Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
3,848
datasets available to search
ShareScore release 0.9.0
Dataset results
3,848 results for “sources”
Source-Sink Dynamics of Garlic Mustard Invasion at Harvard Forest 2004-2006
We are investigating the role of disturbance history, source-sink metapopulation dynamics, and genetic adaptation to different canopy environments in the invasion success of Alliaria petiolata in the New England understory. In these studies, we employ a combination of population-level experiments, metapopulation studies, and landscape-level historical analyses. We have initiated a statewide survey of forested locations to determine the presence and absence of garlic mustard with respect to major pathways of invasion via past and present habitat disturbances. More explicit spatial analyses will test whether past sites of open canopy could have acted as corridors for the spread of invasive populations to present locations. This broad analysis will enable us to evaluate the generality of emerging conclusions from our population-level studies for predicting the broad-scale factors controlling the distribution, pattern and abundance of non-native species. At the population-level, we have begun long term demographic modeling of A. petiolata and experiments to test for habitat-specific natural selection on physiological, phenological, and allocational traits in sites with different canopy structure (edge or understory). In long term demographic analyses, we will use matrix population models to determine the relative contributions of subpopulations in different habitats to forest invasion. In ongoing reciprocal transplant studies, we will asses whether the maternal source habitat of a propagule contributes to its germination and survival in contrasting habitats.
Data for Table S10 of the article "Source-to-sink aeolian fluxes from arid landscape dynamics in the Lut Desert"
<p>Exhaustive list of the 227 individual denudation rates in arid areas compiled to estimate median denudation rate and sediment discharge for the internal river system of the Lut watershed.</p>
Dataset related to the manuscript: "An open-source integrated framework for the automation of citation collection and screening in systematic reviews"
<p>Dataset related to the manuscript: “An open-source integrated framework for the automation of citation collection and screening in systematic reviews”, to be used together with the code stored at https://github.com/AD-Papers-Material/BART_SystReviewClassifier to reproduce the results.</p> <p>There are three datasets:<br> - The Record data collected from the online scientific databases;<br> - The session journal which describes the search session, i.e., how many records were collected and from which source, for each query/session pairs.<br> - The session data which is the outcome of the classification and review tasks;</p>
OpenFOAM cases of the paper "Development and validation of an open-source CFD model for the efficiency assessment of data centers"
<p>This dataset contains the<em> underling data</em> for the paper "Development and validation of an open-source CFD model for the efficiency assessment of data centers”, submitted for the consideration and open review in Open Research Europe (ORE).</p> <p><strong>Validation1.tar.xz:</strong> OpenFOAM files and scripts for the simulation of flow and thermal structures in an enclosed environment (Wang and Chen, 2009).</p> <p><em>Wang, Miao; Chen, Qingyan (2009). Assessment of Various Turbulence Models for Transitional Flows in an Enclosed Environment (RP-1271). HVAC&R Research, 15(6), 1099–1119. doi:10.1080/10789669.2009.10390881</em></p> <p><strong>Validation2-kOmegaSSTModel.tar.xz:</strong> OpenFOAM files and scripts for the simulation of forced convection in a room (Zhang et al. 2007) using k-omega SST turbulence model. </p> <p><em>Zhao Zhang, Wei Zhang, Zhiqiang John Zhai & Qingyan Yan Chen (2007) Evaluation of Various Turbulence Models in Predicting Airflow and Turbulence in Enclosed Environments by CFD: Part 2—Comparison with Experimental Data from Literature, HVAC&R Research, 13:6, 871-886, DOI: 10.1080/10789669.2007.10391460</em></p> <p><strong>Validation2-RNGkEpsilonModel.tar.xz:</strong> OpenFOAM files and scripts for the simulation of forced convection in a room (Zhang et al. 2007) using RNG k-epsilon turbulence model. </p> <p><em>Zhao Zhang, Wei Zhang, Zhiqiang John Zhai & Qingyan Yan Chen (2007) Evaluation of Various Turbulence Models in Predicting Airflow and Turbulence in Enclosed Environments by CFD: Part 2—Comparison with Experimental Data from Literature, HVAC&R Research, 13:6, 871-886, DOI: 10.1080/10789669.2007.10391460</em></p> <p><strong>Validation3.tar.xz:</strong> OpenFOAM files and scripts for the simulation of strong natural convection in a model fire room (Murakami et al. 1995).</p> <p><em>Murakami, S., S. Kato, and R. Yoshie. 1995. Measurement of turbulence statistics in a model fire room by LDV. ASHRAE Transactions 101(2):287–301.</em></p> <p><strong>Validation4.tar.xz:</strong> OpenFOAM files and scripts for the simulation of thermal distribution in an open-aisle data center (Abdelmaksoud et al. 2013).</p> <p><em>W.A. Abdelmaksoud, T.Q. Dang, H. Ezzat Khalifa, R.R. Schmidt Improved computational fluid dynamics model for open-aisle air-cooled data center simulations J. Electron. Packag., 135 (2013), pp. 030901-30913</em></p> <p><strong>Results_Validation1.tar.xz:</strong> Simulation results of the Validation case 1.</p> <p><strong>Results_Validation2.tar.xz:</strong> Simulation results of the Validation case 2.</p> <p><strong>Results_Validation3.tar.xz:</strong> Simulation results of the Validation case 3.</p> <p><strong>Results_Validation4.tar.xz:</strong> Simulation results of the Validation case 4.</p> <p><strong>layout.csv:</strong> Input file for the Validation case 4.</p>
Single-cycle, 643-mW average power THz source based on tilted pulse front in lithium niobate
<p>This data set is associated with the aforementioned paper.</p> <p>The data and the Jupyter notebooks (Python) to reproduce the figures in this paper can be downloaded below. To run a Jupyter notebook as a beginner, it is easiest to download and install anaconda, a Python environment that comes with many packages preinstalled and also offers Jupyter lab/notebook. It is available at <a href="https://www.anaconda.com/download" target="_blank" rel="noopener">https://www.anaconda.com/download</a>.</p> <h2>Fig.01</h2> <p><strong>Fig.01_literature_lithium_niobate_sources.csv</strong> contains a summary table of the last decades of published THz power values obtained with lithium niobate in the tilted pulse front geometry. The accompanying jupyter notebook allows to reproduce the figure that was used in the paper.</p> <h2>Fig.03</h2> <p>Each individual data frame (df), which is saved as an HDF file in the .zip file, contains a "power curve" measurement (i.e. measured THz power as a function of the applied pump power). Whenever a parameter is changed, all positions, angles, THz power and cryostat parameters are saved.</p> <ul> <li><strong>x1 </strong>is the position of the last mirror before the transmission grating (parallel to the pump beam direction before the crystal) in [mm]</li> <li><strong>x2 </strong>is the position of the first imaging lens in direction of the pump beam propagation direction before the crystal in [mm]</li> <li><strong>x3 </strong>is the position of the second imaging lens in the direction of the pump beam propagation direction before the crystal in [mm]</li> <li><strong>x4 </strong>is the position of the cryostat in the direction of the pump beam before reaching the crystal in [mm]</li> <li><strong>y0 </strong>is the position of the cryostat in the perpendicular direction of the pump beam before reaching the crystal in [mm]</li> <li><strong>α0 </strong>is the angle of the lambda/2 waveplate that allows the pump power to be varied at the crystal in [°]</li> <li><strong>α1 </strong>is the angle of the last mirror before the grating in [°]</li> <li><strong>α2 </strong>is the angle of the transmission grating in [°]</li> <li><strong>thz_power_W </strong>is the obtained power obtained from the Ophir 3A-P-THz power meter in [W]</li> <li><strong>temperature_setpoint_K </strong>is the LakeShore cryostat controller setpoint in [K]</li> <li><strong>temperature_K </strong>is the temperature read from the sensor on the cooling finger (above the crystal) in [K]</li> <li><strong>heater_output </strong>is the amount of power in [%] delivered to the resistive heating element inside the cryostat. 100% corresponds to about 50 W. Its value is controlled by an internal PID loop of the cryostat controller, which tries to stabilize <strong>temperature_K </strong>to <strong>temperature_setpoint_K</strong></li> <li><strong>pump_power </strong>is the average laser power reaching the crystal in [W]. It was calibrated before obtaining the data set by characterizing the lambda/2 waveplate angle <strong>α0</strong> to the value of an NIR power meter just before the cryostat.</li> <li><strong>repetition_rate</strong> is the repetition rate of the laser in [Hz]</li> </ul> <p>As an example, below is one line (for one pump power) of such a data frame:</p> <table> <tbody> <tr> <td> </td> <th>x1</th> <th>x2</th> <th>x3</th> <th>x4</th> <th>y0</th> <th>α0</th> <th>α1</th> <th>α2</th> <th>thz_power_W</th> <th>temperature_setpoint_K</th> <th>temperature_K</th> <th>heater_output</th> <th>pump_power</th> <th>repetition_rate</th> </tr> <tr> <td>0</td> <td>-12.000005</td> <td>2.500039</td> <td>9.100015</td> <td>-5.0</td> <td>-2.0</td> <td>35.905660</td> <td>25.68</td> <td>-23.3</td> <td>0.006000</td> <td>80.0</td> <td>79.883</td> <td>4.4</td> <td>20.0</td> <td>40000.0</td> </tr> </tbody> </table> <p>10 of such power curves were obtained at 100 kHz and 40 kHz and can be found in the respective zip-file.</p> <p> </p> <p><strong>Literature_Power_Efficiency.zip</strong> contains digitzed power and efficiency values from the following references:</p> <ol> <li>X. Wu, D. Kong, S. Hao, et al., "Generation of 13.9-mJ Terahertz Radiation from Lithium Niobate Materials," Advanced Materials 35, 2208947 (2023).</li> <li> <p>P. L. Kramer, M. K. R. Windeler, K. Mecseki, et al., "Enabling high repetition rate nonlinear THz science with a kilowatt-class sub-100 fs laser source," Opt. Express 28, 16951 (2020).</p> </li> <li> <p>T. Kroh, T. Rohwer, D. Zhang, et al., "Parameter sensitivities in tilted-pulse-front based terahertz setups and their implications for high-energy terahertz source design and optimization," Opt. Express, OE 30, 24186–24206 (2022).</p> </li> <li> <p>B. Zhang, Z. Ma, J. Ma, et al., "1.4-mJ High Energy Terahertz Radiation from Lithium Niobates," Laser & Photonics Reviews 15, 2000295 (2021).</p> </li> </ol> <p> </p> <h2>Fig.04</h2> <p><strong>EOS_dfs.p</strong> is a pickle file, contain electro-optic sampling traces, which are already averaged for various pump powers at 40 kHz repetition rate.</p>
Technical potential of ground-source heat pumps for Western Switzerland
<p>This dataset contains an estimation of the technical potential of shallow ground-source heat pumps (GSHPs) for Western Switzerland, at a spatial resolution of 200 x 200 m<sup>2</sup>. The technical potential is hereby defined as the maximum energy that could be extracted from GSHP systems in case of their dense deployment, such as to <strong>avoid the over-exploitation</strong> of the heat capacity of the ground. We consider GSHPs with <strong>vertical closed-loop borehole heat exchangers</strong> (BHE) installed at depths of 50 - 200 m. The dataset covers around 80,000 property units (parcels) in the Swiss Cantons of Vaud and Geneva, excluding only the areas of the Alps and the Jura mountains.</p> <p>The estimated potential accounts for:</p> <ul> <li>Norms for geothermal installations set by the Swiss Society of Engineers and Architects (SIA 384/6)</li> <li>Thermal interferences between neighbouring boreholes and their impact on the temperature change in the ground</li> <li>Topographic Landscape data to assess the available area for BHE installation</li> </ul> <p>The methodology used to generate the data is described in:</p> <p>Walch, Alina, Nahid Mohajeri, Agust Gudmundsson, and Jean-Louis Scartezzini. ‘Quantifying the Technical Geothermal Potential from Shallow Borehole Heat Exchangers at Regional Scale’. <em>Renewable Energy</em> 165 (2021): 369–80. <a href="https://doi.org/10.1016/j.renene.2020.11.019">https://doi.org/10.1016/j.renene.2020.11.019</a>.</p> <p><strong>Dataset description</strong></p> <p>As the data is targeted to large-scale applications and potential studies, it is shared in the format of <strong>pixels of 200 x 200 m<sup>2</sup></strong>. Upon request it can be provided at different aggregation levels, as it is generated at the resolution of individual building units (parcels). The potential is provided as <strong>annual</strong> <strong>values</strong>, and it can be converted to monthly values using the provided heating degree weights. For each pixel of 200 x 200 m<sup>2</sup>, we provide the following variables:</p> <ul> <li>Annual total technical heat extraction potential (in MWh)</li> <li>Potential heat delivered <em>to buildings </em>(heat pump output), assuming a heat pump performance (COP) of 4.5 (in MWh)</li> <li>Available area for GSHP installation (in m<sup>2</sup>)</li> <li>Number of installed boreholes </li> <li>Average heat extraction rate (in W/m)</li> <li>Average borehole depth (in m)</li> <li>Average borehole spacing within the parcels located in the pixel (in m)</li> <li>Heating degree weights (i.e. heat demand variation) for each month</li> </ul> <p>A description of the metadata is provided in the document <em>gshp_VD_GE_metadata_V1.pdf.</em></p> <p>This work is part of the PhD Thesis of Alina Walch. </p>
Malawi probabilistic seismic hazard analysis (PSHA) using the Malawi Seismogenic Source Model (MSSM). Supplementary Files v1.1
<p>Updated (October 2022) version of supplementary files for running probabilistic seismic hazard analysis (PSHA) MATLAB codes for Malawi. The PSHA codes themselves (v1.0) are available at: https://doi.org/10.5281/zenodo.7265781and the most recent version will be available on GitHub at: https://github.com/jack-williams1/Malawi_PSHA. Note the variables stored here are not stored on GitHub due to the file size.</p> <p>Includes both input files for performing PSHA and output ground motions for plotting PSHA results.</p> <p>Files are:</p> <ul> <li>malawi_Vs30_active.txt: Input USGS slope-based Vs30 values for Malawi (Wald and Allen 2007)</li> <li>EQCAT_comb.mat: MSSM Direct catalog for all possible rupture weightings (stored as MATLAB variable)</li> <li>GM_MSSM_em_20221027: Ground motions for plotting PSHA maps (stored as MATLAB variable)</li> <li>GM_MSSM_20221021.mat: Ground motions needed for plotting PSHA-site analysis figures (stored as MATLAB variable)</li> <li>mssm_comb.mat: Matlab file for combined MSSM Direct and Adapted MSSM catalogs (stored as MATLAB variable)</li> <li>MSSM_Catalog_Adapted_em.mat: Adapated MSSM event catalog (stored as MATLAB variable)</li> <li>syncat_bg.mat: Areal source stochastic event catalog (stored as MATLAB variable)</li> </ul> <p>Further descriptions of these files and how to use them are provided on Github. An open-access manuscript describing the PSHA is available at: </p> <p>Williams J. N., Werner M. J., Goda K., Wedmore L. N. J., De Risi R., Biggs J., Mdala H., Dulanya Z., Fagereng Å, Mphepo F., Chindandali P. (2023). Fault-based probabilistic seismic hazard analysis in regions with low strain rates and a thick seismogenic layer: a case study from Malawi, Geophysical Journal International, Volume 233, Issue 3, June 2023, Pages 2172–2206, <a href="https://doi.org/10.1093/gji/ggad060">https://doi.org/10.1093/gji/ggad060</a></p> <p>Please reference this publication along with this repository when using these data.</p> <p>USGS vs30 value compilation described in:</p> <p>Allen, T. I., and Wald, D. J., 2009, On the use of high-resolution topographic data as a proxy for seismic site conditions (Vs30), Bulletin of the Seismological Society of America, 99, no. 2A, 935-943.</p> <p> </p>
DisVis-based filtering of contacts from co-evolution data (or other sources)
<p>Dataset described in the manuscript: <em>Improving the Quality of Co-evolution Intermolecular Contact Prediction with DisVis</em>Siri Camee van Keulen, Alexandre M.J.J. Bonvin</p> <p>Details about the data set can be found at: https://github.com/haddocking/contact-filtering</p> <p>This archive contains in addition all the models generated with HADDOCK.</p>
Calculated moisture sources for the Yangtse River Valley for past, present and future climate using a Lagrangian moisture source diagnostic
<p>This dataset contains calculated moisture sources for the Yangtse River Valley (110–122°E and 27–33°N, eastern China) for past, present and future climate using a Lagrangian moisture source diagnostic. The dataset comprises gridded monthly moisture source data files and monthly time series files for a Last Glacial Maximum (LGM) simulation and a Pre-Industrial reference simulation (PRE) with CAM5.1 using prescribed sea surface temperatures, and a control simulation (CTL, 2001-2010) and a climate scenario run with representative concentration pathway 6 (RCP, 2061-2070) with the coupled NorESM-1M model. Each file covers a 10-year time period, computed with the Lagrangian moisture source diagnostic WaterSip (Sodemann et al., 2008).</p>
Indicators of Contaminant Sources, PFAS, and Water Quality in Ellerbe Creek and New Hope Creek, NC (2019-2022)
Thousands of chemical contaminants are found in urban stream globally. This is a dataset of water quality measures of (1) compounds that are indicative of specific contaminant sources, (2) common water quality measures [trace metals, major ions, nutrients], and (3) PFAS. Sampling was conducted in Ellerbe Creek and New Hope Creek in the Durham and Orange counties of North Carolina. Biweekly and synoptic sampling was undertaken to explore spatial and temporal variation in water concentrations.
Plant community data at water sources, Mpala Research Centre, Kenya (2015-2017)
Data package contains four datasets of plant measurements taken at Mpala Research Centre, Laikipia County, Kenya from November 2015-September 2017. Additional code for data analysis is also provided as part of the publication `The effects of herbivore aggregations at water sources on savanna plants differ across soil and climate gradients`.
Wisconsin Lake Plants - multi source database of lake plant abundance 1930 - 2004
This data set provides sampling-point by sampling-point macrophyte data for lakes sampled by a number of agencies in Wisconsin. The relational tables in this dataset were originally used to generate plant community tables. This dataset contains detailed and recent data from approximately the 1970s onward. Sampling timing and intensity varied. Table DATSOUR contains sources of data for tables AQUAPLT2 and LAKEHAB. Table AQUAPLT2 gives an estimate of plant density at each sample point. Table MAXDEPLNG has initial lake parameters derived from data in AQUAPLT2 and LAKEHAB Table LAKEHAB contains habitat characteristics at macrophyte sampling locations. Table PLTNAME has species information for plants in tables AQUAPLT2 and LAKESPEC. Table LAKES contains information for lakes included in this dataset. Table COUNTY contains information associated with the counties where the lakes in the AQUAPLT2 dataset and the LAKESPEC dataset are located. . Sampling Frequency: varies Number of sites: 1938
15N (Nitrate and Ammonium) Uptake by Spartina Alterniflora Sourced from Plum Island Estuary, MA and North Inlet, SC
15N uptake measurements were taken from Spartina alterniflora plants grown from seeds originating from six march locations in PLUM Island Estuary (PIE) and North Inlet, SC. All sites were Spartina alterniflora dominated marsh. 15N uptake was measured after 60 minutes in Nitrate, Ammonium, or control treatments. Seeds were collected in Fall 2022, reared Winter-Summer 2023, and uptake measured in Summer 2023.
Genome-Wide DNA Methylation in Peripheral Blood and Long-Term Exposure to Source-Specific Transportation Noise and Air Pollution: The SAPALDIA Study (Supplementary Data)
<p>The zip file contains supplementary data for the publication - Genome-Wide DNA Methylation in Peripheral Blood and Long-Term Exposure to Source-Specific Transportation Noise and Air Pollution: The SAPALDIA Study, accepted for publication in Environmental Health Perspectives (DOI: 10.1289/EHP6174).</p> <p>The description of the files are noted below:</p> <p><strong>1. Readme File for SAPALDIA Noise and Air Pollution EWAS Single Exposure.zip </strong></p> <p>This zip file contains all the results of the association between source-specific transportation noise (aircraft, railway and road traffic), air pollution (NO<sub>2</sub> and PM<sub>2.5</sub>), and genome-wide DNA methylation, derived from multi-exposure models.</p> <p><strong>SAPALDIA_EWAS_SingleExposure_AircraftLden.txt</strong> contains the results for aircraft noise</p> <p><strong>SAPALDIA_EWAS_SingleExposure_RailwayLden.txt</strong> contains the results for railway noise</p> <p><strong>SAPALDIA_EWAS_SingleExposure_RoadtrafficLden.txt</strong> contains the results for road traffic noise</p> <p><strong>SAPALDIA_EWAS_SingleExposure_NO2.txt</strong> contains the results for nitrogen dioxide</p> <p><strong>SAPALDIA_EWAS_SingleExposure_PM25.txt</strong> contains the results for fine particulate matter</p> <p> </p> <p><strong>General footnote for all files:</strong>SAPALDIA: Swiss cohort study on air pollution and lung and heart diseases in adults. CpG: Cytosine-phosphate-Guanine. CHR: chromosome. SE: standard error. Lden: day-evening-night noise level. NO<sub>2</sub>: nitrogen dioxide. PM<sub>2.5</sub>: particulate matter with aerodynamic diameter <2.5 µm. Beta coefficients represent increase or decrease in DNA methylation per 10 dB increase in aircraft, railway or road traffic Lden or 10 µg/m<sup>3</sup> increase in NO<sub>2</sub> or PM<sub>2.5</sub>. All estimates were from single exposure epigenome-wide linear mixed models, with random intercept at the level of participant. Each model was adjusted for age, sex, educational level, area, and neighborhood socio-economic status, greenness index, smoking status and pack years, exposure to passive smoke, consumption of fruits, vegetables and alcohol, nested study, asthma status, survey, source-specific noise truncation indicator (for Lden models) and leukocyte composition. In a preliminary step, DNA methylation β-values were regressed on the Illumina control probe-derived first 30 principal components to correct for correlation structures and technical bias, and residuals of these regressions covering 430,477 CpGs were used as the technical bias-corrected methylation level at the CpG sites.</p> <p>Extreme values of the residuals (lying beyond three times the interquartile range below the first quartile and above the third quartile at each CpG site) were replaced with their corresponding detection threshold value (“modified winsorization”). The “winsorized” data were then used as the dependent variables in the epigenome-wide association study.</p> <p> </p> <p><strong>2. Readme File for SAPALDIA Noise and Air Pollution EWAS Multi Exposure.zip </strong></p> <p>This zip file contains all the results of the association between source-specific transportation noise (aircraft, railway and road traffic), air pollution (NO<sub>2</sub> and PM<sub>2.5</sub>), and genome-wide DNA methylation, derived from multi-exposure models.</p> <p><strong>SAPALDIA_EWAS_MultiExposure_AircraftLden.txt</strong> contains the results for aircraft noise</p> <p><strong>SAPALDIA_EWAS_MultiExposure_RailwayLden.txt</strong> contains the results for railway noise</p> <p><strong>SAPALDIA_EWAS_MultiExposure_RoadtrafficLden.txt</strong> contains the results for road traffic noise</p> <p><strong>SAPALDIA_EWAS_MultiExposure_NO2.txt</strong> contains the results for nitrogen dioxide</p> <p><strong>SAPALDIA_EWAS_MultiExposure_PM25.txt</strong> contains the results for fine particulate matter</p> <p><strong>General table footnotes: </strong>SAPALDIA: Swiss cohort study on air pollution and lung and heart diseases in adults. CpG: Cytosine-phosphate-Guanine. CHR: chromosome. SE: standard error. Lden: day-evening-night noise level. NO<sub>2</sub>: nitrogen dioxide. PM<sub>2.5</sub>: particulate matter with aerodynamic diameter <2.5 µm. Beta coefficients represent increase or decrease in DNA methylation per 10 dB increase in aircraft, railway or road traffic Lden or 10 µg/m<sup>3</sup> increase in NO<sub>2</sub> or PM<sub>2.5</sub>. All estimates were from multi-exposure epigenome-wide linear mixed models, with random intercept at the level of participant, and were adjusted for age, sex, educational level, area, and neighborhood socio-economic status, greenness index, smoking status and pack years, exposure to passive smoke, consumption of fruits, vegetables and alcohol, nested study, asthma status, survey, source-specific noise truncation indicator and leukocyte composition. Multi-exposure models included all five exposures (Aircraft, railway, road traffic Lden and respective truncation indicators, NO<sub>2</sub> and PM<sub>2.5</sub>) at the same time. In a preliminary step, DNA methylation β-values were regressed on the Illumina control probe-derived first 30 principal components to correct for correlation structures and technical bias, and residuals of these regressions covering 430,477 CpGs were used as the technical bias-corrected methylation level at the CpG sites. Extreme values of the residuals (lying beyond three times the interquartile range below the first quartile and above the third quartile at each CpG site) were replaced with their corresponding detection threshold value (“modified winsorization”). The “winsorized” data were then used as the dependent variables in the epigenome-wide association study.</p>
Compact continuum source finding for next generation radio surveys
<p>This is a data set that accompanies the paper "Compact continuum source finding for next generation radio surveys" (2012MNRAS.422.1812H)</p> <p>The image files and source catalogues contained here were used to test the completeness and false detection rate of a number of source finding algorithms including: Aegean, Selavy, Sfind, SExtractor, and IMSAD. These data can be used to assess the performance of future source finding codes, and to verify the the performance of code during development.</p>
Variant, Metabolite and Source Data for: Population genomics uncover loci for trait improvement in the indigenous African cereal tef (Eragrostis tef)
<p>These files contain the variant and metabolome for a collection of 220 tef (<em>Eragrsotis tef)</em> accessions from an ethiopian diversity panel. The accessions were assembled and managed by the Ethiopian Institute of Agricultural Research (EIAR, Ethiopia). The variant data was produced at the John Innes Centre (UK). The metabolome data was produced at Aberystwyth University (UK). These dataset are described in Jones et al. (2024), <em>bioRxiv</em>, https://doi.org/10.1101/2024.09.30.615331. The source data for main figures in the publication are also included.</p> <p>The submission contains</p> <ol> <li>EIAR_filtered.vcf.gz: This is the variant data obtained from alignment of Illumina reads from all 220 teff accessions to the reference assembly of tef (Dabbi). Low quality variants were filtered out. This variant data was used for constructing the phylogenetic relationship between the accessions. The samples names corresponds to the DNA code in Supplementary Table S10 (Jones et al, 2024).</li> <li>pooled_EIAR_filtered.vcf.gz: After the phylogentic analysis described above, reads from accessions that were found to be genetically redundant were pooled before variant calling. This file was used for the SNP GWAS analysis. The samples names corresponds to the DNA code in Supplementary Table S10 (Jones et al, 2024).</li> <li> Metabolite_Profile.xlxs (source data for Figure 5): This file contains m/z feature intensities from untargeted metabolite fingerprinting using Flow Infusion Electrospray High-resolution Mass Spectrometry (FIE-HRMS). The sample names contains a combination of Location code and Plot number in Supplementary Table S10 e.g AT plot 1, CD plot 1, DZ plot 1, where AT, CD and DZ represent Alem Tena, Chefe Donsa and Debre Zeit, respectively. The data was used for the partial least squares discriminant analysis and differentially accumulated metabolites analysis presented in Figure 5.</li> <li>Source data: Numerical source data for graphs and charts in Figures 3 - 7.</li> <li>Tsedey TT2 Sequence from Improved Assembly: The 4A and 4B sequences around the TT2 orthologue in tef from the improved PacBio-based chromosome-scale assembly of tef. These sequences were used for plotting the LTR Copia alignments presented in Supplementary Figure 9. We thank Corteva for pre-publication access to this improved Tsedey genome assembly.</li> </ol>
AN OPEN-SOURCE, THREE-DIMENSIONAL GROWTH MODEL OF THE MANDIBLE
<p>This repository contains all geometrical data and metadata belonging to the paper AN OPEN-SOURCE, THREE-DIMENSIONAL GROWTH MODEL OF THE MANDIBLE by the MAGIC Amsterdam research consortium. The following contents are uploaded:</p><p><strong>shapeVectors_original.csv</strong> | shape vectors of the original data<br><strong>shapeVectors_rescaled.csv</strong> | shape vectors of the rescaled data<br>678 x 62589 matrices where the rows are samples and the columns are shape vectors. The shape vectors are formatted<i> [x1, x2, x3, ..., y1, y2, y3, ..., z1, z2, z3, ...].</i></p><p><strong>PCA_coeff_original.csv</strong> | principal component coefficients of the original data<br><strong>PCA_coeff_rescaled.csv</strong> | principal component coefficients of the rescaled data<br>62589 x 677 matrices where each row of these matrices is a variable (x-, y-, or z-coordinate of a vertex) and each column is a principal component.</p><p><strong>PCA_score_original.csv</strong> | principal component scores of the original data<br><strong>PCA_score_rescaled.csv</strong> | principal component scores of the rescaled data<br>678 x 677 matrices where rows correspond to samples and columns correspond to principal components.</p><p><strong>PCA_latent_original.csv</strong> | principal component variances of the original data<br><strong>PCA_latent_rescaled.csv</strong> | principal component variances of the rescaled data<br>677 x 1 vectors where each element is an eigenvalue of a principal component.</p><p><strong>PCA_mu_original.csv</strong> | mean of the original data<br><strong>PCA_mu_rescaled.csv</strong> | mean of the rescaled data<br>1 x 62589 vectors that represent the average shape vector. All (centered) data can be reconstructed as follows: <i>shapeVectors = PCA_score * PCA_coeff' + PCA_mu.</i></p><p><strong>PCA_standardDeviations_original.csv</strong> | standard deviations of each sample for each principal component of the original data.<br><strong>PCA_standardDeviations_rescaled.csv</strong> | standard deviations of each sample for each principal component of the rescaled data.<br>677 x 678 matrices where the rows are principal components and the columns are samples. The standard deviations were calculated as follows: <i>PCA_standardDeviations = PCA_score' ./ sqrt(PCA_latent).</i></p><p><strong>metadata.csv</strong> | This matrix contains the age in years (first column) and biological sex (second column, 1 = male and 2 = female) for all samples (rows).</p><p><strong>connectivityList.csv</strong> | This matrix defines the mesh of the 3D model of the mandible. The vector in each row represents which vertices define a triangle. Indexing starts at 0, so for use in e.g. Matlab, add 1 to all elements.</p>
Enhanced Biosafety of the Sleeping Beauty Transposon System by Using mRNA as Source of Transposase to Efficiently and Stably Transfect Retinal Pigment Epithelial Cells
<p>Raw data of the publication "Enhanced Biosafety of the Sleeping Beauty Transposon System by Using mRNA as Source of Transposase to Efficiently and Stably Transfect Retinal Pigment Epithelial Cells".</p> <p>Abstract: Neovascular age-related macular degeneration (nvAMD) is characterized by choroidal<br> neovascularization (CNV), which leads to retinal pigment epithelial (RPE) cell and photoreceptor<br> degeneration and blindness if untreated. Since blood vessel growth is mediated by endothelial cell<br> growth factors, including vascular endothelial growth factor (VEGF), treatment consists of repeated,<br> often monthly, intravitreal injections of anti-angiogenic biopharmaceuticals. Frequent injections are<br> costly and present logistic difficulties; therefore, our laboratories are developing a cell-based gene<br> therapy based on autologous RPE cells transfected ex vivo with the pigment epithelium derived factor<br> (PEDF), which is the most potent natural antagonist of VEGF. Gene delivery and long-term expression<br> of the transgene are enabled by the use of the non-viral Sleeping Beauty (SB100X) transposon system<br> that is introduced into the cells by electroporation. The transposase may have a cytotoxic effect and a<br> low risk of remobilization of the transposon if supplied in the form of DNA. Here, we investigated<br> the use of the SB100X transposase delivered as mRNA and showed that ARPE-19 cells as well as<br> primary human RPE cells were successfully transfected with the Venus or the PEDF gene, followed<br> by stable transgene expression. In human RPE cells, secretion of recombinant PEDF could be detected<br> in cell culture up to one year. Non-viral ex vivo transfection using SB100X-mRNA in combination<br> with electroporation increases the biosafety of our gene therapeutic approach to treat nvAMD while<br> ensuring high transfection efficiency and long-term transgene expression in RPE cells.</p>
Open-source traffic and CO2 emission dataset for commercial aviation
<p>This record is a global open-source passenger air traffic dataset primarily dedicated to the research community. <br>It gives a seating capacity available on each origin-destination route for a given year, 2019, and the associated aircraft and airline when this information is available. </p> <p>Context on the original work is given in the related articles (<a href="https://doi.org/10.59490/joas.2024.7365">https://doi.org/10.59490/joas.2024.7365,</a> <a href="https://doi.org/10.59490/joas.2023.7201">https://doi.org/10.59490/joas.2023.7201)</a> and on the associated GitHub page (<a href="https://github.com/AeroMAPS/AeroSCOPE/">https://github.com/AeroMAPS/AeroSCOPE/</a>).<br>A simple data exploration interface will be available at <a href="www.aeromaps.eu/aeroscope">www.aeromaps.eu/aeroscope.</a><br>The dataset was created by aggregating various available open-source databases with limited geographical coverage. It was then completed using a route database created by parsing Wikipedia and Wikidata, on which the traffic volume was estimated using a machine learning algorithm (XGBoost) trained using traffic and socio-economical data.<br> </p> <h4><br><strong>1- DISCLAIMER</strong></h4> <p><br>The dataset was gathered to allow highly aggregated analyses of the air traffic, at the continental or country levels. At the route level, the accuracy is limited as mentioned in the associated article and improper usage could lead to erroneous analyses. </p> <p>Although all sources used are open to everyone, the Eurocontrol database is only freely available to academic researchers. It is used in this dataset in a very aggregated way and under several levels of abstraction. As a result, it is not distributed in its original format as specified in the contract of use.</p> <p>As a general rule, we decline any responsibility for any use that is contrary to the terms and conditions of the various sources that are used. In case of commercial use of the database, please contact us in advance.</p> <h4><br><strong>2- DESCRIPTION</strong></h4> <p>Each data entry represents an (Origin-Destination-Operator-Aircraft type) tuple.</p> <p><em>Please </em>refer<em> to </em>the<em> support article for more details (see above).</em></p> <p>The dataset contains the following columns:</p> <ul> <li>"First column" : index</li> <li><strong>airline_iata : </strong>IATA code of the operator in nominal cases. An ICAO -> IATA code conversion was performed for some sources, and the ICAO code was kept if no match was found.</li> <li><strong>acft_icao : </strong>ICAO code of the aircraft type</li> <li><strong>acft_class : </strong>Aircraft class identifier, own classification. <ul> <li>WB: Wide Body</li> <li>NB: Narrow Body</li> <li>RJ: Regional Jet</li> <li>PJ: Private Jet</li> <li>TP: Turbo Propeller</li> <li>PP: Piston Propeller</li> <li>HE: Helicopter</li> <li>OTHER</li> </ul> </li> <li><strong>seymour_proxy: </strong>Aircraft code for Seymour Surrogate (https://doi.org/10.1016/j.trd.2020.102528), own classification to derive proxy aircraft when nominal aircraft type unavailable in the aircraft performance model.</li> <li><strong>source: </strong>Original data source for the record, before compilation and enrichment. <ul> <li>ANAC: Brasilian Civil Aviation Authorities</li> <li>AUS Stats: Australian Civil Aviation Authorities</li> <li>BTS: US Bureau of Transportation Statistics T100</li> <li>Estimation: Own model, estimation on Wikipedia-parsed route database</li> <li>Eurocontrol: Aggregation and enrichment of R&D database</li> <li>OpenSky</li> <li>World Bank</li> </ul> </li> <li><strong>seats: </strong>Number of seats available for the data entry, AFTER airport residual scaling</li> <li><strong>n_flights: </strong>Number of flights of the data entry, when available</li> <li><strong>iata_departure</strong>, <strong>iata_arrival : </strong>IATA code of the origin and destination airports. Some BTS inhouse identifiers could remain but it is marginal.</li> <li><strong>departure_lon</strong><em>, </em><strong>departure_lat</strong><em>, </em><strong>arrival_lon</strong><em>, </em><strong>arrival_lat : </strong>Origin and destination coordinates, could be NaN if the IATA identifier is erroneous</li> <li><strong>departure_country, arrival_country</strong>: Origin and destination country ISO2 code. <strong>WARNING: </strong>disable NA (Namibia) as default NaN at import</li> <li><strong>departure_continent, arrival_continent: </strong>Origin and destination continent code. <strong>WARNING: </strong>disable NA (North America) as default NaN at import</li> <li><strong>seats_no_est_scaling: </strong>Number of seats available for the data entry, BEFORE airport residual scaling</li> <li><strong>distance_km: </strong>Flight distance (km)</li> <li><strong>ask: </strong>Available Seat Kilometres</li> <li><strong>rpk: </strong>Revenue Passenger Kilometres (simple calculation from ASK using IATA average load factor)</li> <li><strong>fuel_burn_seymour: </strong>Fuel burn <em>per flight</em> (kg) when seymour proxy available</li> <li><strong>fuel_burn: </strong>Total fuel burn of the data entry (kg)</li> <li><strong>co2: </strong>Total CO2 emissions of the data entry (kg)</li> <li><strong>domestic: </strong>Domestic/international boolean (Domestic=1, International=0)</li> </ul> <p> </p> <h4><strong>3- Citation</strong></h4> <p>Please cite the support paper instead of the dataset itself. </p> <blockquote> <p>Salgas, A., Sun, J., Delbecq, S., Planès, T., & Lafforgue, G. (2024). Compilation and Applications of an Open-Source Dataset on Global Air Traffic Flows and Carbon Emissions. <em>Journal of Open Aviation Science</em>. <a href="https://doi.org/10.59490/joas.2024.7365">https://doi.org/10.59490/joas.2023.7201</a></p> </blockquote>
MarFERReT: an open-source, version-controlled reference library of marine microbial eukaryote functional genes
<p>Metatranscriptomics generates large volumes of sequence data about transcribed genes in natural environments. Taxonomic annotation of these datasets depends on availability of curated reference sequences. For marine microbial eukaryotes, current reference libraries are limited by gaps in sequenced organism diversity and barriers to updating libraries with new sequence data, resulting in taxonomic annotation of only about half of eukaryotic environmental transcripts. Here, we introduce version 1.0 of the Marine Functional EukaRyotic Reference Taxa (MarFERReT), an updated marine microbial eukaryotic sequence library with a version-controlled framework designed for taxonomic annotation of eukaryotic metatranscriptomes. We gathered 902 marine eukaryote genomes and transcriptomes from multiple sources and assessed these candidate entries for sequence quality and cross-contamination issues, selecting 800 validated entries for inclusion in the library. MarFERReT v1 contains reference sequences from 800 marine eukaryotic genomes and transcriptomes, covering 453 species- and strain-level taxa, totaling nearly 28 million protein sequences with associated NCBI and PR2 Taxonomy identifiers and Pfam functional annotations. An accompanying MarFERReT project repository hosts containerized build scripts, documentation on installation and use case examples, and information on new versions of MarFERReT.<br><br>MarFERReT is linked to a code repository hosting containerized build scripts, documentation on installation and use case examples, and information on new versions of MarFERReT here: <a href="https://github.com/armbrustlab/marferret">https://github.com/armbrustlab/marferret</a></p> <p>The raw source data for the 902 candidate entries considered for MarFERReT v1.1.1, including the 800 accepted entries, are available for download from their respective online locations. The source URL for each of the entries is listed here in MarFERReT.v1.1.1.entry_curation.csv, and detailed instructions and code for downloading the raw sequence data from source are available in the MarFERReT code repository (<a href="https://github.com/armbrustlab/marferret/blob/main/docs/process_clean_marmicrodb.log.sh">link</a>). </p> <p>This repository release contains MarFERReT database files from the v1.1.1 MarFERReT release using the following MarFERReT library build scripts: <strong>assemble_marferret.sh</strong>, <strong>pfam_annotate.sh</strong>, and <strong>build_diamond_db.sh</strong><br><br>The following MarFERReT data products are available in this repository:</p> <p><strong>MarFERReT.v1.1.1.metadata.csv</strong><br>This CSV file contains descriptors of each of the 902 database entries, including data source, taxonomy, and sequence descriptors. Data fields are as follows:</p> <ol> <li><strong>entry_id</strong>: Unique MarFERReT sequence entry identifier.</li> <li><strong>accepted: </strong>Acceptance into the final MarFERReT build (Y/N). The Y/N values can be adjusted to customize the final build output according to user-specific needs.</li> <li><strong>marferret_name</strong>: A human and machine friendly string derived from the NCBI Taxonomy organism name; maintaining strain-level designation wherever possible.</li> <li><strong>tax_id</strong>: The NCBI Taxonomy ID (taxID).</li> <li><strong>pr2_accession</strong>: Best-matching PR2 accession ID associated with entry</li> <li><strong>pr2_rank</strong>: The lowest shared rank between the entry and the pr2_accession</li> <li><strong>pr2_taxonomy</strong>: PR2 Taxonomy classification scheme of the pr2_accession</li> <li><strong>data_type</strong>: Type of sequence data; transcriptome shotgun assemblies (TSA), gene models from assembled genomes (genome), and single-cell amplified genomes (SAG) or transcriptomes (SAT).</li> <li><strong>data_source</strong>: Online location of sequence data; the Zenodo data repository (<a href="../">Zenodo</a>), the datadryad.org repository (<a href="http://datadryad.org/">datadryad.org</a>), MMETSP re-assemblies on Zenodo (MMETSP)17, NCBI GenBank (<a href="https://www.ncbi.nlm.nih.gov/genbank/">NCBI</a>), JGI Phycocosm (<a href="https://phycocosm.jgi.doe.gov/phycocosm/home">JGI-Phycocosm</a>), the TARA Oceans portal on Genoscope (<a href="http://www.genoscope.cns.fr/tara/">TARA</a>), or entries from the Roscoff Culture Collection through the METdb database repository (<a href="https://metdb.sb-roscoff.fr/metdb/">METdb</a>).</li> <li><strong>source_link</strong>: URL where the original sequence data and/or metadata was collected.</li> <li><strong>pub_year</strong>: Year of data release or publication of linked reference.</li> <li><strong>ref_link</strong>: Pubmed URL directs to the published reference for entry, if available.</li> <li><strong>ref_doi</strong>: DOI of entry data from source, if available.</li> <li><strong>source_filename</strong>: Name of the original sequence file name from the data source.</li> <li><strong>seq_type</strong>: Entry sequence data retrieved in nucleotide (nt) or amino acid (aa) alphabets.</li> <li><strong>n_seqs_raw</strong>: Number of sequences in the original sequence file.</li> <li><strong>source_name:</strong> Full organism name from entry source</li> <li><strong>original_taxID</strong>: Original NCBI taxID from entry data source metadata, if available</li> <li><strong>alias:</strong> Additional identifiers for the entry, if available</li> </ol> <p><br><strong>MarFERReT.v1.1.1.curation.csv</strong><br>This CSV file contains curation and quality-control information on the 902 candidate entries considered for incorporation into MarFERReT v1, including curated NCBI Taxonomy IDs and entry validation statistics. Data fields are as follows:</p> <ol> <li><strong>entry_id:</strong> Unique MarFERReT sequence entry identifier</li> <li><strong>marferret_name: </strong>Organism name in human and machine friendly format, including additional NCBI taxonomy strain identifiers if available.</li> <li><strong>tax_id</strong>: Verified NCBI taxID used in MarFERReT</li> <li><strong>taxID_status</strong>: Status of the final NCBI taxID (Assigned, Updated, or Unchanged)</li> <li><strong>taxID_notes</strong>: Notes on the original_taxID</li> <li><strong>n_seqs_raw</strong>: Number of sequences in the original sequence file</li> <li><strong>n_pfams</strong>: Number of Pfam domains identified in protein sequences</li> <li><strong>qc_flag</strong>: Early validation quality control flags for the following: LOW_SEQS; less than 1,200 raw sequences; LOW_PFAMS; less than 500 Pfam domain annotations.</li> <li><strong>flag_Lasek</strong>: Flag notes from Lasek-Nesselquist and Johnson (2019); contains the flag 'FLAG_LASEK' indicating ciliate samples reported as contaminated in this study.</li> <li><strong>VV_contam_pct</strong>: Estimated contamination reported for MMETSP entries in Van Vlierberghe et al., (2021).</li> <li><strong>flag_VanVlierberghe: </strong>Flag for a high level of estimated contamination, from 'flag_VanVlierberghe' values over 50%: FLAG_VV.</li> <li><strong>rp63_npfams</strong>: Number of ribosomal protein Pfam domains out of 63 total.</li> <li><strong>rp63_contam_pct</strong>: Percent of total ribosomal protein sequences with an inferred taxonomic identity in any lineage other than the recorded identity, as described in the Technical Validation section from analysis of 63 Pfam ribosomal protein domains.</li> <li><strong>flag_rp63</strong>: Flag for a high level of estimated contamination, from 'rp63_contam_pct' values over 50%: FLAG_RP63.</li> <li><strong>flag_sum: </strong>Count of the number of flag columns (`qc_flag`, `flag_Lasek`, `flag_VanVlierberghe`, and `flag_rp63`). All entries with one or more flag are nominally rejected ('accepted' = N); entries without any flags are validated and accepted ('accepted' = Y).</li> <li><strong>accepted: </strong>Acceptance into the final MarFERReT build (Y or N).</li> </ol> <p> </p> <p><strong>MarFERReT.v1.1.1.proteins.faa.gz</strong><br>This Gzip-compressed FASTA file contains the 27,951,013 final translated and clustered protein sequences for all 800 accepted MarFERReT entries. The sequence defline contains the unique identifier for the sequence and its reference (mftX, where 'X' is a ten-digit integer value). </p> <p> </p> <p><strong>MarFERReT.v1.1.1.taxonomies.tab.gz</strong><br>This Gzip-compressed tab-separated file is formatted for interoperability with the DIAMOND protein alignment tool commonly used for downstream analyses and contains some columns without any data. Each row contains an entry for one of the MarFERReT protein sequences in MarFERReT.v1.proteins.faa.gz. Note that 'accession.version' and 'taxid' are populated columns while 'accession' and 'gi' have NA values; the latter columns are required for back-compatibility as input for the DIAMOND alignment software and LCA analysis. </p> <p>The columns in this file contain the following information:</p> <ol> <li><strong>accession</strong>: (NA)</li> <li><strong>accession.version</strong>: The unique MarFERReT sequence identifier ('mftX').</li> <li><strong>taxid</strong>: The NCBI Taxonomy ID associated with this reference sequence.</li> <li><strong>gi</strong>: (NA).</li> </ol> <p> </p> <p><strong>MarFERReT.v1.1.1.proteins_info.tab.gz</strong><br>This Gzip-compressed tab-separated file contains a row for each final MarFERReT protein sequence with the following columns:</p> <ol> <li><strong>aa_id</strong>: the unique identifier for each MarFERReT protein sequence.</li> <li><strong>entry_id</strong>: The unique numeric identifier for each MarFERReT entry.</li> <li><strong>source_defline</strong>: The original, unformatted sequence identifier</li> </ol> <p> </p> <p><strong>MarFERReT.v1.1.1.best_pfam_annotations.csv.gz<br></strong>This Gzip-compressed CSV file contains the best-scoring Pfam annotation for intra-species clustered protein sequences from the 800 validated MarFERReT entries; derived from the hmmsearch annotations against Pfam 34.0 functional domains. This file contains the following fields:</p> <ol> <li><strong>aa_id</strong>: The unique MarFERReT protein sequence ID ('mftX').</li> <li><strong>pfam_name</strong>: The shorthand Pfam protein family name.</li> <li><strong>pfam_id</strong>: The Pfam identifier.</li> <li><strong>pfam_eval</strong>: hmm profile match e-value score</li> <li><strong>pfam_score:</strong> hmm profile match bitscore</li> </ol> <p><br><strong>MarFERReT.v1.1.1.dmnd</strong><br>This binary file is the indexed database of the MarFERReT protein library with embedded NCBI taxonomic information generated by the DIAMOND makedb tool using the build_diamond_db.sh script from the MarFERReT /scripts/ library. This can be used as the reference DIAMOND database for annotating environment sequences from eukaryotic metatranscriptomes. <br><br></p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.