Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
424
datasets available to search
ShareScore release 0.9.0
Dataset results
424 results for “Compilation”
A database of MMS bow shock crossings compiled using machine learning
<p>We use a machine learning approach to automatically identify shock crossings from the Magnetospheric Multiscale (MMS) spacecraft. We compile a database of 2797 crossings including various spacecraft related and shock related parameters for each event. Furthermore, for each event we provide an overview plot containing key parameters of the shock crossing.</p> <p>A Technical report detailing the content of the database can be found at the DOI: http://dx.doi.org/10.1029/2022JA030454</p>
Intermediate data files from the compilation of Economy-wide Material Flow Accounts for the Domestic Extraction of abiotic materials
<p>These files represent a selection of intermediary files from the compilation of material flow accounts on Domestic Extraction (DE) of abiotic materials. The main output of this compilation has been integrated in the UNEP IRP Global Material Flow Database (GMFD).</p> <p>These files include input files (e.g. IDs for data harmonization, or factors for data conversion), as well as output files (e.g. supplementary information, or detailed data accounts before aggregation and integration into the GMFD).</p> <p>Please note: These files are published exclusively for the purpose of making this information available to interested parties in a transparent and orderly fashion, in particular to other research projects who may have use for it. Therefore, these files are not associated with any publication and have not been adjusted or formatted with regard to any publications standards, i.e. they are uploaded exactly as they have been processed in the respective R Github repository of the underlying data compilation.</p> <p>The following description attempts to give a short overview of the respective types of files and their contents. For more detailed information on the data compilation, please refer to chapters 6, 8, and 10 in the <a href="https://resourcepanel.org/sites/default/files/irp_technical_annex_global_material_flows_database.pdf">technical report</a> of the GMFD.</p> <p> </p> <p><strong>Main data output (i.e. detailed material flow accounts)</strong></p> <p><em>DE_met_min_fos_CCC_2021-11-04.csv</em>: Data aggregated to the official categories (CCC/TCCC) used for integration into the GMFD. With IDs, without names (e.g. for materials and countries).</p> <p><em>DE_met_min_fos_CCC_with_names_2021-11-04.csv</em>: Same as above, but with names.</p> <p><em>DE_met_min_fos_Detailed_2021-11-04.csv</em>: Detailed accounts, as compiled, before final aggregation. With IDs, without names.</p> <p><em>DE_met_min_fos_Detailed_with_names_2021-11-04.csv</em>: Same as above, but with names.</p> <p> </p> <p><em>DE_met_min_fos_Detailed_2022-05-24.csv: </em>Slightly revised version from May 2022. But not included in current GMFD version.</p> <p> </p> <p><strong>ID and concordance tables</strong></p> <p><em>ccc_vs_mat_ids.csv:</em> Concordance table for allocation of detailed material accounts to aggregated CCC accounts.</p> <p><em>country_ids.csv:</em> General ID table for country IDs and names</p> <p><em>estimated_ids.csv:</em> Material IDs which are assigned during application of ore estimation factors.</p> <p><em>geo_exist.csv:</em> Table for consistent geographic adjustment of data for specific countries which have disintegrated over time (not including regions like Germany, Yemen, Ethiopia, Sudan, which all have been dealt with individually if necessary).</p> <p><em>material_ids.csv:</em> General ID table for material IDs and names.</p> <p><em>source_country_ids.csv:</em> Concordance table for country IDs (i.e. for allocation of harmonized IDs to source namings/IDs)</p> <p><em>source_material_ids.csv:</em> Concordance table for material IDs (i.e. for allocation of harmonized IDs to source namings/IDs)</p> <p><em>source_unit_ids.csv:</em> Concordance table for unit IDs (i.e. for allocation of harmonized IDs to source namings/IDs)</p> <p><em>unit_ids.csv:</em> General ID table for unit IDs and names.</p> <p> </p> <p><strong>Conversion factors</strong></p> <p><em>conversion_elements.csv:</em> Factors applied for conversion from metal compounds to elemental metals.</p> <p><em>conversion_factors_units.csv:</em> Factors applied for unit conversions.</p> <p> </p> <p><strong>Ore estimation factors</strong></p> <ul> <li>These factors represent content-to-ore ratios (i.e. "tonnes of extracted ore/mineral per ton of produced content”).</li> </ul> <p><em>all_integrated_est_fac_2021-11-04.csv:</em> Factors for estimation of metal and mineral ores (all which have been compiled/integrated)</p> <p><em>applied_est_fac_2021-11-04.csv:</em> Factors for estimation of metal and mineral ores (which have actually been applied)</p> <p><em>average_metal_prices_1990-2020.csv:</em> Metal prices applied in the compilation of estimation ratios.</p> <p><em>raw_metal_to_ore_ratios_fineprint.csv:</em> Raw metal-to-ore ratios derived from FINEPRINT mining data.</p> <p><em>raw_metal_to_ore_ratios_snl.csv:</em> Raw metal-to-ore ratios derived from SNL mining data.</p> <p> </p> <p><strong>Intermediate files from data processing</strong></p> <p><em>all_interm_conv_integr_2021-12-13.csv:</em> Harmonized, converted (units & elemental metals), integrated (i.e. without double counting). Before any estimations (ores & construction minerals) and before final cleaning/formatting.</p> <p><em>wmd_bgs_usgs_interm_harmonized_2021-12-13.csv:</em> All harmonized raw data from WMD/BGS/USGS, before any further processing (i.e. with double counting).</p> <p> </p> <p><strong>Comparison data (for mining accounts from FINEPRINT project)</strong></p> <p><em>data_for_comparison_2022-05-24.csv:</em> Data set specifically compiled for comparison with accounts on production of mines from the FINEPRINT project. Has been applied for verification of data published in <em>"Jasansky et al. (2022) An open database on global coal and metal mining"</em>.</p> <ul> <li>Includes all available types of materials which have been reported (e.g. ores and metals and metal compounds) <ul> <li>meaning: also materials which are associated with each other, for example iron ore and iron (and would therefore have been selectively integrated/excluded in the GMFD compilation).</li> </ul> </li> <li>Excludes any double counting for the exact same material from different data sources</li> <li>Harmonized IDs, converted units</li> <li>Includes metal compounds approximated/converted from reported elemental metals (where possible)</li> </ul> <p> </p> <p><strong>materialflows.net</strong></p> <p><em>data_sunburst_material_profiles_20220607.csv:</em> Full detail of data underlying the Sunburst visualization "Global Domestic Extraction in 2019, by material group" in section "Raw Material Profiles" on materialsflow.net. However, for all available years. Includes data on Domestic Extraction of biomass. Includes detail for CCC, MFA13+, MFA4+</p> <p> </p> <p><strong>Outliers</strong></p> <p><em>overview_adjusted_outliers_2021-12-22.csv:</em> Overview of outliers which were adjusted during final formatting.</p> <p> </p> <p><strong>Other</strong></p> <p><em>approximation_tailings_detailed_2021-09-26.csv:</em> Estimation of tailings based on reported amounts of ores/minerals and their respective contents.</p> <p> </p>
Compiled soil N2O emission data from Eastern China forests
<p>We searched the Thomson Reuters Web of Science Core Collection and China National Knowledge Infrastructure (CNKI) theses database for literature published before January 1, 2020, using the terms “forest” AND “greenhouse gas” OR “N<sub>2</sub>O” OR “nitrous oxide” in the titles, abstracts, and keywords; this returned 5,948 records and 590 records, respectively. Considering the different preferences for terminology (e.g., nutrient addition, simulated N deposition), we refined our search manually. The criteria for our refinement were as follows: (a) N addition experiment was conducted in eastern China forests with detailed records of N addition levels; (b) N<sub>2</sub>O flux was observed using the static chamber method so that data from different sites were comparable. A total of 58 papers from 38 research sites met our criteria. </p> <p>We also collected the data of soil N<sub>2</sub>O emission rates under natural conditions. The data were collected from literature that met the following criteria: (a) N<sub>2</sub>O fluxes were observed in eastern China forests using the static chamber method; (b) no manipulated experiments were conducted, and no nitrogen was added except for natural N deposition. There were 42 papers and theses that met these criteria, covering 30 different sites. </p> <p>From the abovementioned papers and theses, the soil N<sub>2</sub>O emission rates and the auxiliary information (coordinates, ecoregion type, natural N deposition rate, multi-year mean annual temperature, multi-year mean annual precipitation, and soil texture) were compiled (ds01).</p> <p>To obtain the soil N<sub>2</sub>O emission rates of Eastern China forests on grid level (10km × 10km), another dataset (ds02) on the environmental factors (ecoregion type, natural N deposition rate, multi-year mean annual temperature, multi-year mean annual precipitation, and soil texture) of each grid was compiled from the spatial datasets, including the Chinese vegetation ecoregion map from the Resources, Environmental Sciences, and Data Center at the Chinese Academy of Sciences (<a href="https://www.resdc.cn/">https://www.resdc.cn/data.aspx?DATAID=133</a>), as well as the soil texture raster data (<a href="https://www.resdc.cn/data.aspx?DATAID=260">https://www.resdc.cn/data.aspx?DATAID=260</a>), mean annual temperature raster data (1995–2015; <a href="https://www.resdc.cn/data.aspx?DATAID=228">https://www.resdc.cn/data.aspx?DATAID=228</a>), and mean annual precipitation raster data (1995–2015; <a href="https://www.resdc.cn/data.aspx?DATAID=229">https://www.resdc.cn/data.aspx?DATAID=229</a>). Land use and land cover raster data (2000–2015) were obtained from the Project on Big Earth Data Science Engineering (<a href="https://data.casearth.cn/sdo/detail/5c19a5650600cf2a3c557aae">https://data.casearth.cn/sdo/detail/5c19a5650600cf2a3c557aae</a>). Finally, annual N deposition rate and multi-year mean N deposition (MAN) data were derived from a N deposition dataset from mainland China previously constructed by our group. </p>
Understanding Sampling Bias in the Global Heat Flow Compilation
<p>Geothermal heat flow measurements, including calculated weights and geological, tectonic and topographic settings.</p> <p>Paper in review (July 14th, 2022)</p>
Compiled executables and their resulting Binary Ninja database files
<p>2,972 amd64 executables compiled with gcc and clang and their resulting Binary Ninja Database Files (.bndb). This dataset was made available for the conference paper "What Makefile? Detecting Compiler Information Using The Code Property Graph Without Source". Analysis of the 2,972 executables took 23 hours, 59 minutes, and 1 second on Binary Ninja build 3.1.3561-dev on macOS with an Intel Xeon W-3235. Analysis was done serially but without limiting the amount of threads working on analysis at a single time. CPU utilization averaged 290.2%.</p> <p>Executables compiled by Davide Pizzolotto and Katsuro Inoue of Osaka University and originally available on <a href="https://doi.org/10.5281/zenodo.4659370">Zenodo</a>. Please cite them if you use this work.</p>
Ixodes Scapularis monitoring data compiled from 7 studies
<p>This dataset lists 289 blacklegged tick population datasets from 7 studies that record abundance. These datasets were found by inputing keywords <em>Ixodes Scapularis</em> and <em>tick </em>in data repositories including Long Term Ecological Research data portal, National Ecological Observatory Network data portal, Google Datasets, Data Dryad, and Data One. The types of tick data recorded from these studies include density (number per square meter for example), proportion of ticks, count of ticks found on people. The locations of the datasets range from New York, New Jersey, Iowa, Massachusetts, and Connecticut, and range from 9 to 24 years in length. These datasets vary in that some record different life stages, geographic scope (county/town/plot), sampling technique (dragging/surveying), and different study length. The impact of these study factors on study results is analyzed in our research.</p> <p>Funding:</p> <p>RMC is supported by the National Institute of General Medical Sciences of the National Institutes of the Health under Award Number R25GM122672. CAB, JP, and KSW are supported by the Office of Advanced Cyberinfrastructure in the National Science Foundation under Award Number #1838807. The content is solely the responsibility of the authors and does not necessarily represent the official views of the National Institutes of Health or the National Science Foundation.</p>
Compiled dataset and posterior results for NEMo analyses on Acacia tolerance to aridity and salinity
<p>The dataset contains the compiled data for the aridity and salinity conditions at <em>Acacia</em> species' presence and absence locations, factors on sampling bias, as well as <em>Acacia</em> phylogeny. The data file is the input to the NEMo analyses. The dataset also contains the result files from the analyses.</p>
Compilation of Fe and ligand data along the West coast of the United States
<p>Compilation of Fe and ligand data along the West coast of the United States (Compiled by Anh Le-Duy Pham from current available literature as of April 2024)</p> <p> </p> <p><span>USWC_Iron_Ligand_Compilation.xlsx: an excel file containing all the combined dFe and ligand data</span></p> <p><span>USWC_IronLiganddata.mat: a MATLAB file containing all the combined dFe and ligand data</span></p>
The compiled 8-year dataset (2012-2019) consisting of weekly river water quality indicators (CODMn, DO, NH3-N and PH ) in majors 10 sub-basin of Yangtze river based on imputation of machine learning
<p>Water quality is significantly affected by global climate change and human activities, with diverse critical factors shaping its state in rivers and lakes. In the study, we utilized four indicators to characterize water quality: the physical water quality parameters included dissolved oxygen (DO, mg/L) and PH, while the chemical water quality parameters encompassed chemical oxygen demand (CODMn, mg/L) and ammonia nitrogen (NH3-N, mg/L). This study establishes weekly water quality models for typical 10 sub-basins along the Yangtze River using machine learning methods, which incorporate the impacts of hydro-meteorological and anthropogenic factors.These 10 sub-basins represent the principal tributaries of the Yangtze River basin and include Dongting Lake, the upper Han River, the lower Han River, the Jialing River, the Jinsha River, the Li River, the Min River, Poyang Lake, the Xiang River, and the Yuan River. This data collection was performed by National Environmental Monitoring Centre (http://www.cnemc.cn/sssj/szzdjczb/index_1.shtml). The water quality indicators discussed in this study are assessed in accordance with the national standard GB 3838-2002. Please refer to the paper for details.</p>
A compilation of beryllium-isotope, element, and grainsize data from sediments sampled from Prydz Bay and beneath Amery Ice Shelf, East Antarctica
<p>All tables are included in a single .xlsx file across three sheets. Each sheet includes sample information data: expedition and sample location information, reference to corresponding method section in text, and a reference to the source of the method employed for different procedures, or the reference to source data. Footnotes are used where necessary to explain a component of a table.</p> <p><strong>Supplementary Table 1:</strong> All beryllium data used for Sequential, Grainsize, Partial, and Total experiments described in text. 10Be concentration and corresponding 1-sigma (10^8 at/g), 9Be concentration and corresponding 1-sigma (10^15 at/g), and the 10Be/9Be ratio and corresponding 1-sigma (10^-8 at/at).</p> <p><strong>Supplementary Table 2: </strong>Element concentrations (µg/g) from samples across open marine and sub-ice shelf environments and their resultant enrichment factors (EF). Enrichment factors calculated using in text Equation 1. Estimated crustal abundance and ratio displayed below the data table.</p> <p><strong>Supplementary Table 3: </strong> Grainsize of samples used in this study. </p> <p> </p> <p>This research was supported by the Australian Research Council Special Research Initiative, Australian Centre for Excellence in Antarctic Science (Project Number SR200100008).</p>
PISCO: The PMAS/PPak Integral field Supernova hosts COmpilation
<p>Dataproducts from PISCO. More details at https://github.com/lgalbany/pisco</p>
Less is More: Exploiting the Standard Compiler Optimization Levels for Better Performance and Energy Consumption
<p>The data used to generate the graphs in Figures 2 and 3 of the paper "Less is More: Exploiting the Standard Compiler Optimization Levels for Better Performance and Energy Consumption" at SCOPES'18. The data includes power consumption, code size, and running time, indexed by an ID that identifies the combination of compiler options. Data is stored in a CSV file that identifies the specific benchmark, organized into a zip file that identifies the chip architecture.</p> <p>See README.txt for details.</p>
Compilation of Moon internal structure models and seismic event locations presented in Garcia et al., Space Science Review, 2019 study
<p>Compilation of of previously published internal structure model of the Moon in "named discontinuities" seismological format + 3 new models associated to the above mentioned study + previously published Moon quake location estimates by various authors.</p> <p>The study, the internal structure models and quake locations compilation were performed by the ISSI international research team on Moon Seismology and internal structure described here: http://www.issibern.ch/teams/internstructmoon/</p>
Compilation of hydraulic models for the study of the spatial averaging on flow laws
<p><strong>1.Summary</strong></p> <p>Datasets used in the article written by Ernesto Rodríguez, Michael Durand and Renato Prata de Moraes Frasson entitled “Observing rivers with varying spatial scales”.</p> <p><strong>2.File description</strong></p> <p>Data will be contained in one NetCDF file per river. The file contains the following groups and variables:</p> <p><strong>/River_Info/</strong></p> <p>Name: River name, data type: char</p> <p>QWBM: Mean annual discharge from the water balance model WBMsed (Cohen et al., 2014)</p> <p>rch_bnd: Reach boundaries measured in meters from the upstream end of the model</p> <p>gdrch: Reaches used in the study. Used to exclude small reaches defined around low-head dams and other obstacles where Manning’s equation should not be applied.</p> <p><strong>/XS_Timeseries/</strong></p> <p>t: Time measured in days since the first day or “0-January-0000” for cases when specific dates were available. Dimension: 1,time step.</p> <p>Z: Bed elevation in meters. Dimension: Cross-section, time step.</p> <p>xs_rch: Reach number for each cross-section. Dimension: Cross-section,1.</p> <p>X: Flow distance measured from the most upstream end of the model to the cross-section (meters). Dimension: Cross-section, 1.</p> <p>W: River width in meters. Dimension: Cross-section, time step.</p> <p>Q: Discharge (m<sup>3</sup>/s). Dimension: Cross-section, time step.</p> <p>H: Water surface elevation in meters. Dimension: Cross-section, time step.</p> <p>A: Cross-sectional area of flow in m<sup>2</sup>. Dimension: Cross-section, time step.</p> <p>P: Wetted perimeter in meters. Dimension: Cross-section, time step.</p> <p>n: Manning’s roughness. Dimension: Cross-section, time step.</p> <p><strong>/Reach_Timeseries/</strong></p> <p>t: Time measured in days since the first day or “0-January-0000” for cases when specific dates were available. Dimension: 1,time step.</p> <p>W: Reach averaged river width in meters. Dimension: Reach, time step.</p> <p>Q: Reach averaged discharge (m<sup>3</sup>/s). Dimension: Reach, time step.</p> <p>H: Reach averaged water surface elevation in meters. Dimension: Reach, time step.</p> <p>S: Reach averaged water surface slope in meters per meter. Reach, time step.</p> <p>A: Reach averaged area of flow in m<sup>2</sup>. Dimension: Reach, time step.</p> <p>P: Reach averaged wetted perimeter in meters. Not available for all rivers. Fill value: NaN. Dimension: Reach, time step.</p> <p><strong>References</strong></p> <p>Cohen, S., A. J. Kettner, and J. P. M. Syvitski (2014), Global suspended sediment and water discharge dynamics between 1960 and 2010: Continental trends and intra-basin sensitivity, <em>Glob. Planet. Change</em>, <em>115</em>, 44-58, doi: <a href="https://doi.org/10.1016/j.gloplacha.2014.01.011">10.1016/j.gloplacha.2014.01.011</a>.</p> <p>Rodríguez, E., Durand, M. T., & Frasson, R. P. d. M. (2020). Observing rivers with varying spatial scales. Water Resources Research. doi: <a href="https://doi.org/10.1029/2019WR026476">10.1029/2019WR026476 </a></p>
Figure 1 in Bat species diversity from Reserva Ecológica de Guapiaçu, Rio de Janeiro, Brazil: a compilation of two decades of sampling
Figure 1. Location of Reserva Ecológica de Guapiaçu (red/white circle) in the context of the Atlantic Forest remnants of Rio de Janeiro (green), Southeastern Brazil. Shapefile of forest coverage is from the Brazilian NGO SOS Mata Atlântica database.
A Compiled Archaeobiological Dataset for Central Asia's Chalcolithic through Bronze Age: macrobotanical and zooarchaeological data transcribed, standardized, and summarized from original publications
<p>This dataset contains archaeobotnaical and zooarchaeological data that have been compiled, transcribed, standardized, and summarized from original published data sources. Original data publications are given herein as a List of References (Microsoft Word file). These publications appeared between 1960-2022, presenting data in various formats, in various languages, and in scientific works that included journals, books, and conference proceedings - .</p> <p>Data have been compiled and are given in two Microsoft Excel files, one corresponding to archaeobotanical data and one to zooarchaeological data. On the first tab (worksheet) of each of these files, original publication sources are given in a summarized reference (Author, Year, Table/Figure Number) that corresponds to the full bibliographic reference in the accompanying List of References file. This first tab (worksheet) also summarizes additional information on the archaeological context of each dataset, collection and analysis methods (when reported), and the availabilty of other relevant and/or corresponding datasets. The remaining tabs (worksheets) in each file, organized alphabetically by author last name, tabularize the original data in a standardized format; these data are compiled and transcribed as necessary from the various formatting of original data sources, though they keep the original reported species names, table ordering, and numerical data. In the cases where totals were obviously erroneous or superceded by later or additional analyses, transcription notes have been offered directly in the worksheet.</p> <p>A fair number of these data sources are now out of print, have no digital distribution, and are otherwise difficult to access physically and/or linguistically. Accordingly, the sole aim in compiling these data here is to facilitate their widespread availablity, proper citation, and increased use within the international community of archaeological scientists working in Central Asia in the present day. Other scholars are encouraged to utilize these compiled datasets for foundational regional data and extended analyses, and to add to and expand these datasets going forward.</p>
Fig. 34 in Compilation of personal tributes to William Roy Branch (1946-2018): a loving husband and father, a good friend, and a mentor
Fig. 34. Ozzy Osbourne meets Elvis Presley. My all-time favourite photo of Bill and I (Photo: John Measey).
Fig. 35 in Compilation of personal tributes to William Roy Branch (1946-2018): a loving husband and father, a good friend, and a mentor
Fig. 35. Bill emerging from swamp in Gabon (2002) after checking funnel traps that he had set with the hope of catching an African Parachanna (Photo: Marius Burger).
Fig. 20 in Compilation of personal tributes to William Roy Branch (1946-2018): a loving husband and father, a good friend, and a mentor
Fig. 20. Bill Branch and Mark-Oliver Rödel in July 2018 in Bill's home in Port Elizabeth (Photo: Mark-Oliver Rödel).
Fig. 16 in Compilation of personal tributes to William Roy Branch (1946-2018): a loving husband and father, a good friend, and a mentor
Fig. 16. Bill Branch discussing the finer points of the day's photographic record of collections with colleagues and students, Lagoa Carumbo, May 2012. (Photo: Brian Huntley).
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.