Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
1,079
datasets available to search
ShareScore release 0.9.0
Dataset results
1,079 results for “source data”
Organisation for Economic Co-operation and Development (OECD) data for Antalya (Turkey), Antwerp (Belgium), Cork (Ireland), Thessaloniki (Greece) (source: OECD)
<p>The data have been collected via the official OECD Application Programming Interface (API)<strong> </strong>and<strong> </strong>includes the following indicators:</p> <ul> <li>EmpPlaRes - Employment at place of residence</li> <li>LfPartRa - Labour Force and Participation rate</li> <li>UnemReg - Unemployment in regions </li> <li>RegGdpTL2 - Regional Gross Domestic Product (Large regions TL2)</li> <li>GDPLT3 - Gross Domestic Product (Small regions TL3)</li> <li>RegEmIndu - Regional Employment by industry (ISIC rev 4)</li> <li>RegGVAWorker - Regional GVA per worker</li> <li>RegIncPC - Regional income per capita</li> </ul> <p>Source: https://data.oecd.org/api/</p>
RDF version of the data from Choi, JS., Ha, M.K., Trinh, T.X. et al. Towards a generalized toxicity prediction model for oxide nanomaterials using integrated data from different sources. Sci Rep 8, 6110 (2018)
<p>RDF version of the data from Choi, JS., Ha, M.K., Trinh, T.X. et al. Towards a generalized toxicity prediction model for oxide nanomaterials using integrated data from different sources. Sci Rep 8, 6110 (2018)</p>
Raw and analyzed data for manuscript: "An open-source surface barrier discharge plasma pretreatment for reduced cracking of outdoor wood coatings"
<p><strong>Highlights:</strong></p> <ul> <li>Surface barrier discharges are an affordable and available plasma technology for industrial, laboratory and home-workshop applications.</li> <li>Plasma pretreatments had no impact on the appearance of different protective wood coating for outdoor usage.</li> <li>The weathering performance of outdoor wood coatings improved by plasma, showing less cracks and less biotic factors.</li> </ul>
Source Data for Manuscript "Sodium salicylate improves detection of amplitude-modulated sound in mice"
<p>This repository contains the source data for our papers <strong>Sodium salicylate improves detection of amplitude-modulated sound in mice </strong>(van den Berg*, Wong*, Houtak, Williamson, Borst). The code to generate figure panels can be found in our github repository at https://github.com/aaronbwong/salicylateonam</p>
Bubble/Foam Simulations for Malej et al. 2023, source codes, input files, matlab files, data files
<p><i>.F are source codes, *.m are matlab scripts for analysis and postprocessing, .txt are data files including bathymetry and data from sensitivity tests</i></p>
Source data for publication "Bacteria use exogenous peptidoglycan as a danger signal to trigger biofilm formation"
<p><strong>This dataset contains the source data for the figures in the following publication: </strong></p> <p><strong>Title: </strong>Bacteria use exogenous peptidoglycan as a danger signal to trigger biofilm formation</p> <p><strong>Authors: </strong>Sanika Vaidya, Dibya Saha, Daniel K.H. Rode, Gabriel Torrens, Mads F. Hansen, Praveen K. Singh, Eric Jelli, Kazuki Nosho, Hannah Jeckel, Stephan Göttig, Felipe Cava, Knut Drescher</p> <p><strong>Journal: </strong>Nature Microbiology, 2025</p> <p> </p> <p><strong>Description of the dataset: </strong></p> <p>The data is organized by figures in the publication receferenced above. For each figure, there is a XLSX-file that contains the processed data and a ZIP-archive that contains the raw image data. The ZIP-archive also contains readme documents with more detailed descriptions for every set of raw image data. </p> <p>Example: For Figure 1 in the main text of the publication, there following data are available:</p> <ul> <li>raw data: Figure_01.zip</li> <li>processed data: Processed_Source_Data_Figure1.xlsx</li> </ul> <p>Similarly, XLSX-files and ZIP-archives are available for the main text Figures 1-6. For the Extended Data (ED) Figures 1-10, there are XLSX-files available that present the data shown in each figure. Only some of the Extended Data Figures present results based on image data - therefore raw image data ZIP-archives are only available for ED Figures 3, 6, 7, 8, 9, 10.</p>
Source Data and ambient ozone dataset generated in "Substantially underestimated global health risks of current ozone pollution"
<p>Existing assessments might have underappreciated ozone-related health impacts worldwide. Here our study assesses current global ozone pollution using the high-resolution (0.05°) estimation from a geo-ensemble learning model, with key focuses on population exposure and all-cause mortality burden. Our model demonstrates strong performance, achieving a mean bias of less than -1.5 parts per billion against in-situ measurements. We estimate that 66.2% of the global population is exposed to excess ozone for short term (> 30 days per year), and 94.2% suffers from long-term exposure. Furthermore, severe ozone exposure levels are observed in Cropland areas, particularly over Asia. Importantly, the all-cause ozone-attributable deaths significantly surpass previous recognition from specific diseases worldwide. Notably, mid-latitude Asia (30°N) and the western United States show high mortality burden, contributing substantially to global ozone-attributable deaths. Our study highlights current significant global ozone-related health risks and may benefit the ozone-exposed population in the future.</p>
Tabular Data for "A Comprehensive Catalog of UVIT Observations I: Catalog Description and First Release of Source Catalog (UVIT DR1)"
<p>The first comprehensive catalog of UVIT includes the sources from the observations between 2016 and 2017. </p> <p>The catalog is formatted according to the machine-readable format used by the AAS Journals and CDS/VizieR. Specific information on the structure of MRT files can be found at:</p> <p><span> </span>AAS: <a href="https://journals.aas.org/mrt-overview/"><span>https://journals.aas.org/mrt-overview/</span></a></p> <p><span> </span>CDS: <a href="http://cds.u-strasbg.fr/doc/catstd.htx"><span>http://cds.u-strasbg.fr/doc/catstd.htx</span></a></p> <p><span> </span>These files can be read in python using the astropy package<span> </span>or with the most recent version of TOPCAT (> Version 4.8)</p>
Machine learning code and dataset for "Nowcasting thunderstorm hazards using machine learning: the impact of data sources on performance"
<p>This repository contains the code and dataset for the paper:</p> <p>Nowcasting thunderstorm hazards using machine learning: the impact of data sources on performance, Natural Hazards and Earth System Sciences, 2022, <a href="https://doi.org/10.5194/nhess-2021-171">https://doi.org/10.5194/nhess-2021-171</a></p> <p>The GitHub code repository at <a href="https://github.com/meteoswiss-mdr/ts-nowcast-datasources">https://github.com/meteoswiss-mdr/ts-nowcast-datasources</a> may contain a more up-to-date version of the code if bug fixes etc. have been necessary. The file <a href="https://zenodo.org/api/files/41faa1b7-17f6-4a75-be09-7743426ef13c/ts-nowcast-datasources-publication.zip">ts-nowcast-datasources-publication.zip</a> in this Zenodo release contains the status of the GitHub repository at the time of the publication of the paper.</p> <p>For instructions for using the data, please see the <a href="https://github.com/meteoswiss-mdr/ts-nowcast-datasources">code repository</a>.</p> <p> </p>
Data Sources for the World Atlas of late Quaternary Foraminiferal Oxygen and Carbon Isotope Ratios 2021
<p>A tabulated text file containing all data sources used for the World Atlas of late Quaternary Foraminiferal Oxygen and Carbon Isotope Ratios 2021 (WA_Foraminiferal_Isotopes_2021), https://doi.org/10.1594/PANGAEA.936747 (Mulitza et al. 2021)</p>
Single-crystal X-ray diffractometry data for a sample of [Cu(HF₂)(pyrazine)₂]PF₆ collected on beamline I19-2 at Diamond Light Source
<p>Single-crystal X-ray diffractometry data for a sample of [Cu(HF₂)(pyrazine)₂]PF₆.</p> <p>These data were collected at Diamond Light Source, on beamline I19 (experiments hutch 2), on 2022-01-30, and are particularly useful for testing data reduction routines. They are known to produce good merging statistics and final structure refinement.</p> <p>The sample was prepared as follows:<br> Ammonium hexafluorophosphate (NH₄PF₆) (0.310 g, 1.9 mmol), ammonium hydrogen difluoride ((NH₄)HF₂) (0.109 g, 1.9 mmol) and pyrazine (C₄H₄N₂) (0.300 g, 3.7 mmol) were dissolved in 5 mL of deionised water. The obtained colourless solution was slowly added to a blue solution of copper(II) nitrate prepared by dissolving copper(II) nitrate hemipentahydrate (Cu(NO₃)₂ · 2.5(H₂O)) (0.425 g, 1.8 mmol) in 5 mL of deionised water. The solutions were mixed in a plastic beaker at room temperature. The formation of blue crystals of [Cu(HF₂)(pyrazine)₂]PF₆ on the side of the beaker started after few seconds and continued for about 24 hours during which the sealed beaker was not moved.</p> <p>The sample was measured at room temperature and the illuminating beam had a wavelength of 0.4859 Å (25.52 keV).</p> <p>Beamline I19-2 at Diamond Light Source, a four-circle κ-geometry diffractometer (see <a href="https://onlinelibrary.wiley.com/doi/10.1107/97809553602060000936">[Kern 2019]</a>) with an undulator source, is described in <a href="https://doi.org/10.1107/S0909049512008801">[Nowell 2012]</a> but has since been upgraded to use a Dectris Eiger2 X 4M CdTe hybrid photon counting detector. The data are written in the <a href="https://manual.nexusformat.org/classes/applications/NXmx.html">NXmx variant</a> of the <a href="https://www.nexusformat.org/">NeXus format</a>, and so include metadata with a functionally complete description of the diffractometer.</p> <p>Inventory of data:</p> <ul> <li><strong><code>01_CuHF2pyz2PF6b_Phi.tar.xz</code></strong><br> A single 1750-image 350° φ rotation scan from -175° to 175° with 0.2° rotation per image, an exposure time of 0.1 s per image, ω = -90°, κ = 0° and 2θ = 0°.</li> <li><strong><code>02_CuHF2pyz2PF6b_2T.tar.xz</code></strong><br> A single 1750-image 350° φ rotation scan from -175° to 175° with 0.2° rotation per image, an exposure time of 0.1 s per image, ω = -90°, κ = 0° and 2θ = 20°.</li> <li><strong><code>03_CuHF2pyz2PF6b_P_O.tar.xz</code></strong><br> Two sequential rotation scans: <ul> <li><strong><code>CuHF2pyz2PF6b_P_O_01.nxs</code></strong><br> A 1750-image 350° φ scan from -175° to 175° with ω = -90°, κ = 0° and 2θ = 0°.</li> <li><strong><code>CuHF2pyz2PF6b_P_O_02.nxs</code></strong><br> A 600-image 120° ω scan from -125° to -5° with φ = -90°, κ = 45° and 2θ = 0°.</li> </ul> Both scans had 0.2° rotation per image and an exposure time of 0.1 s per image.</li> </ul> <p>The same sample was used for all these measurements. Throughout, the sample-to-detector distance was 85 mm and the beam was attenuated to 0.2% of its full intensity.</p> <p>For each rotation scan, the data comprise a single top-level NXmx-format NeXus file named <code><filename>.nxs</code>, one or more image files named <code><filename>_00000n.h5</code>, where <code>n</code> is a numeral, and a single detector metadata file named <code><filename>_meta.h5</code>. The NeXus file contains an HDF5 virtual data set that links to the data in the image file(s), and several HDF5 external links to data in the detector metadata file.</p> <p>For internal reference of Diamond Light Source staff, these data were collected as part of commissioning visit CM31144-1. Some file names and corresponding HDF5 link targets have been altered from their original names for consistency with the file contents.</p>
First Three-dimensional Quantification of Planktic Food Chain lower levels (Copepods) for the Ross Sea region Marine Protected Area (RSRMPA), Antarctica: Using FAIR-inspired legacy data with Machine Learning, and Open Source GIS
<p>This dataset is relative to the paper entitled: "First Three-dimensional Quantification of Planktic Food Chain lower levels (Copepods) for the Ross Sea region Marine Protected Area (RSRMPA), Antarctica: Using FAIR-inspired legacy data with Machine Learning, and Open Source GIS" publishing in journal Diversity (MPDI).</p> <p>Abstract:</p> <p>Zooplankton is a fundamental group in all aquatic ecosystems located the base of the food chain. It forms a link between the lower trophic levels with secondary consumers and shows marked fluctuations of populations with environmental change, especially reacting to heating and water acidification. At sea copepod crustaceans account for app. 70% in abundance of zooplankton and are a target of monitoring activities in key areas such as the Southern Ocean. In this study we have used FAIR-inspired legacy data (dating back to the ‘80s) collected in the Ross Sea by the Italian National Antarctic Program in GBIF.org. Together with other open-access GIS data sources and tools it allows generating, for the first time, three-dimensional predictive distribution maps for twenty-six copepod species. These predictive maps were obtained by applying machine learning techniques to grey literature data, which were visualized in open-source GIS platforms. In a Species Distribution Modeling (SDM) framework we used machine learning with three types of algorithms (TreeNet, RandomForest and Ensemble) to analyze the presence and absence of copepods at different areas and depth classes in function of environmental descriptors obtained from the Polar Macroscope Layers present in Quantartica. The models allow for the first time to map-predict the food chain in quantitative terms showing the relative index of occurrence (RIO) and identified the presence for each copepod species analyzed in the Ross Sea. Our results show marked geographical preferences that vary with species and trophic strategy. This study demonstrates that machine learning is a successful method in accurately predicting Antarctic copepod presence, also providing useful data to orient future sampling and management of wildlife and conservation.</p>
Data analysis source code and measurement data of chemosensor salt-responsiveness
<p>Dataset with measurement data of salt-responsiveness of macrocyclic chemosensors and the Python source code for data analysis.</p>
Sharing research data and findings relevant to the novel coronavirus (COVID-19) outbreak - Literature sources
<p>The spreadsheet in the present dataset (CSV format) includes the sources considered during the literature review stage for the report: From intent to impact: Investigating the effects of open sharing commitments. Please note that not all sources in this deposit have been referenced in the above-mentioned report and that the report may include additional sources</p>
Data from article "Is the Atlantic a Source for Decadal Predictability of Sea-Level Rise in Venice?"
<p>Data from the article Zanchettin D., et al.: Is the Atlantic a Source for Decadal Predictability of Sea-Level Rise in Venice?, Earth and Space Science, article number 2022EA002494</p> <p>The dataset contains:</p> <p>- annual time series of October-March average of relative sea level in Venice corrected for vertical land movement (VLMcorrectedRSL) with associated standard error of the mean (LMcorrectedRSL_SEM) for the period 1873-2019;</p> <p>- annual time series of estimate of subsidence in Venice (Subsidence) for the period 1873-2019</p> <p>- modeled state of Venice sea level with associated uncertainty, provided as mean (delta_mean), 1st percentile (delta_1_percentile) and 99th percentile (delta_99_percentile), for the period 1873-2019</p> <p>- modeled local stochastic trend of Venice sea level with associated uncertainty, provided as mean (beta_mean), 1st percentile (beta_1_percentile) and 99th percentile (beta_99_percentile), for the period 1873-2019</p>
Waveform data for centroid moment tensor solutions presented in publication "Bayesian seismic source inversion with a 3-D Earth model of the Japanese islands"
<p>The dataset includes waveform data for centroid moment tensor solutions inferred using Hamiltonian Monte Carlo and a 3-D Earth model in the Japanese islands. The data are provided as Green's strains at the maximum-likelihood location (indicated in the title of each text file) for all study events inverted at different periods. Inversion period is also indicated in the title. All the data are filtered between 15 s and 80 s. Additionally we provide a Python code to obtain displacement from strains given a moment tensor.</p>
How do Google News' top 100 sources visually represent the data centres' energy footprint?
<p><strong>By querying "data centres' energy footprint" on Google News in incognito mode, the candidate has selected and mapped the top 100 results according to the ranking on May 15, 2022. </strong></p>
Data from "Source characterization of the declared North Korean Nuclear Tests from regional distance coda wave spectral ratios"
<p>Results of the coda spectral ratio analysis presented in Delbridge et al. (2022).</p> <p>network_average_ratios.csv - a csv file which contains each of the network average ratios calculated from all channels and stations for each event pair and component.</p>
Source code and data accompanying "Man gave names to all those animals"
<p>This package provides source code and data for the two blog post on animal names ("Man gave names to all those animals..."), which were published in the blog "The genealogical world of phylogenetic networks" (http://phylonetworks.blogspot.de/2017/10/man-gave-names-to-all-those-animals.html). This upload also contains the data from Gerhard Jäger's paper titled "Support for linguistic macrofamilies from weighted sequence alignment" (PNAS, DOI: 10.1073/pnas.1500331112), generously provided by the author.</p>
Open-source DGGS comparison data supplement
<p>A DGGS is a type of spatial reference system that partitions the globe into many individual, evenly spaced, and well-aligned cells to encode location. We calculated normalized area and compactness of cell geometries for 5 open-source DGGS implementations - Uber H3, Google S2, RiskAware OpenEAGGR, rHEALPix by Landcare Research New Zealand, HEALPix by NASA Jet Propulsion Labs, and DGGRID by Southern Oregon University - to evaluate their suitability for a global-level statistical data cube.</p> <p>This repository contains all generated data and statistics.</p> <ul> <li>EAGGR doesn't seem to have a predefined logic of hierarchical cell resolutions for ISEA3H</li> <li>EAGGR doesn't seem to have a region filling algorithm available, neither for ISEA4T nor ISEA3H</li> <li>rHEALPix is pure Python (with Numpy/Scipy support), but cell generation/conversion is slower than the other C/C++ based implementations</li> <li>DGGRID is a commandline tool and can predominantly only be used to generate a grid and fill with sampling data, the Python API is only a wrapper</li> <li>healpy is a Python package to handle pixelated data on the sphere. It is based on the Hierarchical Equal Area isoLatitude Pixelization (HEALPix) scheme and bundles the HEALPix C++ library.</li> </ul> <p>Kmoch et. al (2022). Area and Shape Distortions in Open-Source Discrete Global Grid Systems. <strong><em>Big Earth Data</em></strong></p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.