Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

1,943

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

1,943 results for “machine learning”

Learn how ShareScore rates datasets ↗
dryad40/100

Machine learning can be as good as maximum likelihood when reconstructing phylogenetic trees and determining the best evolutionary model on four taxon alignments

Open the record for dataset details and reuse information.

publicJun 2023View details →
dryad40/100

Training data from: Machine learning predicts which rivers, streams, and wetlands the Clean Water Act regulates

Open the record for dataset details and reuse information.

publicDec 2023View details →
dryad40/100

Data from: SLICE-MSI: A machine learning interface for system suitability testing of mass spectrometry imaging platforms

Open the record for dataset details and reuse information.

publicJan 2025View details →
dryad40/100

Isotope mixing scenarios and machine learning model in: To what extent are the source mixing models accurate: evaluation of the model accuracy and guidelines for the site-specific model selection

Open the record for dataset details and reuse information.

publicNov 2024View details →
dryad40/100

Data from: TC-GEN: Data-driven tropical cyclone downscaling using machine learning-based high-resolution weather model

Open the record for dataset details and reuse information.

publicSep 2024View details →
dryad40/100

Machine learning reveals that climate, geography, and cultural drift all predict bird song variation in coastal Zonotrichia leucophrys

Open the record for dataset details and reuse information.

publicDec 2023View details →
dryad40/100

Machine learning analysis of wing venation patterns accurately identifies Sarcophagidae, Calliphoridae and Muscidae fly species

Open the record for dataset details and reuse information.

publicJul 2023View details →
dryad40/100

Vocalization data and scripts to model reindeer rut activity using on-animal acoustic recorders and machine learning

Open the record for dataset details and reuse information.

publicMay 2024View details →
dryad40/100

Assessing predictive performance of supervised machine learning algorithms for a diamond pricing model

Open the record for dataset details and reuse information.

publicMay 2023View details →
dryad40/100

Data for: Morphological species delimitation in the Western Pond Turtle (Actinemys): Can machine learning methods aid in cryptic species identification?

Open the record for dataset details and reuse information.

publicMar 2024View details →
dryad40/100

Prediction data from: Machine learning predicts which rivers, streams, and wetlands the Clean Water Act regulates

Open the record for dataset details and reuse information.

publicDec 2023View details →
zenodo36/100

Predicting antimicrobial resistance in Pseudomonas aeruginosa with machine learning-enabled molecular diagnostics

<p>Datasets for manuscript &quot;Predicting antimicrobial resistance in Pseudomonas aeruginosa with machine learning-enabled molecular diagnostics&quot;</p> <p><strong>Metadata.zip</strong></p> <ol> <li><strong>phenotypes.txt:&nbsp;</strong>tabular file containing binary resistance phenotypes based on CLSI guidelines, where the rows are the isolates and the columns correspond to&nbsp;different drugs. Resistance : 1, susceptibility: 0, missing: intermediate resistant</li> </ol> <p><strong>Features_gpa_exp_snps.zip &nbsp;</strong></p> <p>We provide the processed molecular data as Numpy compressed files (npz.). You can use the Numpy load method to read in these tables https://docs.scipy.org/doc/numpy/reference/generated/numpy.load.htm. The row (strains_list) and column labels (feature_lists) are stored separately.</p> <ol> <li><strong>genexp</strong>: gene expression table directory <ul> <li>genexp_feature_vect.npz: The feature matrix in the numpy format</li> <li>genexp_feature_list.txt: The columns of the&nbsp;feature matrix (features)</li> <li>genexp_strains_list.txt:&nbsp;The rows of the&nbsp;feature matrix (isolates)</li> </ul> </li> <li><strong>gpa:&nbsp;</strong>gene presence/absence table directory <ul> <li>gpa_feature_vect.npz: The feature matrix in the numpy format</li> <li>gpa_feature_list.txt: The columns of the&nbsp;feature matrix (features)</li> <li>gpa_strains_list.txt:&nbsp;The rows of the&nbsp;feature matrix (isolates)</li> </ul> </li> <li><strong>snps:&nbsp;</strong>SNPs table directory <ul> <li>snps_feature_vect.npz: The feature matrix in the numpy format</li> <li>snps_feature_list.txt: The columns of the&nbsp;feature matrix (features)</li> <li>snps_strains_list.txt:&nbsp;The rows of the&nbsp;feature matrix (isolates)</li> </ul> </li> </ol>

opencc-bySep 2019View details →
zenodo36/100

Machine learning estimates of eddy covariance carbon flux in a scrub in the Mexican highland

<p>Arid and semi-arid ecosystems contain relatively high species diversity and are subject to intense use, in particular extensive cattle grazing, which has favoured the expansion and encroachment of perennial thorny shrubs into the grasslands, thus decreasing the value of the rangeland. However, these environments have been shown to positively impact global carbon dynamics. Machine learning and remote sensing had enhanced our knowledge about carbon dynamics, but they need to be further developed and adapted to particular analysis. We measured the net ecosystem exchange of C (NEE) with the Eddy Covariance (EC) method and estimated GPP in a thorny scrub at Bernal in Mexico. We tested the agreement between EC estimates and remotely sensed GPP estimates from MODIS, and also with two alternative modelling methods: ordinary least squares multiple regression (OLS) or ensembles of machine learning algorithms (EML). The variables used as predictors were Moderate Resolution Spectroradiometer (MODIS) spectral bands, vegetation indices and products, as well as gridded environmental variables. The Bernal site was a carbon sink despite it was overgrazed, the average NEE during fifteen months of 2017 and 2018 was -0.78 g C m<sup>-2</sup> d<sup>-1</sup> and the flux was negative or neutral during the measured months. The probability of agreement (&theta;s) represented the agreement between observed and estimated values of GPP across the range of measurement. According to the mean value of &theta;s, agreement was higher for the EML (0.6) followed by OLS (0.5) and then MODIS (0.24). This graphic metric was more informative than r<sup>2</sup> (0.98, 0.67, 0.58 respectively) to evaluate the model performance. This was particularly true for MODIS because the maximum &theta;s of 4.3 was for measurements of 0.8 g C m<sup>-2</sup> d<sup>-1</sup> and then decreased steadily below 1 &theta;s for measurements above 6.5 g C m<sup>-2</sup> d<sup>-1 </sup>for this scrub vegetation. In the case of EML and OLS the &theta;s was stable across the range of measurement. We used an EML for the Ameriflux site US-SRM, which is similar in vegetation and climate, to predict GPP at Bernal, but &theta;s was low (0.16) indicating the local specificity of this model. Although cacti were an important component of the vegetation, the night time flux was characterized by positive NEE, suggesting that the photosynthetic dark-cycle flux of cacti was lower than ecosystem respiration. The discrepancy between MODIS and EC GPP estimates stresses the need to understand the limitations of both methods.</p>

opencc-by-4.0Jan 2020View details →
zenodo36/100

Processed Saccharomyces cerevisiae transcriptomics and genomics data for machine learning

<p><strong>Genomic data including open reading frame (ORF) boundaries of Saccharomyces cerevisiae C288 was obtained from the Saccharomyces Genome Database (<a href="https://www.yeastgenome.org/">https://www.yeastgenome.org/</a>)&nbsp;(<a href="http://paperpile.com/b/QmuOBv/gg2yy">Cherry, J. M. et al. Saccharomyces Genome Database: the genomics resource of budding yeast. Nucleic Acids Res. 40, D700&ndash;5 (2012)</a>) and published data (<a href="http://paperpile.com/b/QmuOBv/g4EbQ">Xu, Z. et al. Bidirectional promoters generate pervasive transcription in yeast. Nature 457, 1033&ndash;1037 (2009)</a>,&nbsp;<a href="http://paperpile.com/b/QmuOBv/uJqAw">Nagalakshmi, U. et al. The transcriptional landscape of the yeast genome defined by RNA sequencing. Science 320, 1344&ndash;1349 (2008)</a>).&nbsp;</strong><strong>Coding regions were extracted based on ORF boundaries and codon frequencies were normalized to probabilities. Processed raw RNA sequencing Star counts were obtained from the Digital Expression Explorer V2 database (<a href="http://dee2.io/index.html">http://dee2.io/index.html</a>) (<a href="http://paperpile.com/b/QmuOBv/LCGD">Ziemann, M., Kaspi, A. &amp; El-Osta, A. Digital expression explorer 2: a repository of uniformly processed RNA sequencing data. GigaScience vol. 8 (2019)</a>) and filtered for experiments that passed quality control. Raw mRNA data were transformed to transcripts per million (TPM) counts&nbsp;and genes with zero mRNA output (TPM &lt; 5) were removed. Prior to modeling, the mRNA counts were Box-Cox transformed.</strong></p>

opencc-by-sa-4.0Feb 2020View details →
zenodo36/100

A Machine-Learning-Based Global Atmospheric Forecast Model

<p>Data used in &quot;A Machine-Learning-Based Global Atmospheric Forecast Model&quot; 2020. Included in this dataset is the machine learning predictions and the truth for the year&#39;s worth of simulated forecast.&nbsp;</p>

opencc-by-4.0Mar 2020View details →
zenodo36/100

A machine learning framework for computationally expensive transient models

<p>The following dataset contains DEM simulation data on multiple material properties and operational parameters and their impact on uniform mixing time. Unifrom mixing time is calculated when segregation index reaches 1.1. Segregation index is defined based on the paper:&nbsp;Marigo, M., Cairns, D. L., Davies, M., Ingram, A. &amp; Stitt, E. H. A numerical comparison of mixing efficiencies of solids in a cylindrical vessel subject to a range of motions. <em>Powder Technol.</em> <strong>217</strong>, 540&ndash;547 (2012).&nbsp;</p> <p>The dataset was used for training ML model as described in paper:&nbsp;<a href="https://arxiv.org/abs/1907.05928">https://arxiv.org/abs/1907.05928</a></p>

opencc-by-4.0Jul 2019View details →
zenodo36/100

Replication code and data for: "Machine Learning Predicts Large Scale Declines in Native Plant Phylogenetic Diversity."

<p>Replication code and data for the paper: &quot;Machine Learning Predicts Large Scale Declines in Native Plant Phylogenetic Diversity.&quot; The following files are included in this repository:</p> <p>1) R scripts (numbered 0 through 9) include replication code for data analysis</p> <p>2) Datasets (6 zip folders) contain the data analyzed in&nbsp;the R scripts</p> <p>&nbsp;</p> <p>&nbsp;</p>

opencc-by-4.0Apr 2020View details →
zenodo36/100

Machine-learning-long-term-wind-power-time-series

<p>In this research work we assess how time-series generated by machine learning models (MLM) compare to Renewables.ninja in terms of their ability to replicate the characteristics of observed nationally aggregated wind power generation for Germany. With this archive, we want to make three machine learning derived time-series available in the Feather format for everyone interested.</p>

opencc-by-4.0Mar 2020View details →
zenodo36/100

Understanding the Nature of System-Related Issues in Machine Learning Frameworks: An Exploratory Study

<p>Modern systems are built using development frameworks. The infrastructure provided by these frameworks have a major impact on how the developed system executes, how configurations are managed, how it is tested, and how and where it is deployed. Machine learning (ML) systems have revolutionized multiple industries and come with different kinds of frameworks. Naturally, the issues that manifest in such systems may differ as well---as may the behavior of developers correcting those issues. We are interested in characterizing the types of system-related issues---issues impacting performance, memory and resource usage, and other quality attributes---that emerge in machine learning frameworks, and how they differ from those in traditional frameworks. To this end, we have conducted a large-scale exploratory study analyzing real-world system-related issues from 10 popular machine learning frameworks.<br> <br> Our findings offer a number of interesting observations, with implications for the development of machine learning systems, including differences in the frequency of occurrence of certain issue types, observations regarding the impact of debate and time on issue correction, and differences in the specialization of developers. &nbsp;We hope that this exploratory study will enable developers to improve their expectations, plan for risk, and allocate resources accordingly when making use of the tools provided by these frameworks to develop ML-based systems.</p>

opencc-by-4.0May 2020View details →
zenodo36/100

Image Dataset for 'Digitally deconstructing leaves in 3D using X-ray microcomputed tomography and machine learning'

<p>Dataset used in the manuscript &#39;Digitally Deconstructing Leaves in 3D Using X-ray microcomputed Tomography and Machine Learning&#39;. Please cite the paper presenting this dataset:</p> <p><strong>Citation:</strong> Th&eacute;roux-Rancourt, G., M. R. Jenkins, C. R. Brodersen, A. McElrone, E. J. Forrestel, and J. M. Earles. 2020. Digitally deconstructing leaves in 3D using X-ray microcomputed tomography<strong> </strong>and machine learning. <em>Applications in Plant Sciences</em> 8(7): .</p> <p>&nbsp;</p> <p><strong>Description of the dataset</strong></p> <p>A &#39;Cabernet Sauvignon&#39; grapevine (<em>Vitis vinifera</em> L.) leaf from a plant of the BOKU experimental vineyard in Tulln, Austria, was scanned using microCT at the Swiss Light Source. The original reconstructions of the scans are using the gridrec (<a href="https://zenodo.org/api/files/bbca544a-15d0-40f3-8cc9-a3ee3c08fd7e/Gridrec_reconstruction_downsized.zip?versionId=28d98982-f69d-4eac-9dfa-efcc89c6823c">Gridrec_reconstruction_downsized.zip</a>) and the paganin, or phase-contrast, algortithm (<a href="https://zenodo.org/api/files/bbca544a-15d0-40f3-8cc9-a3ee3c08fd7e/Phase_contrast_reconstruction_downsized.zip?versionId=bef3260d-2865-4c9b-b5e1-e692edefb691">Phase_contrast_reconstruction_downsized.zip</a>). To facilitate automated segmentation, the size of the image in the <em>x </em>and <em>y</em> dimensions have been halved, so that the size of the pixels is 0.325 &micro;m in those dimensions, but 0.1625 &micro;m in the <em>z</em> (slices) dimension.</p> <p>A binary image segmenting the leaf cells and the airspace for each gridrec and phase-contrast stacks are created, and both are combined together (<a href="https://zenodo.org/api/files/bbca544a-15d0-40f3-8cc9-a3ee3c08fd7e/Binary_stack_for_local_thickness.zip?versionId=165e3938-b490-4e56-9c8e-a2084cb39d49">Binary_stack_for_local_thickness.zip</a>), a map of the local thickness is created (<a href="https://zenodo.org/api/files/bbca544a-15d0-40f3-8cc9-a3ee3c08fd7e/Local_thickness_map.zip?versionId=ce0a7dc7-5e3f-44a4-8881-cf84b6efd87c">Local_thickness_map.zip</a>). This map gives information on the largest diameter of the pixels labeled as cells in the binary stack.</p> <p>Hand-labeled slices or ground truths were drawn on the following slices:&nbsp;80, 140, 200, 260, 340, 400, 440, 540, 620, 740, 800, 860, 940, 1060, 1140, 1240, 1300, 1400, 1480, 1540, 1600, 1690, 1740, 1840 (<a href="https://zenodo.org/api/files/bbca544a-15d0-40f3-8cc9-a3ee3c08fd7e/Hand_labelled_slices.tif?versionId=a21a13ac-fa47-4ef8-a903-ecc433787184">Hand_labelled_slices.tif</a>).</p> <p>Using the hand-labeled slices and the different images, a random-forest model was trained, which allowed to automatically segment the remaining slices of the stack (<a href="https://zenodo.org/api/files/bbca544a-15d0-40f3-8cc9-a3ee3c08fd7e/Fullstack_Prediction_Example-6_training_slices-6_testing_slices.zip?versionId=02b69e65-da85-492e-9b72-9b2b3ccd085f">Fullstack_Prediction_Example-6_training_slices-6...</a>).</p> <p>The source code for the segmentation program is available <a href="https://github.com/plant-microct-tools/leaf-traits-microct/tree/master">here</a>, and the source code for the testing used in the paper is available <a href="https://github.com/plant-microct-tools/leaf-traits-microct/tree/nb-slices-eval">here</a>.</p>

opencc-by-4.0Mar 2020View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record