Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
1,943
datasets available to search
ShareScore release 0.9.0
Dataset results
1,943 results for “machine learning”
Machine learning can be as good as maximum likelihood when reconstructing phylogenetic trees and determining the best evolutionary model on four taxon alignments
Open the record for dataset details and reuse information.
Training data from: Machine learning predicts which rivers, streams, and wetlands the Clean Water Act regulates
Open the record for dataset details and reuse information.
Data from: SLICE-MSI: A machine learning interface for system suitability testing of mass spectrometry imaging platforms
Open the record for dataset details and reuse information.
Isotope mixing scenarios and machine learning model in: To what extent are the source mixing models accurate: evaluation of the model accuracy and guidelines for the site-specific model selection
Open the record for dataset details and reuse information.
Data from: TC-GEN: Data-driven tropical cyclone downscaling using machine learning-based high-resolution weather model
Open the record for dataset details and reuse information.
Machine learning reveals that climate, geography, and cultural drift all predict bird song variation in coastal Zonotrichia leucophrys
Open the record for dataset details and reuse information.
Machine learning analysis of wing venation patterns accurately identifies Sarcophagidae, Calliphoridae and Muscidae fly species
Open the record for dataset details and reuse information.
Vocalization data and scripts to model reindeer rut activity using on-animal acoustic recorders and machine learning
Open the record for dataset details and reuse information.
Assessing predictive performance of supervised machine learning algorithms for a diamond pricing model
Open the record for dataset details and reuse information.
Data for: Morphological species delimitation in the Western Pond Turtle (Actinemys): Can machine learning methods aid in cryptic species identification?
Open the record for dataset details and reuse information.
Prediction data from: Machine learning predicts which rivers, streams, and wetlands the Clean Water Act regulates
Open the record for dataset details and reuse information.
Predicting antimicrobial resistance in Pseudomonas aeruginosa with machine learning-enabled molecular diagnostics
<p>Datasets for manuscript "Predicting antimicrobial resistance in Pseudomonas aeruginosa with machine learning-enabled molecular diagnostics"</p> <p><strong>Metadata.zip</strong></p> <ol> <li><strong>phenotypes.txt: </strong>tabular file containing binary resistance phenotypes based on CLSI guidelines, where the rows are the isolates and the columns correspond to different drugs. Resistance : 1, susceptibility: 0, missing: intermediate resistant</li> </ol> <p><strong>Features_gpa_exp_snps.zip </strong></p> <p>We provide the processed molecular data as Numpy compressed files (npz.). You can use the Numpy load method to read in these tables https://docs.scipy.org/doc/numpy/reference/generated/numpy.load.htm. The row (strains_list) and column labels (feature_lists) are stored separately.</p> <ol> <li><strong>genexp</strong>: gene expression table directory <ul> <li>genexp_feature_vect.npz: The feature matrix in the numpy format</li> <li>genexp_feature_list.txt: The columns of the feature matrix (features)</li> <li>genexp_strains_list.txt: The rows of the feature matrix (isolates)</li> </ul> </li> <li><strong>gpa: </strong>gene presence/absence table directory <ul> <li>gpa_feature_vect.npz: The feature matrix in the numpy format</li> <li>gpa_feature_list.txt: The columns of the feature matrix (features)</li> <li>gpa_strains_list.txt: The rows of the feature matrix (isolates)</li> </ul> </li> <li><strong>snps: </strong>SNPs table directory <ul> <li>snps_feature_vect.npz: The feature matrix in the numpy format</li> <li>snps_feature_list.txt: The columns of the feature matrix (features)</li> <li>snps_strains_list.txt: The rows of the feature matrix (isolates)</li> </ul> </li> </ol>
Machine learning estimates of eddy covariance carbon flux in a scrub in the Mexican highland
<p>Arid and semi-arid ecosystems contain relatively high species diversity and are subject to intense use, in particular extensive cattle grazing, which has favoured the expansion and encroachment of perennial thorny shrubs into the grasslands, thus decreasing the value of the rangeland. However, these environments have been shown to positively impact global carbon dynamics. Machine learning and remote sensing had enhanced our knowledge about carbon dynamics, but they need to be further developed and adapted to particular analysis. We measured the net ecosystem exchange of C (NEE) with the Eddy Covariance (EC) method and estimated GPP in a thorny scrub at Bernal in Mexico. We tested the agreement between EC estimates and remotely sensed GPP estimates from MODIS, and also with two alternative modelling methods: ordinary least squares multiple regression (OLS) or ensembles of machine learning algorithms (EML). The variables used as predictors were Moderate Resolution Spectroradiometer (MODIS) spectral bands, vegetation indices and products, as well as gridded environmental variables. The Bernal site was a carbon sink despite it was overgrazed, the average NEE during fifteen months of 2017 and 2018 was -0.78 g C m<sup>-2</sup> d<sup>-1</sup> and the flux was negative or neutral during the measured months. The probability of agreement (θs) represented the agreement between observed and estimated values of GPP across the range of measurement. According to the mean value of θs, agreement was higher for the EML (0.6) followed by OLS (0.5) and then MODIS (0.24). This graphic metric was more informative than r<sup>2</sup> (0.98, 0.67, 0.58 respectively) to evaluate the model performance. This was particularly true for MODIS because the maximum θs of 4.3 was for measurements of 0.8 g C m<sup>-2</sup> d<sup>-1</sup> and then decreased steadily below 1 θs for measurements above 6.5 g C m<sup>-2</sup> d<sup>-1 </sup>for this scrub vegetation. In the case of EML and OLS the θs was stable across the range of measurement. We used an EML for the Ameriflux site US-SRM, which is similar in vegetation and climate, to predict GPP at Bernal, but θs was low (0.16) indicating the local specificity of this model. Although cacti were an important component of the vegetation, the night time flux was characterized by positive NEE, suggesting that the photosynthetic dark-cycle flux of cacti was lower than ecosystem respiration. The discrepancy between MODIS and EC GPP estimates stresses the need to understand the limitations of both methods.</p>
Processed Saccharomyces cerevisiae transcriptomics and genomics data for machine learning
<p><strong>Genomic data including open reading frame (ORF) boundaries of Saccharomyces cerevisiae C288 was obtained from the Saccharomyces Genome Database (<a href="https://www.yeastgenome.org/">https://www.yeastgenome.org/</a>) (<a href="http://paperpile.com/b/QmuOBv/gg2yy">Cherry, J. M. et al. Saccharomyces Genome Database: the genomics resource of budding yeast. Nucleic Acids Res. 40, D700–5 (2012)</a>) and published data (<a href="http://paperpile.com/b/QmuOBv/g4EbQ">Xu, Z. et al. Bidirectional promoters generate pervasive transcription in yeast. Nature 457, 1033–1037 (2009)</a>, <a href="http://paperpile.com/b/QmuOBv/uJqAw">Nagalakshmi, U. et al. The transcriptional landscape of the yeast genome defined by RNA sequencing. Science 320, 1344–1349 (2008)</a>). </strong><strong>Coding regions were extracted based on ORF boundaries and codon frequencies were normalized to probabilities. Processed raw RNA sequencing Star counts were obtained from the Digital Expression Explorer V2 database (<a href="http://dee2.io/index.html">http://dee2.io/index.html</a>) (<a href="http://paperpile.com/b/QmuOBv/LCGD">Ziemann, M., Kaspi, A. & El-Osta, A. Digital expression explorer 2: a repository of uniformly processed RNA sequencing data. GigaScience vol. 8 (2019)</a>) and filtered for experiments that passed quality control. Raw mRNA data were transformed to transcripts per million (TPM) counts and genes with zero mRNA output (TPM < 5) were removed. Prior to modeling, the mRNA counts were Box-Cox transformed.</strong></p>
A Machine-Learning-Based Global Atmospheric Forecast Model
<p>Data used in "A Machine-Learning-Based Global Atmospheric Forecast Model" 2020. Included in this dataset is the machine learning predictions and the truth for the year's worth of simulated forecast. </p>
A machine learning framework for computationally expensive transient models
<p>The following dataset contains DEM simulation data on multiple material properties and operational parameters and their impact on uniform mixing time. Unifrom mixing time is calculated when segregation index reaches 1.1. Segregation index is defined based on the paper: Marigo, M., Cairns, D. L., Davies, M., Ingram, A. & Stitt, E. H. A numerical comparison of mixing efficiencies of solids in a cylindrical vessel subject to a range of motions. <em>Powder Technol.</em> <strong>217</strong>, 540–547 (2012). </p> <p>The dataset was used for training ML model as described in paper: <a href="https://arxiv.org/abs/1907.05928">https://arxiv.org/abs/1907.05928</a></p>
Replication code and data for: "Machine Learning Predicts Large Scale Declines in Native Plant Phylogenetic Diversity."
<p>Replication code and data for the paper: "Machine Learning Predicts Large Scale Declines in Native Plant Phylogenetic Diversity." The following files are included in this repository:</p> <p>1) R scripts (numbered 0 through 9) include replication code for data analysis</p> <p>2) Datasets (6 zip folders) contain the data analyzed in the R scripts</p> <p> </p> <p> </p>
Machine-learning-long-term-wind-power-time-series
<p>In this research work we assess how time-series generated by machine learning models (MLM) compare to Renewables.ninja in terms of their ability to replicate the characteristics of observed nationally aggregated wind power generation for Germany. With this archive, we want to make three machine learning derived time-series available in the Feather format for everyone interested.</p>
Understanding the Nature of System-Related Issues in Machine Learning Frameworks: An Exploratory Study
<p>Modern systems are built using development frameworks. The infrastructure provided by these frameworks have a major impact on how the developed system executes, how configurations are managed, how it is tested, and how and where it is deployed. Machine learning (ML) systems have revolutionized multiple industries and come with different kinds of frameworks. Naturally, the issues that manifest in such systems may differ as well---as may the behavior of developers correcting those issues. We are interested in characterizing the types of system-related issues---issues impacting performance, memory and resource usage, and other quality attributes---that emerge in machine learning frameworks, and how they differ from those in traditional frameworks. To this end, we have conducted a large-scale exploratory study analyzing real-world system-related issues from 10 popular machine learning frameworks.<br> <br> Our findings offer a number of interesting observations, with implications for the development of machine learning systems, including differences in the frequency of occurrence of certain issue types, observations regarding the impact of debate and time on issue correction, and differences in the specialization of developers. We hope that this exploratory study will enable developers to improve their expectations, plan for risk, and allocate resources accordingly when making use of the tools provided by these frameworks to develop ML-based systems.</p>
Image Dataset for 'Digitally deconstructing leaves in 3D using X-ray microcomputed tomography and machine learning'
<p>Dataset used in the manuscript 'Digitally Deconstructing Leaves in 3D Using X-ray microcomputed Tomography and Machine Learning'. Please cite the paper presenting this dataset:</p> <p><strong>Citation:</strong> Théroux-Rancourt, G., M. R. Jenkins, C. R. Brodersen, A. McElrone, E. J. Forrestel, and J. M. Earles. 2020. Digitally deconstructing leaves in 3D using X-ray microcomputed tomography<strong> </strong>and machine learning. <em>Applications in Plant Sciences</em> 8(7): .</p> <p> </p> <p><strong>Description of the dataset</strong></p> <p>A 'Cabernet Sauvignon' grapevine (<em>Vitis vinifera</em> L.) leaf from a plant of the BOKU experimental vineyard in Tulln, Austria, was scanned using microCT at the Swiss Light Source. The original reconstructions of the scans are using the gridrec (<a href="https://zenodo.org/api/files/bbca544a-15d0-40f3-8cc9-a3ee3c08fd7e/Gridrec_reconstruction_downsized.zip?versionId=28d98982-f69d-4eac-9dfa-efcc89c6823c">Gridrec_reconstruction_downsized.zip</a>) and the paganin, or phase-contrast, algortithm (<a href="https://zenodo.org/api/files/bbca544a-15d0-40f3-8cc9-a3ee3c08fd7e/Phase_contrast_reconstruction_downsized.zip?versionId=bef3260d-2865-4c9b-b5e1-e692edefb691">Phase_contrast_reconstruction_downsized.zip</a>). To facilitate automated segmentation, the size of the image in the <em>x </em>and <em>y</em> dimensions have been halved, so that the size of the pixels is 0.325 µm in those dimensions, but 0.1625 µm in the <em>z</em> (slices) dimension.</p> <p>A binary image segmenting the leaf cells and the airspace for each gridrec and phase-contrast stacks are created, and both are combined together (<a href="https://zenodo.org/api/files/bbca544a-15d0-40f3-8cc9-a3ee3c08fd7e/Binary_stack_for_local_thickness.zip?versionId=165e3938-b490-4e56-9c8e-a2084cb39d49">Binary_stack_for_local_thickness.zip</a>), a map of the local thickness is created (<a href="https://zenodo.org/api/files/bbca544a-15d0-40f3-8cc9-a3ee3c08fd7e/Local_thickness_map.zip?versionId=ce0a7dc7-5e3f-44a4-8881-cf84b6efd87c">Local_thickness_map.zip</a>). This map gives information on the largest diameter of the pixels labeled as cells in the binary stack.</p> <p>Hand-labeled slices or ground truths were drawn on the following slices: 80, 140, 200, 260, 340, 400, 440, 540, 620, 740, 800, 860, 940, 1060, 1140, 1240, 1300, 1400, 1480, 1540, 1600, 1690, 1740, 1840 (<a href="https://zenodo.org/api/files/bbca544a-15d0-40f3-8cc9-a3ee3c08fd7e/Hand_labelled_slices.tif?versionId=a21a13ac-fa47-4ef8-a903-ecc433787184">Hand_labelled_slices.tif</a>).</p> <p>Using the hand-labeled slices and the different images, a random-forest model was trained, which allowed to automatically segment the remaining slices of the stack (<a href="https://zenodo.org/api/files/bbca544a-15d0-40f3-8cc9-a3ee3c08fd7e/Fullstack_Prediction_Example-6_training_slices-6_testing_slices.zip?versionId=02b69e65-da85-492e-9b72-9b2b3ccd085f">Fullstack_Prediction_Example-6_training_slices-6...</a>).</p> <p>The source code for the segmentation program is available <a href="https://github.com/plant-microct-tools/leaf-traits-microct/tree/master">here</a>, and the source code for the testing used in the paper is available <a href="https://github.com/plant-microct-tools/leaf-traits-microct/tree/nb-slices-eval">here</a>.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.