Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

1,943

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

1,943 results for “machine learning”

Learn how ShareScore rates datasets ↗
dryad32/100

Passively Addressed Robotic Morphing Surface (PARMS) based on machine learning

Open the record for dataset details and reuse information.

publicJun 2023View details →
dryad32/100

Data for: Theropod dinosaur diversity of the lower English Wealden: analysis of a tooth-based fauna from the Wadhurst Clay Formation (Lower Cretaceous: Valanginian) via phylogenetic, discriminant and machine learning methods

Open the record for dataset details and reuse information.

publicDec 2024View details →
dryad32/100

Automated detection of lameness in sheep using machine learning approaches: novel insights into behavioural differences among lame and non-lame sheep

Open the record for dataset details and reuse information.

publicDec 2019View details →
dryad32/100

Data from: ClinicNet: machine learning for personalized order set recommendations

Open the record for dataset details and reuse information.

publicJun 2020View details →
dryad32/100

Data for: Estimating causal effects with machine learning: A guide for ecologists

Open the record for dataset details and reuse information.

publicOct 2025View details →
dryad32/100

Black box attack on machine learning assisted wide area monitoring and protection systems

Open the record for dataset details and reuse information.

publicApr 2021View details →
dryad32/100

Data for training AMSR2-CNN and its corresponding machine learning algorithm

Open the record for dataset details and reuse information.

publicSep 2023View details →
dryad32/100

Data from: Biogeographic and anthropogenic correlates of Aleutian Islands plant diversity: a machine-learning approach

Open the record for dataset details and reuse information.

publicSep 2018View details →
dryad32/100

Machine learning to extract muscle fascicle length changes from dynamic ultrasound images in real-time

Open the record for dataset details and reuse information.

publicJun 2021View details →
dryad32/100

Training dataset for Nile delta shoreline change prediction until 2050 using machine learning

Open the record for dataset details and reuse information.

publicSep 2025View details →
dryad32/100

Data from: Olfactory testing in Parkinson’s disease & REM behavior disorder: a machine learning approach

Open the record for dataset details and reuse information.

publicJan 2022View details →
dryad32/100

The choices we make and the impacts they have: Machine learning and species delimitation in North American box turtles (Terrapene spp.)

Open the record for dataset details and reuse information.

publicDec 2020View details →
dryad32/100

Benchmarking parametric and machine learning models for genomic prediction of complex traits

Open the record for dataset details and reuse information.

publicOct 2019View details →
dryad32/100

Calibration of probability predictions from machine-learning and statistical models

Open the record for dataset details and reuse information.

publicMar 2020View details →
dryad32/100

Precursor recommendation for inorganic synthesis by machine learning materials similarity from scientific literature

Open the record for dataset details and reuse information.

publicMay 2023View details →
dryad32/100

Data and code for: Explainable machine learning revealed the conditional response of biogenic isoprene to the changes in environmental factors at an urban site in the Yangtze River Delta, China

Open the record for dataset details and reuse information.

publicMay 2022View details →
zenodo28/100

Monitoring forest health using hyperspectral imagery: Does feature selection improve the performance of machine-learning techniques?

<p>This is a research compendium (RC) for the publication</p> <blockquote> <p>Monitoring forest health using hyperspectral imagery: Does feature selection improve the performance of machine-learning techniques?</p> </blockquote> <p>Code, figures, appendices and the manuscript can be found in the corresponding <a href="https://github.com/pat-s/2019-feature-selection">GitHub repository</a>.</p> <p>This RC is a static snapshot at the time of submission. The GitHub repository holds the latest version and may see changes after the publication was accepted.</p> <p><strong>Data sources and description</strong></p> <ul> <li><em>aoi.gpkg</em><strong>:</strong> Area of interest for downloading Sentinel-2 images. <em>Not used in the publication<strong>. </strong></em>Source:<em><strong> </strong></em>Custom.</li> <li><em>forest_mask.gpkg</em>: A forest/non-forest mask of the Basque Country. <em>Not used in the publication</em>. Source:<em><strong> </strong></em>Custom.</li> <li><em>hyperspectral.zip: </em>Hyperspectral remote sensing data used to extract reflectance values on the tree level. Source:<em><strong> </strong></em>Custom.</li> <li><em>plot-locations.gpkg: </em>Spatial location of the plots used in the study. Source:<em><strong> </strong></em>Custom.</li> <li><em>tree-in-situ-data-corrected.zip: </em>Corrected in-situ data containing defoliation information on the tree level. A correction of the spatial location was applied by the creators of the data. Source:<em><strong> </strong></em>Custom.</li> <li><em>tree-in-situ-data.zip: </em>First version of in-situ data containing defoliation information on the tree level.<em> Not used in the publication. </em>Source:<em><strong> </strong></em>Custom.</li> </ul> <p><strong>Licenses</strong></p> <p>All files are licensed under CC BY 4.0.</p>

opencc-by-4.0Jan 2020View details →
zenodo28/100

Detection of COVID-19 Infection from Routine Blood Exams with Machine Learning: a Feasibility Study

<p>This upload consists of the dataset employed in the publication &quot;Detection of COVID-19 Infection from Routine Blood Exams with Machine Learning: a Feasibility Study&quot;. The paper was accepted for publication at Journal of Medical Systems (Springer), and a pre-print version is also available on MedRXiv and has been attached to the upload for further reference. If you decide to use or reference this dataset (or the related work) please cite the journal version (details will be added as soon as available).</p> <p>The dataset&nbsp;consists of&nbsp;280 records of patients admitted to the San Raffaele Hospital (Milan, Italy), annotated with a collection of&nbsp; hematochemical values from routine blood exams (namely: white blood cells counts, and the platelets, CRP, AST, ALT, GGT, ALP, LDH plasma levels), and a target variable that describes COVID-19 positivity/negativity (in the target column, class 2 and class 1 can both be treated as COVID-19 positive patients)</p> <p>&nbsp;</p> <p>The abstract of the supporting publication follows:</p> <p>&nbsp;</p> <p><strong>Background</strong> - The COVID-19 pandemia due to the SARS-CoV-2 coronavirus, in its first 4 months since its outbreak, has to date reached more than 200 countries worldwide with more than 2 million confirmed cases (probably a much higher number of infected), and almost 200,000 deaths. Amplification of viral RNA by (real time) reverse transcription polymerase chain reaction (rRT-PCR) is the current gold standard test for confirmation of infection, although it presents known shortcomings: long turnaround times (3-4 hours to generate results), potential shortage of reagents, false-negative rates as large as 15-20%, the need for certified laboratories, expensive equipment and trained personnel. Thus there is a need for alternative, faster, less expensive and more accessible tests.</p> <p><strong>Material and methods</strong> - We developed two machine learning classification models using hematochemical values from routine blood exams (namely: white blood cells counts, and the platelets, CRP, AST, ALT, GGT, ALP, LDH plasma levels) drawn from 279 patients who, after being admitted to the San Raffaele Hospital (Milan, Italy) emergency-room with COVID-19 symptoms, were screened with the rRT-PCR test performed on respiratory tract specimens. Of these patients, 177 resulted positive, whereas 102 received a negative response.</p> <p><strong>Results</strong> - We have developed two machine learning models, to discriminate between patients who are either positive or negative to the SARS-CoV-2: their accuracy ranges between 82% and 86%, and sensitivity between 92% e 95%, so comparably well with respect to the gold standard. We also developed an interpretable Decision Tree model as a simple decision aid for clinician interpreting blood tests (even off-line) for COVID-19 suspect cases.</p> <p><strong>Discussion</strong> - This study demonstrated the feasibility and clinical soundness of using blood tests analysis and machine learning as an alternative to rRT-PCR for identifying COVID-19 positive patients. This is especially useful in those countries, like developing ones, suffering from shortages of rRT-PCR reagents and specialized laboratories. We made available a Web-based tool for clinical reference and evaluation. This tool is available at https://covid19-blood-ml.herokuapp.com.</p>

opencc-by-2.0Apr 2020View details →
zenodo28/100

Development and validation of a machine learning model for use as an automated artificial intelligence tool to predict mortality risk in patients with COVID-19

<p><strong>Background</strong></p> <p>New York City quickly became an epicenter of the COVID-19 pandemic. Due to a sudden and massive increase in patients during COVID-19 pandemic, healthcare providers incurred an exponential increase in workload which created a strain on the staff and limited resources. As this is a new infection, predictors of morbidity and mortality are not well characterized.</p> <p><strong>Methods</strong></p> <p>We developed a prediction model to predict patients at risk for mortality using only laboratory, vital and demographic information readily available in the electronic health record on more than 3000 hospital admissions with COVID-19. A variable importance algorithm was used for interpretability and understanding of performance and predictors.</p> <p><strong>Findings</strong></p> <p>We built a model with 84-97% accuracy to identify predictors and patients with high risk of mortality, and developed an automated artificial intelligence (AI) notification tool that does not require manual calculation by the busy clinician. Oximetry, respirations, blood urea nitrogen, lymphocyte percent, calcium, troponin and neutrophil percentage were important features and key ranges were identified that contributed to a 50% increase in patients&rsquo; mortality prediction score. With an increasing negative predictive value (NPV) starting 0.90 after the second day of admission, we are able more confidently able identify likely survivors. This study serves as a use case of a model with visualizations to aide clinicians with a better understanding of the model and predictors of mortality. Additionally, an example of the operationalization of the model via an AI notification tool is illustrated.</p>

opencc-by-4.0Jun 2020View details →
zenodo28/100

Machine learning based root stress calculation model for gears with progressive curved path of contact

<p><em>A machine learning based model for calculation of root stress in gears with a progressive curved path of contact is proposed. The prediction model is used as a surrogate model to the finite element analysis simulations, from which the data was collected.&nbsp;The results of the validation of all the methods are presented. Further validation was done with new simulations.</em></p>

opencc-by-4.0Dec 2019View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record