Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
1,943
datasets available to search
ShareScore release 0.9.0
Dataset results
1,943 results for “machine learning”
Passively Addressed Robotic Morphing Surface (PARMS) based on machine learning
Open the record for dataset details and reuse information.
Data for: Theropod dinosaur diversity of the lower English Wealden: analysis of a tooth-based fauna from the Wadhurst Clay Formation (Lower Cretaceous: Valanginian) via phylogenetic, discriminant and machine learning methods
Open the record for dataset details and reuse information.
Automated detection of lameness in sheep using machine learning approaches: novel insights into behavioural differences among lame and non-lame sheep
Open the record for dataset details and reuse information.
Data from: ClinicNet: machine learning for personalized order set recommendations
Open the record for dataset details and reuse information.
Data for: Estimating causal effects with machine learning: A guide for ecologists
Open the record for dataset details and reuse information.
Black box attack on machine learning assisted wide area monitoring and protection systems
Open the record for dataset details and reuse information.
Data for training AMSR2-CNN and its corresponding machine learning algorithm
Open the record for dataset details and reuse information.
Data from: Biogeographic and anthropogenic correlates of Aleutian Islands plant diversity: a machine-learning approach
Open the record for dataset details and reuse information.
Machine learning to extract muscle fascicle length changes from dynamic ultrasound images in real-time
Open the record for dataset details and reuse information.
Training dataset for Nile delta shoreline change prediction until 2050 using machine learning
Open the record for dataset details and reuse information.
Data from: Olfactory testing in Parkinson’s disease & REM behavior disorder: a machine learning approach
Open the record for dataset details and reuse information.
The choices we make and the impacts they have: Machine learning and species delimitation in North American box turtles (Terrapene spp.)
Open the record for dataset details and reuse information.
Benchmarking parametric and machine learning models for genomic prediction of complex traits
Open the record for dataset details and reuse information.
Calibration of probability predictions from machine-learning and statistical models
Open the record for dataset details and reuse information.
Precursor recommendation for inorganic synthesis by machine learning materials similarity from scientific literature
Open the record for dataset details and reuse information.
Data and code for: Explainable machine learning revealed the conditional response of biogenic isoprene to the changes in environmental factors at an urban site in the Yangtze River Delta, China
Open the record for dataset details and reuse information.
Monitoring forest health using hyperspectral imagery: Does feature selection improve the performance of machine-learning techniques?
<p>This is a research compendium (RC) for the publication</p> <blockquote> <p>Monitoring forest health using hyperspectral imagery: Does feature selection improve the performance of machine-learning techniques?</p> </blockquote> <p>Code, figures, appendices and the manuscript can be found in the corresponding <a href="https://github.com/pat-s/2019-feature-selection">GitHub repository</a>.</p> <p>This RC is a static snapshot at the time of submission. The GitHub repository holds the latest version and may see changes after the publication was accepted.</p> <p><strong>Data sources and description</strong></p> <ul> <li><em>aoi.gpkg</em><strong>:</strong> Area of interest for downloading Sentinel-2 images. <em>Not used in the publication<strong>. </strong></em>Source:<em><strong> </strong></em>Custom.</li> <li><em>forest_mask.gpkg</em>: A forest/non-forest mask of the Basque Country. <em>Not used in the publication</em>. Source:<em><strong> </strong></em>Custom.</li> <li><em>hyperspectral.zip: </em>Hyperspectral remote sensing data used to extract reflectance values on the tree level. Source:<em><strong> </strong></em>Custom.</li> <li><em>plot-locations.gpkg: </em>Spatial location of the plots used in the study. Source:<em><strong> </strong></em>Custom.</li> <li><em>tree-in-situ-data-corrected.zip: </em>Corrected in-situ data containing defoliation information on the tree level. A correction of the spatial location was applied by the creators of the data. Source:<em><strong> </strong></em>Custom.</li> <li><em>tree-in-situ-data.zip: </em>First version of in-situ data containing defoliation information on the tree level.<em> Not used in the publication. </em>Source:<em><strong> </strong></em>Custom.</li> </ul> <p><strong>Licenses</strong></p> <p>All files are licensed under CC BY 4.0.</p>
Detection of COVID-19 Infection from Routine Blood Exams with Machine Learning: a Feasibility Study
<p>This upload consists of the dataset employed in the publication "Detection of COVID-19 Infection from Routine Blood Exams with Machine Learning: a Feasibility Study". The paper was accepted for publication at Journal of Medical Systems (Springer), and a pre-print version is also available on MedRXiv and has been attached to the upload for further reference. If you decide to use or reference this dataset (or the related work) please cite the journal version (details will be added as soon as available).</p> <p>The dataset consists of 280 records of patients admitted to the San Raffaele Hospital (Milan, Italy), annotated with a collection of hematochemical values from routine blood exams (namely: white blood cells counts, and the platelets, CRP, AST, ALT, GGT, ALP, LDH plasma levels), and a target variable that describes COVID-19 positivity/negativity (in the target column, class 2 and class 1 can both be treated as COVID-19 positive patients)</p> <p> </p> <p>The abstract of the supporting publication follows:</p> <p> </p> <p><strong>Background</strong> - The COVID-19 pandemia due to the SARS-CoV-2 coronavirus, in its first 4 months since its outbreak, has to date reached more than 200 countries worldwide with more than 2 million confirmed cases (probably a much higher number of infected), and almost 200,000 deaths. Amplification of viral RNA by (real time) reverse transcription polymerase chain reaction (rRT-PCR) is the current gold standard test for confirmation of infection, although it presents known shortcomings: long turnaround times (3-4 hours to generate results), potential shortage of reagents, false-negative rates as large as 15-20%, the need for certified laboratories, expensive equipment and trained personnel. Thus there is a need for alternative, faster, less expensive and more accessible tests.</p> <p><strong>Material and methods</strong> - We developed two machine learning classification models using hematochemical values from routine blood exams (namely: white blood cells counts, and the platelets, CRP, AST, ALT, GGT, ALP, LDH plasma levels) drawn from 279 patients who, after being admitted to the San Raffaele Hospital (Milan, Italy) emergency-room with COVID-19 symptoms, were screened with the rRT-PCR test performed on respiratory tract specimens. Of these patients, 177 resulted positive, whereas 102 received a negative response.</p> <p><strong>Results</strong> - We have developed two machine learning models, to discriminate between patients who are either positive or negative to the SARS-CoV-2: their accuracy ranges between 82% and 86%, and sensitivity between 92% e 95%, so comparably well with respect to the gold standard. We also developed an interpretable Decision Tree model as a simple decision aid for clinician interpreting blood tests (even off-line) for COVID-19 suspect cases.</p> <p><strong>Discussion</strong> - This study demonstrated the feasibility and clinical soundness of using blood tests analysis and machine learning as an alternative to rRT-PCR for identifying COVID-19 positive patients. This is especially useful in those countries, like developing ones, suffering from shortages of rRT-PCR reagents and specialized laboratories. We made available a Web-based tool for clinical reference and evaluation. This tool is available at https://covid19-blood-ml.herokuapp.com.</p>
Development and validation of a machine learning model for use as an automated artificial intelligence tool to predict mortality risk in patients with COVID-19
<p><strong>Background</strong></p> <p>New York City quickly became an epicenter of the COVID-19 pandemic. Due to a sudden and massive increase in patients during COVID-19 pandemic, healthcare providers incurred an exponential increase in workload which created a strain on the staff and limited resources. As this is a new infection, predictors of morbidity and mortality are not well characterized.</p> <p><strong>Methods</strong></p> <p>We developed a prediction model to predict patients at risk for mortality using only laboratory, vital and demographic information readily available in the electronic health record on more than 3000 hospital admissions with COVID-19. A variable importance algorithm was used for interpretability and understanding of performance and predictors.</p> <p><strong>Findings</strong></p> <p>We built a model with 84-97% accuracy to identify predictors and patients with high risk of mortality, and developed an automated artificial intelligence (AI) notification tool that does not require manual calculation by the busy clinician. Oximetry, respirations, blood urea nitrogen, lymphocyte percent, calcium, troponin and neutrophil percentage were important features and key ranges were identified that contributed to a 50% increase in patients’ mortality prediction score. With an increasing negative predictive value (NPV) starting 0.90 after the second day of admission, we are able more confidently able identify likely survivors. This study serves as a use case of a model with visualizations to aide clinicians with a better understanding of the model and predictors of mortality. Additionally, an example of the operationalization of the model via an AI notification tool is illustrated.</p>
Machine learning based root stress calculation model for gears with progressive curved path of contact
<p><em>A machine learning based model for calculation of root stress in gears with a progressive curved path of contact is proposed. The prediction model is used as a surrogate model to the finite element analysis simulations, from which the data was collected. The results of the validation of all the methods are presented. Further validation was done with new simulations.</em></p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.