Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

1,943

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

1,943 results for “machine learning”

Learn how ShareScore rates datasets ↗
zenodo36/100

Machine learning reactive mixing dataset-5

<p>Reactive-mixing dataset-2&nbsp;for machine learning analyses</p>

opencc-by-4.0Jan 2021View details →
zenodo36/100

Machine learning reactive mixing dataset-12

<p>Reactive-mixing dataset-3.3&nbsp;for machine learning analyses</p>

opencc-by-4.0Jan 2021View details →
zenodo36/100

Machine learning reactive mixing dataset-10

<p>Reactive-mixing dataset-3.1&nbsp;for machine learning analyses</p>

opencc-by-4.0Jan 2021View details →
zenodo36/100

Machine learning reactive mixing dataset-16

<p>Reactive-mixing dataset-4.3&nbsp;for machine learning analyses</p>

opencc-by-4.0Jan 2021View details →
zenodo36/100

Machine learning reactive mixing dataset-15

<p>Reactive-mixing dataset-4.2&nbsp;for machine learning analyses</p>

opencc-by-4.0Jan 2021View details →
zenodo36/100

Machine learning reactive mixing dataset-13

<p>Reactive-mixing dataset-4&nbsp;for machine learning analyses</p>

opencc-by-4.0Jan 2021View details →
zenodo36/100

Machine learning reactive mixing dataset-21

<p>Reactive-mixing dataset-6.1 for machine learning analyses</p>

opencc-by-4.0Jan 2021View details →
zenodo36/100

Machine learning reactive mixing dataset-14

<p>Reactive-mixing dataset-4.1 for machine learning analyses</p>

opencc-by-4.0Jan 2021View details →
zenodo36/100

Machine learning reactive mixing dataset-11

<p>Reactive-mixing dataset-3.2&nbsp;for machine learning analyses</p>

opencc-by-4.0Jan 2021View details →
zenodo36/100

Machine learning reactive mixing dataset-20

<p>Reactive-mixing dataset-6&nbsp;for machine learning analyses</p>

opencc-by-4.0Jan 2021View details →
zenodo36/100

Machine learning reactive mixing dataset-18

<p>Reactive-mixing dataset-5.1 for machine learning analyses</p>

opencc-by-4.0Jan 2021View details →
zenodo36/100

Machine learning reactive mixing dataset-19

<p>Reactive-mixing dataset-5.2&nbsp;for machine learning analyses</p>

opencc-by-4.0Jan 2021View details →
zenodo36/100

Machine learning reactive mixing dataset-17

<p>Reactive-mixing dataset-5&nbsp;for machine learning analyses</p>

opencc-by-4.0Jan 2021View details →
zenodo36/100

Machine learning reactive mixing dataset-22

<p>Reactive-mixing dataset-6.2&nbsp;for machine learning analyses</p>

opencc-by-4.0Jan 2021View details →
zenodo36/100

Dataset for Black Tea Fermentation Detection based on Image Processing and Machine Learning Techniques

<p>This is a dataset on black tea fermentation. The dataset contains black tea fermentation conditions and images. The fermentation conditions captured are: temperature, humidty and time. The images belong to black tea as they underwent the fermentation process. The dataset was collected in Sisibo tea factory, Kenya in July and August 2020.</p>

opencc-by-4.0Jan 2021View details →
zenodo36/100

Data from Understanding X-ray spectroscopy of carbonaceous materials by combining experiments, density functional theory and machine learning. Parts I and II.

<p>Understanding X-ray spectroscopy of carbonaceous materials by combining experiments, density functional theory, and machine learning; Parts I and II.</p> <p>This data-set is published in Refs. [1-2] and it is now made openly accessible. The data-set consists of computational X-ray spectroscopy fingerprints of plain and functionalized amorphous carbon. This data can be used in interpretation of experimental spectroscopy data (XAS and XPS). The spectra are averages of certain atomic environments, that are described in the publications. Standard deviation is included in the third column. Please, feel free to use the data-set, and if you do so, remember to cite Refs. [1-2] and this source. If there are any questions, please contact the corresponding author.&nbsp;</p> <p>&nbsp;</p> <p>[1] A. Aarva, V. L. Deringer, S. Sainio, T. Laurila, andM. A. Caro, &ldquo;Understanding X-ray spectroscopy of carbonaceous materials by combining experiments, density functional theory, and machine learning. Part I: Fingerprint spectra,&rdquo; Chem. Mater. 31, 9243&ndash;9255 (2019).</p> <p>[2] A. Aarva, V. L. Deringer, S. Sainio, T. Laurila, andM. A. Caro, &ldquo;Understanding X-ray spectroscopy of carbonaceous materials by combining experiments, density functional theory, and machine learning. Part II: Quantitative fitting of spectra,&rdquo; Chem. Mater. 31, 9256&ndash;9267(2019).</p> <p>&nbsp;</p> <p>Funding and resources for the work are acknowledged as follows:</p> <p>Funding from the Academy of Finland (project no.285526) and the computational resources provided for this project by CSC &ndash; IT Center for Science are gratefully acknowledged. M. A. C. acknowledges personal funding from the Academy of Finland under project no. 310574.V. L. D. acknowledges a Leverhulme Early Career Fellowship and support from the Isaac Newton Trust. M. A. C.and V. L. D. are grateful for travelling support from the HPC-Europa3 program under the auspices of the European Union&rsquo;s Horizon 2020 framework (grant agreement no. 730897). Use of the Stanford Synchrotron Radiation Lightsource, SLAC National Accelerator Laboratory, is supported by the U.S. Department of Energy, Office of Science, Office of Basic Energy Sciences under contract no. DE-AC02-76SF00515. S. S. acknowledges personal funding from Instrumentarium Science Foundation and the Walter Ahlstr&ouml;m Foundation.</p> <p>&nbsp;</p>

opencc-by-4.0Jan 2021View details →
zenodo36/100

Exploratory analysis using machine learning of predictive factors for falls in persons with type 2 diabetes: A Longitudinal Study

<p>The risk of falls in elderly individuals with diabetes was reported to be 1.5 - 3 times higher than in those without diabetes. However, it is not clear what risk factors are strongly related to falls in those with diabetes. In this study, we aimed to investigate the status of falls and to identify important risk factors for falls in persons with type 2 diabetes (T2D) including the non-elderly. Participants were 316 persons with T2D who were admitted to the University of Tsukuba Hospital for treatment of diabetes. They were assessed for medical history, laboratory data and physical capabilities during the hospitalization and were given a questionnaire on falls one year after discharge. Two different statistical models, logistic regression and random forest classifier, were used to investigate important predictors of falls. The response rate to the survey was 72%; of the 226 respondents, there were 129 males and 97 females (median age 62 years). The fall rate during the first year after discharge was 19% and increased with age; fall rates were 17% for those &lt;60 years, 20% for those aged 60 &ndash; 69 years and 24% for those &ge;70 years. Logistic regression revealed that knee extension strength (&beta;= -0.698, P = 0.002), fasting C-peptide (F-CPR) level (&beta;= 0.492, P = 0.009) and dorsiflexion strength (&beta;= -0.432, P = 0.047) were independent predictors of falls. The random forest classifier placed knee extension strength (covariate importance = 0.304), grip strength (0.234), F-CPR level (0.232) and dorsiflexion strength (0.230) in the top 4 important variables for falls. The rate of falls in persons with T2D was high even in middle age. Lower extremity muscle weakness as well as elevated F-CPR levels and reduced grip strength were shown to be important risk factors for falls in T2D.</p>

opencc-by-4.0Jun 2021View details →
dryad36/100

Data from: A demonstration of unsupervised machine learning in species delimitation

One major challenge to delimiting species with genetic data is successfully differentiating population structure from species-level divergence, an issue exacerbated in taxa inhabiting naturally fragmented habitats. Many fields of science are now using machine learning, and in evolutionary biology supervised machine learning has recently been used to infer species boundaries. These supervised methods require training data with associated labels. Conversely, unsupervised machine learning (UML) uses inherent data structure and does not require user-specified training labels, potentially providing more objectivity in species delimitation. Here we demonstrate the utility of three UML approaches (random forests, variational autoencoders, t-distributed stochastic neighbor embedding) for species delimitation in an arachnid taxon with high population genetic structure (Opiliones, Laniatores, Metanonychus). We find that UML approaches successfully cluster samples according to species-level divergences and not high levels of population structure, while model-based validation methods severely over-split putative species. UML offers intuitive data visualization in two-dimensional space, the ability to accommodate various data types, and has potential in many areas of systematic and evolutionary biology. We argue that machine learning methods are ideally suited for species delimitation and may perform well in many natural systems and across taxa with diverse biological characteristics.

opencc-zeroJul 2019View details →
dryad36/100

Data from: Unsupervised machine learning reveals mimicry complexes in bumble bees occur along a perceptual continuum

Müllerian mimicry theory states that frequency dependent selection should favour geographic convergence of harmful species onto a shared colour pattern. As such, mimetic patterns are commonly circumscribed into discrete mimicry complexes each containing a predominant phenotype. Outside a few examples in butterflies, the location of transition zones between mimicry complexes and the factors driving mimicry zones has rarely been examined. To infer the patterns and processes of Müllerian mimicry, we integrate large-scale data on the geographic distribution of colour patterns of social bumble bees across the contiguous United States and use these to quantify colour pattern mimicry using an innovative, unsupervised machine learning approach based on computer vision. Our data suggest that bumble bees exhibit geographically clustered, but sometimes imperfect colour patterns and that mimicry patterns gradually transition spatially, rather than exhibit discrete boundaries. Additionally, examination of colour pattern transition zones of three comimicking, polymorphic species, where active selection is driving phenotype frequencies, revealed their transition zones to differ in location within a broad region of poor mimicry. Potential factors influencing mimicry transition zone dynamics are discussed.

opencc-zeroAug 2019View details →
zenodo36/100

Classification of word levels with usage frequency, expert opinions and machine learning

<p>This dataset includes classification of English words according to CEFR language levels. It can be used in various educational applications including determining levels of text that is appropriate for students learning English.&nbsp;</p> <p>For each word, part-of-speech, the word lemma and usage frequency is provided. For words that have no survey results, a machine learning based methodology is used to predict levels. These predictions are also included as a separate file. This data is released as part of the&nbsp;submission process to British Journal of Educational Technology Special Issue on Open Data.</p> <p>The included readme.pdf file contains a detailed&nbsp;description of data.&nbsp;</p>

opencc-by-4.0Oct 2014View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record