Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

1,943

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

1,943 results for “machine learning”

Learn how ShareScore rates datasets ↗
zenodo36/100

Comparing Machine Learning Classifiers and Linear/Logistic Regression to Explore the Relationship between Hand Dimensions and Demographic Characteristics

<p>-----------------------------------------------------------------------------------------------------------------</p> <p>Data for "<strong>Comparing Machine Learning Classifiers and Linear/Logistic Regression to Explore the Relationship between Hand Dimensions and Demographic Characteristics</strong>" (PLOSONE)</p> <p>Oscar Miguel-Hurtado<sup>1</sup>, Richard Guest<sup>1</sup>, Sarah V. Stevenage<sup>2</sup>,Greg J. Neil<sup>2</sup>,Sue Black<sup>3</sup><br>  </p> <ul> <li><sup>1</sup> School of Engineering and Digital Arts, University of Kent, Canterbury, UK</li> <li><sup>2</sup> Department of Psychology, University of Southampton, Southampton, UK</li> <li><sup>3</sup> Centre for Anatomy and Human Identification, University of Dundee, Dundee, UK</li> </ul> <p>-----------------------------------------------------------------------------------------------------------------</p> <p>For more information please contact: O.Miguel-Hurtado-98@kent.ac.uk (Oscar Miguel)</p> <p>-----------------------------------------------------------------------------------------------------------------</p> <p>The zip contains right and left hand geometry images  from 112 participants. The images were captured using a Nikon D200 SLR camera (format: jpg, size: 3504x2336 pixels), with both the palm of the hand and camera facing downwards. Participants placed each hand on an acetate sheet with a series of positioning pegs.</p> <p>-----------------------------------------------------------------------------------------------------------------</p> <p>The excel contains a series of length measurements (based on the underlying skeleton of the hand) manually extracted (see Figure 1 for details) along with demographic information from the participants: sex (male or female), height (in cm), weight (in kg) and foot size (in UK sizes).</p>

opencc-by-nc-4.0Oct 2016View details →
zenodo36/100

Biological and environmental data for a study on transferability of statistical and machine learning models using North Sea Macrozoobenthos

<p>General</p> <p>Data documented here are not the product of our research but was scraped from various sources and processed - so no genuine reupload. This collection is a contribution to reproduceable reseach. All datasets are given in "RData" binary format</p> <p> </p> <p>Data description</p> <p>majornorthseabenthos </p> <p>This is macrozoobenthos data as data frame scraped from the GBIF repository (gbif.org). Species are  Corbula gibba, Tellina fabula, Turritella communis, Euspira pulchella, Corystes cassive- launus, Upogebia deltaura, Lanice conchilega, Nephtys hombergii, Echinocardium cordatum, and Amphiura filiformis. data was postprocessed to have only single occurrence fon the approxinatel 1x1 km grid used for this study. Also, occurrences closer than 5 km  close to shore were removed - including occurrences on land.</p> <p> </p> <p>Predictors</p> <p>A SpatialPixelsDataFrame in EPSG 4326 with five layers: Median grain size in micrometers, mud content in percent (both MUDAB database), water depth in meters above MSL (Weatherall et al, 2015), modelled average bottom shear stress from waves in N/sqrm (The Wamdi Group, 1988) and climatologival average winter bottom water temperature in deg. C (Stips et al, 2004).</p> <p> </p> <p> </p> <p>References</p> <p>Stips A, Bolding K, Pohlmann T, Burchard H (2004) Simulating the temporal and spatial dy- namics of the North Sea using the new model GETM (general estuarine transport model). Ocean Dynamics 54(2):266–283</p> <p>The Wamdi Group (1988) The WAM model-a third generation ocean wave prediction model. Journal of Physical Oceanography 18(12):1775–1810</p> <p>Weatherall P, Marks K, Jakobsson M, Schmitt T, Tani S, Arndt JE, Rovere M, Chayes D, Ferrini V, Wigley R (2015) A new digital bathymetric model of the world’s oceans. Earth and Space Science 2(8):331–345</p> <p> </p>

opencc-by-4.0Apr 2017View details →
zenodo36/100

Data mining-based machine learning methods for improving hydrological data: a case study of salinity field in the Western Arctic Ocean

<p><span><span>Salinity variations in Arctic Ocean determine the strength of stratification,</span> <span>ocean circulation, and biogeochemical cycles.&nbsp;Therefore, a<span>ccurate </span>salinity product is of great significance for our study of the Arctic Ocean. The mean density structure and wind-driven surface circulation of the Arctic Ocean are largely dominated by the anti-cyclonic Beaufort Gyre in the Canadian Basin, along with the Transpolar Drift<span> (Hall </span><span>et al.,2022)</span>. We focus on the salinity in Western Arctic Ocean. Multiple machine learning methods were used to reconstruct annual salinity product in the Western Arctic Ocean temporal for the period 2003-2022.</span></span></p>

opencc-by-4.0Apr 2024View details →
zenodo36/100

Data for "A multimodal machine learning fused global 0.1° daily evapotranspiration dataset from 1950-2022" (2000-2024)

<p>The data contains simulation results from 2000-2024, 25 years total.</p> <p>You can access the remaining part of the dataset via Qingchen Xu and Lu Li (2025) using the following reference:</p> <p>Qingchen Xu, &amp; Lu Li. (2025). Data for "A multimodal machine learning fused global 0.1&deg; daily evapotranspiration dataset from 1950-2022" (1950-1974) [Data set]. Zenodo. https://doi.org/10.5281/zenodo.15671245</p> <p>Qingchen Xu, &amp; Lu Li. (2025). Data for "A multimodal machine learning fused global 0.1&deg; daily evapotranspiration dataset from 1950-2022" (1975-1999) [Data set]. Zenodo. https://doi.org/10.5281/zenodo.15671253</p>

opencc-by-4.0Apr 2024View details →
zenodo36/100

Data for "A multimodal machine learning fused global 0.1° daily evapotranspiration dataset from 1950-2022" (1975-1999)

<p>The data contains simulation results from 1975-1999, 25 years total.</p> <p>You can access the remaining part of the dataset via Qingchen Xu and Lu Li (2025) using the following reference:</p> <p>Qingchen Xu, &amp; Lu Li. (2025). Data for "A multimodal machine learning fused global 0.1&deg; daily evapotranspiration dataset from 1950-2022" (1950-1974) [Data set]. Zenodo. https://doi.org/10.5281/zenodo.15671245</p> <p>Qingchen Xu, &amp; Lu Li. (2025). Data for "A multimodal machine learning fused global 0.1&deg; daily evapotranspiration dataset from 1950-2022" (2000-2024) [Data set]. Zenodo. https://doi.org/10.5281/zenodo.15671254</p>

opencc-by-4.0Apr 2024View details →
zenodo36/100

Data for "A multimodal machine learning fused global 0.1° daily evapotranspiration dataset from 1950-2022" (1950-1974)

<p>The data contains simulation results from 1950-1974, 25 years total.</p> <p>You can access the remaining part of the dataset via Qingchen Xu and Lu Li (2025) using the following reference:<br>Qingchen Xu, &amp; Lu Li. (2025). Data for "A multimodal machine learning fused global 0.1&deg; daily evapotranspiration dataset from 1950-2022" (1975-1999) [Data set]. Zenodo. https://doi.org/10.5281/zenodo.15671253</p> <p>Qingchen Xu, &amp; Lu Li. (2025). Data for "A multimodal machine learning fused global 0.1&deg; daily evapotranspiration dataset from 1950-2022" (2000-2024) [Data set]. Zenodo. https://doi.org/10.5281/zenodo.15671254</p>

opencc-by-4.0Apr 2024View details →
zenodo36/100

Global snow water equivalent product derived from machine learning model trained with in situ measurement data

<p>This dataset is a global snow water equivalent dataset using machine learning trained with in-situ measurements. The temporal resolution of the SWEML product is daily, and the spatial resolution is 0.25˚ (approximately 25km). It covers latitudes of 90S to 90N and longitudes of 180W to 180E with global scales, excluding Antarctica. The dataset is provided in NetCDF format, organized by year. Each year contains daily SWE data, including leap days in leap years.</p>

opencc-by-4.0Jul 2024View details →
zenodo36/100

Machine learning refinement of in situ images acquired by low electron dose LC-TEM

<p>This is a dataset of images acquired in vacuum condition and corresponding images in a solution listed in the text file list_dataset.txt.&nbsp;</p><p>&nbsp;</p><p>The code is available at <a href="https://github.com/hiroyasukatsuno/Machine-learning-refinement-of-images-acquired-by-LC-TEM">this website</a>:</p><p>https://github.com/hiroyasukatsuno/Machine-learning-refinement-of-images-acquired-by-LC-TEM/</p><p><br>Equipment of TEM:</p><ul><li>field-emission gun (JEM-2100F, JEOL, Tokyo)</li><li>OneView IS (Gatan, Inc., Pleasanton, CA, USA)</li></ul><p>&nbsp;</p>

opencc-by-4.0Oct 2023View details →
zenodo36/100

POMFinder: Identifying polyoxometalate cluster structures from pair distribution function data using explainable machine learning

<p>Databases that were used to train the POMFinder ML model.</p>

opencc-by-4.0Oct 2023View details →
dryad36/100

Data archive for: Exploring the use of machine learning to improve vertical profiles of temperature and moisture

<p>Vertical profiles of temperature and dewpoint are useful in predicting deep convection that leads to severe weather that threatens property and lives. Currently, forecasters rely on observations from radiosonde launches and numerical weather prediction (NWP) models. Radiosonde observations are, however, temporally and spatially sparse, and NWP models contain inherent errors that influence short-term predictions of high-impact events. This work explores using machine learning (ML) to postprocess NWP model forecasts, combining them with satellite data to improve vertical profiles of temperature and dewpoint. We focus on different ML architectures, loss functions, and input features to optimize predictions. Because we are predicting vertical profiles at 256 levels in the atmosphere, this work provides a unique perspective at using ML for 1-D tasks. Compared to baseline profiles from the Rapid Refresh (RAP), ML predictions offer the largest improvement for dewpoint, particularly in the mid- and upper-atmosphere.  emperature improvements are modest, but CAPE values are improved by up to 40%. Feature importance analyses indicate that the ML models are primarily improving incoming RAP biases. While additional model and satellite data offer some improvement to the predictions, architecture choice is more important than feature selection in fine-tuning the results. Our proposed deep residual UNet performs the best by leveraging spatial context from the input RAP profiles; however, the results are remarkably robust across model architecture. Further, uncertainty estimates for every level are well-calibrated and can provide useful information to forecasters.</p>

opencc-zeroOct 2023View details →
zenodo36/100

IMBALANCED MACHINE LEARNING CLASSIFICATION MODELS FOR REMOVAL BIOSIMILAR DRUGS AND INCREASED ACTIVITY IN PATIENTS WITH RHEUMATIC DISEASES

<p>Objective: Predict long-term disease worsening and the removal of biosimilar medication in patients with rheumatic diseases.</p><p>Methodology: Observational, retrospective, and descriptive study. Review of a database of patients with immune-mediated inflammatory rheumatic diseases. Disease worsening and removing biosimilars are imbalanced variables, that require using imbalanced machine learning models selected based on their superior f1-scores and great accuracy. Previously, we selected the most important variables using mutual information tests.</p><p>Results: The best imbalanced machine learning models to predict disease worsening and the removal of the biosimilar obtained f1-scores of 0.52 and 0.63, respectively. Both models are decision trees. In the first one, two important factors are switching of biosimilar and age, and in the second, the relevant variables are optimization and the value of the initial CRP.&nbsp;</p><p>Conclusions: Biosimilar drugs do not always work well for rheumatic diseases. We obtained two imbalanced machine learning models to detect those cases, where the drug should be removed or where the activity of the disease increases from low to high. Our decision trees use variables, such as age or switching, not considered in previous studies.</p>

opencc-by-4.0Nov 2022View details →
dryad36/100

NLP and machine learning to measure peace from news media

<p>"Hate speech" can mobilize violence and destruction.  What are the characteristics of "peace speech" that reflect and support the social processes that maintain peace?  In this study we used a data driven, machine learning approach to identify the words most associated with lower-peace versus higher-peace countries. Logistic regression and random forest classifiers were trained using five respected, traditional peace indices: Global Peace Index, Positive Peace Index, World Happiness Index, Fragile States Index, and Human Development Index. The feature inputs into the machine learning model were the word frequencies from the news media in each country and the output classifications were the level of peace in that country.  The machine learning model was successful in properly classifying the level of peace from the news media in a country (both accuracy and F1: 96% - 100%). We also used that trained machine model to create a machine learning peace index that measured the level of peace in countries, including countries not in the training set, which correlated with the average of those five traditional peace indices (r-squared = 0.8349). Using the random forest feature importance method we found that the words in news media in lower-peace countries were characterized by words related to government, order, control and fear (such as government, state, law, security and court), while higher-peace countries were characterized by an increased prevalence of words related to optimism for the future and fun (such as time, like, home, believe and game).</p>

opencc-zeroNov 2023View details →
zenodo36/100

Data Release for "Revisiting the evidence for precession in GW200129 with machine learning noise mitigation"

<p>Cleaned gravitational-wave data frame for the Livingston interferometer around GW200129 using NLSub, a machine-learning algorithm. For more details, see the publication on ArXiv: <a href="https://arxiv.org/abs/2311.09921">https://arxiv.org/abs/2311.09921</a>.</p><p>The data frame can be loaded in Python using e.g. GWPy TimeSeries class:</p><blockquote><p>from gwpy.timeseries import TimeSeries<br>tseries = TimeSeries.read('L-L1_DCS-CALIB_STRAIN_CLEAN_C01_NLSUB_P2300358_v4-1264314077-4078.hdf5')</p></blockquote><p>&nbsp;</p>

opencc-by-4.0Nov 2023View details →
zenodo36/100

Gauging Ambient Environmental Carbon Dioxide Concentration Solely Using Biometric Observations: A Machine Learning Approach

<p>Data and code in form of Jupyter Notebook to accompany an unpublished paper with the title " Gauging Ambient Environmental Carbon Dioxide Concentration Solely Using Biometric Observations: A Machine Learning Approach". This work makes use of biometric variables of a participant to estimate the inhaled carbon dioxide in microenvironments and understand various physiological and cognitive responses. &nbsp;</p><p>Github link: <a href="https://github.com/mi3nts/Estimate-CO2">mi3nts/Estimate-CO2: Data and code to estimate inhaled CO2 (github.com)</a></p>

opencc-by-4.0Nov 2023View details →
zenodo36/100

Automated bio-AFM generation of large mechanome data set and their analysis by machine learning to classify prostatic cell lines_Training base 100 PC3-GFP

Open the record for dataset details and reuse information.

opencc-by-sa-4.0Nov 2023View details →
zenodo36/100

Combined bioinformatics and machine learning methodologies reveal prognosis-related ceRNA network and propose ABCA8, CAT, and CXCL12 as independent protective factors against osteosarcoma

<p><strong>Supplementary Table 1</strong>. Basic traits of the seven microarray datasets from the Gene Expression Omnibus and The Cancer Genome Atlas.</p><p><strong>Supplementary Table 2</strong>. Basic characteristics of the nine differentially expressed circRNAs</p><p><strong>Supplementary Table 3</strong>. Index of concordance (C-index) and variance inflation factor (VIF) of ABCA8, CXCL12, and CAT.</p><p><strong>Supplementary Table 4.</strong> LASSO and cox analysis of ceRNA with coef/se(coef) &lt; 0.01 and P-value of proportional hazards assumption (PH) &gt; 0.05.</p><p><strong>Supplementary Table 5</strong>. Robust rank aggregation analysis of ABCA8, CXCL12, and CAT. LogFC in the four datasets and RRA score of the three RNAs.</p><p><strong>Supplementary Figure 1.</strong> Competitive endogenous RNA in osteosarcoma.</p><p><strong>Supplementary Figure 2.</strong> Protein–protein interaction (PPI) network of genes in the competitive endogenous RNA network</p><p><strong>Supplementary Figure 3.</strong> Proportional hazards assumption (left) and linearity assumption (right).</p><p><strong>Supplementary Figure 4.</strong> Survival analysis for competitive endogenous RNA in osteosarcoma.</p>

opencc-by-4.0Nov 2023View details →
zenodo36/100

Dataset for "Mo-Si alloys studied by atomistic computer simulations using a novel machine-learning interatomic potential: Thermodynamics and interface phenomena"

<p>This dataset was used to fit a general purpose machine-learning interatomic potential for Mo-Si alloys based on the Atomic Cluster Expansion (ACE) formalism. It supports the paper "Mo-Si alloys studied by atomistic computer simulations using a novel machine-learning interatomic potential: Thermodynamics and interface phenomena".</p>

opencc-by-4.0Nov 2023View details →
zenodo36/100

INFLAMeR: a machine learning algorithm based on large-scale perturbation screening identified new lncRNAs regulating differentiation and survival of leukaemia cells

Open the record for dataset details and reuse information.

opencc-by-4.0Jul 2023View details →
dryad36/100

Machine learning driven self-discovery of the robot body morphology

<p>Conventionally, the kinematic structure of a robot is assumed to be known and data from external measuring devices are used mainly for calibration. We take an agent-centric perspective to explore whether a robot could learn its body structure by relying on scarce knowledge and depending only on unorganized proprioceptive signals. To achieve this, we analyze a mutual-information-based representation of the relationships between the proprioceptive signals, which we call proprioceptive information graphs (pi-graph), and use it to look for connections that reflect the underlying mechanical topology of the robot. We then use the inferred topology to guide the search for the morphology of the robot; i.e. the location and orientation of its joints. Results from different robots show that the correct topology and morphology can be effectively inferred from their pi-graph, regardless of the number of links and body configuration.</p>

opencc-zeroDec 2023View details →
zenodo36/100

Subseasonal to Seasonal (S2S) Prediction Algorithms using Hybrid Machine Learning Techniques

<p>&lt; S2S dataset.zip &gt;</p><p>1.ECMWF observations/hindcast realizations</p><ul><li>hindcast-like-observations_2000-2019_biweekly_deterministic.zarr</li><li>forecast-like-observations_2020_biweekly_deterministic.zarr</li><li>ecmwf_hindcast-input_2000-2019_biweekly_deterministic.zarr</li><li>ecmwf_forecast-input_2020_biweekly_deterministic.zarr</li><li>hindcast-like-observations_2000-2019_biweekly_tercile-edges.nc</li></ul><p>2. External variables</p><ul><li>"nino" folder -&gt; nino12.long.anom.data, nino34.long.anom.data : El Niño data</li><li>"Oscillation" folder<ul><li>-&gt; ersst.v5.pdo.dat.text : PDO (Pacific Decadal Oscillation)</li><li>-&gt; norm.nao.monthly.b5001.current.ascii.table.txt : NAO (North Atlantic Oscillation)</li><li>-&gt; qbo.dat : QBO (Quasi Biennial Oscillation)</li></ul></li><li>"great_lake" folder -&gt; N_seaice_extent_daily_v3.0 : Great lakes ice cover</li><li>observed-solar-cycle-indices.json : Sunspot cycles (two variables: original value and smoothed value)</li></ul><p>3. Region.txt : Region and its bound</p><p>4. Biweekly historical statistics data</p><ul><li>biw_stat_w34 folder -&gt; data (mean, standard deviation, median, skewness, kurtosis)&nbsp;for Week 3-4</li><li>biw_stat_w56 folder -&gt; data (mean, standard deviation, median, skewness, kurtosis)&nbsp;for Week 5-6</li></ul><p>&nbsp;</p><p>&lt; ML_code.zip &gt;</p><ul><li>ML codes for training, testing, and calculating RPSS based on Python3</li><li>Check run_val.sh and run_2020.sh&nbsp;</li></ul><p>&nbsp;</p>

opencc-by-4.0Dec 2023View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record