Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

1,075

datasets available to search

ShareScore release 0.7.1

Reset

Dataset results

1,075 results for “ML”

Learn how ShareScore rates datasets ↗
zenodo36/100

Regional seismicity (ML≥1.0) from 2008 to 2022 for the Haiyuan fault system

<p>This data is the regional seismicity (ML&ge;1.0) from 2008 to 2022 from the article &quot;Strain Accumulation and Release on the Haiyuan Fault System from Joint Analysis of InSAR, GPS and Seismological Observations&quot;.</p>

opencc-by-4.0Jul 2022View details →
zenodo36/100

Forecasting 24-hour-averaged PM2.5concentration in the Aburrá Valley using tree-based ML models, global forecasts, and satellite information: Dataset

<p>Data necessary for the training and evaluating the 24-hourly-averaged PM2.5 forecast over 19 stations within the Aburr&aacute; Valley, Colombia,&nbsp;is included here.</p>

opencc-by-4.0Sep 2022View details →
zenodo36/100

Water, acetonitrile, and methanol MD simulations driven by many-body ML potentials

<p>Input, output, and trajectories of molecular dynamics (MD) simulations of water, acetonitrile, and methanol. Simulations were driven by many-body machine learning (mbML) potentials including explicit 1-, 2-, and 3-body contributions. <a href="https://keithgroup.github.io/mbGDML/">GDML</a>, <a href="https://libatoms.github.io/GAP/">GAP</a>, and <a href="https://schnetpack.readthedocs.io/en/stable/">SchNet</a>&nbsp;models are provided in a <a href="https://doi.org/10.5281/zenodo.7112163">separate repository</a>. All simulations were performed in the <a href="https://wiki.fysik.dtu.dk/ase/">atomic simulation environment (ASE)</a>.&nbsp;Analyses including radial distribution function (rdf) curves are provided <a href="https://github.com/keithgroup/mbgdml-h2o-meoh-mecn">here</a>.</p> <p><strong>Manifest</strong></p> <p>The following simulations are included in this repository for each solvent.</p> <ul> <li>1 ps hexamer NVE MD simulation driven by MP2/def2-TZVP (in ORCA v4.2.0), mbGDML, mbGAP, mbSchNet, and GFN2-xTB started with the same positions and velocities. Velocities were initialized at 298.15 K with a Maxwell-Boltzmann distribution.</li> <li>Periodic NVT MD simulation at 298.15 K for 10 or 30 ps with a 1 fs time step driven by mbGDML. These simulations contained <ul> <li>58 or <strong>137</strong> water molecules,</li> <li><strong>67</strong> or 122 acetonitrile molecules,</li> <li><strong>61</strong> methanol molecules.</li> </ul> </li> </ul> <p>Systems that are not bolded were used for testing purposes.</p>

opencc-by-4.0Sep 2022View details →
zenodo36/100

Monitoring ML Systems: Challenges, Solutions and Metrics from a Practitioners' Perspective

<p>This is the dataset for the paper &quot;Monitoring ML Systems: Challenges, Solutions and Metrics from a&nbsp;Practitioners&#39; Perspective&quot;. The dataset is recorded in an MS Excel file which contains the following Excel sheets, and the description of each sheet is briefly presented below.</p> <p>(1)&nbsp;<strong>Selected Projects (GitHub)</strong> contain the 15 selected ML projects with the URL of each project.</p> <p>(2)&nbsp;<strong>Raw Data (GitHub)</strong> contains the information about the randomly selected 600 issues, such as issue titles, issue links, issue id.</p> <p>(3) <strong>Raw Data (SO) </strong>contains the information about the randomly selected 2088 SO posts (part of the 2174 SO posts), such as post titles, post link, post id, and open date.</p> <p>(4)&nbsp;<strong>MLOps (SO) </strong>contains the information about the 86 MLOps SO posts (part of the 2174 SO posts), such as post titles, post link, post id, and open date.</p> <p>(5)&nbsp;<strong>Identified Challenges (GitHub) </strong>contain the list of categories, subcategories, and codes of identified challenges from GitHub issues.</p> <p>(6)&nbsp;<strong>Identified Solutions (GitHub) </strong>contain the list of categories, subcategories, and codes of identified solutions from GitHub issues.</p> <p>(7)&nbsp;<strong>Identified Metrics (GitHub) </strong>contain the list of identified metrics from GitHub issues.</p> <p>(8)&nbsp;<strong>Identified Challenges (SO) </strong>contain the list of categories, subcategories, and codes of identified challenges from SO posts.</p> <p>(9)&nbsp;<strong>Identified Solutions (SO) </strong>contain the list of categories, subcategories, and codes of identified solutions from SO posts.</p> <p>(10)&nbsp;<strong>Identified Metrics (SO) </strong>contain the lists of identified metrics from SO posts.</p> <p>(11)&nbsp;<strong>Identified Challenges (Interview)</strong> contain the list of categories, subcategories, and codes of identified challenges from interviews.</p> <p>(12)&nbsp;<strong>Identified Solutions (Interview) </strong>contain the list of categories, subcategories, and codes of identified solutions from interviews.</p> <p>(13)&nbsp;<strong>Identified Metrics (Interview) </strong>contain the list of identified metrics from interviews.</p> <p>(14)&nbsp;<strong>Identified Challenges (Final) </strong>contain the taxonomy of the final challenges identified from GitHub issues, SO posts, and interviews.</p> <p>(15)&nbsp;<strong>Identified Solutions (Final) </strong>contain the<strong> </strong>taxonomy of the final solutions identified from GitHub issues, SO posts, and interviews.</p> <p>(16)&nbsp;<strong>Identified Metrics (Final) </strong>contain the<strong> </strong>final list of metrics identified from GitHub issues, SO posts, and interviews.</p>

opencc-by-4.0Feb 2023View details →
zenodo36/100

Dataset for "On Developing an ML-Based Approach for the Automatic Characterization of Behavioral Phenotypes for Dairy Cows Relevant to Thermotolerance"

<p>This dataset consists of 3,421 videos filmed at T &amp; K Dairy in Snyder, TX, over a 24-hour period on March 12-13, 2023.&nbsp; These videos were then used to train, validate, and evaluate a computer vision algorithm that is capable of automatically identifying cows using their coat patterns.</p>

opencc-by-4.0Jun 2024View details →
zenodo36/100

MLSBench: A Synthesizable Dataset of HLS Designs to Support ML Based Design Flows

<p>HLS dataset for training of ML models for post route prediction of parameters. &quot;HLS Benchmarks.zip&quot; are in C and C++; &quot;SystemC Version.zip&quot; are in SystemC.</p>

opencc-by-4.0Sep 2019View details →
zenodo36/100

Biologics repositioning: ml-SOM analysis raw results

<p><strong>Summary</strong></p> <p>This dataset contains raw result files for multiple-layer SOM (ml-SOM) repositioning analysis of infliximab and brodalumab.&nbsp;</p> <p><strong>Infliximab_results.zip</strong>&nbsp;archive contains&nbsp;the results of ml-SOM repositioning analysis of infliximab (T1) as a potential therapeutics for ulcerative colitis (D1), Crohn&rsquo;s disease (D2), COPD (R1) and sarcoidosis (R2).</p> <p>Data is organized in the archive as follows:</p> <ol> <li>inflixmab_UC_CD_COPD_SARC021018_results+.RData - R data file that contains ml-SOM environment</li> <li>folder &quot;inflixmab_UC_CD_COPD_SARC021018_results+ - Results&quot; - comprises of various PDF and CSV files representing the a&nbsp;ml-SOM analysis results. It also contains the separate&nbsp;layer-level SOM analysis results for all datasets: <ul> <li>folder &quot;1&quot; - SOM analysis of&nbsp;infliximab dataset</li> <li>folder &quot;2&quot; - SOM analysis for&nbsp;ulcerative colitis and Crohn&#39;s disease dataset</li> <li>folder &quot;3&quot; - SOM analysis for COPD&nbsp;dataset</li> <li>folder &quot;4&quot; - SOM analysis for&nbsp;sarcoidosis dataset.</li> </ul> </li> </ol> <p><strong>Brodalumab_results.zip</strong> archive contains&nbsp;the results of ml-SOM repositioning analysis of brodalumab (T1) as a potential therapeutics for psoriasis (D1), Crohn&rsquo;s disease (R1),&nbsp;and systemic juvenile idiopathic arthritis (R2).</p> <p>Data is organized in the archive as follows:</p> <ol> <li>brodalumab_CD_SJIA_301218_results+.RData - R data file that contains ml-SOM environment</li> <li>folder &quot;brodalumab_CD_SJIA_301218_results+ - Results&quot; - comprises of various PDF and CSV files representing the a&nbsp;ml-SOM analysis results. It also contains the separate&nbsp;layer-level SOM analysis results for all datasets: <ul> <li>folder &quot;1&quot; - SOM analysis of&nbsp;brodalumab and psoriasis dataset</li> <li>folder &quot;2&quot; - SOM analysis for Crohn&#39;s disease dataset</li> <li>folder &quot;3&quot; - SOM analysis for&nbsp;systemic juvenile idiopathic arthritis&nbsp;dataset</li> </ul> </li> </ol> <p>For detailed instructions on browsing the results and their interpretation please refer to the oposSOM package manual [1], as well as original publications [2-4].&nbsp;</p> <p><strong>References</strong></p> <ol> <li>Henry Loeffler-Wirth, Hoang Thanh Le and Martin KalcheroposSOM.Comprehensive analysis of transcriptome data.&nbsp;DOI: <a href="https://doi.org/doi:10.18129/B9.bioc.oposSOM">10.18129/B9.bioc.oposSOM</a>&nbsp;</li> <li>L&ouml;ffler-Wirth H, Kalcher M, Binder H.&nbsp;oposSOM: R-package for high-dimensional portraying of genome-wide expression landscapes on bioconductor.Bioinformatics. 2015 Oct 1;31(19):3225-7. DOI: 10.1093/bioinformatics/btv342. Epub 2015 Jun 10.</li> <li>Wirth H, von Bergen M, Binder H.&nbsp;Mining SOM expression portraits: feature selection and integrating concepts of molecular function.&nbsp;BioData Min. 2012 Oct 8;5(1):18. DOI: 10.1186/1756-0381-5-18.</li> <li>Wirth H, L&ouml;ffler M, von Bergen M, Binder H.&nbsp;Expression cartography of human tissues using self organizing maps.&nbsp;BMC Bioinformatics. 2011 Jul 27;12:306. DOI: 10.1186/1471-2105-12-306.</li> </ol> <p>&nbsp;</p>

opencc-by-4.0Sep 2019View details →
zenodo36/100

OWL2VecOA Resources for Bio-ML 2023

<p>1. The repository&nbsp;<a href="../api/records/13309009/draft/files/omim2ordo_exp_results.zip/content" target="_blank" rel="noopener noreferrer">omim2ordo_exp_results.zip</a> contains results of applying our extended&nbsp; OWL2VecOA method to biomedical ontology alignments, specifically focusing on the alignment between OMIM and ORDO.&nbsp;</p> <p>The specifications are: walk depth = 3, embedding size =100, iteration =70, walker iteration k=20.&nbsp;</p> <p>The alignment process utilized a combined approach, integrating results from two well-established ontology matching systems: AML and LogMap. Specifically the following input configurations were used:</p> <ul> <li>Train.tsv (from BIO-ML Track 2023) combined with the intersection of AML and LogMap alignments</li> <li>Train.tsv combined with the union of AML and LogMap alignments</li> <li>Train.tsv combined with LogMap alignments (Logmapping)</li> <li>Train.tsv combined with LogMap alignments (Anchor Mappings)</li> <li>Train.tsv combined with LogMap alignments (OverEstimation Mappings)</li> <li>Train.tsv only</li> </ul> <p>2. The repository "<a href="../api/records/13309009/draft/files/owl2vecstart_initres_2&amp;3.zip/content" target="_blank" rel="noopener noreferrer">owl2vecstart_initres_2&amp;3.zip</a>" contains the results of applying the initial version of the OWL2VecStar method to biomedical ontology alignments 2023 : OMIM-ORDO (o2o), NCIT-DOID (ncit2doid), SNOMED-NCIT-N (s2nn), and SNOMED-NCIT-PHARMA (sn2p) with walk depths of 2 and 3. Similarly, the repository "<a href="../api/records/13309009/draft/files/owl2vecstar_initres_4&amp;5.zip/content" target="_blank" rel="noopener noreferrer">owl2vecstar_initres_4&amp;5.zip</a>" contains analogous results, but with walk depths of 4 and 5&nbsp;&nbsp;</p> <p>3. The repository "<a href="../api/records/13309009/draft/files/owl2vecOA_results_2&amp;3.zip/content" target="_blank" rel="noopener noreferrer">owl2vecOA_results_2&amp;3.zip</a>" contains the results of applying our extended OWL2VecOA method to the BIO-ML&nbsp; datasets, utilizing walk depths of 2 and 3.</p> <p>The results package contains three key components of each input data: Embedding file, Cosine Similarity Scores file and Euclidian Distance Scores file. The embedding files can be used for various ML downstream tasks, while the similarity and distance scores provide direct measures of entity relatedness, potentially useful for ontology alignment, entity matching, or other biomedical informatics applications.&nbsp;</p>

opencc-by-4.0Aug 2024View details →
zenodo36/100

Dataset for "Property-based Testing within ML Projects: an Empirical Study" (ICSME NIER 2024)

<p>The dataset for the ICSME NIER 2024 paper "Property-based Testing within ML Projects: an Empirical Study". Descriptions of each column can be found in the readme.md file.</p>

opencc-by-4.0Aug 2024View details →
zenodo36/100

ML-Optimized QKD Frequency Assignment for Efficient Quantum-Classical Coexistence in Multi-Band EONs

<p>Abstract: Quantum key distribution (QKD) represents a cutting-edge technology that ensures unbreakable security. Coexisting quantum and classical signals on a multi-band (O+E+S+C+L-band) system offer a viable solution for secure, high-rate networks amidst growing classical traffic and address quantum signal sensitivity. In this study, we assume a dynamic classical traffic load and varying configurations of classical channels (CChs). Considering the varying behavior of Secure Key Rate (SKR) under different classical conditions, solving the integral noise equations are crucial for optimizing QKD implementation and enhancing resource efficiency. The complexity and time-consuming nature of this process challenge infrastructure providers in determining the optimal quantum channel (QCh) frequency in real time. To tackle these challenges, we propose a machine learning (ML) algorithm. By leveraging ML, QKD can be implemented efficiently, optimizing resource utilization while significantly reducing computation and processing time in dynamic classical traffic. We implement three ML algorithms at various fiber intervals, all of which estimate the optimal frequency for QCh with 99\% accuracy and perform computations on average in 0.09 seconds, which is significantly faster compared to integral computational methods that have a mean time of 637 seconds.<br><br>Information: In this file, the Excel sheet contains data for each fiber interval, including inputs such as fiber length in each interval, the overall classical loading factor percentage, the C-band loading factor percentage, the L-band loading factor percentage, the highest active classical frequency (which serves as input to the machine learning model), and the QCh frequency that resulted in the highest SKR.</p>

opencc-by-4.0Aug 2024View details →
zenodo36/100

Float+SOCAT sampling masks for ML reconstruction of surface ocean pCO2 using the Large Ensemble Testbed

<p>Here we provide sampling masks used in the study "The importance of adding unbiased Argo observations to the ocean carbon observing system" (Heimdal &amp; McKinley, 2024, Scientific Reports). In this paper, we reconstruct surface ocean pCO2 using the Large Ensemble Testbed (Gloege et al., 2021, https://doi.org/10.1029/2020GB006788) and the pCO2-Residual method (Bennington et al., 2022, https://doi.org/10.1029/2021MS002960). We provide 2 different sampling masks used in the experiments presented in Heimdal &amp; McKinley (2024). These masks represent two different float sampling schemes (+SOCAT) including 500 floats, corresponding to historical Argo float observations (https://fleetmonitoring.euro-argo.eu/dashboardpatterns) and potential optimized float sampling (following Chamberlain et al., 2023, <a href="https://doi.org/10.1175/JTECH-D-22-0093.1" target="_blank" rel="noopener">https://doi.org/10.1175/JTECH-D-22-0093.1</a>).&nbsp;</p>

opencc-by-4.0Aug 2024View details →
zenodo36/100

GS_EgoExo_Plaster Turning on Wheel_ML_EN 1 - audio_ambient

<p>GS_EgoExo_Plaster Turning on Wheel_ML_EN 1 - audio_ambient</p>

opencc-by-4.0Oct 2024View details →
zenodo36/100

GS_EgoExo_Plaster Turning on Wheel_ML_EN 1 - audio_contact

<p>GS_EgoExo_Plaster Turning on Wheel_ML_EN 1 - audio_contact</p>

opencc-by-4.0Oct 2024View details →
zenodo36/100

ACSAC_ML_SCA_EVA Artifacts

<p>The artifacts for ACSAC 2024 paper <em><span><span>R+R: Demystifying ML-Assisted Side-Channel Analysis Framework: A Case of Image Reconstruction</span></span></em></p>

opencc-by-4.0Oct 2024View details →
zenodo36/100

ML scripts and data for Buzacott et al. "Drivers and annual totals of methane emissions from Dutch peatlands (2024)"

<div> <div>This repository includes the machine learning (ML) scripts, data, and output for the article "Drivers and annual totals of methane emissions from Dutch peatlands" submitted to Global Change Biology. The scripts utilise the ML FCH4 gapfilling framework described in Irvin et al. (2021) (https://doi.org/10.1016/j.agrformet.2021.108528) which is available at https://github.com/stanfordmlgroup/methane-gapfill-ml and needed to run the scripts.</div> </div>

opencc-by-4.0Nov 2024View details →
zenodo36/100

ReqExp: BERT-based ML Model for Extracting Software Requirements

<p>Datasets that were used during experiments in ReqExp project.</p>

opencc-by-4.0Jun 2021View details →
zenodo36/100

Results: Guided Metamorphic Testing for Software Engineering ML

<p>This package holds results from &quot;Searching for Quality: Guided Metamorphic Testing for Software Engineering ML&quot; as submitted to GECCO 2023.&nbsp;</p> <p>The results are directly gathered from running the (separately) provided replication package.&nbsp;</p> <p>Note: Due to a slight oversight, the postfixes &quot;min&quot; and &quot;max&quot; are switched, i.E. the F1-min as named in the file actually maximizes.<br> This is corrected later in the evaluation.&nbsp;</p>

opencc-by-4.0Nov 2022View details →
zenodo36/100

Enriched CONLLU Ancora for ML training

<p>This is an enriched version for Machine Learning purposes of the CONLLU adaptation of AnCora corpus .</p> <p>This version of the corpus was developed by BSC TeMU as part of the AINA project, and has been used to do multi-task learning for the Catalan language Spacy 3.4 models.</p> <p><strong>Versi&oacute; enriquida de l&#39;adaptaci&oacute; del corpus AnCora al format CONLLU orientada a l&#39;aprenentatge autom&agrave;tic.</strong></p> <p><strong>Aquesta versi&oacute; del corpus ha estat desenvolupada per BSC TeMU com a part del projecte Aina, i s&#39;ha fet servir per a l&#39;entrenament multitasca dels models Spacy 3.0 per al catal&agrave;.</strong></p> <p>&nbsp;</p>

opencc-by-4.0Jun 2021View details →
zenodo36/100

Fibroblasts reaction to 1 and 3 µM ML-7

<p>Healthy (HF), scar (SF) and Dupuytren (DF) fibroblasts videos showing how cells reacted to 1 and 3 &micro;M ML-7 addition.</p>

opencc-by-4.0Nov 2022View details →
zenodo36/100

ML Research pdf2img

<p>Machine Learning related papers: pdf2img</p> <p>Used at https://github.com/stepp1/research-app/</p>

opencc-by-4.0Feb 2023View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record