Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
1,075
datasets available to search
ShareScore release 0.7.1
Dataset results
1,075 results for “ML”
Regional seismicity (ML≥1.0) from 2008 to 2022 for the Haiyuan fault system
<p>This data is the regional seismicity (ML≥1.0) from 2008 to 2022 from the article "Strain Accumulation and Release on the Haiyuan Fault System from Joint Analysis of InSAR, GPS and Seismological Observations".</p>
Forecasting 24-hour-averaged PM2.5concentration in the Aburrá Valley using tree-based ML models, global forecasts, and satellite information: Dataset
<p>Data necessary for the training and evaluating the 24-hourly-averaged PM2.5 forecast over 19 stations within the Aburrá Valley, Colombia, is included here.</p>
Water, acetonitrile, and methanol MD simulations driven by many-body ML potentials
<p>Input, output, and trajectories of molecular dynamics (MD) simulations of water, acetonitrile, and methanol. Simulations were driven by many-body machine learning (mbML) potentials including explicit 1-, 2-, and 3-body contributions. <a href="https://keithgroup.github.io/mbGDML/">GDML</a>, <a href="https://libatoms.github.io/GAP/">GAP</a>, and <a href="https://schnetpack.readthedocs.io/en/stable/">SchNet</a> models are provided in a <a href="https://doi.org/10.5281/zenodo.7112163">separate repository</a>. All simulations were performed in the <a href="https://wiki.fysik.dtu.dk/ase/">atomic simulation environment (ASE)</a>. Analyses including radial distribution function (rdf) curves are provided <a href="https://github.com/keithgroup/mbgdml-h2o-meoh-mecn">here</a>.</p> <p><strong>Manifest</strong></p> <p>The following simulations are included in this repository for each solvent.</p> <ul> <li>1 ps hexamer NVE MD simulation driven by MP2/def2-TZVP (in ORCA v4.2.0), mbGDML, mbGAP, mbSchNet, and GFN2-xTB started with the same positions and velocities. Velocities were initialized at 298.15 K with a Maxwell-Boltzmann distribution.</li> <li>Periodic NVT MD simulation at 298.15 K for 10 or 30 ps with a 1 fs time step driven by mbGDML. These simulations contained <ul> <li>58 or <strong>137</strong> water molecules,</li> <li><strong>67</strong> or 122 acetonitrile molecules,</li> <li><strong>61</strong> methanol molecules.</li> </ul> </li> </ul> <p>Systems that are not bolded were used for testing purposes.</p>
Monitoring ML Systems: Challenges, Solutions and Metrics from a Practitioners' Perspective
<p>This is the dataset for the paper "Monitoring ML Systems: Challenges, Solutions and Metrics from a Practitioners' Perspective". The dataset is recorded in an MS Excel file which contains the following Excel sheets, and the description of each sheet is briefly presented below.</p> <p>(1) <strong>Selected Projects (GitHub)</strong> contain the 15 selected ML projects with the URL of each project.</p> <p>(2) <strong>Raw Data (GitHub)</strong> contains the information about the randomly selected 600 issues, such as issue titles, issue links, issue id.</p> <p>(3) <strong>Raw Data (SO) </strong>contains the information about the randomly selected 2088 SO posts (part of the 2174 SO posts), such as post titles, post link, post id, and open date.</p> <p>(4) <strong>MLOps (SO) </strong>contains the information about the 86 MLOps SO posts (part of the 2174 SO posts), such as post titles, post link, post id, and open date.</p> <p>(5) <strong>Identified Challenges (GitHub) </strong>contain the list of categories, subcategories, and codes of identified challenges from GitHub issues.</p> <p>(6) <strong>Identified Solutions (GitHub) </strong>contain the list of categories, subcategories, and codes of identified solutions from GitHub issues.</p> <p>(7) <strong>Identified Metrics (GitHub) </strong>contain the list of identified metrics from GitHub issues.</p> <p>(8) <strong>Identified Challenges (SO) </strong>contain the list of categories, subcategories, and codes of identified challenges from SO posts.</p> <p>(9) <strong>Identified Solutions (SO) </strong>contain the list of categories, subcategories, and codes of identified solutions from SO posts.</p> <p>(10) <strong>Identified Metrics (SO) </strong>contain the lists of identified metrics from SO posts.</p> <p>(11) <strong>Identified Challenges (Interview)</strong> contain the list of categories, subcategories, and codes of identified challenges from interviews.</p> <p>(12) <strong>Identified Solutions (Interview) </strong>contain the list of categories, subcategories, and codes of identified solutions from interviews.</p> <p>(13) <strong>Identified Metrics (Interview) </strong>contain the list of identified metrics from interviews.</p> <p>(14) <strong>Identified Challenges (Final) </strong>contain the taxonomy of the final challenges identified from GitHub issues, SO posts, and interviews.</p> <p>(15) <strong>Identified Solutions (Final) </strong>contain the<strong> </strong>taxonomy of the final solutions identified from GitHub issues, SO posts, and interviews.</p> <p>(16) <strong>Identified Metrics (Final) </strong>contain the<strong> </strong>final list of metrics identified from GitHub issues, SO posts, and interviews.</p>
Dataset for "On Developing an ML-Based Approach for the Automatic Characterization of Behavioral Phenotypes for Dairy Cows Relevant to Thermotolerance"
<p>This dataset consists of 3,421 videos filmed at T & K Dairy in Snyder, TX, over a 24-hour period on March 12-13, 2023. These videos were then used to train, validate, and evaluate a computer vision algorithm that is capable of automatically identifying cows using their coat patterns.</p>
MLSBench: A Synthesizable Dataset of HLS Designs to Support ML Based Design Flows
<p>HLS dataset for training of ML models for post route prediction of parameters. "HLS Benchmarks.zip" are in C and C++; "SystemC Version.zip" are in SystemC.</p>
Biologics repositioning: ml-SOM analysis raw results
<p><strong>Summary</strong></p> <p>This dataset contains raw result files for multiple-layer SOM (ml-SOM) repositioning analysis of infliximab and brodalumab. </p> <p><strong>Infliximab_results.zip</strong> archive contains the results of ml-SOM repositioning analysis of infliximab (T1) as a potential therapeutics for ulcerative colitis (D1), Crohn’s disease (D2), COPD (R1) and sarcoidosis (R2).</p> <p>Data is organized in the archive as follows:</p> <ol> <li>inflixmab_UC_CD_COPD_SARC021018_results+.RData - R data file that contains ml-SOM environment</li> <li>folder "inflixmab_UC_CD_COPD_SARC021018_results+ - Results" - comprises of various PDF and CSV files representing the a ml-SOM analysis results. It also contains the separate layer-level SOM analysis results for all datasets: <ul> <li>folder "1" - SOM analysis of infliximab dataset</li> <li>folder "2" - SOM analysis for ulcerative colitis and Crohn's disease dataset</li> <li>folder "3" - SOM analysis for COPD dataset</li> <li>folder "4" - SOM analysis for sarcoidosis dataset.</li> </ul> </li> </ol> <p><strong>Brodalumab_results.zip</strong> archive contains the results of ml-SOM repositioning analysis of brodalumab (T1) as a potential therapeutics for psoriasis (D1), Crohn’s disease (R1), and systemic juvenile idiopathic arthritis (R2).</p> <p>Data is organized in the archive as follows:</p> <ol> <li>brodalumab_CD_SJIA_301218_results+.RData - R data file that contains ml-SOM environment</li> <li>folder "brodalumab_CD_SJIA_301218_results+ - Results" - comprises of various PDF and CSV files representing the a ml-SOM analysis results. It also contains the separate layer-level SOM analysis results for all datasets: <ul> <li>folder "1" - SOM analysis of brodalumab and psoriasis dataset</li> <li>folder "2" - SOM analysis for Crohn's disease dataset</li> <li>folder "3" - SOM analysis for systemic juvenile idiopathic arthritis dataset</li> </ul> </li> </ol> <p>For detailed instructions on browsing the results and their interpretation please refer to the oposSOM package manual [1], as well as original publications [2-4]. </p> <p><strong>References</strong></p> <ol> <li>Henry Loeffler-Wirth, Hoang Thanh Le and Martin KalcheroposSOM.Comprehensive analysis of transcriptome data. DOI: <a href="https://doi.org/doi:10.18129/B9.bioc.oposSOM">10.18129/B9.bioc.oposSOM</a> </li> <li>Löffler-Wirth H, Kalcher M, Binder H. oposSOM: R-package for high-dimensional portraying of genome-wide expression landscapes on bioconductor.Bioinformatics. 2015 Oct 1;31(19):3225-7. DOI: 10.1093/bioinformatics/btv342. Epub 2015 Jun 10.</li> <li>Wirth H, von Bergen M, Binder H. Mining SOM expression portraits: feature selection and integrating concepts of molecular function. BioData Min. 2012 Oct 8;5(1):18. DOI: 10.1186/1756-0381-5-18.</li> <li>Wirth H, Löffler M, von Bergen M, Binder H. Expression cartography of human tissues using self organizing maps. BMC Bioinformatics. 2011 Jul 27;12:306. DOI: 10.1186/1471-2105-12-306.</li> </ol> <p> </p>
OWL2VecOA Resources for Bio-ML 2023
<p>1. The repository <a href="../api/records/13309009/draft/files/omim2ordo_exp_results.zip/content" target="_blank" rel="noopener noreferrer">omim2ordo_exp_results.zip</a> contains results of applying our extended OWL2VecOA method to biomedical ontology alignments, specifically focusing on the alignment between OMIM and ORDO. </p> <p>The specifications are: walk depth = 3, embedding size =100, iteration =70, walker iteration k=20. </p> <p>The alignment process utilized a combined approach, integrating results from two well-established ontology matching systems: AML and LogMap. Specifically the following input configurations were used:</p> <ul> <li>Train.tsv (from BIO-ML Track 2023) combined with the intersection of AML and LogMap alignments</li> <li>Train.tsv combined with the union of AML and LogMap alignments</li> <li>Train.tsv combined with LogMap alignments (Logmapping)</li> <li>Train.tsv combined with LogMap alignments (Anchor Mappings)</li> <li>Train.tsv combined with LogMap alignments (OverEstimation Mappings)</li> <li>Train.tsv only</li> </ul> <p>2. The repository "<a href="../api/records/13309009/draft/files/owl2vecstart_initres_2&3.zip/content" target="_blank" rel="noopener noreferrer">owl2vecstart_initres_2&3.zip</a>" contains the results of applying the initial version of the OWL2VecStar method to biomedical ontology alignments 2023 : OMIM-ORDO (o2o), NCIT-DOID (ncit2doid), SNOMED-NCIT-N (s2nn), and SNOMED-NCIT-PHARMA (sn2p) with walk depths of 2 and 3. Similarly, the repository "<a href="../api/records/13309009/draft/files/owl2vecstar_initres_4&5.zip/content" target="_blank" rel="noopener noreferrer">owl2vecstar_initres_4&5.zip</a>" contains analogous results, but with walk depths of 4 and 5 </p> <p>3. The repository "<a href="../api/records/13309009/draft/files/owl2vecOA_results_2&3.zip/content" target="_blank" rel="noopener noreferrer">owl2vecOA_results_2&3.zip</a>" contains the results of applying our extended OWL2VecOA method to the BIO-ML datasets, utilizing walk depths of 2 and 3.</p> <p>The results package contains three key components of each input data: Embedding file, Cosine Similarity Scores file and Euclidian Distance Scores file. The embedding files can be used for various ML downstream tasks, while the similarity and distance scores provide direct measures of entity relatedness, potentially useful for ontology alignment, entity matching, or other biomedical informatics applications. </p>
Dataset for "Property-based Testing within ML Projects: an Empirical Study" (ICSME NIER 2024)
<p>The dataset for the ICSME NIER 2024 paper "Property-based Testing within ML Projects: an Empirical Study". Descriptions of each column can be found in the readme.md file.</p>
ML-Optimized QKD Frequency Assignment for Efficient Quantum-Classical Coexistence in Multi-Band EONs
<p>Abstract: Quantum key distribution (QKD) represents a cutting-edge technology that ensures unbreakable security. Coexisting quantum and classical signals on a multi-band (O+E+S+C+L-band) system offer a viable solution for secure, high-rate networks amidst growing classical traffic and address quantum signal sensitivity. In this study, we assume a dynamic classical traffic load and varying configurations of classical channels (CChs). Considering the varying behavior of Secure Key Rate (SKR) under different classical conditions, solving the integral noise equations are crucial for optimizing QKD implementation and enhancing resource efficiency. The complexity and time-consuming nature of this process challenge infrastructure providers in determining the optimal quantum channel (QCh) frequency in real time. To tackle these challenges, we propose a machine learning (ML) algorithm. By leveraging ML, QKD can be implemented efficiently, optimizing resource utilization while significantly reducing computation and processing time in dynamic classical traffic. We implement three ML algorithms at various fiber intervals, all of which estimate the optimal frequency for QCh with 99\% accuracy and perform computations on average in 0.09 seconds, which is significantly faster compared to integral computational methods that have a mean time of 637 seconds.<br><br>Information: In this file, the Excel sheet contains data for each fiber interval, including inputs such as fiber length in each interval, the overall classical loading factor percentage, the C-band loading factor percentage, the L-band loading factor percentage, the highest active classical frequency (which serves as input to the machine learning model), and the QCh frequency that resulted in the highest SKR.</p>
Float+SOCAT sampling masks for ML reconstruction of surface ocean pCO2 using the Large Ensemble Testbed
<p>Here we provide sampling masks used in the study "The importance of adding unbiased Argo observations to the ocean carbon observing system" (Heimdal & McKinley, 2024, Scientific Reports). In this paper, we reconstruct surface ocean pCO2 using the Large Ensemble Testbed (Gloege et al., 2021, https://doi.org/10.1029/2020GB006788) and the pCO2-Residual method (Bennington et al., 2022, https://doi.org/10.1029/2021MS002960). We provide 2 different sampling masks used in the experiments presented in Heimdal & McKinley (2024). These masks represent two different float sampling schemes (+SOCAT) including 500 floats, corresponding to historical Argo float observations (https://fleetmonitoring.euro-argo.eu/dashboardpatterns) and potential optimized float sampling (following Chamberlain et al., 2023, <a href="https://doi.org/10.1175/JTECH-D-22-0093.1" target="_blank" rel="noopener">https://doi.org/10.1175/JTECH-D-22-0093.1</a>). </p>
GS_EgoExo_Plaster Turning on Wheel_ML_EN 1 - audio_ambient
<p>GS_EgoExo_Plaster Turning on Wheel_ML_EN 1 - audio_ambient</p>
GS_EgoExo_Plaster Turning on Wheel_ML_EN 1 - audio_contact
<p>GS_EgoExo_Plaster Turning on Wheel_ML_EN 1 - audio_contact</p>
ACSAC_ML_SCA_EVA Artifacts
<p>The artifacts for ACSAC 2024 paper <em><span><span>R+R: Demystifying ML-Assisted Side-Channel Analysis Framework: A Case of Image Reconstruction</span></span></em></p>
ML scripts and data for Buzacott et al. "Drivers and annual totals of methane emissions from Dutch peatlands (2024)"
<div> <div>This repository includes the machine learning (ML) scripts, data, and output for the article "Drivers and annual totals of methane emissions from Dutch peatlands" submitted to Global Change Biology. The scripts utilise the ML FCH4 gapfilling framework described in Irvin et al. (2021) (https://doi.org/10.1016/j.agrformet.2021.108528) which is available at https://github.com/stanfordmlgroup/methane-gapfill-ml and needed to run the scripts.</div> </div>
ReqExp: BERT-based ML Model for Extracting Software Requirements
<p>Datasets that were used during experiments in ReqExp project.</p>
Results: Guided Metamorphic Testing for Software Engineering ML
<p>This package holds results from "Searching for Quality: Guided Metamorphic Testing for Software Engineering ML" as submitted to GECCO 2023. </p> <p>The results are directly gathered from running the (separately) provided replication package. </p> <p>Note: Due to a slight oversight, the postfixes "min" and "max" are switched, i.E. the F1-min as named in the file actually maximizes.<br> This is corrected later in the evaluation. </p>
Enriched CONLLU Ancora for ML training
<p>This is an enriched version for Machine Learning purposes of the CONLLU adaptation of AnCora corpus .</p> <p>This version of the corpus was developed by BSC TeMU as part of the AINA project, and has been used to do multi-task learning for the Catalan language Spacy 3.4 models.</p> <p><strong>Versió enriquida de l'adaptació del corpus AnCora al format CONLLU orientada a l'aprenentatge automàtic.</strong></p> <p><strong>Aquesta versió del corpus ha estat desenvolupada per BSC TeMU com a part del projecte Aina, i s'ha fet servir per a l'entrenament multitasca dels models Spacy 3.0 per al català.</strong></p> <p> </p>
Fibroblasts reaction to 1 and 3 µM ML-7
<p>Healthy (HF), scar (SF) and Dupuytren (DF) fibroblasts videos showing how cells reacted to 1 and 3 µM ML-7 addition.</p>
ML Research pdf2img
<p>Machine Learning related papers: pdf2img</p> <p>Used at https://github.com/stepp1/research-app/</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.