Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
1,943
datasets available to search
ShareScore release 0.9.0
Dataset results
1,943 results for “machine learning”
Data for: ToadFishFinder classifier model v4: A catalog of oyster toadfish (Opsanus tau) calls for machine learning
<p>This data repository contains labeled passive underwater acoustic data used to train and test the machine-learning model of Bohnenstiehl (in prep – 2023), <span>Automated cataloging of oyster toadfish (<em>Opsanus</em> <em>tau</em>) calls using template matching and machine learning</span>. The software accompanying this paper is known as ToadFishFinder, and the classifier model presented in the paper is v4. It consists of more than 10000 labeled toadfish and 10000 labeled other signals. Labeled spectrogram images are provided, along with pressure-corrected waveforms (micro-Pascals) sampled at 24 kHz. Each waveform sample is 1350 ms long. The center 850 ms of these waveform segments represent the portion of the signal used in training and testing the classifier model. Waveform data are provided in multiple formats: 1) MATLAB (.mat) files containing the 'boatwhistle' and 'other' waveforms stored in column format, and 2) individual .wav files, each containing a labeled waveform example. Codes are provided to demonstrate how these .wav files can be read into MATLAB and PYTHON. These labeled data can be used to re-train the ToadFishFinder model or develop alternative classifiers. </p>
Machine Learning-based Energy Optimisation in Smart City Internet of Things
<p>Dataset for the paper Machine Learning-based Energy Optimisation in Smart City Internet of Things accepted for publication at The First International Workshop on the Integration between Distributed Machine Learning and the Internet of Things, ACM MobiHoc 2023.</p> <p>The dataset is collected from a real-world deployment of environmental sensors in the city of Bern, Switzerland. Our proposed approach can be applied to determine the tradeoff between the accuracy of temperature measurements and reducing the energy consumption for a single sensor; hence, without loss of generality, the evaluation is conducted on a dataset from a single sensor. Overall, we acquired 3697 measurements, each long 138 seconds. To correct the measurements, we set the maximum ventilation duration of 138 seconds, during which the multivariate time series of humidity and temperature sensor values are recorded together with their corresponding timestamps. The sensor values are recorded at a fixed frequency.</p> <p>From this raw data, we created the training and test sets through data augmentation to simulate time series of different lengths. Namely, for each measurement, we generated 136 samples with the increasing length of measurement time-series, padding the residual time-series length with zeros until reaching a time-series length of 137.</p> <p>We released the source code and trained models on the following GitHub repository https://www.github.com/ricsamikwa/ml-iot-smartcitytemp</p>
Handful of Pixels - machine learning data
<p>Training and testing data within the context of the "Handful of Pixels" course, and in particular the Land-Use and Land-Cover mapping chapter. The data included samples 300 random locations across 10 land cover classes from the data described in Fritz et al. 2017, and downloaded through the AppEEARS API using the {appeears} R package (Hufkens et al. 2023). A 80% split is executed on this larger dataset with 240 locations retained for training, while the remainder is used for testing purposes. For the testing data the input data is shared, the labels are withheld (stored in a closed release of this archive, and accessible on reasonable request). This data can be used within the context of small demonstration machine learning exercises or competitions.</p> <p><strong>Data structure</strong></p> <p>The data contains all seven (7) bands of the MODIS MCD43A4 data product for the year 2012. Band names are indicated in full. In addition MODIS MOD11A2 daytime land surface temperature (LST) data is provided, where band names only contain the date (YYYY-MM-DD) of acquisition. Additional indices can be calculated from these band combinations if so desired.</p> <p>Data is provided in compressed serialized R rds files, and can be read into R as follows:</p> <pre><code>df <- readRDS("training_data.rds")</code></pre> <p> </p> <p><strong>References</strong></p> <p>Fritz, Steffen, Linda See, Christoph Perger, Ian McCallum, Christian Schill, Dmitry Schepaschenko, Martina Duerauer, et al. “A Global Dataset of Crowdsourced Land Cover and Land Use Reference Data.” <em>Scientific Data</em> 4, no. 1 (June 13, 2017): 170075. <a href="https://doi.org/10.1038/sdata.2017.75">https://doi.org/10.1038/sdata.2017.75</a>.</p> <p>Koen Hufkens. (2023). bluegreen-labs/appeears: appeears: an interface to the NASA AppEEARS API (v1.0). Zenodo. https://doi.org/10.5281/zenodo.7958270</p>
Automatic Identification of Kidney Cell Types in scRNA-seq and snRNA-seq Data Using Machine Learning Algorithms - Datasets
<p>Datasets for reproducibility of the results found in Automatic Identification of Kidney Cell Types in scRNA-seq and snRNA-seq Data Using Machine Learning Algorithms. This study utilized data from the following 4 journals:</p> <p>Lake, B.B. et al. A single-nucleus RNA-sequencing pipeline to decipher the molecular anatomy and pathophysiology of human kidneys. Nat Commun 10, 2832 (2019).</p> <p>Liao, J., Yu, Z., Chen, Y. et al. Single-cell RNA sequencing of human kidney. Sci Data 7, 4 (2020).</p> <p>Menon, R. et al. Single cell transcriptomics identifies focal segmental glomerulosclerosis remission endothelial biomarker. JCI Insight 5, e133267 (2020).</p> <p>Wu, H. et al. Single-cell transcriptomics of a human kidney allograft biopsy specimen defines a diverse inflammatory response. J Am Soc Nephrol 29: 2069–2080 (2018).</p> <p>Young, M. D. et al. Single-cell transcriptomes from human kidneys reveal the cellular identity of renal tumors. Science 361, 594–599 (2018).</p>
Data for DRExM³L: Drug REpurposing using eXplainable Machine Learning and Mechanistic Models of signal transduction
<p>(DREM³L) Drug REpurposing using Mechanistic Models of signal transduction and Machine Learning </p>
WaivOps RTRO-DRM: Open Audio Resources for Machine Learning in Music
<p><strong>WaivOps RTRO-DRM Dataset</strong></p> <p>RTRO-DRM is an open audio dataset composed of a series of drum recordings in the style of 1980s electronic music. The dataset comprises 2138 raw, unedited audio clips recorded in uncompressed stereo WAV format. These recordings were curated using an internal drum sample dataset and MIDI files sourced from a code-based music generation system, along with a MIDI transformer model trained on more than 30,000 MIDI files. The files primarily consist of recordings that may not meet conventional audio quality standards but can still be valuable for a range of applications and research projects.</p> <p><strong>Dataset</strong></p> <p>The primary objective of this dataset is to provide accessible content for machine learning applications in music and audio research. Some potential use cases for this dataset include tempo detection and classification, drum rhythm analysis, audio-to-MIDI conversion, source separation, automated mixing, music information retrieval, AI music generation, sound design, and signal processing.</p> <p>Specifications</p> <ul> <li>2138 audio loops (4.3 hours)</li> <li>24-bit WAV format</li> <li>BPM labeled</li> <li>Tempo range: 100-145bpm</li> <li>Variational drum patterns</li> <li>Electronic drum machine sound (circa 1980s)</li> </ul> <p><strong>License</strong></p> <p>This dataset was compiled by WaivOps, a crowdsourced music project managed by the sound label company Patchbanks. All recordings have been compiled by verified sources for copyright clearance.</p> <p>The RTRO-DRM dataset is licensed under Creative Commons Attribution 4.0 International <a href="https://creativecommons.org/licenses/by/4.0/">(CC BY 4.0)</a>.</p> <p><strong>Additional Info</strong></p> <p>For audio examples or more information about this dataset, please refer to the <a href="https://github.com/patchbanks/WaivOps-RTRO-DRM">GitHub repository</a>.</p>
Replication package for the paper: "Machine Learning for the Identification and Classification of Technical Debt Types on StackOverflow Discussions"
<p>This is the replication package for the article "Machine Learning for the Identification and Classification of Technical Debt Types on StackOverflow Discussions". The article was published in the Research Track of the third Brazilian Workshop on Intelligent Software Engineering (ISE'23).</p> <p>The replication package consists of 8 files:<br> 1) dataset.csv, 2) code_anayses.ipynb and 3) example_test_balanced.csv and the others are results of word cloud generation.</p> <p>In dataset.csv, we provide the data for future replications.</p> <p>In code_anayses.ipynb, we provide the code we use to arrive at the results.</p> <p>In example_test_balanced.csv, we provide an example input dataset for training the models.</p> <p>For future references in this article, please contact lead author Eliakim Gama, or one of the co-authors.</p>
Soil texture dataset from the publication: "Machine learning applied for Antarctic soil mapping: Spatial prediction of soil texture for Maritime Antarctica and Northern Antarctic Peninsula'
<p>Clay, silt and sand distribution in Antarctic soils modeled and predicted through Machine Learning approaches, legacy soil data and environmental covariates. The coefficient of variation and quantile data represent the spatial uncertainty of the predictions. For more information about the methodology used, users are referred to the article: </p> <p>Siqueira, R.G., Moquedace, C.M., Francelino, M.R., Schaefer, C.E.G.R., Fernandes-Filho, E.I., 2023. Machine learning applied for Antarctic soil mapping: Spatial prediction of soil texture for Maritime Antarctica and Northern Antarctic Peninsula. Geoderma 432, 116405. https://doi.org/10.1016/j.geoderma.2023.116405</p> <p>The .zip file has the following folders:</p> <p>1) soil_texture_antarctica: soil texture information containing clay, silt and sand contents</p> <p>2) soil_texture_coefficient_variation: uncertainty from the coefficient of variation of the soil texture prediction</p> <p>3) soil_texture_prediction_interval: uncertainty from the prediction interval 90% (Q95% - Q5%) of the soil texture prediction</p> <p>4) soil_texture_quantile05: quantile 5% of the soil texture prediction</p> <p>5) soil_texture_quantile95: quantile 95% of the soil texture prediction</p>
Rating results obtained during a review of original articles on radiomics and machine learning for outcome prediction based on PET
<p>This upload provides Open Data associated with the publication "Methodological evaluation of original articles on radiomics and machine learning for outcome prediction based on positron emission tomography (PET)" by Rogasch JMM <em>et al.</em> (2023).</p> <p>The upload contains the item-by-item results of rating for all criteria and all 100 original articles. PubMed IDs are also included.</p> <p>Furthermore, a description of all variable names and how the rating categories were encoded in the data tables can be found in the PDF file "ML_prediction_Dictionary_2023_08_27.pdf".</p>
Soil chemistry dataset from the work "Modelling and prediction of major soil chemical properties with Random Forest: machine learning as tool to understand soil-environment relationships in Antarctica"
<p>Bases sum, H+Al (potential acidity), pH, phosphorous, remaining P (P-rem), sodium and total organic carbon distribution in Antarctic soils modeled and predicted through Machine Learning approaches, legacy soil data and environmental covariates. The quantile and prediction interval data represent the spatial uncertainty of the predictions.</p> <p>As soon as the work "Modelling and prediction of major soil chemical properties with Random Forest: machine learning as tool to understand soil-environment relationships in Antarctica" is published, the paper will be cited here. </p> <p>The .zip file contains the following folders:</p> <p>1) soil_chemistry_antarctica: data containing the soil chemical attributes distribution</p> <p>2) soil_chemistry_prediction_interval: uncertainty from the prediction interval 90% (Q95% - Q5%) of the soil attributes prediction</p> <p>4) soil_texture_quantile05: quantile 5% of the soil attributes prediction</p> <p>5) soil_texture_quantile95: quantile 95% of the soil attributes prediction</p>
Data and codes: Who is calling? Optimising source identification from marmoset vocalisations with hierarchical machine learning classifiers
<p>Data and codes that accompany the article titled "Who is calling? Optimising source identification from marmoset vocalisations with hierarchical machine learning classifiers".</p>
Data for Progress on Climate Action: a Multilingual Machine Learning Analysis of the Global Stocktake
<p>Data to go with our submission to Climatic Change titled "Progress on Climate Action: a Multilingual Machine Learning Analysis of the Global Stocktake".</p> <p>Dataset contains the embeddings (.zip with pickles) as well as the associated document items (idem), the most-closely associated keywords and paragraphs per topic in the final model (.xlsx), the reduced 2d embeddings with all selected paragraphs (.csv utf-8 encoded), as well as an overview with the meta-data per source (.csv utf-8 encoded).</p>
State-of-the-Art Review on the Aspects of Martensitic Alloys Studied via Machine Learning
<p>Description</p> <p>The dataset for the review paper titled "State-of-the-Art Review on the Aspects of Martensitic Alloys Studied via Machine Learning" consists of the four files with the names (i) alloy_names.csv, (ii) machine_learning_methods.csv, (iii) nomenclature.csv, and (iv) ptmc_terminologies.csv.</p> <p><strong>(i) alloy_names.csv </strong>: This file presents the summarized list of alloys' names which have been discussed in the review paper. The list thus provides the names of the alloys for which data-driven studies have been attempted to explore one of the effects - martensitic transformation, phase transformation or shape memory effect. </p> <p><strong>(ii) machine_learning_methods.csv</strong> : The machine learning methods that have been discussed in the review paper in relation to the simulation, modeling or prediction tasks in martensitic alloys are listed in this file. This csv file conssits of three columns. The first column "Methods" lists the names of the machine learning methods whereas the second column "Purpose" briefly reveals the objective of the use of the named machine learning method. The final column "Reference" provides the information about the original work (source) from which the data is obtained. </p> <p><strong>(iii) nomenclature.csv </strong>: This file lists all of the acronyms utilized in the review paper, and provides their corresponding full forms. </p> <p><strong> (iv) ptmc_terminologies.csv</strong> : One of the major theories considered significant in the study of martensitic alloys and shape memory effects is phenomenological theory of martensite crystallography (PTMC). The review paper discusses this theory. The different concepts that might be helpful in understanding PTMC , have been assembled in the form of terminologies. </p>
WaivOps WRLD-LP: Open Audio Resources for Machine Learning in Music
<p><strong>WRLD-LP Dataset</strong></p> <p>WRLD-LP is an open audio dataset comprised of a series of symbolic drum recordings in the genres of world percussion music. The dataset includes 3,162 audio loops recorded in uncompressed stereo WAV format. The compositions were generated with an internal sample dataset played with note-dense MIDI drum files from a code-based music generation system, along with a MIDI transformer model trained on more than 30,000 MIDI files.</p> <p><strong>Dataset</strong></p> <p>The primary objective of this dataset is to provide accessible content for machine learning applications in music and audio research. Some potential use cases for this dataset include tempo detection and classification, drum rhythm analysis, audio-to-MIDI conversion, source separation, automated mixing, music information retrieval, AI music generation, sound design, and signal processing.</p> <p><strong>Specifications</strong></p> <ul> <li>3,162 audio loops (7.3 hours)</li> <li>24-bit WAV format</li> <li>BPM labeled</li> <li>Tempo range: 100-130bpm</li> <li>Expressive percussion drumming</li> <li>Mixed rhythms of world music</li> </ul> <p><strong>License</strong></p> <p>This dataset was compiled by WaivOps, a crowdsourced music project managed by the sound label company Patchbanks. All recordings have been compiled by verified sources for copyright clearance.</p> <p>The WRLD-LP dataset is licensed under Creative Commons Attribution 4.0 International <a href="https://creativecommons.org/licenses/by/4.0/">(CC BY 4.0)</a>.</p> <p><strong>Additional Info</strong></p> <p>For audio examples or more information about this dataset, please refer to the <a href="https://github.com/patchbanks/WaivOps-WRLD-LP">GitHub repository</a>.</p>
Image for the dataset in "Extraction of stratigraphic exposures on visible images using a supervised machine learning technique"
<p>This is the original and hand-masked image for the dataset used in a research paper "Extraction of stratigraphic exposures on visible images using a supervised machine learning technique".</p> <p>The content is</p> <ul> <li>Original images with hand-masked images (original_images_NOGUCHIandShoji.zip) <ul> <li>training/* : original images used for the training dataset generation (60 files)</li> <li>training_masks/* : hand-masked images for the training dataset generation (60 files)</li> <li>validation/* : original images used for the training dataset generation (10 files)</li> <li>validation_masks/* : hand-masked images for the validation dataset generation (10 files)</li> <li>test/* : original images used as the test data (5 files)</li> <li>test_masks/* : hand-masked images used as the test data (5 files).</li> </ul> </li> </ul> <p>Note that original images include images obtained using <em>google-image-download</em>, a Python script published on GitHub (<a href="https://github.com/Joeclinton1/google-images-download/tree/patch-1">https://github.com/Joeclinton1/google-images-download/tree/patch-1</a>, Copyright © 2015-2019 Hardik Vasa). The whole images we obtained by <em>google-image-download</em> were labeled as noncommercial reuse with modification.</p> <p>For more details, please refer to a research paper "Extraction of stratigraphic exposures on visible images using a supervised machine learning technique".</p> <p>Correspondence: Rina Noguchi (r-noguchi@env.sc.niigata-u.ac.jp)</p>
Libraries generated in: Using Machine Learning to Predict the Antibacterial Activity of Ruthenium Complexes
<p>Libraries generated in the manuscript: "<strong>Using Machine Learning to Predict the Antibacterial Activity of Ruthenium Complexes"</strong>. The libraries can be generated locally by running the code provided on <a href="https://github.com/TheFreiLab/RutheniumML">GitHub</a>, but are also provided here free to download.</p>
Simulated datasets for detector and particle flow reconstruction: CLIC detector, hit-based data, machine learning format
<p>Derived from https://zenodo.org/record/8260741, prepared in a machine-learning friendly TFDS format, ready to be used with https://zenodo.org/record/8397954.</p> <ul> <li>clic_edm_ttbar_hits_pf10k.tar: ee -> ttbar, center of mass energy at 380 GeV, 10k events</li> <li>clic_edm_qq_hits_pf10k.tar: ee -> Z* -> qqbar, center of mass energy at 380 GeV, 10k events</li> </ul> <p><strong>Contents</strong></p> <p>Each .tar file contains the dataset in the <a href="https://github.com/tensorflow/datasets">tensorflow-datasets</a> (minimum version v4.9.1), <a href="https://github.com/google/array_record">array_record</a> format.</p> <p><strong>Dataset semantics</strong></p> <p>Each dataset consists of events that can be iterated over using the tensorflow-datasets library in either tensorflow or pytorch. Each event has the following information available:</p> <ul> <li>X: the reconstruction input features, i.e. tracks and calorimeter hits</li> <li>ygen: the ground truth particles with the features ["PDG", "charge", "pt", "eta", "sin_phi", "cos_phi", "energy", "jet_idx"], with "jet_idx" corresponding to the gen-jet assignment of this particle</li> <li>ycand: the baseline Pandora PF particles with the features ["PDG", "charge", "pt", "eta", "sin_phi", "cos_phi", "energy", "jet_idx"], with "jet_idx" corresponding to the gen-jet assignment of this particle</li> </ul> <p>The full semantics, including the list of features for X, are available at https://github.com/jpata/particleflow/blob/v1.6/mlpf/heptfds/clic_pf_edm4hep_hits/utils_edm.py.</p>
CCClim - A machine-learning powered cloud class climatology
<p>CCClim is based on cloud property retrievals from the European Space Agency's (ESA) Cloud\_cci dataset, adding relative occurrences of eight major cloud types as defined by the World Meteorological Organization (WMO) at 1° resolution. </p><p>The cloud types are predicted using a two stage machine learning framework, in which a 1 km pixel-level classifier is followed up with a grid-box scale Random Forest regression model.</p><p>CCClim's global coverage being almost gapless from 1982 to 2016 allows for performing process-oriented analyses of clouds on a climatological time scale. Similarly, the moderate spatial and temporal resolutions make it a lightweight dataset while enabling straightforward comparison to climate models.</p><p>The compressed tarball contains 35 netCDF files, each covering one calendar year. Each file provides daily averages of nine cloud-related variables and the nine classes (eight cloud types+undetermined) as per 1° grid box fractional amounts.</p><p>Cloud-related variables:</p><ul><li>cloud water path</li><li>ice water path</li><li>liquid water path</li><li>cloud optical depth</li><li>effective liquid droplet radius at cloud top</li><li>effective ice particle radius at cloud top</li><li>cloud top pressure</li><li>surface temperature</li><li>cloud area fraction</li></ul><p>Cloud types:</p><ul><li>Ci: Cirrus/Cirrostratus</li><li>As: Altostratus</li><li>Ac: Altocumulus</li><li>St: Stratus</li><li>Sc: Stratocumulus</li><li>Cu: Cumulus</li><li>Ns: Nimbostratus</li><li>Dc: Deep convective</li></ul>
Data related to the publication "Efficient molecular dynamics simulations of deep eutectic solvents with first-principles accuracy using machine learning interatomic potentials"
<p>The training data sets, the trained machine learning models, and input scripts for the training and molecular dynamics simulations.</p>
Bee Tracker – an open-source machine-learning based video analysis software for the assessment of nesting and foraging performance of cavity-nesting solitary bees
Open the record for dataset details and reuse information.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.