Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

558

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

558 results for “Training Data”

Learn how ShareScore rates datasets ↗
zenodo32/100

HagesLab/Absorber_NN - Sample Training Data and Models

Open the record for dataset details and reuse information.

opencc-by-4.0Apr 2024View details →
zenodo32/100

AIMD training data for H2 adsorption on MoxCy facets for machine learning model

<p>This data is used to train the MACE model for modeling hydrogen dissociation on molybdenum carbide surfaces.</p>

opencc-by-4.0Nov 2024View details →
zenodo32/100

Code and Training Data for "Cascaded Machine Learning of Soil Moisture and Salinity Prediction in Estuarine Wetlands based on In-situ Internet of Things Monitoring"

Open the record for dataset details and reuse information.

opencc-by-4.0Nov 2024View details →
zenodo32/100

ATELIER. Self-Labelling dynamic Loss function training data

<p>This dataset contains all the statistics collected during the training of multiple Neural Networks, which uses the following architecture technologies:</p> <p>Siamese Neural Networks, Dynamic loss function, modified during the training through Reinforcement learning and a 'Curriculum-Learning' cycle in between training cycles. The goal of the Models is to correctly identify anomalies in the QoE of real-time videos through the network KPIs.</p> <p>The dataset contains 3 type of experiment with different levels of model complexity.</p> <p>Please refer to the associated github repository and the published paper for a detailed description of both the dataset and the infrastructure that generated the dataset.<br>Journal paper title: "ATELIER: Service Tailored and Limited-Trust Network Analytics Using Cooperative Learning"</p>

openbsd-3-clause-clearApr 2024View details →
zenodo32/100

SynapseNet Training Data

<p>This dataset contains room-temperature single-axis TEM tomograms from Schaffer collateral and mossy fiber synapses in organotypic hippocampal slices.&nbsp;<br>The tomograms were published in the two studies [1, 2]. The data was re-used for training deep neural networks to segment different synaptic structures in electron micrographs in [3].&nbsp;<br>For the tomograms, organotypic slices were prepared from the hippocampi of neonatal mice according to the interface protocol55 and vitrified after 28 days in vitro in culture medium supplemented with 20% (w/v) bovine serum albumin using an HPM100 (Leica) high-pressure freezing device. The dataset also contains 23 tomograms resulting from chemically-fixed material, which were also published in (Maus et al., 2020). For these tomograms, wild-type animals at postnatal day 28 were transcardially perfused under deep anesthesia, first with 0.9% sodium chloride solution, and then one of two fixatives (Fixative 1: Ice-cold 4% paraformaldehyde, 2.5% glutaraldehyde in 0.1 M phosphate buffer16; Fixative 2: 37&deg; C 2% paraformaldehyde, 2.5% glutaraldehyde, 2 mM CaCl2, in 0.1 M cacodylate buffer56). Brains were rinsed and sectioned coronally through the dorsal hippocampus in an ice-cold 0.1 M phosphate buffer using a VT 1200S vibratome (Leica) (step size 100 &micro;m; amplitude 1.5 mm, speed 0.1 mm/sec). Hippocampal CA3 subregions were excised using a 1.5 mm diameter biopsy punch and high-pressure frozen on the same day in 20% (w/v) bovine serum albumin using an HPM100 (Leica) high-pressure freezing device. For both sample preparations, automated freeze-substitution was performed. Tomograms were collected using a 200 kV JEM-2100 (JEOL) transmission electron microscope equipped with an 11 MP Orius SC1000 CCD camera (Gatan). Tilt-series (tilt range +/- 60&deg;; 1&deg; angular increments) were acquired at 30 000x magnification using SerialEM58. Tomographic reconstructions were generated using weighted back-projection with etomo.<br>The data is organized into two different subfolders for data with annotations for "vesicles" and "active_zones". Each of these subfolders is further subdivided into "train" and "test" folders, which contain<br>tomograms for the two different sample preparations in "chemical_fixation" and "single_axis_tem".<br>Each tomogram and the corresponding annotation is stored as a hdf5 file, containing the following internal datasets:<br>- raw: The tomogram data.<br>- labels/vesicles: Annotations for the synaptic vesicles, annotated with IMOD, further postprocessed and then exported to instance masks. (for tomograms in "vesicles")<br>- labels/AZ: Annotations for the active zone, annotated with IMOD and exported to binary masks.</p> <p>[1] Imig et al., The Morphological and Molecular Nature of Synaptic Vesicle Priming at Presynaptic Active Zones, Neuron, 2014, DOI:10.1016/j.neuron.2014.10.009<br>[2] Maus et al., Ultrastructural Correlates of Presynaptic Functional Heterogeneity in Hippocampal Synapses, Cell Reports, 2020, DOI: 10.1016/j.celrep.2020.02.083<br>[3] Muth, Moschref et al., 2024, Preprint to be published</p>

opencc-by-4.0Nov 2024View details →
zenodo32/100

Annotated data of simultaneous broadband radio and optical emission of meteor trains imaged by LOFAR / AARTFAAC and CAMS

<p>This data set contains simultaneous 30 - 60 MHz LOFAR / AARTFAAC12 radio observations and CAMS low-light video observations of +4 to -10 magnitude meteors at the peak of the Perseid meteor shower on August 12/13, 2020. 204 meteor trains were imaged in both the radio and optical domain.</p>

opencc-by-4.0Nov 2021View details →
zenodo32/100

Pre-trained DNN model data for pruning example code

<p>Pre-trained DNN model datasets for example codes of&nbsp;neural network pruning.</p> <p>Example pruning codes are published in &quot;https://github.com/FujitsuLaboratories/CAC/tree/main/cac/pruning&quot;.</p>

opencc-zeroNov 2021View details →
zenodo32/100

Data set from: Robot-Assisted Gait Training in Patients with Multiple Sclerosis: A Randomized Controlled Crossover Trial.

<p>Data set from the paper&nbsp;&quot;Robot-Assisted Gait Training in Patients with Multiple Sclerosis: A Randomized Controlled Crossover Trial&quot;&nbsp;&nbsp;doi:&nbsp;<a href="https://dx.doi.org/10.3390%2Fmedicina57070713">10.3390/medicina57070713</a></p>

opencc-by-4.0Dec 2021View details →
zenodo32/100

AHI-CALIOP Collocated Data for Training and Validation of Cloud Masking Neural Networks

<p>Collocated data between AHI at 2km resolution (nadir)&nbsp;and CALIOP 1km cloud product v4.20 used for training and validating cloud identification neural networks. The main training and validation data from 2019 is stored in monthly directories, whilst the collocated dataset used to compare the NN, JMA and BoM cloud mask performances is the file &quot;superdf.h5&quot;. All collocated data is stored as .h5 files and was built using the Python Pandas package. In&nbsp;this archive, the data has been&nbsp;stored&nbsp;as compressed directories for each month or as a single compressed file in the case of &quot;superdf.h5&quot; using tar with bzip2 compression or just bzip2 compression respectively.</p>

opencc-by-4.0Dec 2021View details →
dryad32/100

Linkage of hospital records and death certificates by a search engine and machine learning: training and test set data

<p>INTRODUCTION: Vital status is of central importance to hospital clinical research. However, hospital information systems record only in-hospital death information. Recently, the French government released a publicly available dataset containing death-certificate data for over 25 million individuals. The objective of this study was to link French death certificates to the Bordeaux University Hospital records to complete the vital status information.</p> <p>MATERIALS AND METHODS: Our linkage strategy was composed of a search engine to reduce the number of comparisons and machine-learning algorithms. The overall pipeline was evaluated by assembling a file containing 3,565 in-hospital deaths and 15,000 alive persons.</p> <p>RESULTS: The recall and precision of our linkage strategy were 97.5% and 99.97% for the upper threshold and 99.4% and 98.9% for the lower threshold, respectively.</p> <p>CONCLUSION: In this article, we demonstrated the feasibility of accurately linking hospital records with death certificates using a search engine and machine learning.</p>

opencc-zeroJan 2022View details →
zenodo32/100

Training data for the "Metabarcoding/eDNA through Obitools" tutorial

<p>Training dataset for Obitools tutorial using Galaxy.</p>

opencc-by-4.0Jan 2022View details →
zenodo32/100

Mothra training data

<p>Training data related to the <a href="https://github.com/machine-shop/mothra">mothra</a> project.</p>

opencc-by-4.0Jul 2021View details →
zenodo32/100

Toward General Speech Restoration With VoiceFixer - Training Data

<p>This data set contains speech data and noise data for the paper:&nbsp;VoiceFixer: Toward General Speech Restoration with Neural Vocoder.</p> <p>If you found this dataset helpful, please consider citing:&nbsp;</p> <blockquote> <pre>@article{liu2021voicefixer, title={VoiceFixer: Toward General Speech Restoration with Neural Vocoder}, author={Liu, Haohe and Kong, Qiuqiang and Tian, Qiao and Zhao, Yan and Wang, DeLiang and Huang, Chuanzeng and Wang, Yuxuan}, journal={arXiv preprint arXiv:2109.13731}, year={2021} }</pre> </blockquote>

opencc-by-4.0Sep 2021View details →
zenodo32/100

ONT Training Data (WHO AFRO/SANBI RCEGSB Training)

<p>Oxford Nanopore (ONT) reads for training</p>

opencc-by-4.0Feb 2022View details →
zenodo32/100

Machine Learning Potentials for Metal-Organic Frameworks with Thermodynamic Transferability: training data

<p>This dataset contains&nbsp;potential energies, forces, and virial stress for a large set of reference configurations for UiO-66(Zr) and MIL-53(Al), computed at the PBE-D3 level using CP2K 7.1. The basis set contained both TZVP Gaussian basis functions as well as plane waves (cutoff energy 800 Ry for UiO-66(Zr) and 900 Ry for MIL-53(Al)). The sampling of the Brillouin zone was restricted to the gamma point.</p>

opencc-by-4.0Mar 2022View details →
zenodo32/100

Multilingual CoNaLa Datset, train data

<p>Training datasets used in the Multilingual CoNaLa benchmark experimentation.&nbsp;</p>

opencc-by-4.0Mar 2022View details →
zenodo32/100

Data set - The design and evaluation of an integrated training load and injury/illness surveillance system in competitive swimming.

<p>Dataset for paper as per title above.&nbsp;</p>

opencc-by-4.0Jun 2022View details →
zenodo32/100

Training data - Hydraulic Head Probabilistic MLP-NN

<p>Training data for paper on hydraulic head predictions using an MLP-NN.</p>

opencc-by-4.0Jul 2022View details →
dryad32/100

Data from: Is there a benefit for anesthesiologists of adding difficult airway scenarios for learning fiberoptic intubation skills using virtual reality training? A randomized controlled study

<p><strong><span>Introduction</span></strong><span><strong>:</strong> Fiberoptic intubation for a difficult airway requires significant experience. Traditionally only normal airways were available for high fidelity bronchoscopy simulators. It is not clear if training on difficult airways offers an advantage over training on normal airways. This study investigates the added value of difficult airway scenarios during virtual reality fiberoptic intubation training.</span></p> <p><span><strong>Methods:</strong> </span><span>A prospective multicentric randomized study was conducted 2019 to 2020, among 86 inexperienced anesthesia residents, fellows and staff. Two groups were compared: Group N (control, n=43) first trained on a normal airway and Group D (n=43) first trained on a normal, followed by three difficult airways. All were then tested by comparing their Global Rating Scores (GRS) on 5 scenarios (1 normal and 4 difficult airways).</span></p> <p><span><strong>Results:</strong> </span><span>The final evaluation GRS score for the normal airway testing scenario was significantly higher for group N than group D: median score 76% (IQR 56.5 - 90) versus 58% (IQR 51.5 - 69, p = 0.0039), but there was no difference in GRS scores for the difficult intubation testing scenarios. </span></p> <p><span><strong>Conclusions:</strong> </span><span>A single exposure to each of 3 different difficult airway scenarios did not lead to better fiberoptic intubation skills on previously unseen difficult airways, when compared to multiple exposures to a normal airway scenario. This finding may be due to the learning curve of approximately 5-10 exposures to a specific airway scenario required to reach proficiency. </span></p>

opencc-zeroJul 2022View details →
zenodo32/100

Training data for 'Maximum Likelihood Phylogeny Reconstruction'' (Galaxy Training Material)

<p>This data is used for Galaxy Training Network (GTN) training &#39;Maximum Likelihood Phylogeny Reconstruction&#39;. It consists of 173 amino acid alignments of orthologs found in chromosome 5 of four strains of S. cerevisiae. Original sequence data (https://zenodo.org/record/6610704) was processed in Galaxy following GTN &#39;Preparing genomic data for phylogeny reconstruction&#39; training (10.48546/workflowhub.workflow.359.1) to generate alignments of orthologs.</p>

opencc-by-4.0Jul 2022View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record