Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

558

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

558 results for “Training Data”

Learn how ShareScore rates datasets ↗
dryad36/100

Training and test data for: Not getting in too deep: A practical deep learning approach to routine crystallisation image classification

<p>These data were used to classify crystallisation experiments in Milne et al., (<a href="https://doi.org/10.1101/2022.09.28.509868">https://doi.org/10.1101/2022.09.28.509868</a>). Here, four of the most widely-used convolutional deep-learning network architectures that can be implemented without the need for extensive computational resources were compared. It was shown that the classifiers have different strengths that can be combined to provide an ensemble classifier achieving a classification accuracy comparable to that obtained by a large consortium initiative (Bruno et al. PLOS one, 13(6), 2018). Eight classes were used to rank the experimental outcomes, thereby providing detailed information that can be used with routine crystallography experiments to automatically identify crystal formation for drug discovery and pave the way for further exploration of the relationship between crystal formation and crystallisation conditions.</p>

opencc-zeroJan 2023View details →
zenodo36/100

Galaxy Training Data for "End-to-End Tissue Microarray Image Analysis with Galaxy-ME"

<p>This dataset provides the inputs&nbsp;used in the Galaxy Training Network (GTN) training &#39;End-to-End Tissue Microarray Image Analysis with Galaxy-ME&#39;. The tutorial demonstrates how to use the Galaxy-ME tool suite for primary image processing, data analysis, and interactive visualization&nbsp;of multiple tissue imaging datasets. Original data was published by <a href="https://pubmed.ncbi.nlm.nih.gov/34824477/">Schapiro <em>et al</em></a>.</p>

opencc-by-4.0Feb 2023View details →
zenodo36/100

Training, Validation and Test Sets for paper 'A Little Data goes a Long Way: Automating Seismic Phase Arrival Picking at Nabro Volcano with Transfer Learning'

<p>Training, Validation and Test Data for model presented in&nbsp;paper &#39;A Little Data Goes A Long Way: Automating Seismic Phase Arrival Picking at Nabro Volcano with Transfer Learning&#39;, submitted to Journal of Geophysical Research: Solid Earth.</p> <p>Files:</p> <p>- train_events_2498.h5 = training set of seismic waveforms (events with P-/S-wave labelled arrivals only, i.e., no noise waveforms)</p> <p>- train_events_2498.pkl = event training set metadata (UTC P-/S-wave phase arrival times)</p> <p>- train_noise_2498.h5 = training set of seismic waveforms (noise sections only, i.e., no event waveforms)</p> <p>- train_noise_2498.pkl = noise training set metadata (UTC time&nbsp;for training noise waveforms)</p> <p>- val_events.h5 = validation set of seismic waveforms (events with P-/S-wave labelled arrivals only, i.e., no noise waveforms)</p> <p>- val_events.pkl = event validation set metadata (UTC P-/S-wave phase arrival times)</p> <p>- val_noise.h5 = validation&nbsp;set of seismic waveforms (noise sections only, i.e., no event waveforms)</p> <p>- val_noise.pkl = noise validation set metadata (UTC time&nbsp;for validation noise waveforms)</p> <p>- test.h5 = test&nbsp;set of seismic waveforms (events and noise)</p> <p>- test_events.pkl = event test set metadata (UTC P-/S-wave phase arrival times for test event waveforms)</p> <p>- test_noise.pkl = noise test set metadata (UTC time for test noise waveforms)</p> <p>- nabro_2011-247.mseed = 24 hours seismic data from Nabro Urgency Array (2011-09-04), saved in mseed format (e.g., can be read with obspy)</p> <p>- nabro_2011-269.mseed = 24 hours seismic data from Nabro Urgency Array (2011-09-26), saved in mseed format (e.g., can be read with obspy)</p> <p>&nbsp;</p> <p>Further details and code for reading and using&nbsp;these files can be found at the GitHub repo for this paper:&nbsp;<a href="https://github.com/sachalapins/U-GPD">https://github.com/sachalapins/U-GPD</a></p> <p>&nbsp;</p>

opencc-by-4.0Feb 2021View details →
dryad36/100

Data from: Cadaveric emergency cricothyrotomy training for non-surgeons using a bronchoscopy-enhanced curriculum

<p>Emergency cricothyrotomy training for non-surgeons providing critical care is important as rare "cannot intubate or oxygenate events" may occur multiple times in a provider's career when surgical expertise is not immediately available. However, such training is highly variable and often infrequent so enhancing these experiences is important.</p> <p>Our study was performed after implementing a program to train non-surgeon providers on cadaveric donors. Standard training with an instructional video and live coaching was enhanced by bronchoscopic visualization of the trachea allowing participants to review their technique after performing scalpel and Seldinger-technique procedures, and to review their colleagues' technique on live video. Feasibility was measured through assessing helpfulness for trainees, cost, setup time, quality of images, and operator needs. Footage from the bronchoscopy recordings was analyzed to assess puncture-to-tube time and errors. Participants submitted pre- and post-session surveys assessing their levels of experience and gauging their confidence and anxiety with cricothyrotomies.</p> <p>We found our training program met feasibility criteria for low costs, setup time, and operator needs. Furthermore, all 24 participants rated the cadaveric session as helpful and demonstrated efficient technique by puncture-to-tube times. Bronchoscopy videos revealed that sharp instrument puncturing of the posterior tracheal wall was common and that improper tube placement occurred but was uncommon. Bronchoscopic enhancement was rated as quite or extremely helpful for visualizing the trachea and to assess depth of instrumentation. There was a significant increase in confidence and decrease in performance anxiety after the session.</p> <p>Our findings confirm that supplementing cadaveric emergent cricothyrotomy training programs with tracheal bronchoscopy is feasible, helpful to trainees, meets prior documented times for efficient technique, and detects technical errors that would have been missed in a standard training program. Bronchoscopic enhancement is a valuable addition to cricothyrotomy cadaveric training programs and may help avoid real-life complications.</p>

opencc-zeroMar 2023View details →
zenodo36/100

Machine-guided path sampling to discover mechanisms of molecular self-organization (Training and validation data)

<p>Training and validation data for the Nature Computational Science manuscript &quot;Machine-guided path sampling to discover mechanisms of molecular self-organization&quot;</p>

opencc-by-4.0Mar 2023View details →
zenodo36/100

TARDIS configuration and emulator weights and training data for "1991T-Like Type Ia Supernovae as an Extension of the Normal Population"

<p>This dataset contains two archives of data related to the paper &quot;1991T-Like Type Ia Supernovae as an Extension of the Normal Population&quot;<br> <br> The first dataset, <a href="https://zenodo.org/api/files/de696fe0-3280-44f2-8ef4-975b92fad260/TARDIS_Emulator_Config.tar.gz">TARDIS_Emulator_Config.tar.gz </a>, contains the atomic data used to run TARDIS and a template configuration file from which samples are generated including the flags for the physics implementation used.</p> <p>The second dataset, InferenceScripts.tar.gz, contains the trained probabilistic neural network, the training/validation data (Under NNData), and scripts used to train the model and load and evaluate the model.&nbsp; Scripts that perform inference on spectra, as well as a folder of observed spectra (Under CorrectedSpectra), are included as well.&nbsp; A conda environment yaml file is included to rebuild the Python environment required to run all of the scripts.&nbsp; For questions please email John O&#39;Brien.</p>

opencc-by-4.0Apr 2023View details →
dryad36/100

Raw data of: Spelling training "Errorless learning" using Tablet PC in elementary schools

<p>Mobile devices and multimedia content have recently played a central role in conveying information and ideas. In addition, these can also be used effectively for individualized learning situations. Successful mastery of writing and reading skills is considered a key competence for academic and professional success. Particular difficulties in spelling often show a high degree of persistence without appropriate intervention. Based on a BMBF-funded research project, a tablet-based spelling training program with handwritten input and direct feedback on word correctness was developed and applied. The data sets are raw data showing the development of children's spelling skills. They also include the writing movement and pressure of each letter of all children. The data sets indicate that the children's spelling improved as a result of the individualized tablet application.</p>

opencc-zeroApr 2023View details →
zenodo36/100

Training Data for "Metatranscriptomics analysis using microbiome RNASeq data"

<p>Microbiomes play a critical role in host health, disease, and the environment.. Functional microbiome analysis which estimates the functional groups expressed by microbial community enables researchers to look beyond taxonomic composition and correlation with the condition under study. Using microbial community RNA-Seq data and subsequent metatranscriptomics workflows to elucidate the functional complement of the microbiome is gaining interest in the field.<br> This&nbsp;tutorial from Galaxy training network will introduce researchers to the basic concepts of metatranscriptomics data analysis. It takes in paired-end datasets of raw shotgun sequences (in FastQ format) as an input and:</p> <ol> <li>preprocess</li> <li>extract and analyze the community structure (taxonomic information)</li> <li>extract and analyze the community functions (functional information)</li> <li>combine taxonomic and functional information to offer insights into taxonomic contribution to a function or functions expressed by a particular taxonomy.</li> </ol> <p>The dataset used in the tutorial comes from a time-serie analysis of a microbial community inside a bioreactor (Kunath et al, ISME, 2018). Only the data for one time point (1st) and a biological replicate (A) is analyzed here, after having been trimmed out the original file for the purpose of saving time and resources.</p>

opencc-by-4.0May 2023View details →
zenodo36/100

FunMap feature data file used for model training

<p>The provided file is a TSV (Tab-Separated Values, gzipped) file that serves as input for training a machine learning model. The file structure consists of rows representing gene pairs and columns containing 16 Spearman correlation coefficients and 16 mutual ranks. The Spearman correlation coefficients capture the strength and direction of the relationship between gene pairs. Additionally, the mutual ranks are derived from the correlation coefficients and are utilized in the training process of the model.</p>

opencc-by-4.0May 2023View details →
zenodo36/100

Training data and test data sets for simultaneous inversion of velocity density based on U-T

<p>Here are the&nbsp;training and testing data sets involved in the numerical experiments in the article that has been submitted to the journal &ldquo;Journal of Geophysical Research: Solid Earth&rdquo;, named &ldquo;Joint Model and Data-Driven Simultaneous Inversion of Velocity and Density&rdquo;: Marmousi model. Each dataset consists of two parts: a training dataset and a testing dataset. Both training and testing data sets contain three parts: seismic data, velocity model and density model.</p>

opencc-by-4.0May 2023View details →
zenodo36/100

U-T training and test data for LayerFault model

<p>Here are the training and testing data sets involved in the numerical experiments in the article that has been submitted to the journal &ldquo;Journal of Geophysical Research: Solid Earth&rdquo;, named &ldquo;Joint Model and Data-Driven Simultaneous Inversion of Velocity and Density&rdquo;:&nbsp; LayerFault model.&nbsp;Each dataset consists of two parts: a training dataset and a testing dataset. Both training and testing data sets contain three parts: seismic data, velocity model and density model.</p>

opencc-by-4.0May 2023View details →
zenodo36/100

Anndata object of 10x Mouse Brain 5k data set for scDGD training

<p>This data from 10x (5k Adult Mouse Brain Nuclei Isolated with Chromium Nuclei Isolation Kit, Single Cell Gene Expression Dataset by Cell Ranger 7.0.0, 2022) is comprised of 7377 cells from the adult mouse brain with 32285 features. Cell type annotations were approximated using CellTpyist with the `Developing_Mouse_Brain` reference model and majority voting. This resulted in 7 distinct cell types.</p>

opencc-by-4.0May 2023View details →
zenodo36/100

Training data for the Singapore Butterfly Classifier used in the XPRIZE Rainforest semifinals

<p>Training data based on&nbsp;Lepidoptera observation data from&nbsp;Singapore with images&nbsp;attached from the Global Biodiversity Index Facility (GBIF) (de Vries H, Lemmens M (2023). Observation.org, Nature data from around the World. Observation.org. Occurrence dataset https://doi.org/10.15468/5nilie accessed via GBIF.org on 2023-04-17; Nature data from around the World, Observation.org, https://doi.org/10.15468/5nilie accessed via GBIF.org on 2023-04-17; iNaturalist Research-grade Observations, iNaturalist.org, https://doi.org/10.15468/ab3s5x; Earth Guardians Weekly Feed, Occurrence dataset <a href="https://doi.org/10.15468/slqqt8">https://doi.org/10.15468/slqqt8</a>)</p> <p>Annotations were produced with a generic butterfly detector trained on annotations available at https://www.kaggle.com/datasets/mistag/arthropod-taxonomy-orders-object-detection-dataset, accessed on 2023-05-08). Annotations were produced with the generic butterfly detector while&nbsp;labels assigned using the specied ID&nbsp;assigned in the GBIF&nbsp;observation.</p> <p>Training data contains 230 species. These were the ones that contained species ID in relevant columns, contained&nbsp;50+ images, and received predictions with the generic butterfly detector.</p> <p>&nbsp;</p> <p>Dataset: Observation.org, Nature data from around the World&nbsp;<br> Rights as supplied: http://creativecommons.org/licenses/by-nc/4.0/legalcode<br> Dataset: iNaturalist Research-grade Observations&nbsp;<br> Rights as supplied: http://creativecommons.org/licenses/by-nc/4.0/legalcode<br> Dataset: Earth Guardians Weekly Feed&nbsp;<br> Rights as supplied: http://creativecommons.org/licenses/by-nc/4.0/legalcode</p> <p>&nbsp;</p> <p>&nbsp;</p>

opencc-by-4.0Jun 2023View details →
zenodo36/100

TM-Vec training data

<p>Protein/Domain pairs and corresponding TM-scores used for training TM-Vec models</p> <p>Training data for TM-Vec models used in: &ldquo;TM-Vec: template modeling vectors for fast homology detection and alignment&rdquo; by Hamamsy et al. Preprint: https://www.biorxiv.org/content/10.1101/2022.07.25.501437v1</p>

opencc-by-4.0Jun 2023View details →
zenodo36/100

Unlocking the Potential of Health Data: A Distributed Analysis Approach based on Personal Health Train Infrastructure

<p>While there is a great availability of medical datasets, they are usually focused on a specific research question. This is useful for making experiments transparent and reproducible, however, these datasets can be more efficiently used in other kinds of analyses, where it not for data privacy issues. The Personal Health Train (PHT) provides a distributed analysis infrastructure that follows the FAIR principles and gives control to the data owners (providers) about how their data are used by scientists or other users (consumers).</p>

opencc-by-4.0Jul 2023View details →
zenodo36/100

choderalab/refit-espaloma: Preprocessed data to train espaloma-0.3

<p>This is a collection of preprocessed data&nbsp;used to train <strong>espaloma-0.3</strong>, which is&nbsp;published in:</p> <p>Kenichiro Takaba, Iv&aacute;n Pulido,&nbsp;Pavan Kumar Behara, Mike Henry, Hugo MacDermott Opeskin, John D. Chodera, Yuanqing Wang.&nbsp;&quot;Machine-learned molecular mechanics force field for the simulation of protein-ligand systems and beyond&quot; (<a href="https://arxiv.org/abs/2307.07085">arXiv:2307.07085</a>)</p> <p>The provided data is compatible with the data generated at&nbsp;<a href="https://github.com/choderalab/refit-espaloma/tree/main/openff-default/02-train/merge-data">https://github.com/choderalab/refit-espaloma/tree/main/openff-default/02-train/merge-data</a>.</p>

opencc-by-4.0Jul 2023View details →
zenodo36/100

DADA2 formatted Silva SSU taxonomic training data (Silva version 138.1) with emended description of the genus Lactobacillus Beijerinck 1901

<p>These training fasta files are derived from the Silva 138.1 prokaryotic SSU taxonomic training data formatted for DADA2 (from <a href="https://zenodo.org/record/4587955">https://zenodo.org/record/4587955</a>). The species assignment file contains changes in species names according to <a href="https://doi.org/10.1099/ijsem.0.004107">Zheng et al. 2020</a>&nbsp;based on data from <a href="https://github.com/swuyts/lactotax/tree/master">Lactotax</a> (file <a href="https://github.com/swuyts/lactotax/raw/master/data/2023_05_30.xlsx">2023_05_30.xlsx</a>). The script in R for making changes in species names is in the silva.R file.</p> <p><br> Please cite one or both of the Silva references, the DADA2 paper (reference below), and the Zenodo record for this specific version in your Methods or published source code to record the specific taxonomic database files used in your analysis.</p>

opencc-by-4.0Jul 2023View details →
zenodo36/100

Covid training data

<p><strong>Vous disposez pour ce TP d&rsquo;&eacute;chantillons d&rsquo;&eacute;couvillonnage issus des voies respiratoires sup&eacute;rieures (&eacute;couvillon nasal, s&eacute;cr&eacute;tions) chez des patients pr&eacute;sentant un syndrome respiratoire s&eacute;v&egrave;re. Ces &eacute;chantillons ont &eacute;t&eacute; r&eacute;alis&eacute;s entre le 24 janvier et le 24 mars 2020 en France par l&rsquo;Institut Pasteur (<a href="https://doi.org/10.1101/2020.04.24.059576">https://doi.org/10.1101/2020.04.24.059576</a>). Apr&egrave;s extraction des ARN pr&eacute;sents dans l&rsquo;&eacute;chantillon, ceux-ci ont &eacute;t&eacute; convertis en ADN double brin. Les librairies g&eacute;n&eacute;r&eacute;es ont &eacute;t&eacute; s&eacute;quenc&eacute;es par s&eacute;quen&ccedil;age Illumina NextSeq500 (2x150)</strong></p>

opencc-by-4.0Aug 2023View details →
dryad36/100

Training and test data with scripts for simulation-trained deep learning and likelihood-based phylogeography comparisons

Open the record for dataset details and reuse information.

publicJan 2024View details →
dryad36/100

Training and test data for: Not getting in too deep: A practical deep learning approach to routine crystallisation image classification

Open the record for dataset details and reuse information.

publicJan 2023View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record