Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

558

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

558 results for “Training Data”

Learn how ShareScore rates datasets ↗
zenodo40/100

Data for Project 'Test-Retest Reliability and Validity of vagally-mediated Heart Rate Variability to Monitor Internal Training Load in Older Adults: A within-subjects (repeated-measures) randomized study'

<p>Data for Project &#39;Test-Retest Reliability and Validity of vagally-mediated Heart Rate Variability to Monitor Internal Training Load in Older Adults: A within-subjects (repeated-measures) randomized study&#39; consisting of (1)&nbsp;the original and complete dataset (&#39;Data_Brain-IT-Reliability-of-HRV-during-Exergaming_for-publication&#39;; and (2)&nbsp;a corresponding README file including (a) general information, (b) data and file overview, (c) sharing and access information, (d) methodological information, and (e) data-specific information.</p>

opencc-by-4.0May 2023View details →
zenodo40/100

PhageHostLearn - training data and cluster analysis

<p>These data comprise the processed phage RBP and&nbsp;<em>Klebsiella&nbsp;</em>K-loci sequence data to train our PhageHostLearn system, along with ESM-2 embeddings of the RBPs and loci, as well as results from the cluster analyses of K-loci proteins and RBPs at 50% identity with CD-HIT.</p>

opencc-by-4.0Jun 2023View details →
zenodo40/100

GOBAI-O2 Training Data

<p><strong>About</strong></p> <p>This repository contains quality controlled data obtained from the Biogeochemical Argo database (Argo, 2000) and GLODAP data product (Lauvset et al., 2022) that are used to train machine leaning models to be applied to three-dimensional temperature and salinity maps compiled using the Core Argo data (<a href="https://sio-argo.ucsd.edu/RG_Climatology.html">Roemmich and Gilson, 2009</a>) to produce GOBAI-O2 (Sharp et al., 2022).</p> <p><strong>References</strong></p> <p>Argo (2000). Argo float data and metadata from Global Data Assembly Centre (Argo GDAC). SEANOE.&nbsp;<a href="https://doi.org/10.17882/42182">https://doi.org/10.17882/42182</a>.</p> <p>Lauvset,&nbsp;S. K., Lange,&nbsp;N., Tanhua,&nbsp;T., Bittig,&nbsp;H. C., Olsen,&nbsp;A., Kozyr,&nbsp;A., Alin,&nbsp;S., &Aacute;lvarez,&nbsp;M., Azetsu-Scott,&nbsp;K., Barbero,&nbsp;L., Becker,&nbsp;S., Brown,&nbsp;P. J., Carter,&nbsp;B. R., da Cunha,&nbsp;L. C., Feely,&nbsp;R. A., Hoppema,&nbsp;M., Humphreys,&nbsp;M. P., Ishii,&nbsp;M., Jeansson,&nbsp;E., Jiang,&nbsp;L.‑Q., Jones,&nbsp;S. D., Lo Monaco,&nbsp;C., Murata,&nbsp;A., M&uuml;ller,&nbsp;J. D., P&eacute;rez,&nbsp;F. F., Pfeil,&nbsp;B., Schirnick,&nbsp;C., Steinfeldt,&nbsp;R., Suzuki,&nbsp;T., Tilbrook,&nbsp;B., Ulfsbo,&nbsp;A., Velo,&nbsp;A., Woosley,&nbsp;R. J. and Key,&nbsp;R. M. (2022). GLODAPv2.2022: the latest version of the global interior ocean biogeochemical data product. Earth System Science Data, 14(12), 5543‑5572. <a href="https://doi.org/10.5194/essd-14-5543-2022">https://doi.org/10.5194/essd‑14‑5543‑2022</a>.</p> <p>Roemmich, D. and Gilson, J. (2009). The 2004-2008 mean and annual cycle of temperature, salinity, and steric height in the global ocean from the Argo Program. Progress in Oceanography, 82, 81-100. <a href="https://doi.org/10.1016/j.pocean.2009.03.004">https://doi.org/10.1016/j.pocean.2009.03.004</a>.</p> <p>Sharp, J. D. Fassbender, A. J. Carter, B. R., Johnson, G. C., Schultz, C., and Dunne, J. P. (2022). GOBAI-O<sub>2</sub>: A Global Gridded Monthly Dataset of Ocean Interior Dissolved Oxygen Concentrations Based on Shipboard and Autonomous Observations (NCEI Accession 0259304). v1.0. NOAA National Centers for Environmental Information. Dataset. <a href="https://doi.org/10.25921/z72m-yz67">https://doi.org/10.25921/z72m-yz67</a>.</p>

opencc-by-4.0Mar 2023View details →
zenodo40/100

Training Data for "Creating Quality FAIR assessment reports and draft of Data Papers from EML metadata with MetaShRIMPS"

<p>Training Data for &quot;Training Data for &quot;Creating Quality FAIR assessment reports and draft of Data Papers from EML metadata with MetaShRIMPS&quot;&quot;</p>

opencc-by-4.0Jul 2023View details →
zenodo40/100

Reproduction package for the paper "The Impact of Hard and Easy Negative Training Data on Vulnerability Prediction Performance"

<p>This Reproduction package contains the datasets, code and results for the paper &quot;The Impact of Hard and Easy Negative Training Data on Vulnerability Prediction Performance&quot; for other researchers to use for reproducing or improving our work.&nbsp;</p>

opencc-by-4.0Feb 2023View details →
dryad40/100

Annotated spiral ganglion neuron training data for object detection

Open the record for dataset details and reuse information.

publicJun 2025View details →
dryad40/100

Data from: How many specimens make a sufficient training set for automated three dimensional feature extraction?

Open the record for dataset details and reuse information.

publicMay 2024View details →
dryad40/100

Data from: Postdoctoral T32 training is correlated with obtaining an academic primarily research faculty position

Open the record for dataset details and reuse information.

publicNov 2025View details →
dryad40/100

Data and trained models for: Human-robot facial co-expression

Open the record for dataset details and reuse information.

publicMar 2024View details →
dryad40/100

Training data from: Machine learning predicts which rivers, streams, and wetlands the Clean Water Act regulates

Open the record for dataset details and reuse information.

publicDec 2023View details →
dryad40/100

Data from: Convolutional neural networks trained on internal variability predict forced response of TOA radiation by learning the pattern effect

Open the record for dataset details and reuse information.

publicFeb 2025View details →
dryad40/100

An individual-based model trained on multiple data sources estimates population connectivity and facilitates aggregation of harvest management units

Open the record for dataset details and reuse information.

publicOct 2024View details →
zenodo36/100

Training data for 'Unicycler assembly of SARS-CoV-2 genome with preprocessing to remove human genome reads' tutorial (Galaxy Training Material)

<p>The data here is a copy of the corresponding SRR records in the NCBI SRA. The duplication serves a dual purpose:</p> <ol> <li>as a backup should there be problems connecting to NCBI servers, e.g., during Galaxy user trainings.</li> <li>to illustrate how to obtain raw sequencing data from alternative sources, and to organize the data into the same collection structure in a Galaxy history that is generated by specialized Galaxy SRA download tools.</li> </ol>

opencc-by-4.0Mar 2020View details →
zenodo36/100

360-degree video recording of an outdoor camerawork training session for qualitative data collection

<p>In this equirectangular 360&deg; video clip, a recording of an outdoor camerawork training session is stitched together from the footage taken by a stereoscopic 360&deg; camera with eight lenses. To play the video and spatial sound correctly, use a digital video player that can play equirectangular videos with the YouTube ambiX First Order Ambisonic audio format, eg. VLC or PotPlayer. Please wear headphones.</p> <p>In the camerawork training session, all the participants are playing particular roles in the training session, and each carries a camera. In preparation for the real data collection with a guide, one person is pretending to be a nature guide. She carries a GoPro camera on a gimbal. There is an instructor, who is carrying a single lens 360&deg; camera on a raised extension pole with a separate ambisonic microphone. Two others are filming with a prosumer camcorder and a single lens 360&deg; camera on a lowered extension pole respectively. And a fifth person is filming with a stereoscopic 360&deg; camera and an independent ambisonic microphone on a monopod. In a nutshell, this is a typical team filming arrangement, in which the team needs to attentively yet silently coordinate their joint camerawork. Languages: Danish and English</p>

opencc-by-nc-nd-4.0Oct 2018View details →
zenodo36/100

Data and code for training and evaluating machine learning models for thunderstorm prediction from reanalysis data

<p>FIXED Data and Python code for training and evaluating machine learning models for predicting thunderstorms, associated with the paper:</p> <p>&quot;Evaluation of machine learning classifiers for predicting deep convection&quot;</p> <p>by Peter Ukkonen and Antti M&auml;kel&auml;&nbsp;(to appear in JAMES)</p> <p>The data (preprocessed inputs and outputs)&nbsp;is stored as netCDF files and .mat files which can be loaded with Python.&nbsp;</p>

opencc-by-4.0Nov 2018View details →
zenodo36/100

Generating Physically Sound Training Data for Image Recognition of Additively Manufactured Parts Data and Scripts

<p>The repository contains the data corresponding to the Paper &quot;Generating Physically Sound Training Data for Image Recognition of Additively Manufactured Parts&quot;.</p> <p>Random30, Random50, Random100, Similiar10, Similar30 and Similar50.zip contain the data sets (obj Files).</p> <p>R30_physical_images.zip and sim50_physical_images.zip contain the photos made from the physical components which are used for the evaluation.</p>

opencc-by-4.0Jan 2021View details →
zenodo36/100

Training and Test data for BaCoN

<p>Training and test datasets consisting of matter power spectra for use with the code BAyesian COsmological Network (BaCoN):&nbsp;https://github.com/Mik3M4n/BaCoN</p> <p>The dataset&nbsp;was generated using the code ReACT:&nbsp;https://github.com/nebblu/ReACT</p> <p>See the github repo for additional details.</p> <p>&nbsp;</p> <p>&nbsp;</p>

opencc-by-4.0Dec 2020View details →
zenodo36/100

Training data from Reinforced SciNet

<p><strong>Summary:</strong></p> <p>The results from the training of neural networks in v2 of <a href="https://github.com/HendrikPN/reinforced_scinet">Reinforced SciNet</a>, published partially in v2 of the paper <a href="https://arxiv.org/abs/2001.00593">Operationally meaningful representations of physical systems in neural networks</a>.</p> <p>&nbsp;</p> <p><strong>File description:</strong></p> <p>results.txt - The results from the training during<em> reinforcement learning</em>.</p> <p>results_loss.txt - The loss from the training during <em>representation learning</em>.</p> <p>selection.txt - The noise level of latent neurons during <em>representation learning</em>.</p> <p>&nbsp;</p> <p><strong>Parameters: Reinforcement Learning</strong></p> <p>Server parameters</p> <ul> <li>21 workers, 2 predictors, 1 trainer each</li> <li>3M episodes</li> </ul> <p>Training parameters</p> <ul> <li>glow: 0.1</li> <li>gamma: 0.01</li> <li>softmax: 0.5</li> <li>learning rate: 0.00005</li> <li>reward clipping: 1.0e-7</li> </ul> <p>Network parameters</p> <ul> <li>DPS model:<br> {&#39;env1&#39;: [128, 128, 128, 128, 64, 32],<br> &#39;env2&#39;: [128, 128, 128, 128, 64, 32],<br> &#39;env3&#39;: [128, 128, 128, 128, 64, 32]}</li> </ul> <p>&nbsp;</p> <p><strong>Parameters: Representation Learning</strong></p> <p>Server parameters</p> <ul> <li>21 workers, 2 predictors, 1 trainer each</li> <li>5M episodes</li> </ul> <p>Training parameters</p> <ul> <li>learning rate: 0.0001</li> <li>reward clipping: 1.0e-7</li> <li>selection discount: 0.04</li> <li>minimization discount: 0.02</li> <li>ae discount: 10.0</li> <li>agent discount: 1.</li> <li>reward rescaling: 10</li> <li>predicted actions: 1</li> <li>training data: 200K</li> </ul> <p>Network parameters</p> <ul> <li>Prediction model:<br> {&#39;env1&#39;: [64, 128, 128, 128, 128, 64, 32],<br> &#39;env2&#39;: [64, 128, 128, 128, 128, 64, 32],<br> &#39;env3&#39;: [64, 128, 128, 128, 128, 64, 32]}</li> <li>Encoder model: [128, 128, 64, 32]</li> <li>Decoder model: [32, 64, 128, 128, 128]</li> </ul>

opencc-by-4.0Jan 2021View details →
zenodo36/100

Training data for "From small to large-scale genome comparison", a tutorial for the Galaxy Training Network

<p>This dataset comprises two sequence pairs in FASTA format, one including two mycoplasmas (<em>Hyopneumoniae</em> 232 and 7422) and the other including the first chromosome of two plant genomes (<em>Aegilops tauschii</em> and <em>Triticum aestivum</em>).</p>

opencc-by-4.0Jan 2021View details →
zenodo36/100

Data and Codes for "Explainable Offline-Online Training of Neural Networks for Parameterizations: A 1D Gravity Wave-QBO Testbed in the Small-data Regime" by Pahlavan et al. (2023)

<p>This is part of the code and data related to the paper entitled Explainable Offline-Online Training of Neural Networks for Parameterizations: A 1D Gravity Wave-QBO Testbed in the Small-data Regime, available at https://arxiv.org/abs/2309.09024.</p><p>The original sources of the codes are the v1.0.0 version of open source software EnsembleKalmanProcesses.jl for EKI analysis, accessible at zenodo.org/records/7806813, and the \emph{qbo1d} code for the 1D-QBO model simulations, accessible at github.com/DataWaveProject/qbo1d.git.</p>

opencc-by-4.0Nov 2023View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record