Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
558
datasets available to search
ShareScore release 0.9.0
Dataset results
558 results for “Training Data”
Data for Project 'Test-Retest Reliability and Validity of vagally-mediated Heart Rate Variability to Monitor Internal Training Load in Older Adults: A within-subjects (repeated-measures) randomized study'
<p>Data for Project 'Test-Retest Reliability and Validity of vagally-mediated Heart Rate Variability to Monitor Internal Training Load in Older Adults: A within-subjects (repeated-measures) randomized study' consisting of (1) the original and complete dataset ('Data_Brain-IT-Reliability-of-HRV-during-Exergaming_for-publication'; and (2) a corresponding README file including (a) general information, (b) data and file overview, (c) sharing and access information, (d) methodological information, and (e) data-specific information.</p>
PhageHostLearn - training data and cluster analysis
<p>These data comprise the processed phage RBP and <em>Klebsiella </em>K-loci sequence data to train our PhageHostLearn system, along with ESM-2 embeddings of the RBPs and loci, as well as results from the cluster analyses of K-loci proteins and RBPs at 50% identity with CD-HIT.</p>
GOBAI-O2 Training Data
<p><strong>About</strong></p> <p>This repository contains quality controlled data obtained from the Biogeochemical Argo database (Argo, 2000) and GLODAP data product (Lauvset et al., 2022) that are used to train machine leaning models to be applied to three-dimensional temperature and salinity maps compiled using the Core Argo data (<a href="https://sio-argo.ucsd.edu/RG_Climatology.html">Roemmich and Gilson, 2009</a>) to produce GOBAI-O2 (Sharp et al., 2022).</p> <p><strong>References</strong></p> <p>Argo (2000). Argo float data and metadata from Global Data Assembly Centre (Argo GDAC). SEANOE. <a href="https://doi.org/10.17882/42182">https://doi.org/10.17882/42182</a>.</p> <p>Lauvset, S. K., Lange, N., Tanhua, T., Bittig, H. C., Olsen, A., Kozyr, A., Alin, S., Álvarez, M., Azetsu-Scott, K., Barbero, L., Becker, S., Brown, P. J., Carter, B. R., da Cunha, L. C., Feely, R. A., Hoppema, M., Humphreys, M. P., Ishii, M., Jeansson, E., Jiang, L.‑Q., Jones, S. D., Lo Monaco, C., Murata, A., Müller, J. D., Pérez, F. F., Pfeil, B., Schirnick, C., Steinfeldt, R., Suzuki, T., Tilbrook, B., Ulfsbo, A., Velo, A., Woosley, R. J. and Key, R. M. (2022). GLODAPv2.2022: the latest version of the global interior ocean biogeochemical data product. Earth System Science Data, 14(12), 5543‑5572. <a href="https://doi.org/10.5194/essd-14-5543-2022">https://doi.org/10.5194/essd‑14‑5543‑2022</a>.</p> <p>Roemmich, D. and Gilson, J. (2009). The 2004-2008 mean and annual cycle of temperature, salinity, and steric height in the global ocean from the Argo Program. Progress in Oceanography, 82, 81-100. <a href="https://doi.org/10.1016/j.pocean.2009.03.004">https://doi.org/10.1016/j.pocean.2009.03.004</a>.</p> <p>Sharp, J. D. Fassbender, A. J. Carter, B. R., Johnson, G. C., Schultz, C., and Dunne, J. P. (2022). GOBAI-O<sub>2</sub>: A Global Gridded Monthly Dataset of Ocean Interior Dissolved Oxygen Concentrations Based on Shipboard and Autonomous Observations (NCEI Accession 0259304). v1.0. NOAA National Centers for Environmental Information. Dataset. <a href="https://doi.org/10.25921/z72m-yz67">https://doi.org/10.25921/z72m-yz67</a>.</p>
Training Data for "Creating Quality FAIR assessment reports and draft of Data Papers from EML metadata with MetaShRIMPS"
<p>Training Data for "Training Data for "Creating Quality FAIR assessment reports and draft of Data Papers from EML metadata with MetaShRIMPS""</p>
Reproduction package for the paper "The Impact of Hard and Easy Negative Training Data on Vulnerability Prediction Performance"
<p>This Reproduction package contains the datasets, code and results for the paper "The Impact of Hard and Easy Negative Training Data on Vulnerability Prediction Performance" for other researchers to use for reproducing or improving our work. </p>
Annotated spiral ganglion neuron training data for object detection
Open the record for dataset details and reuse information.
Data from: How many specimens make a sufficient training set for automated three dimensional feature extraction?
Open the record for dataset details and reuse information.
Data from: Postdoctoral T32 training is correlated with obtaining an academic primarily research faculty position
Open the record for dataset details and reuse information.
Data and trained models for: Human-robot facial co-expression
Open the record for dataset details and reuse information.
Training data from: Machine learning predicts which rivers, streams, and wetlands the Clean Water Act regulates
Open the record for dataset details and reuse information.
Data from: Convolutional neural networks trained on internal variability predict forced response of TOA radiation by learning the pattern effect
Open the record for dataset details and reuse information.
An individual-based model trained on multiple data sources estimates population connectivity and facilitates aggregation of harvest management units
Open the record for dataset details and reuse information.
Training data for 'Unicycler assembly of SARS-CoV-2 genome with preprocessing to remove human genome reads' tutorial (Galaxy Training Material)
<p>The data here is a copy of the corresponding SRR records in the NCBI SRA. The duplication serves a dual purpose:</p> <ol> <li>as a backup should there be problems connecting to NCBI servers, e.g., during Galaxy user trainings.</li> <li>to illustrate how to obtain raw sequencing data from alternative sources, and to organize the data into the same collection structure in a Galaxy history that is generated by specialized Galaxy SRA download tools.</li> </ol>
360-degree video recording of an outdoor camerawork training session for qualitative data collection
<p>In this equirectangular 360° video clip, a recording of an outdoor camerawork training session is stitched together from the footage taken by a stereoscopic 360° camera with eight lenses. To play the video and spatial sound correctly, use a digital video player that can play equirectangular videos with the YouTube ambiX First Order Ambisonic audio format, eg. VLC or PotPlayer. Please wear headphones.</p> <p>In the camerawork training session, all the participants are playing particular roles in the training session, and each carries a camera. In preparation for the real data collection with a guide, one person is pretending to be a nature guide. She carries a GoPro camera on a gimbal. There is an instructor, who is carrying a single lens 360° camera on a raised extension pole with a separate ambisonic microphone. Two others are filming with a prosumer camcorder and a single lens 360° camera on a lowered extension pole respectively. And a fifth person is filming with a stereoscopic 360° camera and an independent ambisonic microphone on a monopod. In a nutshell, this is a typical team filming arrangement, in which the team needs to attentively yet silently coordinate their joint camerawork. Languages: Danish and English</p>
Data and code for training and evaluating machine learning models for thunderstorm prediction from reanalysis data
<p>FIXED Data and Python code for training and evaluating machine learning models for predicting thunderstorms, associated with the paper:</p> <p>"Evaluation of machine learning classifiers for predicting deep convection"</p> <p>by Peter Ukkonen and Antti Mäkelä (to appear in JAMES)</p> <p>The data (preprocessed inputs and outputs) is stored as netCDF files and .mat files which can be loaded with Python. </p>
Generating Physically Sound Training Data for Image Recognition of Additively Manufactured Parts Data and Scripts
<p>The repository contains the data corresponding to the Paper "Generating Physically Sound Training Data for Image Recognition of Additively Manufactured Parts".</p> <p>Random30, Random50, Random100, Similiar10, Similar30 and Similar50.zip contain the data sets (obj Files).</p> <p>R30_physical_images.zip and sim50_physical_images.zip contain the photos made from the physical components which are used for the evaluation.</p>
Training and Test data for BaCoN
<p>Training and test datasets consisting of matter power spectra for use with the code BAyesian COsmological Network (BaCoN): https://github.com/Mik3M4n/BaCoN</p> <p>The dataset was generated using the code ReACT: https://github.com/nebblu/ReACT</p> <p>See the github repo for additional details.</p> <p> </p> <p> </p>
Training data from Reinforced SciNet
<p><strong>Summary:</strong></p> <p>The results from the training of neural networks in v2 of <a href="https://github.com/HendrikPN/reinforced_scinet">Reinforced SciNet</a>, published partially in v2 of the paper <a href="https://arxiv.org/abs/2001.00593">Operationally meaningful representations of physical systems in neural networks</a>.</p> <p> </p> <p><strong>File description:</strong></p> <p>results.txt - The results from the training during<em> reinforcement learning</em>.</p> <p>results_loss.txt - The loss from the training during <em>representation learning</em>.</p> <p>selection.txt - The noise level of latent neurons during <em>representation learning</em>.</p> <p> </p> <p><strong>Parameters: Reinforcement Learning</strong></p> <p>Server parameters</p> <ul> <li>21 workers, 2 predictors, 1 trainer each</li> <li>3M episodes</li> </ul> <p>Training parameters</p> <ul> <li>glow: 0.1</li> <li>gamma: 0.01</li> <li>softmax: 0.5</li> <li>learning rate: 0.00005</li> <li>reward clipping: 1.0e-7</li> </ul> <p>Network parameters</p> <ul> <li>DPS model:<br> {'env1': [128, 128, 128, 128, 64, 32],<br> 'env2': [128, 128, 128, 128, 64, 32],<br> 'env3': [128, 128, 128, 128, 64, 32]}</li> </ul> <p> </p> <p><strong>Parameters: Representation Learning</strong></p> <p>Server parameters</p> <ul> <li>21 workers, 2 predictors, 1 trainer each</li> <li>5M episodes</li> </ul> <p>Training parameters</p> <ul> <li>learning rate: 0.0001</li> <li>reward clipping: 1.0e-7</li> <li>selection discount: 0.04</li> <li>minimization discount: 0.02</li> <li>ae discount: 10.0</li> <li>agent discount: 1.</li> <li>reward rescaling: 10</li> <li>predicted actions: 1</li> <li>training data: 200K</li> </ul> <p>Network parameters</p> <ul> <li>Prediction model:<br> {'env1': [64, 128, 128, 128, 128, 64, 32],<br> 'env2': [64, 128, 128, 128, 128, 64, 32],<br> 'env3': [64, 128, 128, 128, 128, 64, 32]}</li> <li>Encoder model: [128, 128, 64, 32]</li> <li>Decoder model: [32, 64, 128, 128, 128]</li> </ul>
Training data for "From small to large-scale genome comparison", a tutorial for the Galaxy Training Network
<p>This dataset comprises two sequence pairs in FASTA format, one including two mycoplasmas (<em>Hyopneumoniae</em> 232 and 7422) and the other including the first chromosome of two plant genomes (<em>Aegilops tauschii</em> and <em>Triticum aestivum</em>).</p>
Data and Codes for "Explainable Offline-Online Training of Neural Networks for Parameterizations: A 1D Gravity Wave-QBO Testbed in the Small-data Regime" by Pahlavan et al. (2023)
<p>This is part of the code and data related to the paper entitled Explainable Offline-Online Training of Neural Networks for Parameterizations: A 1D Gravity Wave-QBO Testbed in the Small-data Regime, available at https://arxiv.org/abs/2309.09024.</p><p>The original sources of the codes are the v1.0.0 version of open source software EnsembleKalmanProcesses.jl for EKI analysis, accessible at zenodo.org/records/7806813, and the \emph{qbo1d} code for the 1D-QBO model simulations, accessible at github.com/DataWaveProject/qbo1d.git.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.