Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
558
datasets available to search
ShareScore release 0.9.0
Dataset results
558 results for “Training Data”
Data from: Increased dopamine release after working-memory updating training: neurochemical correlates of transfer
Open the record for dataset details and reuse information.
Data from: Education research: simulation training for neurology residents on acquiring tPA consent: an educational initiative
Open the record for dataset details and reuse information.
Autoimmune inflammation causes hematopoietic stem cells to generate a trained immunity program inherited by BMDMs [RNAseq_primary_data]
GEO Series GSE267563. Mus musculus. 6 samples. Type: Expression profiling by high throughput sequencing.
E-Predict Training Data Set and Examples
GEO Series GSE2228. Viruses. 56 samples. Type: Expression profiling by array.
Expression data from exercised-trained, myog-deleted adult mice
GEO Series GSE22046. Mus musculus. 6 samples. Type: Expression profiling by array.
Transgenic Eµ-myc mouse lymphoma expression data for training dataset
GEO Series GSE40756. Mus musculus. 39 samples. Type: Expression profiling by array.
Microarray data of osteoblast from sham, ovariectomized mice (OVX) with and without treadmill training
GEO Series GSE111628. Mus musculus. 3 samples. Type: Expression profiling by array.
Expression data from human skeletal muscle in endurance-trained athletes
GEO Series GSE155271. Homo sapiens. 13 samples. Type: Expression profiling by array.
Autoimmune inflammation causes hematopoietic stem cells to generate a trained immunity program inherited by BMDMs [ATACseq_primary_data]
GEO Series GSE267516. Mus musculus. 4 samples. Type: Genome binding/occupancy profiling by high throughput sequencing.
Acetaminophen (APAP) Rat Blood Training Gene Expression Data Set
GEO Series GSE5593. Rattus norvegicus. 68 samples. Type: Expression profiling by array.
Autoimmune inflammation causes hematopoietic stem cells to generate a trained immunity program inherited by BMDMs [ATACseq_secondary_data]
GEO Series GSE267518. Mus musculus. 4 samples. Type: Genome binding/occupancy profiling by high throughput sequencing.
RNA-seq counts to genes (training data)
<p>Dataset for Galaxy counts to genes tutorial</p>
WMT-SLT SRF: Training data for the WMT shared task on sign language translation (videos, subtitles)
<p>These are Standard German daily news (Tagesschau) and Swiss German weather forecast (Meteo) episodes broadcast and interpreted into Swiss German Sign Language by hearing interpreters (among them, children of Deaf adults, CODA) via Swiss National TV (Schweizerisches Radio und Fernsehen, SRF) (<a href="https://www.srf.ch/play/tv/sendung/tagesschau-in-gebaerdensprache?id=c40bed81-b150-0001-2b5a-1e90e100c1c0">https://www.srf.ch/play/tv/sendung/tagesschau-in-gebaerdensprache?id=c40bed81-b150-0001-2b5a-1e90e100c1c0</a>). For a more extended description of the data, visit <a href="https://www.wmt-slt.com/data">https://www.wmt-slt.com/data</a>.<br> <br> </p>
Camera trap data accompanying paper titled: "Deep learning-based ecological analysis of camera trap images is impacted by training data quality and size"
<p>ZIp file that contains two camera trap datasets that support the experiments of the paper title: "Deep learning-based ecological analysis of camera trap images is impacted by training data quality and size". The paper is currently under submission.</p>
Data Donation Model for Inclusive Cardiovascular Prevention Using the TRAIN Health Platform
ClinicalTrials.gov study NCT07238036. IPD Sharing: Not stated. Countries: 0. Publications: 0.
Expression data from baseline and post-endurance training in human PBMCs
GEO Series GSE57999. Homo sapiens. 26 samples. Type: Expression profiling by array.
Training data for the shared task Ideology and Power Identification in Parliamentary Debates - with errors please do not use!
<p><strong>This dataset is an early release with errors. Please do not use this data set. The official dataset for the shared task will be announced on <a href="https://touche.webis.de/clef24/touche24-web/ideology-and-power-identification-in-parliamentary-debates.html">the shared task webpage</a>.</strong></p>
PIMD data for training effective potential incorporating nuclear quantum effects
<p>The dataset contains training data to generate machine-learned effective potentials reproducing correct nuclear quantum statistics. The generation procedure is described in I. Zaporozhets, F. Musil, V. Kapil, & C. Clementi (2024). Accurate nuclear quantum statistics on machine-learned classical effective potentials. <a href="https://arxiv.org/abs/2407.03448">[arXiv: 2407.03448]</a></p> <p>The code required to generate the dataset can be found in the repository: <a href="https://github.com/ClementiGroup/accurate_nuclear_quantum_statistics_on_machine_learned_classical_effective_potentials">cg_nuclear_quantum_statistics</a>.</p> <h2>Dataset structure</h2> <p>The archive contains datasets for four systems: a particle in 1D Morse potential, a single water molecule in a vacuum, a Zundel cation, and a box of 256 water molecules. For the particle in Morse potential, separate .npy files (loadable with numpy in Python) for harmonic coupling (spring) forces and coordinates are given for temperatures 100, 300, and 600 K. For other systems, the data are provided in HDF5 format (see below).</p> <p>Directory structure:</p> <p>CG_quantum_statistics<br>├── 0_morse_potential<br>│ ├── temp_100_spring_forces.npy<br>│ ├── temp_100_total_coordinates.npy<br>│ ├── temp_300_spring_forces.npy<br>│ ├── temp_300_total_coordinates.npy<br>│ ├── temp_600_spring_forces.npy<br>│ └── temp_600_total_coordinates.npy<br>├── 1_h2o_molecule<br>│ └── h2o.h5<br>├── 2_zundel_cation<br>│ └── h5o2+.h5<br>└── 3_bulk_h2o<br> └── h2o_256.h5</p>
Training Eval data for TropiGAT and TropiSEQ
<p><span>Compilation of the training data, comprising the depolymerase sequences, the encoding prophage, the KL type of the infected ancestor and related information.</span></p>
LAPSO PM2.5 in Scotland (training data and the model)
<p>Estimating Near-Surface Concentrations of Major Air Pollutants From Space: A Universal Estimation Framework LAPSO</p> <p>Like many other countries, China is still facing severe air pollution issues after extensive efforts. The difficulties in deriving near-surface concentrations from satellite measurements restrict the application of remote sensing of large-scale surface air quality. Aiming at providing daily accurate near-surface ail pollution estimates (PM2.5, PM10, O3, NO2, SO2, and CO), we propose a robust estimation framework called learning air pollutants from satellite observations (LAPSO). The principle of LAPSO is to derive a nonlinear relationship between surface pollutant concentrations of interest and satellite observations with the aid of meteorological reanalyzes based on deep learning techniques. The LAPSO framework is superior to other algorithms due to its robust retrieval performance, independence from chemical transport models (CTMs), lower hardware requirements, and a user-friendly interface. The retrieval results of LAPSO were in good agreement with ground-level measurements according to extensive cross-validation at 1628 sites ( R2> 0.8 in polluted areas and uncertainty ≪5 μg/m3 for most pollutants) in China. The framework also showed a strong capability to capture the temporal variability of different air pollutants. By comparing with the estimation results from different satellite platforms, TROPOspheric monitoring instrument (TROPOMI) onboard the Sentinel-5P demonstrated marginally better performance for estimating PM2.5. Although the selection of satellite observations did not significantly affect the results of O3 estimation, the number and spatial sampling density of in situ sites imposed large impacts on O3 estimation performance. The success of LAPSO for estimating near-surface concentrations from satellite remote sensing at an enhanced spatiotemporal resolution is expected to serve the continuous and dynamical monitoring of regional and global air pollution.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.