Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
558
datasets available to search
ShareScore release 0.9.0
Dataset results
558 results for “Training Data”
Malwa survey : GIS data SRTM [trains]
<p>Malwa survey : GIS data SRTM [trains] (2009).</p>
PDBscreen with multiple data augmentation strategies suitable for training protein-ligand interaction prediction methods
<p>PDBscreen with multiple data augmentation strategies suitable for training protein-ligand interaction prediction methods.</p> <p>PDBscreen is the training dataset for EquiScore.</p>
Landsat time series classification training data
<p>Data for the paper </p> <p>Hankui K. Zhang, Dong Luo, Zhongbin Li, Classifying raw irregular Landsat time series (CRIT) for large area land cover mapping by adapting Transformer model.</p> <p>It stores the daily raw Landsat ARD annual good quality surface reflectance time series for 1985, 2006 and 2018 for CONUS with 7 land cover classes. Details are in the paper. </p>
Raw motif mapping bedfile data and model training set class probabilities
<p>Leveraging prior viral genome sequencing data to make predictions on whether an unknown, emergent virus harbors a 'phenotype-of-concern' has been a long-sought goal of genomic epidemiology. A predictive phenotype model built from nucleotide-level information alone is challenging with respect to RNA viruses due to the ultra-high intra-sequence variance of their genomes, even within closely related clades. We developed a degenerate k-mer method to accommodate this high intra-sequence variation of RNA virus genomes for modeling frameworks. By leveraging a taxonomy-guided 'group-shuffle-split' cross validation paradigm on complete coronavirus assemblies from prior to October 2018, we trained multiple regularized logistic regression classifiers at the nucleotide k-mer level. We demonstrate the feasibility of this method by finding models accurately predicting withheld SARS-CoV-2 genome sequences as human pathogens and accurately predicting withheld Swine Acute Diarrhea Syndrome coronavirus (SADS-CoV) genome sequences as non-human pathogens. Feature selection using L1 regularization identified several degenerate nucleotide predictor motifs with high model coefficients for the human pathogen class that were present across widely disparate clades of coronaviruses. However, these motifs differed in which genes they were present in, what specific codons were used to encode them, and what the translated amino acid motif was. This emphasizes the importance of a phenetic view of emerging pathogenic RNA viruses, as opposed to the canonical phylogenetic interpretations most commonly used to track and manage viral zoonoses. Applying our model to more recent Orthocoronavirinae genomes deposited since October 2018 yields a novel contextual view of pathogen potential across bat-related, canine-related, porcine-related, and rodent-related coronaviruses and critical adaptations which may have contributed to the emergence of the pandemic SARS-CoV-2 virus. Finally, we discuss the next steps to achieve robust predictive ensembles and the utility of these models (and their associated predictor motifs) to novel biosurveillance protocols that substantially increase the 'pound-for-pound' information content of field-collected sequencing data and make a strong argument for the necessity of routine collection and sequencing of zoonotic viruses. </p>
Dataset from "Synthetic Training Data for Semantic Segmentation of the Environment from UAV Perspective"
<p>This dataset contains the images and ground truth label masks for semantic segmentation created and described in "Hinniger, C.; Rüter, J. Synthetic Training Data for Semantic Segmentation of the Environment from UAV Perspective. Aerospace 2023, 10, 604. https://doi.org/10.3390/aerospace10070604".</p>
animal soup sample data, ground truth dataset, and pre-trained models
<p>- sample data and ground truth files for animal soup tests</p> <p>- pre-trained models</p>
Data for "No Train No Gain: Revisiting Efficient Training Algorithms For Transformer-based Language Models"
<p>Datasets to reproduce the experiments associated with the paper: https://doi.org/10.48550/arXiv.2307.06440</p> <p>The readme contains instructions for how to use them: https://github.com/JeanKaddour/NoTrainNoGain/blob/main/bert/README.md</p> <p>c4-subset-random.tar.bz2 is a subset of the C4 dataset (https://arxiv.org/abs/1910.10683), licensed under ODC-BY 1.0.</p>
Free throw psychological procedure training data
<p>罚球心理程序训练数据</p>
Vascular Positioning System G4 Algorithm ECG Data Collection for Model Training Study
ClinicalTrials.gov study NCT05702515. IPD Sharing: NO. Countries: 1. Publications: 0.
Data from: Adherence to trained standards after a faculty development workshop on “Teaching With Simulated Patients”
Open the record for dataset details and reuse information.
Data from: Using matrix and tensor factorizations for the single-trial analysis of population spike trains
Open the record for dataset details and reuse information.
Data from: Machine learning biogeographic processes from biotic patterns: a new trait-dependent dispersal and diversification model with model choice by simulation-trained discriminant analysis
Open the record for dataset details and reuse information.
Data from: Generalized polyspike train: an EEG biomarker of drug-resistant idiopathic generalized epilepsy
Open the record for dataset details and reuse information.
Data from: An integrated iterative annotation technique for easing neural network training in medical image analysis
Open the record for dataset details and reuse information.
Data from: Genomic selection and association mapping in rice (Oryza sativa): effect of trait genetic architecture, training population composition, marker number and statistical model on accuracy of rice genomic selection in elite, tropical rice breeding lines
Open the record for dataset details and reuse information.
Data from: ASSET: analysis of sequences of synchronous events in massively parallel spike trains
Open the record for dataset details and reuse information.
Data from: Comparison of a newly established emotional stimulus approach to a classical assessment-driven approach in BLS training: a randomised controlled trial
Open the record for dataset details and reuse information.
Data from: Handover training for medical students – a controlled educational trial of a pilot curriculum
Open the record for dataset details and reuse information.
Data from: Predicting classifier performance with limited training data: applications to computer-aided diagnosis in breast and prostate cancer
Open the record for dataset details and reuse information.
Data from: Data reliability in citizen science: learning curve and the effects of training method, volunteer background and experience on identification accuracy of insects visiting ivy flowers
Open the record for dataset details and reuse information.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.