Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
558
datasets available to search
ShareScore release 0.9.0
Dataset results
558 results for “Training Data”
Training data of protein classifier SolubEcoli.pgc and GDP1.pgc
<p>The classifiers Solub.Ecoli.pgc and GDP1.pgc have been built using Pro-Gyan (https://code.google.com/p/pro-gyan/).</p> <p>The file AGGREGATING.fasta contains chaperone-dependent 502 aggregation prone proteins from <em>E.coli</em>. The file SOLUBLE.fasta contains chaperone-independent 475 soluble proteins from <em>E.coli</em>. The file GROEL_C3.fasta contains GroEL-dependent 83 "class 3" proteins from <em>E.coli</em>. The file SOLUBLE.fasta contains chaperone-independent 475 soluble proteins from <em>E.coli</em>.</p>
Training for RD Management: Comparative European Approaches, Raw data & Tidied Data
<p>Raw and tidied data of responses to the Knowledge Exchange survey around Training for Research Data management. The data accompanies the Knowledge Exchange report 'Training for RD Management: Comparative European Approaches' DOI: 10.5281/zenodo.50068</p> <p> </p>
SemEval 2017 Task 3 Subtask E Train data (StackExchange)
<p>This is the train data that was used for Task 3, Subtask E of the SemEval-2017 Shared Task. More information on the task can be found here: http://alt.qcri.org/semeval2017/task3/</p>
Data and trained word2vec model for ``Easy over Hard: A Case Study on Deep Learning''
<p>The data include: training and testing data pairs</p> <p>The word2vec model is pre-trained. </p> <p>More details, please refer to the paper</p>
RDP taxonomic training data formatted for DADA2 (RDP trainset 16/release 11.5)
<p>These DADA2-formatted training fasta files were derived from the Ribosomal Database Project's Training Set 16 and the 11.5 release of the RDP database.</p> <p>These fastas were generated by the following commands (using the dada2 R package version 1.5.1):</p> <blockquote> <p>path <- "~/Desktop/RDP/RDPClassifier_16S_trainsetNo16_rawtrainingdata"<br> dada2:::makeTaxonomyFasta_RDP(file.path(path, "trainset16_022016.fa"), file.path(path, "trainset16_db_taxid.txt"), "~/tax/rdp_train_set_16.fa.gz")</p> <p>dada2:::makeSpeciesFasta_RDP("~/Desktop/RDP/current_Bacteria_unaligned.fa", "~/tax/rdp_species_assignment_16.fa.gz")</p> </blockquote>
Galaxy Training Date for "Analyses of metagenomic data - The global picture"
<p>These training datasets are part of a Galaxy Training Network tutorial that analyzes metagenomic (amplicon and WGS) data. These datasets are extracted of a project studying the Argentinean agricultural pampean soils (https://www.ebi.ac.uk/metagenomics/projects/SRP016633). </p>
RDP LSU taxonomic training data formatted for DADA2 (trainingset 11)
<p>#Format RDP taxonomic training set for DADA2<br> #1 Wrangle the RDP trainingsets and unaligned data into the downloads folder by executing this from a terminal and move the file somewhere with >30GB free<br> wget https://sourceforge.net/projects/rdp-classifier/files/RDP_Classifier_TrainingData/RDPClassifier_fungiLSU_trainsetNo11_rawtrainingdata.zip/download<br> wget http://rdp.cme.msu.edu/download/current_Fungi_unaligned.fa.gz<br> #2 Unzip the trainingset file and replace Us with Ts in the fasta by executing in terminal<br> awk 'NR%2==0 {gsub(/[uU]/,"T"); print} NR%2==1' /media/lauren/96BA-19E6/RDPClassifier_fungiLSU_trainsetNo11_rawtrainingdata/fungiLSU_train_012014.fa > /media/lauren/96BA-19E6/RDPClassifier_fungiLSU_trainsetNo11_rawtrainingdata/fungiLSU_train_012014_lsu_fixed_v2.fa<br> #3 Summon the dada2 pkg<br> library(dada2);packageVersion("dada2")<br> #4 Transform the DADA2 formatted training fastas<br> path<-"/media/lauren/96BA-19E6/RDPClassifier_fungiLSU_trainsetNo11_rawtrainingdata"<br> dada2:::makeTaxonomyFasta_RDP(file.path(path, "fungiLSU_train_012014_lsu_fixed_v2.fa"), file.path(path, "fungiLSU_taxid_012014.txt"),"/media/lauren/96BA-19E6/Upload/RDP_LSU_fixed_train_set_v2.fa",compress=FALSE)<br> dada2:::makeSpeciesFasta_RDP("/media/lauren/96BA-19E6/RDPClassifierLSU/current_Fungi_unaligned.fa", "/media/lauren/96BA-19E6/Upload/rdp_species_assignment_LSU_v2.fa", compress=FALSE)</p> <p>#5 Make the compressed DADA2 formatted training fastas in gz and zip format<br> dada2:::makeTaxonomyFasta_RDP("/media/lauren/96BA-19E6/RDPClassifier_fungiLSU_trainsetNo11_rawtrainingdata/fungiLSU_train_012014_lsu_fixed_v2.fa", file.path(path, "fungiLSU_taxid_012014.txt"),"/media/lauren/96BA-19E6/Upload/RDP_LSU_fixed_train_set_v2.fa.gz",compress=TRUE)<br> dada2:::makeTaxonomyFasta_RDP("/media/lauren/96BA-19E6/RDPClassifier_fungiLSU_trainsetNo11_rawtrainingdata/fungiLSU_train_012014_lsu_fixed_v2.fa", file.path(path, "fungiLSU_taxid_012014.txt"),"/media/lauren/96BA-19E6/Upload/RDP_LSU_fixed_train_set_v2.fa.zip",compress=TRUE)<br> dada2:::makeSpeciesFasta_RDP("/media/lauren/96BA-19E6/RDPClassifierLSU/current_Fungi_unaligned.fa", "/media/lauren/96BA-19E6/Upload/rdp_species_assignment_LSU_v2.fa.gz", compress=TRUE)<br> dada2:::makeSpeciesFasta_RDP("/media/lauren/96BA-19E6/RDPClassifierLSU/current_Fungi_unaligned.fa", "/media/lauren/96BA-19E6/Upload/rdp_species_assignment_LSU_v2.fa.zip", compress=TRUE)</p> <p> </p>
Global snow water equivalent product derived from machine learning model trained with in situ measurement data
<p>This dataset is a global snow water equivalent dataset using machine learning trained with in-situ measurements. The temporal resolution of the SWEML product is daily, and the spatial resolution is 0.25˚ (approximately 25km). It covers latitudes of 90S to 90N and longitudes of 180W to 180E with global scales, excluding Antarctica. The dataset is provided in NetCDF format, organized by year. Each year contains daily SWE data, including leap days in leap years.</p>
Training data for the EPInformer model
<p>Training data for the EPInformer framework deep learning model. The data contains enhancer-gene links of protein coding genes nominated by ABC method (https://github.com/broadinstitute/ABC-Enhancer-Gene-Prediction/tree/master)</p>
Data from: Effectiveness of Online Off-the-Job Training in Attracting Participants and Video-On-Demand Streaming in Improving Work-Life Balance: A Study Focusing on Medical Technologists
<p>The Nara Association of Medical Technologists has introduced online Off-Job Training (Off-JT) starting from FY2020 in response to the COVID-19 pandemic. This study aims to evaluate the online Off-JT, which differs from the traditional face-to-face format. Firstly, we compared the online format's ability to attract participants with the face-to-face format based on the number of training sessions and attendees. Despite having fewer training sessions (40.8% less), the online format had an average attendance of 105.4% higher (39.7 vs. 19.3) than the face-to-face format. To enhance participant convenience, we offered a limited number of live and video-on-demand (VOD) sessions on YouTube, evaluating their usefulness through an online survey focusing on work-life balance (WLB). The survey results showed that 81.9% (458/559) of respondents reported an improvement in WLB. The effect on WLB improvement varied depending on the viewing method, with VOD sessions showing 84.1% (376/447) and live sessions showing 73.2% (82/112). We believe that the increased ability to attract participants in the online Off-JT is mainly due to the elimination of travel burdens through internet-connected devices. The combination of live and VOD sessions on YouTube allowed participants to adjust their viewing time, leading to better allocation of free time and improved WLB. The online Off-JT and VOD delivery have shown to enhance convenience for participants by removing geographical and time constraints, resulting in positive effects.</p>
Data for training SuperdropNet
<p>Dataset for training and testing SuperdropNet. Includes superdroplet simulations in a warm rain scenario which were simulated using McSnow.</p>
Research data on pass-by sound emission measurements from train units for D7.4-Susteren pilot
<p>This dataset compiles sample measurement results for the N-RSD rail pilot from D7.4-Susteren pilot project report available once approved by the EC at https://cordis.europa.eu/project/id/860441/results. The data includes train unit subtypes, pass-by sound levels, speed, weight, wheel flat indication, brake type information and a sound emission rating. </p>
Automated bio-AFM generation of large mechanome data set and their analysis by machine learning to classify prostatic cell lines_Training base 100 PC3-GFP
Open the record for dataset details and reuse information.
Supplementary material for the publication: Deep Learning of Crystalline Defects from TEM images: A Solution for the Problem of "Never Enough Training Data"
Open the record for dataset details and reuse information.
Training and test data for antibody humanness evaluation
<p>### Training and test data for humanness evaluation</p> <p>This data was collected in conjunction with and used for<br>training and testing for Parkinson / Wang et al 2024. The<br>data is organized as follows:</p> <p>- Heavy chain training and multispecies test data (under the heavy chain folder)<br> - The conslidated cAb rep file contains training human sequences<br> - The test sample sequences folder contains fasta files with test sequences for each species<br>- Light chain training and multispecies test data (under the light chain folder)<br> - The conslidated cAb rep file contains training human sequences<br> - The test sample sequences folder contains fasta files with test sequences for each species<br>- Abybank data (under the abybank compiled data folder)<br> - This folder contains separate folders for heavy and light chain<br> - Each subfolder contains test data for a more diverse species set under fasta files for each species<br>- Humanization test data (under the humanization test data folder)<br> - The sequences in the parental.fa file were originally humanized as part of drug discovery programs<br> - The experimental.fa file contains the humanization results<br>- IMGT and ADA data (under the imgt test data folder)<br> - The imgt mab db fa and tsv files contain sequences and species assignments for IMGT mAb DB<br> - The thera ada fa file contains sequences evaluated in the clinic<br> - The Therapeutic ADA txt file contains anti drug antibody results for those antibodies<br>- VDJ statistics (under the vdj_statistics_eval folder)</p> <p>The data was retrieved from the following sources.</p> <p>1. All heavy and light chain training data is from the cAb-Rep database from [Guo et al.](https://pubmed.ncbi.nlm.nih.gov/31649674/)<br>2. All testing data is from the Observed Antibody Space [(OAS) database](https://opig.stats.ox.ac.uk/webapps/oas/)</p> <p>The training and test data show is after filtering for quality. The testing data was additionally randomly sampled to yield a set of 50,000 sequences for each species, then filtered to remove duplicates. The human test data was checked to ensure no overlap with the human training set.</p> <p><br>The IMGT, ADA and humanization test data was retrieved from Prihoda et al. and<br>the associated [Github repo](https://github.com/Merck/BioPhi-2021-publication).</p> <p>See Parkinson et al. 2024 and the associated github repos for more details on how models other than<br>SAM / AntPack were evaluated on this data.</p>
Data set of axle loads and distances for operating passenger and freight trains for development of load models
<p>Train data for moving load models of operating trains (5,007 passenger trains and 139,182 freight trains).</p> <ul> <li>axle distances [m] and axle loads [kN] for all trains included in zip-files as txt-files</li> <li>information on line categories (DIN EN 15528), loadcases and maximum speeds for passenger trains (and assignment to passenger train numbers in previous version of publication (used for development of load model)) in xlsx-file</li> </ul>
Data: Responsiveness and habituation to repeated sound exposures and pulse trains in blue mussels
<p>Data abstract:</p> <p>Time series data on the valve gape behaviour of blue mussels (<em>Mytilus edulis</em>) that were exposured to sound treatments. Here, we provide the valve gape (time series data expressed in proportion open) of all mussels over the course of their trial and the timing of the sound exposures.</p> <p> </p> <p>Paper abstract:</p> <p>Anthropogenic sound has been shown to affect marine animals across taxa. However, bivalves and other invertebrates received limited attention and most studies across taxa focussed on immediate, rather than long-term, effects of sound. Most bivalves adopt a sessile or sedentary lifestyle and are therefore expected to be exposed to the same sounds for long periods or repeatedly. For this reason, bivalves are an especially relevant taxonomic group to study long-term effects of sound. In the current study, we examined whether blue mussels (<em>Mytilus edulis</em>) habituate to repeated sound exposures and whether they recover quicker from a single pulse exposure than from a pulse train. We equipped individual mussels with sensors to monitor valve gape and exposed them to repeated sound playback. We found that mussels responded to sound by partially closing their valves. This response was consistent and repeatable, but decayed over sequential exposures to the same sound stimulus, and was stronger again with exposure to a different sound. This pattern is clear evidence for acoustic habituation in a bivalve. Additionally, we found no differences in the initial response and recovery (time to return to baseline levels) between mussels that were exposed to single pulses and pulse trains. Our results therefore show that mussels are able to habituate to sound and suggest that mussels mostly respond to the onset of a pulse train. Future research is needed to determine whether mussels also habituate in situ to actual anthropogenic sound and whether a lack of a behavioural response also implies that other negative effects are also absent.</p> <p> </p> <p>Paper reference:</p> <p>Hubert, J., Booms, E., Witbaard, R., Slabbekoorn, H. (2022). Responsiveness and habituation to repeated sound exposures and pulse trains in blue mussels. <em>Journal of Experimental Marine Biology and Ecology</em>. 547, 151668. DOI: 10.1016/j.jembe.2021.151668</p>
rime fraction training data set extracted from BAECC
<p>training data set used in Vogl et al. (https://amt.copernicus.org/preprints/amt-2021-137/) to derive rime mass fraction from Doppler cloud radar observations. Extracted from the BAECC data set.</p> <p>rime mass fraction retrieved from PIP data</p> <p>cloud radar observations at Ka- and W-band (ARM KAZR and MWACR)</p> <p>attenuation estimated using the Passive and Active Microwave Remote Sensing Tool (PAMTRA)</p>
Training Data Label Distributions
<p>A comparison of the label distribution of the full training dataset of OGB-PPA and the 80% and 80-99% of the dataset sorted by graph size. </p>
Experimental Data for the paper Training AI to Recognize Realizable Gauss diagrams: the Same Instances Confound AI and Human Mathematicians
<p>This upload contains supplementary materials for the paper <br> Training AI to Recognize Realizable Gauss diagrams: the Same Instances Confound AI and Human Mathematicians, by Abdullah Khan, Alexei Lisitsa and Alexei Vernitski, to appear in Proceedings of ICAART 2022 (14th International Conference on Agents and Artificial Intelligence, 3-5 February, 2022), SCITEPRESS. </p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.