Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
558
datasets available to search
ShareScore release 0.9.0
Dataset results
558 results for “Training Data”
Training data for 'Mapping-by-sequencing' tutorial (Galaxy Training Material)
<p>The data provided here are part of a Galaxy Training Network tutorial that demonstrates mapping-by-sequencing analysis and represent a subsample of the data used in Sun & Schneeberger, 2015 (DOI:10.1007/978-1-4939-2444-8_19).</p>
Co-curricular Quantitative Data Analytics Training Program (CQDATP) for Undergraduate Natural Science Students
<p>While data analytics has increasingly become an essential skill for various professionals in natural science fields, college curricula often cannot be updated frequently to meet the emerging changes of the workforce’s needs. A drastic curricular change is particularly difficult at small colleges with limited availability of faculty. This study proposes a co-curricular training program that enables natural science students to master skills of applying quantitative data analytics on real datasets without requiring a significant curricular change. The training program was developed and implemented as a sequence of in-person one-hour training sessions at a small four-year college. The results show that the approach is effective in improving students’ data analytics skills. This program can be adapted by other small institutions to achieve similar goals.</p> <p> </p> <p>This dataset include the training materials and deidentified student performance and survey responses. </p>
Baltic Sea Region Land Cover Plus - Training and Validation data
<p>Training and validation data used in creating Baltic Sea Region Land Cover Plus (BSRLC+) maps: <a href="https://doi.org/10.5281/zenodo.10653871" target="_blank" rel="noopener">Dataset link</a></p> <ul> <li><strong>landcover_training_data_2006_2018.gpkg</strong>: Points data of consistent land cover from 2006 to 2018</li> <li><strong>crop_training_data_{year}.gpkg</strong>: Points data of crop types derived from <a href="https://doi.org/10.1038/s41597-023-02517-0">EuroCrop dataset </a>in particular year (2019, 2021, 2023)</li> <li><strong>landcover_validation_{year}.gpkg</strong>: Points data of validation data derived from <a href="https://doi.org/10.1038/s41597-020-00675-z">LUCAS points </a>in particular year (2009, 2012, 2015, 2018)</li> <li><strong>Metadata.pdf</strong>: Information of land cover code in each dataset</li> </ul> <p>Version notes:</p> <p>Version 2: Correcting the validation data 2018 and Metadata file</p> <p>Version 1: Original upload</p>
Replication data for: "Effectiveness of iso-inertial resistance training on eccentric and concentric power, physical performance, and risk of falls in physically active middle-older adults: a randomised controlled trial"
<p>Replication data for: "Effectiveness of iso-inertial resistance training on eccentric and concentric power, physical performance, and risk of falls in physically active middle-older adults: a randomised controlled trial"</p> <p>This folder contains 4 files:</p> <p>1) Database that contains the values for concentric and eccentric power measured with both iso-inertial and gravitational systems (Dataset_power.xlsx)</p> <p>2) Database that contains the values for the Short Physical Performance Battery (SPPB) and Get Up and Go (GUG) test (Dataset_SPPB_GUG.xlsx)</p> <p>3) R Software script used to analyse file 1 (Iso-inertial analysis_power.R)<br> <br>4) R Software script used to analyse file 2 (Iso-inertial analysis_SPPB_GUG.R)</p>
Simulations dataset and pre-trained models of "Deep learning in real-time on the astrophysical data obtained from the Čerenkov CTA Observatory" Ph.D. project
<p>Ph.D. project datasets and models release, <br><em>Deep learning in real-time on the astrophysical data obtained from the Čerenkov CTA Observatory.</em></p>
Phase picker models and training data for paper "Deep learning models for regional phase detection on seismic stations in Northern Europe and the European Arctic"
<p>This ZIP file includes tensorflow models for seismic phase detection. Please see how to use these models here: https://github.com/NorwegianSeismicArray/tphasenet</p> <p>The HDF5 files includes waveforms and labels which are part of the training data set (only NORSAR event catalogue and station ARA0).</p>
Data for the paper "PETScML: Second-Order Solvers for Training Regression Problems in Scientific Machine Learning"
Open the record for dataset details and reuse information.
Data for the project 'Feasibility and Effectiveness of a Personalized Home-Based Motor-Cognitive Training Program in Community-Dwelling Older Adults: a Pragmatic Pilot Randomized Controlled Trial'
<p>Data for the project 'Feasibility and Effectiveness of a Personalized Home-Based Motor-Cognitive Training Program in Community-Dwelling Older Adults: a Pragmatic Pilot Randomized Controlled Trial' including four complete datasets with all variables collected in this project and a corresponding README file ('README-File_Data_Feasibility and Effectiveness of a Personalized Home-Based Motor-Cognitive Training Program in Community-Dwelling Older Adults: a Pragmatic Pilot Randomized Controlled Trial.txt'). The latter provides (1) general information, (2) sharing and access information, (3) data and file overview, (4) methodological information, and (5) data-specific information.</p>
Rumsey Train and Validation Data for ICDAR'24 MapText Competition
<p>Data set of 2Kx2K image tiles cropped from maps of the <a href="https://davidrumsey.com">David Rumsey collection</a> for the <a href="https://rrc.cvc.uab.es/?ch=28">ICDAR'24 Competition on Historical Map Text Detection, Recognition, and Linking</a>.</p> <p>Annotations and images follow the format described at the competition website and can be evaluated using the official <a href="https://github.com/icdar-maptext/evaluation">evaluation repository</a> script.</p> <p><strong>Important</strong>: v1.1 fixes an image channel order error, superseding the prior version. v1.2 corrects group links among several annotations. v1.3 strips markup that inadvertently remained in some annotation transcriptions.</p> <table> <tbody> <tr> <td> </td> <td><strong>Train</strong></td> <td><strong>Validation</strong></td> </tr> <tr> <td>Annotations</td> <td><code>rumsey_train.json</code></td> <td><code>rumsey_val.json</code></td> </tr> <tr> <td>Images</td> <td><code>train.zip</code></td> <td><code>val.zip</code></td> </tr> <tr> <td>Files</td> <td><code>rumsey/train/*.png</code></td> <td><code>rumsey/val/*.png</code></td> </tr> <tr> <td>Tiles</td> <td>200</td> <td>40</td> </tr> <tr> <td>Map Sheets</td> <td>196</td> <td>40</td> </tr> <tr> <td>Words</td> <td>34,518</td> <td>5,544</td> </tr> <tr> <td>Label Groups</td> <td>21,205</td> <td>3,502</td> </tr> <tr> <td>Illegible Words</td> <td>1,870</td> <td>313</td> </tr> <tr> <td>Truncated Words</td> <td>3,582</td> <td>628</td> </tr> <tr> <td>Valid Words</td> <td>30,563</td> <td>4,860</td> </tr> </tbody> </table> <p> </p> <p><strong>Annotations</strong>: Copyright 2024 UMN Knowledge Computing Lab, <a href="https://creativecommons.org/licenses/by-nc-sa/4.0/">CC-BY-NC-SA 4.0 International.</a><br><strong>Images</strong>: David Rumsey Map Collection, David Rumsey Map Center, Stanford Libraries. <a href="https://creativecommons.org/licenses/by-nc-sa/3.0/">CC-BY-NC-SA 3.0 Unported</a>.</p>
Weather data (forecast and observation) at 48 locations in France for beginning of 2024 for Machine Learning Training
<p>The data provided data are historical weather measurement and forecast at 48 locations in France and its boundary.</p> <p>Measurements are inside files named MES_YYYY.csv with YYYY is the id code of the station.</p> <p>The file "Station_list.csv" contains the list of the 45 locations with the id code, the name and then the latitude and longitude.</p> <p><br>Forecasts are inside files named XXX_YYYY.csv with YYYY the id code corresponding of the location of the grid ouput close to the associated observation location.<br>XXX is the id of the numerical forecast:<br> "GFS0.25-Complet" for GFS file at 0.25° resolution<br> "LEXIS" for WRF produced by EVEREST project using the LEXIS chain<br> "WRF3KM-Complet" for WRF at 3km resolution produced by NUMTECH<br> "WRF12KM-Complet" for WRF at 12km resolution produced by NUMTECH</p> <p><br>Description of MES-YYYY files:<br>- One line per measurement with hourly resolution<br>- columns are: Date(TU),Temperature2m_degC,WindSpeed10m_m/s,WindDirection10m_m/s<br> Date = date of measurement in TU and format DD/MM/YYYY HH:MM<br> Temperature2m_degC = air temperature at 2m height in °Celsius<br> WindSpeed10m_m/s = wind speed at 10m height in m/s<br> WindDirection10m_deg = wind direction at 10m height in deg. (0 or 360 = wind from north to south, 45°=wind from east to east, ....)<br>If measurement is not available for a specific hour for one parameter, the value "-999" is used.</p> <p>The observation data gocfrom 28/01/2024 00HTU to 17/03/2024 23HTU</p> <p><br>Description of XXX_YYYY forecast files:<br>- One line per forecast with hourly resolution<br>- columns are: First date run (TU),Forecast date,Temperature2m_degC,WindSpeed10m_m/s,WindDirection10m_m/s<br> First date run (TU) = date of start of the forecast in TU and format DD/MM/YYYY HH:MM. HH could be 00 and 12 according to the cycle of forecast start.<br> Forecast date = date of the forecast in TU and format DD/MM/YYYY HH:MM. HH go from 00 to 23. <br> Temperature2m_degC = air temperature at 2m height in °Celsius<br> WindSpeed10m_m/s = wind speed at 10m height in m/s<br> WindDirection10m_deg = wind direction at 10m height in deg. (0 or 360 = wind from north to south, 45°=wind from east to east, ....)<br>If forecast is not available for a specific hour for one parameter, the value "-999" is used.</p> <p>The forecast data go from 28/01/2024 00HTU to 17/03/2024 23HTU</p>
PreTIS2 Positive and Negative Training Data
<p>We retrieved experimental data on translation initiation sites in HEK293T cells from table S2 of (https://doi.org/10.1093/nar/gkab549) that were determined by those authors using a protocol termed TISCA. As we wanted to focus on non-canonical translation, we filtered that data and kept only truncation, extension, uORF, and overlap.uORF translation initiation types that involve either ATG or any of the nine near-cognate start codons. Using the translated amino acid sequences and transcript version IDs provided by (https://doi.org/10.1093/nar/gkab549), we fetched the respective 5'UTR and cDNA nucleotide sequences from ensembl biomart using human gencode GRCh38.p13.</p> <p>Then, we scanned each cDNA sequence in all three open reading frames and identified the longest matching nucleotide sequence that encodes the respective amino acid chain. For extension, uORF, and overlap.uORF initiation sites, the start site had to be inside the 5'UTR. For truncation initiation sites, the start site had to be inside the coding sequence. </p> <p>We also needed to construct negative cases where translation initiation supposedly does not occur. To this aim, we scanned the 5'UTR sequences for all 10 possible start codons. As negatives, we then considered those start codons that were not detected by TISCA, but have enough nucleotides upstream to meet the feature criteria. This resulted in a massive imbalance between both classes, with 17 times as many negatives as positive samples. </p> <p><span>We used as features the identity of the 20 nucleotide positions upstream of a putative start codon, represented as "U", and the 20 nt positions downstream, represented as "D". As each start codon may have its own optimal sequence context, we additionally considered the nature of the particular translation initiation start codon (ATG, CTG, GTG, TTG, AAG, ACG, AGG, ATA, ATC, and ATT). </span></p> <p>The positive and negative datasets can be found below. Whereby, each row is a possible translation initiation site and the columns symbolize the flanking nucleotide for each position as well as the start codon used.</p>
CodonTransformer - Training Data
<p>Dataset used for pretraining and finetuning CodonTransformer.</p> <p>See paper at <a href="https://adibvafa.github.io/CodonTransformer" target="_blank" rel="noopener">CodonTransformer Webpage</a></p>
Training and validation data for MOF-801(Zr) with adsorbed water molecules.
<p>We employ MACE 0.3.5 (github.com/acesuit/mace) to train an ML potential to the extended XYZ file `data.xyz`, which contains atomic geometries and potential energy and force labels. The system is MOF-801(Zr) at various water loadings. Data was generated in an active learning fashion using psiflow (github.com/molmod/psiflow).</p>
Training data for 'Genetic map RADSeq ' tutorial (Galaxy Training Material)
<p>The data provided here are part of a study published by Amores<em> et al.</em> (2011) (<a href="https://www.ncbi.nlm.nih.gov/pmc/articles/PMC3176089/">doi 10.1534/genetics.111.127324</a>), exploiting massively parallel DNA sequencing to develop meiotic maps by genotyping F<sub>1</sub> offspring of a single female and a single male spotted gar (<em>Lepisosteus oculatus</em>).</p>
Training and testing data, associated code and estimators for emulating a convection scheme
<p>Data and code for a random-forest convection scheme associated with the paper:</p> <p>"Using machine learning to parameterize moist convection: potential for modeling of climate, climate change and extreme events"</p> <p>by Paul A. O'Gorman and John G. Dwyer (to appear in JAMES)</p>
Training data for "Machine learning: classification and regression"
<p>The data provided here are part of a Galaxy Training Network tutorial for "Machine learning: classification and regression".</p>
Training material for analysis small RNA-seq data (Galaxy Training Network tutorial)
<p>The data provided here is part of the Galaxy Training Network tutorial for analysis of small RNA-seq (sRNA-seq) data using mirdeep2 and miranda. This dataset is provided by INRA (Le Rheu, France).</p>
Data and code for training neural network parameterizations from an near-global aqua-planet simulation
<p>This commit contains the code, coarse-grained data, processed training data, neural network models, and coupled NN-GCM simulations. It can be extracted by running</p> <pre><code>tar xzf <archive></code></pre> <p>While this archive contains code (it is slightly out of date). This is the up-to-date code: <a href="https://zenodo.org/record/3248586">https://zenodo.org/record/3248586</a></p> <p>Move the "nn", "debiased", and "data" folders from this archive into that code directory.</p> <p> </p> <p> </p>
Trimmed RNASeq pair for the Galaxy Training Network tutorial - "Metatranscriptomics analysis using microbiome RNASeq data"
<p>Functional microbiome analysis which estimates the functional groups expressed by microbial community enables researchers to look beyond taxonomic composition and correlation with the condition under study. Using microbial community RNA-Seq data and subsequent metatranscriptomics workflows to elucidate the functional complement of the microbiome is gaining interest in the field. <br> This Galaxy training network tutorial will introduce researchers to the basic concepts and tools from the published ASaiM workflow (Batut et al, <em>GigaScience</em> (2018), 7 (6),<a href="http://dx.doi.org/10.1093/gigascience/giy057"> http://dx.doi.org/10.1093/gigascience/giy057</a>). </p> <p>The dataset is a trimmed version of one of the time points from a cellulose degradation biogas reactor dataset. The dataset has been trimmed to facilitate running the workflows for this tutorial. Any biological interpretation from the results would be incorrect, due to the trimmed version of the dataset.</p>
Training Data for "Making sense of a newly assembled genome"
<p>The data provided here is part of the Galaxy Training Network tutorial for Making sense of a newly assembled genome. This data was sourced from NCBI on 2019-08-29</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.