Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
558
datasets available to search
ShareScore release 0.9.0
Dataset results
558 results for “Training Data”
Example data for Immcantation training legacy tutorials
<p>`bcr_phylo_tutorial.zip` is used in the <a href="https://immcantation.readthedocs.io/en/4.4.0/tutorials/dowser_tutorial.html">Reconstruction and analysis of B-cell lineage trees from single cell data using Immcantation</a> tutorial.</p> <p>`immcantation-BCR-Seurat-tutorial.zip` is used in the <a href="https://immcantation.readthedocs.io/en/4.4.0/tutorials/BCR_Seurat_tutorial.html">Integration of BCR and GEX data</a> tutorial. </p>
Data set for training a ML model to predict duration of MPI application phases (HPC system) - with previous phase info
<p>This is the data used to train a ML model predicting the duration of MPI application phases, in a HPC system.</p> <p>There are 10 different data sets corresponding to different HPC applications.</p> <p>These data sets contain information regarding the previous MPI call with same ID and type.</p>
Data set for training a ML model to predict duration of MPI application phases (HPC system) - without previous phases info
<p>This is the data used to train a ML model predicting the duration of MPI application phases, in a HPC system.</p> <p>There are 11 different data sets corresponding to different HPC applications.</p> <p>These data sets do not contain information regarding previous MPI calls</p>
m-sized Training and Evaluation Data for Publication "Using Supervised Learning to Classify Metadata of Research Data by Field of Study"
<p>Automated classification of metadata of research data by their discipline(s) of research can be used<br> in scientometric research, by repository service providers, and in the context of research data aggregation<br> services. Openly available metadata of the DataCite index for research data were used to compile a large<br> training and evaluation set comprised of 609,524 records. This is the cleaned and vectorized version with a feature selection of medium size</p>
Training Data - Effects of seasonality and classifier on the accuracy of grazing resource and land degradation maps in a savanna ecosystem
<p>This csv file contains training data (class and coordinates (latitude and longitude)) used in the classification models presented in the paper - Effects of seasonality and classifier on the accuracy of grazing resource and land degradation maps in a savanna ecosystem</p>
Data and code for training and testing a ResMLP model with experience replay for machine-learning physics parameterization
<p>This directory contains the training data and code for training and testing a ResMLP with experience replay for creating a machine-learning physics parameterization for the Community Atmospheric Model. </p> <p>The directory is structured as follows:</p> <p>1. Download training and testing data: https://portal.nersc.gov/archive/home/z/zhangtao/www/hybird_GCM_ML</p> <p>2. Unzip nncam_training.zip</p> <p>nncam_training</p> <p> - models</p> <p> model definition of ResMLP and other models for comparison purposes</p> <p> - dataloader </p> <p> utility scripts to load data into pytorch dataset</p> <p> - training_scripts</p> <p> scripts to train ResMLP model with/without experience replay</p> <p> - offline_test</p> <p> scripts to perform offline test (Table 2, Figure 2)</p> <p>3. Unzip nncam_coupling.zip</p> <p>nncam_srcmods</p> <p> - SourceMods</p> <p> SourceMods to be used with CAM modules for coupling with neural network</p> <p> - otherfiles</p> <p> additional configuration files to setup and run SPCAM with neural network</p> <p> - pythonfiles</p> <p> python scripts to run neural network and couple with CAM</p> <p> - ClimAnalysis</p> <p> - paper_plots.ipynb</p> <p> scripts to produce online evaluation figures (Figure 1, Figure 3-10)</p> <p> </p>
Raw Data Training Decision Tree Rice Disease
<p>this data is data set for training decision tree in research rice plant disease.</p>
Pre-processed data and trained model weight in ADAF
<h1>Pre-proccesd data</h1> <p>The pre-proccesd data consists of input-target pairs. The inputs include surface weather observations within a 3-hour window, GOES-16 satellite imagery within a 3-hour window, HRRR forecast, and topography. The target is a combination of RTMA and surface weather observations. The table below summarizes the input and target datasets utilized in this study. All data were regularized to grids of size 512 $\times$ 1280 with a spatial resolution of 0.05 $\times$ 0.05 $^\circ$. </p> <table> <tbody> <tr> <td> </td> <td><strong>Dataset</strong></td> <td><strong>Source</strong></td> <td><strong>Time window</strong></td> <td><strong>Variables/Bands</strong></td> </tr> <tr> <td><strong>Input</strong></td> <td>Surface weather observations</td> <td>WeatherReal-Synoptic (Jin et al., 2024)</td> <td>3 hours</td> <td>Q, T2M, U10, V10</td> </tr> <tr> <td><strong>Input</strong></td> <td>Satellite imagery</td> <td>GOES-16 (Tan et al., 2019)</td> <td>3 hours</td> <td>0.64, 3.9, 7.3, 11.2 $\mu m$</td> </tr> <tr> <td><strong>Input</strong></td> <td>Background</td> <td>HRRR forecast (Dowell et al., 2022)</td> <td>N/A</td> <td>Q, T2M, U10, V10</td> </tr> <tr> <td><strong>Input</strong></td> <td>Topography</td> <td>ERA5 (Hersbach et al., 2019)</td> <td>N/A</td> <td>Geopotential</td> </tr> <tr> <td><strong>Target</strong></td> <td>Analysis</td> <td>RTMA (Pondeca et al., 2011)</td> <td>N/A</td> <td>Q, T2M, U10, V10</td> </tr> <tr> <td><strong>Target</strong></td> <td>Surface weather observations</td> <td>WeatherReal-Synoptic (Jin et al., 2024)</td> <td>N/A</td> <td>Q, T2M, U10, V10</td> </tr> </tbody> </table> <p><a href="https://zenodo.org/api/records/14020879/draft/files/2022-10-01_06.nc/content" target="_blank" rel="noopener noreferrer">2022-10-01_06.nc</a> is a sample of pre-proccesd data. The vairables in this file contain the input-target pairs mentioned above.</p> <div> <p>A sample file contains the following variables:</p> <table> <tbody> <tr> <td><strong>Variable</strong></td> <td><strong>Decription</strong></td> <td><strong>Dimension</strong></td> </tr> <tr> <td>z</td> <td>Topography, normalized</td> <td>[lat, lon]</td> </tr> <tr> <td>rtma_t</td> <td>T2M from RTMA, normalized</td> <td>[lat, lon]</td> </tr> <tr> <td>rtma_q</td> <td>Q from RTMA, normalized</td> <td>[lat, lon]</td> </tr> <tr> <td>rtma_u10</td> <td>U10 from RTMA, normalized</td> <td>[lat, lon]</td> </tr> <tr> <td>rtma_v10</td> <td>V10 from RTMA, normalized</td> <td>[lat, lon]</td> </tr> <tr> <td>sta_t</td> <td>T2M from station's observation, 0 means non-station, normalized</td> <td>[obs_time_window, lat, lon]</td> </tr> <tr> <td>sta_q</td> <td>Q from station's observation, 0 means non-station, normalized</td> <td>[obs_time_window, lat, lon]</td> </tr> <tr> <td>sta_u10</td> <td>U10 from station's observation, 0 means non-station, normalized</td> <td>[obs_time_window, lat, lon]</td> </tr> <tr> <td>sta_v10</td> <td>V10 from station's observation, 0 means non-station, normalized</td> <td>[obs_time_window, lat, lon]</td> </tr> <tr> <td>CMI02</td> <td>ABI Band 2: visible (red), normalized</td> <td>[obs_time_window, lat, lon]</td> </tr> <tr> <td>CMI07</td> <td>ABI Band 7: shortwave infrared, normalized</td> <td>[obs_time_window, lat, lon]</td> </tr> <tr> <td>CMI10</td> <td>ABI Band 10: low-level water vapor, normalized</td> <td>[obs_time_window, lat, lon]</td> </tr> <tr> <td>CMI14</td> <td>ABI Bands 14: longwave infrared, normalized</td> <td>[obs_time_window, lat, lon]</td> </tr> <tr> <td>hrrr_t</td> <td>T2M from HRRR 1-hour forecast</td> <td>[lat, lon]</td> </tr> <tr> <td>hrrr_q</td> <td>Q from HRRR 1-hour forecast</td> <td>[lat, lon]</td> </tr> <tr> <td>hrrr_u_10</td> <td>U10 from HRRR 1-hour forecast</td> <td>[lat, lon]</td> </tr> <tr> <td>hrrr_v_10</td> <td>V10 from HRRR 1-hour forecast</td> <td>[lat, lon]</td> </tr> </tbody> </table> <p> </p> </div> <h1>Pre-comuted normalization statistics</h1> <p><a href="https://zenodo.org/api/records/14020879/draft/files/stats.csv/content" target="_blank" rel="noopener noreferrer">stats.csv</a> is pre-comuted normalization statistics.</p> <h1>Pre-trained model weights</h1> <p><a href="https://zenodo.org/api/records/14020879/draft/files/best_ckpt.tar/content" target="_blank" rel="noopener noreferrer">best_ckpt.tar</a> is the pre-trained model weights.</p> <h1>References</h1> <ol> <li>Jin, W. et al. WeatherReal: A Benchmark Based on In-Situ Observations for Evaluating Weather Models. (2024).</li> <li>Dowell, D. et al. The High-Resolution Rapid Refresh (HRRR): An Hourly Updating Convection-Allowing Forecast Model. Part I: Motivation and System Description. Weather and Forecasting 37, (2022).</li> <li>Tan, B., Dellomo, J., Wolfe, R. & Reth, A. GOES-16 and GOES-17 ABI INR assessment. in Earth Observing Systems XXIV vol. 11127 290–301 (SPIE, 2019).</li> <li>Hersbach, H. et al. ERA5 monthly averaged data on single levels from 1979 to present. Copernicus Climate Change Service (C3S) Climate Data Store (CDS) 10, 252–266 (2019).</li> <li>Pondeca, M. S. F. V. D. et al. The Real-Time Mesoscale Analysis at NOAA’s National Centers for Environmental Prediction: Current Status and Development. Weather and Forecasting 26, 593–612 (2011).</li> </ol>
Data from: Increased dopamine release after working-memory updating training: neurochemical correlates of transfer
Previous work demonstrates that working-memory (WM) updating training results in improved performance on a letter-memory criterion task, transfers to an untrained n-back task, and increases striatal dopamine (DA) activity during the criterion task. Here, we sought to replicate and extend these findings by also examining neurochemical correlates of transfer. Four positron emission tomography (PET) scans using the radioligand raclopride were performed. Two of these assessed DAD2 binding (letter memory; n-back) before 5 weeks of updating training, and the same two scans were performed post training. Key findings were (a) pronounced training-related behavioral gains in the letter-memory criterion task, (b) altered striatal DAD2 binding potential after training during letter-memory performance, suggesting training-induced increases in DA release, and (c) increased striatal DA activity also during the n-back transfer task after the intervention, but no concomitant behavioral transfer. The fact that the training-related DA alterations during the transfer task were not accompanied by behavioral transfer suggests that increased DA release may be a necessary, but not sufficient, condition for behavioral transfer to occur.
Data set to train a natural language classifier able to differentiate between 15 topics relevant to biodiversity informatics
<p><strong>Scope and size</strong><br> This data set is used to train a natural language processing classifier. The classifier shall be able to differentiate between 15 topics relevant to biodiversity informatics. The list of relevant topics was adapted from Searls (2012).</p> <p>The data set was split into training data, testing data (for tweaking and unit-testing the classifier) and validation data. Each data set is stored as PDF files in a separate directory.</p> <ul> <li>Training data (5494 pages)</li> <li>Test data (977 pages)</li> <li>Validation data (215 pages)</li> </ul> <p><strong>Data sources and licenses</strong><br> Details about the licenses for each data set can be found in the corresponding directories.</p> <ul> <li>Training data was compiled from MIT OpenCourseWare resources provided by MIT under a Creative Commons BY-NC-SA License.</li> <li>Testing data was compiled from MIT OpenCourseWare exams, provided by MIT under a Creative Commons License BY-NC-SA.</li> <li>Validation data was compiled from Wikipedia, provided under a Creative Commons License by Wikipedia editors and contributors.</li> </ul> <p><strong>Topic references</strong><br> Each topic references one or more MIT OpenCourseWare courses:</p> <ul> <li>Algorithms (Demaine, and Devadas, 2011)</li> <li>Artificial Intelligence (Winston, 2010)</li> <li>Building Dynamic Websites (Abelson, and Greenspun, 2003)</li> <li>Computational Biology (Kellis, 2015)</li> <li>Computer Graphics (Matusik, and Durand, 2012)</li> <li>Computer Science and Programming (Bell, Grimson, and Guttag, 2016)</li> <li>Databases (Madden, Morris, Stonebraker, and Curino, 2010)</li> <li>Data Structures (Demaine, 2012)</li> <li>Digital Image Processing (Clifford, Fisher, Greenberg, and Wells, 2007; Golland, 2005)</li> <li>Machine Learning (Singh, Jaakkola, and Mohammad, 2006)</li> <li>Machine Structures (Morris, and Madden, 2009)</li> <li>Natural Language Processing (Berwick, 2003; Collins, and Barzilay, 2005)</li> <li>Parallel Computing (Edelman, 2011)</li> <li>Software Engineering (Jackson, and Devadas, 2005)</li> <li>Structure and Interpretation of Computer Programs (Miller, and Goldman, 2016).</li> </ul> <p><strong>References</strong></p> <p>Harold Abelson, and Philip Greenspun. 6.171 Software Engineering for Web Applications. Fall 2003. Massachusetts Institute of Technology: MIT OpenCourseWare, https://ocw.mit.edu. License: Creative Commons BY-NC-SA.</p> <p>Ana Bell, Eric Grimson, and John Guttag. 6.0001 Introduction to Computer Science and Programming in Python. Fall 2016. Massachusetts Institute of Technology: MIT OpenCourseWare, https://ocw.mit.edu. License: Creative Commons BY-NC-SA.</p> <p>Robert Berwick. 6.863J Natural Language and the Computer Representation of Knowledge. Spring 2003. Massachusetts Institute of Technology: MIT OpenCourseWare, https://ocw.mit.edu. License: Creative Commons BY-NC-SA.</p> <p>Gari Clifford, John Fisher, Julie Greenberg, and William Wells. HST.582J Biomedical Signal and Image Processing. Spring 2007. Massachusetts Institute of Technology: MIT OpenCourseWare, https://ocw.mit.edu. License: Creative Commons BY-NC-SA.</p> <p>Michael Collins, and Regina Barzilay. 6.864 Advanced Natural Language Processing. Fall 2005. Massachusetts Institute of Technology: MIT OpenCourseWare, https://ocw.mit.edu. License: Creative Commons BY-NC-SA.</p> <p>Erik Demaine. 6.851 Advanced Data Structures. Spring 2012. Massachusetts Institute of Technology: MIT OpenCourseWare, https://ocw.mit.edu. License: Creative Commons BY-NC-SA.</p> <p>Erik Demaine, and Srini Devadas. 6.006 Introduction to Algorithms. Fall 2011. Massachusetts Institute of Technology: MIT OpenCourseWare, https://ocw.mit.edu. License: Creative Commons BY-NC-SA.</p> <p>Alan Edelman. 18.337J Parallel Computing. Fall 2011. Massachusetts Institute of Technology: MIT OpenCourseWare, https://ocw.mit.edu. License: Creative Commons BY-NC-SA.</p> <p>Polina Golland. 6.881 Representation and Modeling for Image Analysis. Spring 2005. Massachusetts Institute of Technology: MIT OpenCourseWare, https://ocw.mit.edu. License: Creative Commons BY-NC-SA.</p> <p>Daniel Jackson, and Srini Devadas. 6.170 Laboratory in Software Engineering. Fall 2005. Massachusetts Institute of Technology: MIT OpenCourseWare, https://ocw.mit.edu. License: Creative Commons BY-NC-SA.</p> <p>Manolis Kellis. 6.047 Computational Biology. Fall 2015. Massachusetts Institute of Technology: MIT OpenCourseWare, https://ocw.mit.edu. License: Creative Commons BY-NC-SA.</p> <p>Robert Miller, and Max Goldman. 6.005 Software Construction. Spring 2016. Massachusetts Institute of Technology: MIT OpenCourseWare, https://ocw.mit.edu. License: Creative Commons BY-NC-SA.</p> <p>Samuel Madden, Robert Morris, Michael Stonebraker, and Carlo Curino. 6.830 Database Systems. Fall 2010. Massachusetts Institute of Technology: MIT OpenCourseWare, https://ocw.mit.edu. License: Creative Commons BY-NC-SA.</p> <p>Wojciech Matusik, and Frédo Durand. 6.837 Computer Graphics. Fall 2012. Massachusetts Institute of Technology: MIT OpenCourseWare, https://ocw.mit.edu. License: Creative Commons BY-NC-SA.</p> <p>Morris, Robert, and Madden, Samuel. 6.033 Computer System Engineering. Spring 2009. Massachusetts Institute of Technology: MIT OpenCourseWare, http://hdl.handle.net/1721.1/118791. License: Creative Commons BY-NC-SA.</p> <p>David B. Searls. An online bioinformatics curriculum. 2012. PLoS computational biology, 8(9), p.e1002632.</p> <p>Rohit Singh, Tommi Jaakkola, and Ali Mohammad. 6.867 Machine Learning. Fall 2006. Massachusetts Institute of Technology: MIT OpenCourseWare, https://ocw.mit.edu. License: Creative Commons BY-NC-SA.</p> <p>Patrick Winston. 6.034 Artificial Intelligence. Fall 2010. Massachusetts Institute of Technology: MIT OpenCourseWare, https://ocw.mit.edu. License: Creative Commons BY-NC-SA.</p> <p> </p>
Data and trained models for "Fourier Ring Correlation and anisotropic kernel density estimation improve deep learning based SMLM reconstruction of microtubules"
<p>Data and trained models for "Fourier Ring Correlation and anisotropic kernel density estimation improve deep learning based SMLM reconstruction of microtubules", https://github.com/CIA-CCTB/FRCnet</p>
Inspiration From Eye-tracking Data: Investigating the Impact of Combining Specific Environmental Features and Power Mobility Training
ClinicalTrials.gov study NCT06928077. IPD Sharing: Not stated. Countries: 1. Publications: 0.
Measuring the Quality of Surgical Care and Setting Benchmarks for Training Using Intuitive Data Recorder Technology
ClinicalTrials.gov study NCT04647188. IPD Sharing: NO. Countries: 1. Publications: 0.
Training and Support Programme on Data-driven Quality Development for Swiss Long-Term Care Facilities (NIP-Q-UPGRADE Subaim 2.6)
ClinicalTrials.gov study NCT07068009. IPD Sharing: NO. Countries: 1. Publications: 0.
Training Data Collection & AI Development
ClinicalTrials.gov study NCT05378854. IPD Sharing: YES. Countries: 1. Publications: 0.
LEOPARD Training and Validation Data Collection Study
ClinicalTrials.gov study NCT06675604. IPD Sharing: NO. Countries: 7. Publications: 0.
Data from: Effects of a systematically offered social and preventive medicine consultation on training and health attitudes of young people not in employment, education or training (NEETs): an interventional study in France
Open the record for dataset details and reuse information.
Data from: Effects of working-memory training on striatal dopamine release
Open the record for dataset details and reuse information.
Data from: A structured training program for health workers in intravenous treatment with fluids and antibiotics in nursing homes: a modified stepped-wedge cluster-randomised trial to reduce hospital admissions
Open the record for dataset details and reuse information.
Data from: Rehabilitative skilled forelimb training enhances axonal remodeling in the corticospinal pathway but not the brainstem-spinal pathways after photothrombotic stroke in the primary motor cortex
Open the record for dataset details and reuse information.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.