Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

558

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

558 results for “Training Data”

Learn how ShareScore rates datasets ↗
zenodo40/100

Mapping Tree Species Fractions in Temperate Mixed Forests Using Sentinel-2 Time Series and Synthetically Mixed Training Data

<p>This dataset contains the latest version of a selection of result data of the paper "Mapping Tree Species Fractions in Temperate Mixed Forests Using Sentinel-2 Time Series and Synthetically Mixed Training Data" (DOI: https://doi.org/10.1016/j.rse.2025.114740 )</p> <p>The dataset contains:</p> <ol> <li>A geopackage of training points of pure tree species</li> <li>The resulting 12-band tree species fraction map of Rhineland-Palatinate</li> <li>HSV-colored map of dominant tree species. For information which tree species are represented by the different colors, refer to the Supplemental in the original paper.</li> <li>CSV-table of predicted and reference propotion of the tree species in the validation polygon (the original polygon data can not be published due to data privacy regulations)&nbsp;</li> </ol> <p>&nbsp;</p>

opengpl-3.0-or-laterOct 2024View details →
zenodo40/100

Training and test data for retrievals based on HATPRO observations during MOSAiC

<p>The dataset consists of one netCDF file that contains the entire training and test data for the retrieval of temperature (ta) and humidity (hua) profiles, integrated water vapour (prw), and liquid water path (clwvi) from brightness temperatures (tb) measured by a HATPRO (humidity and temperature profiler). A regression with quadratic terms, except for the boundary layer scan which is confined&nbsp;to linear terms, is performed to derive these meteorological variables. The trained retrieval is applied on the HATPRO observations gathered onboard the research vessel Polarstern during the&nbsp;Multidisciplinary&nbsp;drifting Observatory for the Study of Arctic Climate (MOSAiC) expedition.&nbsp;For the data to be&nbsp;specialized on Arctic conditions they are based on Ny-&Aring;lesund radiosonde observations. An IDL-based radiative transfer model has been used to simulate brightness temperatures for each radiosonde. Several elevation angles (ele) are given in the file&nbsp;because the HATPRO radiometer performs elevation scans in between zenith scans to increase the resolution of temperature profiles in the atmospheric boundary layer.</p>

opencc-by-4.0Nov 2021View details →
zenodo40/100

Audio samples from generative models trained on the TIMIT speech data.

<p>This is a posting of audio snippets to accompany the paper&nbsp;&quot;Benchmarking Generative Latent Variable&nbsp;Models for Speech&quot;.</p> <p>The snippets include samples and reconstructions.&nbsp;All samples are completely unconditional and utilise only the prior&nbsp;internal representations learned by the model.&nbsp;Reconstructions are computed from a given test audio snippet by first encoding it to a learned representation and then decoding that&nbsp;to a reconstruction of the audio.</p> <p>All models are trained on the TIMIT speech dataset (<a href="https://catalog.ldc.upenn.edu/LDC93s1">https://catalog.ldc.upenn.edu/LDC93s1</a>). Some snippets are from models&nbsp;trained at different temporal resolutions denoted by `s1` and `s64`. We refer to the paper for details.</p> <p>The files include:</p> <ul> <li>`clockwork-vae-s64-reconstruction-*` <ul> <li>Four reconstructions using a&nbsp;two-layered Clockwork VAE trained with temporal resolution s=64.</li> </ul> </li> <li>`clockwork-vae-s64-sample-*` <ul> <li>Four samples from the prior of a Clockwork VAE trained with temporal resolution s=64.</li> </ul> </li> <li>`original-*` <ul> <li>Four original samples from TIMIT corresponding in pairs to the reconstructions.</li> </ul> </li> <li>`vrnn-s64-sample-*` <ul> <li>Two samples from the prior of a VRNN trained with temporal resolution s=64.</li> </ul> </li> <li>`vrnn-s1-sample-*` <ul> <li>Two samples from the prior of a VRNN trained with temporal resolution s=1.</li> </ul> </li> <li>`srnn-s64-sample-*` <ul> <li>Two samples from the prior of a SRNN trained with temporal resolution s=64.</li> </ul> </li> <li>`srnn-s1-sample-*` <ul> <li>Two samples from the prior of a SRNN trained with temporal resolution s=1.</li> </ul> </li> <li>`wavenet-s64-sample-*` <ul> <li>Four samples from a WaveNet trained with temporal resolution s=1.</li> </ul> </li> <li>`wavenet-s1-sample-*` <ul> <li>Two samples from a WaveNet trained with temporal resolution s=64.</li> </ul> </li> </ul>

opencc-by-4.0Nov 2021View details →
zenodo40/100

Bening Flow Data Train (D1)

<p>Train dataset only contains benign traffic, this dataset has been collected using Netflow and implementing a sampling rate of 1 packet out of each 1000 to generate the flows, simulating the conditions of the RedCayle&rsquo;s routers.</p> <p>The traffic has been generated using three Python scripts. The first one simulates email sending using SMTP protocol. The second script simulates SSH connections. The third script simulates a user browsing the internet using different search engines and different protocols like HTTP and HTTPS.</p>

opencc-by-4.0Dec 2021View details →
zenodo40/100

Training and test data for retrievals based on MiRAC-P observations during MOSAiC

<p>The dataset consists of one netCDF file that contains the entire training and test data for the retrieval of integrated water vapour (prw) from brightness temperatures (tb) measured by the MiRAC-P (microwave radiometer for Arctic clouds, aka. LHUMPRO-243-340). A neural network retrieval has been developed to derive the prw. The trained retrieval is applied on the MiRAC-P observations gathered onboard the research vessel Polarstern during the Multidisciplinary drifting Observatory for the Study of Arctic Climate (MOSAiC) expedition. For the data to be specialized on Arctic conditions they are based on ERA-Interim reanalysis. An IDL-based radiative transfer model has been used to simulate brightness temperatures. The elevation angle (ele) is always 90&deg; because the MiRAC-P performed zenith scans only throughout the MOSAiC campaign.</p>

opencc-by-4.0Nov 2021View details →
zenodo40/100

Netflow data with sampling 500 for training (D2)

<p>NetFlow traffic generated using DOROTHEA (DOcker-based fRamework fOr gaTHering nEtflow trAffic) NetFlow is a network protocol developed by Cisco for the collection and monitoring of network traffic flow data generated. A flow is defined as a unidirectional sequence of packets with some common properties that pass through a network device.</p> <p>NetFlow flows have been captured with sampling 500 at the packet level. A sampling means that 1 out of every X packets is selected to be flow while the rest of the packets are not valued.</p> <p>The version of NetFlow used to build the datasets is 5.</p>

opencc-by-4.0Jan 2022View details →
zenodo40/100

Netflow data with sampling 250 for training (D1)

<p>NetFlow traffic generated using DOROTHEA (DOcker-based fRamework fOr gaTHering nEtflow trAffic) NetFlow is a network protocol developed by Cisco for the collection and monitoring of network traffic flow data generated. A flow is defined as a unidirectional sequence of packets with some common properties that pass through a network device.</p> <p>NetFlow flows have been captured with sampling 250 at the packet level. A sampling means that 1 out of every X packets is selected to be flow while the rest of the packets are not valued.</p> <p>The version of NetFlow used to build the datasets is 5.</p>

opencc-by-4.0Jan 2022View details →
zenodo40/100

Netflow data with sampling 1000 for training (D3)

<p>NetFlow traffic generated using DOROTHEA (DOcker-based fRamework fOr gaTHering nEtflow trAffic) NetFlow is a network protocol developed by Cisco for the collection and monitoring of network traffic flow data generated. A flow is defined as a unidirectional sequence of packets with some common properties that pass through a network device.</p> <p>NetFlow flows have been captured with sampling 1000 at the packet level. A sampling means that 1 out of every X packets is selected to be flow while the rest of the packets are not valued.</p> <p>The version of NetFlow used to build the datasets is 5.</p>

opencc-by-4.0Jan 2022View details →
zenodo40/100

Data set - Practical teacher-training program on STEAM activity planning

<p>Data set of answers of 14 Brazilian teachers who took part in a&nbsp;practical teacher-training program on STEAM activity planning.&nbsp;</p>

opencc-by-4.0Feb 2022View details →
zenodo40/100

Training data for the "Biodiversity data exploration" Galaxy-E tutorial

<p>Dataset sample from Reef life survey initiative https://reeflifesurvey.com/ to serve as a training set for &quot;Biodiversity data exploration&quot; tutorial for Galaxy and notably Galaxy for ecology initiative</p>

opencc-by-4.0Jan 2022View details →
zenodo40/100

Galaxy Training Data for "Designing plasmids encoding predicted pathways by using the BASIC assembly method"

<p>This dataset provides the data needed for the Galaxy BASIC assembly workflow training tutorial (<a href="https://galaxy-synbiocad.org">https://galaxy-synbiocad.org</a>). This workflow provides a pathway to design plasmids encoding predicted metabolic pathways using the BASIC assembly method (<a href="https://doi.org/10.1021/sb500356d">https://doi.org/10.1021/sb500356d</a>). It generates scripts allowing the automatic construction of these plasmids using an Opentrons liquid handling robot. After downloading these scripts on a computer connected to an Opentrons (<a href="https://opentrons.com">https://opentrons.com</a>), the user can perform the automatic construction of the plasmids on the bench.</p> <p>The content of the dataset is as follows:</p> <ul> <li> <p>an SBML file modeling a heterologous pathway producing lycopene such as those produced by the Pathway Analysis Workflow (<a href="https://galaxy-synbiocad.org">https://galaxy-synbiocad.org</a>).</p> </li> <li> <p>a CSV file listing the parts to be used (linkers, backbone and promoters) in the constructions.</p> </li> <li> <p>two YAML files providing two examples of settings, i.e. providing the identifiers of the laboratory equipment and the parameters of the DNA robot.</p> </li> </ul>

opencc-by-4.0Feb 2022View details →
zenodo40/100

In silico data-set used to train the PGNNIV

<p>The folder contains the data used to train PGNNIVs to unravel the go or grow behaviour of glioblastoma under 4 different parametric models. These data consist on the solutions (in time and space)&nbsp;of eleven simulations with different oxygen boundary conditions.</p> <p>Each folder is named as DATA_&quot;ModelName&quot;, where &quot;ModelName&quot; can be: &quot;Sigmoid&quot;, &quot;ReLU&quot;, &quot;MichaelisMenten&quot; or &quot;Heaviside&quot;. Inside each folder, the multidimensional arrays for input and output data for the network training can be found. These&nbsp;arrays have dimension [nExp,TimeStep,x,field], where:</p> <ul> <li>nExp = 11&nbsp; and corresponds to the number of different configurations or experiments simulated.</li> <li>TimeStep = 1000 and correspond to the different temporal frames where the solution is given.</li> <li>x = 51 and corresponds to the different spatial points where the solution is given.</li> <li>field = 2 and correspond to the different solution fields (1: cells, 2: oxygen).</li> </ul>

opencc-by-4.0Mar 2022View details →
zenodo40/100

Training and Testing Data, Associated Code, and WRF Code for ML-based nonhydrostatic alternative scheme in dynamical core of atmosphere

<p>Data and codes for a nonhydrostatic alternative scheme (NAS) in dynamical core of atmosphere based on machine learning.</p> <p>In this new version, the&nbsp;randomly sampled&nbsp;training data samples testing data samples from nonhydrostatic simulations in WRF baraclinic wave test&nbsp;are provided. They are processed&nbsp;into a new data structure, which can be directly utilized in training and testing.&nbsp;</p> <p>Follow the instructions in README.txt and download the training and testing data, and the associated codes.</p> <p>Here we provide 3 parts of data and codes:</p> <p>1, Training and testing data from WRF;</p> <p>2, Training and testing codes for two machine learning emulators:&nbsp;machine learning and neural network</p> <p>3, WRF application.</p>

opencc-by-4.0Mar 2022View details →
zenodo40/100

Training data for submitted paper "Wildfire Danger Prediction and Understanding with Deep Learning"

<p>Training data for submitted paper &quot;Wildfire Danger Prediction and Understanding with Deep Learning&quot;. To run the code in https://github.com/Orion-AI-Lab/wildfire_forecasting</p>

opencc-by-4.0May 2022View details →
zenodo40/100

Weather data (forecast and observation) at three locations in France over 2021 for Machine Learning Training

<p>The data provided data are historical weather measurement and forecast at three location in France.</p> <p>Measurements are inside files named OBS_xxx</p> <p>Forecasts are inside files names YYY_xxx, with YYY is the name of the forecast simultion (GFS0.25, WRF12km or WRF3KM).</p> <p>In the two cases, xxx is the name of the site (Site 1, Site2 or Site3).</p> <p><br> <strong>Description of OBS_xxx files:</strong><br> - One line per measurement with hourly resolution<br> - columns are: Date(TU),Temperature2m_degC,WindSpeed10m_m/s,WindDirection10m_m/s<br> &nbsp;&nbsp; &nbsp;Date = date of measurement in TU and format DD/MM/YYYY HH:MM<br> &nbsp;&nbsp; &nbsp;Temperature2m_degC = air temperature at 2m height in &deg;Celsius<br> &nbsp;&nbsp; &nbsp;WindSpeed10m_m/s = wind speed at 10m height in m/s<br> &nbsp;&nbsp; &nbsp;WindDirection10m_deg = wind direction at 10m height in deg. (0 or 360 = wind from north to south, 45&deg;=wind from east to east, ....)<br> If measurement is not available for a specific hour for one parameter, the value &quot;-999&quot; is used.</p> <p>The observation data go:<br> from 16/04/2021 00H&nbsp;<br> to 31/01/2022 23H</p> <p><br> <strong>Description of YYY_xxx files:</strong><br> - One line per forecast with hourly resolution<br> - columns are: First date run (TU),forecast hour,Temperature2m_degC,WindSpeed10m_m/s,WindDirection10m_m/s<br> &nbsp;&nbsp; &nbsp;First date run (TU) = date of start of the forecast in TU and format DD/MM/YYYY HH:MM. HH could be 00 and 12 according to the cycle of forecast start.<br> &nbsp;&nbsp; &nbsp;forecast hour = forecast hour from the start of the forecast date. 00 = forecast for &quot;first date run&quot;. 01 = forecast for &quot;First date run&quot; + 1 hour. .... 95 = forecast for &quot;First date run&quot; + 95 hours.<br> &nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;For GFS0.25, forecast hour go from 00 to 95<br> &nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;For WRF12km, forecast hour go from 00 to 95<br> &nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;For WRF3m, forecast hour go from 00 to 95<br> &nbsp;&nbsp; &nbsp;Temperature2m_degC = air temperature at 2m height in &deg;Celsius<br> &nbsp;&nbsp; &nbsp;WindSpeed10m_m/s = wind speed at 10m height in m/s<br> &nbsp;&nbsp; &nbsp;WindDirection10m_deg = wind direction at 10m height in deg. (0 or 360 = wind from north to south, 45&deg;=wind from east to east, ....)<br> If measurement is not available for a specific hour for one parameter, the value &quot;-999&quot; is used.</p> <p>The forecast data go:<br> from 13/04/2021 00H + 72H = first forecast for the 16/04/2021 00H<br> to 31/01/2022 12H + 11H = last forecast for the 31/01/2022 23H</p>

opencc-by-4.0May 2022View details →
zenodo40/100

MItosis DOmain Generalization Challenge 2022 (MICCAI MIDOG 2022), Training data set (PNG version)

<p>This is the training dataset of the MItosis DOmain Generalization (MIDOG) challenge 2022, held in conjunction with MICCAI 2022. Please find the structured challenge description at 10.5281/zenodo.6362337.</p> <p>The training set consists of 405 tumor cases in total across six tumor types:</p> <ul> <li>Canine Lung Cancer (44 cases, scanned with 3DHistech Pannoramic Scan II)</li> <li>Human Breast Cancer (150 cases, scanned using three scanners, part of MIDOG2021 dataset)</li> <li>Canine Lymphoma (55 cases, scanned with 3DHistech Pannoramic Scan II)</li> <li>Human neuroendocrine tumor (55 cases, scanned with Hamamatsu NanoZoomer XR)</li> <li>Canine Cutaneous Mast Cell Tumor (50 cases, scanned with Aperio ScanScope CS2)</li> <li>Human melanoma (51 cases, scanned with Hamamatsu NanoZoomer XR) (no labels provided)</li> </ul> <p>From each WSI, a trained pathologist selected an area of 2mm&sup2; corresponding to approximately 10 high power fields, according to the grading scheme of Elston and Ellis. We cropped this area and provide it as PNG files in this data set due to restrictions in data set size on zenodo. Each file includes the resolution (in dots per inch, DPI) of the original scanned images.</p> <p>The training set contains 9501 mitotic figures (MF) and 11051 hard examples (non-mitotic figures). All annotations are provided in MS COCO JSON format and as SQLITE database (SlideRunner format).</p>

opencc-by-4.0Apr 2022View details →
zenodo40/100

Dataset: Applications of Educational Data Mining and Learning Analytics on Data From Cybersecurity Training

<p>This repository contains supplementary materials for the following journal paper:</p> <p>Valdemar &Scaron;v&aacute;bensk&yacute;, Jan&nbsp;Vykopal, Pavel&nbsp;Čeleda, Lydia&nbsp;Kraus.<br> <em>Applications of Educational Data Mining and Learning Analytics on Data From Cybersecurity Training.</em><br> In Springer Education and Information Technologies. 2022.<br> <a href="https://doi.org/10.1007/s10639-022-11093-6">https://doi.org/10.1007/s10639-022-11093-6</a></p> <p>Preprint available at:&nbsp;<a href="https://arxiv.org/abs/2307.08582">https://arxiv.org/abs/2307.08582</a></p> <ul> </ul> <p><strong>How to cite</strong></p> <p>If you use or build upon the materials, please use the BibTeX entry below to cite the original paper (not only this web link).</p> <pre><code>@article{Svabensky2022applications, author = {\v{S}v\'{a}bensk\'{y}, Valdemar and Vykopal, Jan and \v{C}eleda, Pavel and Kraus, Lydia}, title = {{Applications of Educational Data Mining and Learning Analytics on Data From Cybersecurity Training}}, journal = {Education and Information Technologies}, publisher = {Springer}, volume = {27}, year = {2022}, issn = {1360-2357}, url = {https://doi.org/10.1007/s10639-022-11093-6}, doi = {10.1007/s10639-022-11093-6}, }</code></pre> <p><strong>Attached content</strong></p> <p>The files included in the ZIP archive are:</p> <ul> <li>`All-discovered-papers.bib` -- a BibTeX export of the Mendeley database of all considered papers discovered by the automated search.</li> <li>`Candidate-papers-reviewer1.bib` -- a BibTeX export of the Mendeley database of the candidate papers suggested by the first investigator.</li> <li>`Candidate-papers-reviewer2.bib` -- a BibTeX export of the Mendeley database of the candidate papers suggested by the second investigator.</li> <li>`Selected-papers.bib` -- a BibTeX export of the Mendeley database of the 35 papers selected for the literature review.</li> <li>`Selected-papers.xlsx` -- an Excel spreadsheet with the extracted information about the selected papers.</li> <li>`Selected-papers.csv` -- a CSV equivalent of the Excel spreadsheet.</li> </ul>

opencc-by-4.0May 2022View details →
zenodo40/100

Galaxy Training Data for "Evaluating and ranking a set of pathways based on multiple metrics"

<pre>This dataset provides the inputs needed for the Galaxy Pathway Analysis workflow training tutorial (<a href="https://galaxy-synbiocad.org">https://galaxy-synbiocad.org</a>). This workflow asseses the performance of predicted pathways by computing 4 criteria (target product flux, thermodynamic feasibility, pathway length, and enzyme availability). A score inform the user about the best candidate pathways to produce a compound of interest. The generated output is a collection of scored and ranked heterologous pathways. The content of the dataset is as follows: - A set of pathways provided in the SBML format (Systems Biology Markup Language) to be ranked, modeling heterologous pathways such as those outputted by the RetroSynthesis workflow (<a href="https://galaxy-synbiocad.org">https://galaxy-synbiocad.org</a>). - The GEM (Genome-scale metabolic models) which is a formalized representation of the metabolism of the host organism (the model is E. coli iML1515), provided in the SBML format.</pre>

opencc-by-4.0Jun 2022View details →
zenodo40/100

Data from Finnish Research Data Management Training Survey 2020-2021

<p>This is the survey data used in The Finnish Research Data Management Training Survey 2020-2021. The survey was sent to 74 Finnish research organizations of which 36 responded. The aim of the report was to gain a deeper understanding of what kind of research data management (RDM) training activities are provided by different Finnish organizations.</p> <p>&nbsp;</p>

opencc-by-4.0Jun 2022View details →
zenodo40/100

Training and validation data used to produce the pre-trained model for the TomoTwin paper.

<p>This datasets represents the training and validation data that was used to produce the pre-trained model for the TomoTwin paper. Please see 10.5281/zenodo.6637357 for the raw tomograms.</p>

opencc-by-4.0Jun 2022View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record