Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

558

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

558 results for “Training Data”

Learn how ShareScore rates datasets ↗
zenodo36/100

l-sized Training and Evaluation Data for Publication "Using Supervised Learning to Classify Metadata of Research Data by Field of Study"

<p>Automated classification of metadata of research data by their discipline(s) of research can be used in scientometric research, by repository service providers, and in the context of research data aggregation services. Openly available metadata of the DataCite index for research data were used to compile a large training and evaluation set comprised of 609,524 records. This is the cleaned and vectorized version with a feature selection of large size.</p>

opencc-by-4.0Oct 2019View details →
zenodo36/100

s-sized Training and Evaluation Data for Publication "Using Supervised Learning to Classify Metadata of Research Data by Field of Study"

<p>Automated classification of metadata of research data by their discipline(s) of research can be used in scientometric research, by repository service providers, and in the context of research data aggregation services. Openly available metadata of the DataCite index for research data were used to compile a large training and evaluation set comprised of 609,524 records. This is the cleaned and vectorized version with a feature selection of small size.</p>

opencc-by-4.0Oct 2019View details →
zenodo36/100

Training data for 'Beacon' tutorial (Galaxy Training Material)

<p>The data files are from the 1000 Genomes Project (1000HG) and GDC database. These datasets will be utilized in the Galaxy training session titled "Working with Beacon V2: A Comprehensive Guide to Creating, Uploading, and Searching for Variants with Beacons and Querying the University of Bradford GDC Beacon Database for Copy Number Variants (CNVs)." This training aims to equip participants with the skills necessary to construct Beacons, prepare and transform data into Beacon-compatible formats, seamlessly import data, and proficiently query Beacons for genetic variants. The provided data sets are integral for hands-on practice and will guide users through working with Beacon V2.</p>

opencc-by-4.0Jan 2024View details →
zenodo36/100

Training data composition determines machine learning generalization and biological rule discovery

<p>Github: https://github.com/csi-greifflab/negative-class-optimization</p> <p>Preprint: https://www.biorxiv.org/content/10.1101/2024.06.17.599333v1</p>

opencc-by-4.0May 2024View details →
zenodo36/100

Training and Testing Data for AP-SVM

<p>The files in here contain training and testing data for the AP-SVM data cleaning model, including datasets curated for leakage and sacrifice studies. Raw and digital signal processed files are included</p>

opencc-zeroSep 2024View details →
zenodo36/100

Cane Toad Acoustic Classifier Audio Training Data

<h3><strong>Cane Toad Audio Dataset for Machine Learning Classifier Development</strong></h3> <h3><strong>Description:</strong></h3> <p>This dataset was created as part of a study aimed at developing a machine learning classifier to detect the advertisement calls of the cane toad (<em>Rhinella marina</em>) using BirdNET. The dataset comprises 3-second audio snippets that capture a variety of sounds, including cane toad vocalizations, calls from spectrally overlapping species, environmental noises, and unidentified sounds. These labelled sound data were collected from various Australian Acoustic Observatory' recording sites, covering a broad range of geographic locations and environmental conditions in Australia.</p> <h3><strong>Sound Classes:</strong></h3> <ul> <li>Background</li> <li>Canis lupus dingo (Dingo)</li> <li>Centropus phasianinus (Pheasant Coucal)</li> <li>Cyclorana australis (Water Holding Frog)</li> <li>Cyclorana cryptotis (Hidden Ear Frog)</li> <li>Cyclorana novaehollandiae (New Holland Frog)</li> <li>Dacelo novaeguineae (Kookaburra)</li> <li>Ninox boobook (Southern Boobook)</li> <li>Notaden melanoscaphus (Northern Spadefoot Toad)</li> <li>Rhinella marina (Cane Toad) &ndash; Processed with a low-pass filter to remove frequencies above 1300 Hz, minimizing interference from co-occurring sounds.</li> <li>Unidentified Sounds</li> </ul> <h3><strong>Use and Applications:</strong></h3> <p>This dataset is valuable for training machine learning models focused on the acoustic detection of cane toads, could be useful for researchers and professionals working in bioacoustics, machine learning, ecological monitoring and invasive species management.</p>

opencc-by-4.0Sep 2024View details →
zenodo36/100

Training data for IoNNo model

Open the record for dataset details and reuse information.

opencc-by-4.0Sep 2024View details →
zenodo36/100

Data from P300-based Neurofeedback Training for Attention Enhancement with 4 EEG Electrodes

<p>The dataset contains EEG and behavioral data of 10 participants who completed 11 runs (i.e. copy-spelled 11 words) in a P300 speller task, as well as a random dot motion (RDM) task and questionnaires in a single experimental session. The data comes from a study that uses a modified version of the protocol described here:</p> <p>Noble SC,&nbsp;Woods E,&nbsp;Ward T,&nbsp;Ringwood JV. &ldquo;Adaptive P300-Based Brain-Computer Interface for Attention Training: Protocol for a Randomized Controlled Trial.&rdquo;&nbsp;<em>JMIR Res Protoc</em>&nbsp;2023, 12:e46135, doi:&nbsp;<a href="https://doi.org/10.2196/46135">10.2196/46135</a></p> <p>These are the differences to the protocol above:</p> <ul> <li>More but shorter words in the P300 speller (see below)</li> <li>Only 4 electrodes were used (Pz, POz, P7 and P8)</li> <li>Task difficulty adaptation is by iterative learning control (ILC) only</li> </ul> <p>Each participant folder contains:</p> <ul> <li>[xxx]-raw.[xxx] &ndash; unprocessed EEG signals from 4 electrodes for all 11 P300 speller runs in Openvibe (.ov) and Matlab (.mat) file formats, see details of the runs below</li> <li>classifier.cfg - LDA classifier weights</li> <li>log.txt - contains start and end time of the experiment, and performance in the P300 speller and RDM tasks</li> </ul> <p>The file &ldquo;Questionnaire scores.csv&rdquo; contains the responses to the questionnaire described in the experimental protocol and the NASA Task Load Index (TLX) for all participants.</p> <p>The .ov and .mat files contain data from the following runs:</p> <table> <tbody> <tr> <th>Filename</th> <th>Word to be copy-spelled</th> <th>Number of flashes per row and column</th> <th>Feedback given to participant</th> </tr> <tr> <td>calibration-signal1</td> <td>THE</td> <td>12</td> <td>no</td> </tr> <tr> <td>calibration-signal2</td> <td>QUICK</td> <td>12</td> <td>no</td> </tr> <tr> <td>calibration-signals</td> <td>Concatenation of calibration-signal1 and calibration-signal2</td> <td>&nbsp;</td> <td>&nbsp;</td> </tr> <tr> <td>eval</td> <td>DOG</td> <td>12</td> <td>yes</td> </tr> <tr> <td>training-run-1</td> <td>WIZARD</td> <td>10</td> <td>yes</td> </tr> <tr> <td>training-run-2</td> <td>HUMBLE</td> <td>varying</td> <td>yes</td> </tr> <tr> <td>training-run-3</td> <td>JOKERS</td> <td>varying</td> <td>yes</td> </tr> <tr> <td>training-run-4</td> <td>UNLOCK</td> <td>varying</td> <td>yes</td> </tr> <tr> <td>training-run-5</td> <td>THRIVE</td> <td>varying</td> <td>yes</td> </tr> <tr> <td>training-run-6</td> <td>JUNGLE</td> <td>varying</td> <td>yes</td> </tr> <tr> <td>training-run-7</td> <td>SHADOW</td> <td>varying</td> <td>yes</td> </tr> <tr> <td>training-run-8</td> <td>FROZEN</td> <td>varying</td> <td>yes</td> </tr> </tbody> </table> <p>&nbsp;</p> <p>This research is supported by the Irish Research Council under project ID GOIPG/2020/692 and Science Foundation Ireland under grant number 12/RC/2289_P2.</p>

opencc-by-4.0Sep 2024View details →
zenodo36/100

Weekly supervised Multilingual Data Set to train Named Entity Recognition for Symptom Extraction

<p>Data Sets were generated using the Weakly Supervised NER pipeline (https://github.com/HUMADEX/Weekly-Supervised-NER-pipline) to train the symptom extraction NER models.&nbsp;</p> <p><strong>Supported Languages and dataset locations for the specific language:</strong></p> <p>&nbsp; &nbsp; English (base language): https://huggingface.co/HUMADEX/english_medical_ner<br>&nbsp; &nbsp; German: https://huggingface.co/HUMADEX/german_medical_ner<br>&nbsp; &nbsp; Italian: https://huggingface.co/HUMADEX/italian_medical_ner<br>&nbsp; &nbsp; Spanish: https://huggingface.co/HUMADEX/spanish_medical_ner<br>&nbsp; &nbsp; Greek: https://huggingface.co/HUMADEX/german_medical_ner<br>&nbsp; &nbsp; Slovenian: https://huggingface.co/HUMADEX/slovenian_medical_ner<br>&nbsp; &nbsp; Polish: https://huggingface.co/HUMADEX/polish_medical_ner<br>&nbsp; &nbsp; Portuguese: https://huggingface.co/HUMADEX/portugese_medical_ner</p> <p>&nbsp;</p> <p><strong>Dataset Building&nbsp;</strong></p> <ul> <li>Data Integration and Preprocessing</li> <li>Data Cleaning</li> <li>Annotation with Stanza's i2b2 Clinical Model&nbsp;</li> <li>Translation into the targeted language</li> <li>Word Alignment&nbsp;</li> <li>Data Augmentation&nbsp;</li> </ul> <p><strong>Acknowledgement</strong><br>This dataset had been created as part of joint research of HUMADEX research group (https://www.linkedin.com/company/101563689/) and has received funding by the European Union Horizon Europe Research and Innovation Program project SMILE (grant number 101080923) and Marie Skłodowska-Curie Actions (MSCA) Doctoral Networks, project BosomShield ((rant number 101073222). Responsibility for the information and views expressed herein lies entirely with the authors.</p> <p><strong>Authors:</strong><br>dr. Izidor Mlakar, Rigona Sallauka, dr. Umut Arioz, dr. Matej Rojc</p> <p><strong>Please cite as:</strong></p> <p><span>Article title: Weakly-Supervised Multilingual Medical NER For Symptom Extraction For Low-Resource Languages</span><br><span>Doi: 10.20944/preprints202504.1356.v1</span><br><span>Website:&nbsp;</span><a title="https://www.preprints.org/manuscript/202504.1356/v1" href="https://www.preprints.org/manuscript/202504.1356/v1">https://www.preprints.org/manuscript/202504.1356/v1</a></p>

opencc-by-4.0Oct 2024View details →
zenodo36/100

Training data for "Identification of allelic variants in SARS-CoV-2 from deep sequencing reads"

<p>Effectively monitoring global infectious disease crises, such as the COVID-19 pandemic, requires capacity to generate and analyze large volumes of sequencing data in near real time. These data have proven essential for monitoring the emergence and spread of new variants, and for understanding the evolutionary dynamics of the virus.</p> <p>Two sequencing platforms in combination with several established library preparation strategies are predominantly used to generate SARS-CoV-2 sequence data. However, data alone do not equal knowledge: they need to be analyzed. The Galaxy community developed analysis workflows to support the <strong>identification of allelic variants (AVs) in SARS-CoV-2 from deep sequencing reads</strong>.</p> <p>These workflows allow one to identify AVs and lineages in SARS-CoV-2 genomes with variant allele frequencies ranging from 5% to 100% (i.e., they detect variants with intermediate frequencies as well.</p> <p>In this tutorial we will see how to run these workflows for the different types of input data:</p> <ul> <li>Single end data derived from Illumina-based RNAseq experiments</li> <li>Paired end data derived from Illumina-based RNAseq experiments</li> <li>Paired-end data generated with Illumina-based Ampliconic (ARTIC) protocols</li> <li>ONT fastq files generated with Oxford nanopore (ONT)-based Ampliconic (ARTIC) protocols</li> </ul> <p>To illustrate the tutorial, we took some example datasets (paired-end data generated with Illumina-based Ampliconic (ARTIC) protocols) from COG-UK, the COVID-19 Genomics UK Consortium.</p>

opencc-by-4.0Jun 2021View details →
zenodo36/100

Training data for the Sei framework sequence model

<p>Training data for the Sei framework deep learning model. The data contains chromatin profiles from the Cistrome Project: <strong>please agree to the terms of usage at the Cistrome Project (http://cistrome.org/db/#/bdown) before downloading.&nbsp;</strong></p>

opencc-by-4.0Jun 2021View details →
zenodo36/100

Training data for MaxQuant and Msstats TMT analysis in Galaxy

<p>The files serve as input and intermediate results for a MaxQuant and MsstatsTMT training on&nbsp;&nbsp;lysine methyl transferase 9 knockdown and control cell proteomics (https://doi.org/10.1186/s12935-020-1141-2) in the Galaxy training network (https://training.galaxyproject.org).</p> <p>Input files: human FASTA protein database for Maxquant. MaxQuant experimental design template, MSstatsTMT annotation file</p> <p>Intermediate result files: MaxQuant protein groups and evidence</p>

opencc-by-4.0Aug 2021View details →
zenodo36/100

Data manuscript: Preventive training does not interfere with mRNA-encoding myosin and collagen expression during pulmonary arterial hypertension

<p>Supporting Information files of manuscript:&nbsp;Preventive training does not interfere with mRNA-encoding myosin and collagen expression during pulmonary arterial hypertension&nbsp;</p> <p>&nbsp;</p> <p>The values behind the means, standard deviations and&nbsp;values used to build graphs;&nbsp;</p>

opencc-by-4.0Aug 2021View details →
zenodo36/100

Training data for MLP - HYDROsim project

<p>Initial release of&nbsp;full data set for training the Multi Layer Perceptron</p>

opencc-by-4.0Sep 2021View details →
zenodo36/100

Experimental data for "A Novel Clinical-Driven Design for Robotic Hand Rehabilitation: Combining Sensory Training, Effortless Setup and Large Range of Motion in a Palmar Device"

<p>Experimental data for the interaction force benchmark test of the PRIDE haptic hand rehabilitation device.</p>

opencc-by-4.0Sep 2021View details →
zenodo36/100

ChIP-seq data for training purposes

<p>Data for the EBAII training and concerning reads extracted from GSE40129 serie (<a href="https://www.ncbi.nlm.nih.gov/geo">NCBI GEO dataset</a>) by human chr11 mapping.</p> <p>siNT_ER_E2_r3_chr11.fastq.gz (from <a href="https://www.ncbi.nlm.nih.gov/geo/query/acc.cgi?acc=GSM986063">GSM986063</a> sample) : ChIP-seq</p> <p>MCF_input_r3_chr21.fastq.gz (from&nbsp;<a href="https://www.ncbi.nlm.nih.gov/geo/query/acc.cgi?acc=GSM986091">GSM986091</a> sample) : input</p> <p>&nbsp;</p> <p>&nbsp;</p>

opencc-by-4.0Oct 2021View details →
zenodo36/100

Experimental data for the study: "Hiding Assistive Robots During Training in Immersive VR Does not Affect Users' Motivation, Presence, Embodiment, and Performance"

<p>The datasets contains the motor performance metrics, the gaze fixation time ratios, and the questionnaire responses for a study involving a motor task with a rehabilitation assistive robot and an immersive virtual reality head-mounted display. The&nbsp;study was performed in the Motor Learning and Neurorehabilitation Laboratory at University of Bern. All data are stored in&nbsp;&ldquo;csv&rdquo; files. The variables inside the files are explained in &ldquo;DataFrameDescription.rtf&rdquo;. For questions, please contact nicolas.wenk@unibe.ch or L.MarchalCrespo@tudelft.nl.</p>

opencc-by-4.0Oct 2021View details →
zenodo36/100

Data: Acoustic disturbance in blue mussels: sound-induced valve closure varies with pulse train speed but does not affect phytoplankton clearance rate

<p>Data abstract:</p> <p>Data on&nbsp;mussels&#39; valve gape behaviour and phytoplankton clearance&nbsp;during sound exposure trials. We provide the raw data, processed data, scripts to process the raw data, make plots, and run the statistics.</p> <p>&nbsp;</p> <p>Paper abstract:</p> <p>Anthropogenic sound has increasingly become part of the marine soundscape and may negatively affect animals across all taxa. Invertebrates, including bivalves, received limited attention even though they make up a significant part of the marine biomass and are very important for higher trophic levels. Behavioural studies are critical to evaluate individual and potentially population-level impact of noise and can be used to compare the effects of different sounds. In the current study, we examined the effect of impulsive sounds with different pulse rates on the valve gape behaviour and phytoplankton clearance rate of blue mussels (<em>Mytilus</em> spp.<em>)</em>. We monitored the mussels&rsquo; valve gape using an electromagnetic valve gape monitor, and their clearance rate using spectrophotometry of phytoplankton densities in the water. We found that the mussels&rsquo; valve gape was positively correlated with their clearance rate, but the sound exposure did not significantly affect the clearance rate or reduce the valve gape of the mussels. They did close their valves upon the onset of a pulse train, but the majority of the individuals recovered to pre-exposure valve gape levels during the exposure. Individuals that were exposed to faster pulse trains returned to their baseline valve gape faster. Our results show that different sound exposures can affect animals differently, which should be taken into account for noise pollution impact assessments and mitigation measures.</p> <p>&nbsp;</p> <p>Paper reference:</p> <p>Hubert, J., Moens, R., Witbaard, R., Slabbekoorn, H. (2022).&nbsp;Acoustic disturbance in blue mussels: sound-induced valve closure varies with pulse train speed but does not affect phytoplankton clearance rate.&nbsp;<em>ICES&nbsp;Journal of Marine Science</em>.&nbsp;DOI:&nbsp;10.1093/icesjms/fsac193.</p>

opencc-by-4.0Oct 2022View details →
zenodo36/100

2D Synthetic Training Data For SyMBac

<p>Synthetic training datasets, used to train models to segment</p> <ul> <li><em>B. subtilis&nbsp;</em>growing in mother machine (100x oil, phase contrast)</li> <li><em>E. coli&nbsp;</em>growing on agar pads (100x oil, phase contrast)</li> <li><em>E. coli&nbsp;</em>streaked onto agar pads (60x air, fluorescence)</li> <li><em>E. coli&nbsp;</em>growing in a microfluidic turbidostat (100x oil, phase contrast)</li> </ul>

opencc-by-4.0May 2022View details →
zenodo36/100

Code and extensive data for training neural networks for radiation, used in "Implementation of a machine-learned gas optics parameterization in the ECMWF Integrated Forecasting System: RRTMGP-NN 2.0""

<p>Data and code used in a paper submitted to JAMES titled :<em>&nbsp;Implementation of a machine-learned gas optics parameterization in the ECMWF Integrated Forecasting System</em></p> <p>1) The files <strong>ml_training_*.7z</strong> contain extensive datasets (in NetCDF format) for training neural network versions of the RRTMGP gas optics scheme as described in the paper. The datasets are read by <a href="https://github.com/peterukk/rte-rrtmgp-nn/blob/main/examples/rrtmgp-nn-training/ml_train.py">ml_train.py.</a></p> <p>2) The ML datasets were in turn generated using the input profiles (in NetCDF format) inside <strong>inputs_to_RRTMGP.zip </strong>by running the Fortran programs <code>rrtmgp_sw_gendata_rfmipstyle.F90 and rrtmgp_lw_gendata_rfmipstyle.F90 </code>in <em>rte-rrtmgp-nn/examples/rrtmgp-nn-training</em>, which call the RRTMGP gas optics scheme, The input profiles contain <strong>millions of columns, hundreds of perturbation experiments (including hypercube-sampled gas concentrations), are derived from several different data sources (including CAMS reanalysis, GCM, and CKDMIP-MMM), and span present-day, preindustrial, and future atmospheric conditions.</strong> They could be used to generate training data for developing emulators of the full RTE+RRTMGP radiation scheme, not just gas optics (see nn_dev on the <a href="https://github.com/peterukk/rte-rrtmgp-nn">RTE+RRTMGP-NN repository on Github</a>, used in a previous paper where different emulation methods were compared)</p> <p>3) The Fortran and Python code used for data generation and NN training are found in<a href="https://github.com/peterukk/rte-rrtmgp-nn/tree/main/examples/rrtmgp-nn-training"> <em>rte-rrtmgp-nn/examples/rrtmgp-nn-training</em> </a>on the main branch on Github; <strong>an archived version is also included here </strong>(<strong>rte-rrtmgp-nn-2.0.zip</strong>). See the readme in the above sub-directory for further information.</p> <p>&nbsp;</p>

opencc-by-4.0May 2022View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record