Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

11,687

datasets available to search

ShareScore release 0.7.1

Reset

Dataset results

11,687 results for “training”

Learn how ShareScore rates datasets ↗
zenodo40/100

Doctoral Studies as part of an Innovative Training Network (ITN): Early Stage Researcher (ESR) experiences - supplemental material & data

<p>Table and data repository for the manuscript &quot;Doctoral Studies as part of an Innovative Training Network (ITN): Early Stage Researcher (ESR) experiences&quot;</p> <p><strong>Supplemental Tables:</strong></p> <ul> <li>table1_ESIT Project Table</li> <li>table2_TIN-ACT Project Table</li> <li>table3_ITN_tinnitus</li> <li>table4_ITN_other</li> <li>table5_Individual PhDs</li> </ul> <p><strong>Individual-level and de-identified survey data (raw data):</strong></p> <ul> <li>raw_data_ITN_tinnitus (survey results from PhDs as part of an ITN with a focus on tinnitus)</li> <li>raw_data_ITN_other (survey results from PhDs associated to ITNs with another focus)</li> <li>raw_data_Individual_Phds (survey results from PhDs not part of an ITN)</li> </ul>

opencc-by-4.0Jan 2021View details →
zenodo40/100

The influence of art expertise and training on emotion and preference ratings for representational and abstract artworks.

<p>Across cultures and throughout recorded history, humans have produced visual art. This raises the question of why people report such an emotional response to artworks and find some works more beautiful or compelling than others. In the current study we investigated the interplay between art expertise, and emotional and preference judgments. Sixty participants (40 novices, 20 art experts) rated a set of 150 abstract artworks and portraits during two occasions: in a laboratory setting and in a museum. Before commencing their second session, half of the art novices received a brief training on stylistic and art historical aspects of abstract art and portraiture. Results showed that art experts rated the artworks higher than novices on aesthetic facets (beauty and wanting), but no group differences were observed on affective evaluations (valence and arousal). The training session made a small effect on ratings of preference compared to the non-trained group of novices. Overall, these findings are consistent with the idea that affective components of art appreciation are less driven by expertise and largely consistent across observers, while more cognitive aspects of aesthetic viewing depend on viewer characteristics such as art expertise.</p>

opencc-zeroAug 2015View details →
zenodo40/100

H2M survey data on commercialisation training needs of Health Researchers

<p>Health-2-Market was a 3-year long Coordination and Support Action, funded by the European Union&rsquo;s Seventh Framework Programme for research, technological development and demonstration (Grant Agreement No 305532). H2M aimed at providing training and individual support to Health / Life Sciences researchers in the process of translating their research results into successful new business ideas.</p> <p>With a view to properly adapting the training offer of the project to the needs of Health / Life Sciences researchers in terms of entrepreneurship and business skill development a Training Needs Analysis (TNA) was conducted. In this context, H2M launched an online survey targeted at Health / Life Sciences researchers who have been involved in EU health projects. In particular, the objectives of the survey were:</p> <ul> <li>To formulate&nbsp; a&nbsp; descriptive&nbsp; understanding&nbsp; of various&nbsp; aspects&nbsp; of&nbsp; commercialisation&nbsp; and&nbsp; training&nbsp; needs&nbsp; of &nbsp;the main target group of the project;</li> <li>To divide this target group into homogeneous sub-groups (clusters) along a number of key characteristics such as demographics, commercialisation attitudes and needs;</li> <li>To understand preferences and importance of different aspects and needs through the analysis of: <ul> <li>Knowledge&nbsp; areas&nbsp; that&nbsp; can&nbsp; influence&nbsp; commercialisation behaviour;</li> <li>Training modalities&nbsp; that&nbsp; have&nbsp; an&nbsp; effect&nbsp; on&nbsp; the&nbsp; intention&nbsp; to participate and /or on the perception of the usefulness of a commercialisation training;</li> <li>Variations identified over different sub-groups.</li> </ul> </li> </ul> <p>The survey was dispatched to a database composed of 7,991 unique contacts of participants in previous health projects, accessed through the Directorate General for Health and Food Safety of the European Commission. The initial aim of at least 50 complete responses was overwhelmingly surpassed: 637 respondents completed the survey in full.</p> <p>The &ldquo;H2M survey data on commercialisation training needs of Health Researchers&rdquo; dataset contains the raw, anonymised data that were collected from these respondents, along with the questionnaire items that were utilised.</p>

opencc-by-nc-4.0Sep 2015View details →
zenodo40/100

Survey data for "Remote Sensing & GIS Training in Ecology and Conservation"

<p>This file provides the raw data of an online survey intended at gathering information regarding remote sensing (RS) and Geographical Information Systems (GIS) for conservation in academic education. The aim was to unfold best practices as well as gaps in teaching methods of remote sensing/GIS, and to help inform how these may be adapted and improved. A total of 73 people answered the survey, which was distributed through closed mailing lists of universities and conservation groups.</p>

opencc-zeroApr 2016View details →
zenodo40/100

Training CNNs with Low-Rank Filters for Efficient Image Classification: Trained Models

<p>Models from experiments referenced in the paper &quot;Training CNNs with Low-Rank Filters for Efficient Image Classification&quot;,&nbsp;https://arxiv.org/abs/1511.06744</p> <p>Model names differ from those in the paper, but the csv files for each set of experiments relates the paper&#39;s name for the model and the real name of the model here:</p> <ul> <li>cifarma.csv: Network-in-Network CIFAR10 Models</li> <li>mitma.csv: MIT Places Models</li> <li>googlenetma.csv: GoogLeNet ILSVRC2012 Models</li> <li>vggma.csv: VGG-11 ILSVRC2012 Models</li> </ul> <p>&nbsp;</p> <p>&nbsp;</p>

opencc-by-nc-4.0May 2016View details →
zenodo40/100

Train-B dataset for ICDAR2017 Competition on Handwritten Text Recognition on the READ Dataset (ICDAR2017 HTR). Batch 1 and Batch 2.

<p>Train-B Dataset.   Dataset of pages without any layout or text line information. The corresponding transcripts are provided at page level with line breaks. It has 10k pages, though for convenience it is divided into two 5k page batches. This information is provided in PAGE format. </p> <p>This dataset is complementary to this other dataset:</p> <p>https://zenodo.org/record/439807#.WOIBZ3WLSkA</p> <p>More information at:</p> <p>https://scriptnet.iit.demokritos.gr/competitions/~icdar2017htr/</p> <p> </p>

opencc-by-4.0Apr 2017View details →
zenodo40/100

Train-A dataset for ICDAR2017 Competition on Handwritten Text Recognition on the READ Dataset (ICDAR2017 HTR)

<p>Train-A Dataset of pages with manually revised baselines and the corresponding transcripts associated to them. This batch is small, 50 pages. Please, keep in mind that only the baselines have been manually corrected, The polygons associated to each line have not been manually reviewed. </p> <p>This dataset is complementary to this other dataset:</p> <p>https://zenodo.org/record/439811#.WOIF9HWLSkA</p> <p>More information at:</p> <p>https://scriptnet.iit.demokritos.gr/competitions/~icdar2017htr/</p>

opencc-by-4.0Apr 2017View details →
zenodo40/100

A rule based Tibetan part-of-speech (POS) tagger for the creation of gold standard training data

<p>This rule based Tibetan part-of-speech (POS) tagger was prepared in the course of the research project &#39;Tibetan in Digital Communication&#39; (2012-2015) hosted at SOAS, University of London and funded by the UK&#39;s Arts and Humanities Research Council (grant code: AH/J00152X/1). For a description of the tag set see Garrett et al. 2014. and Garrett et al. 2015. For a description of the tagger itself see Garrett et al. 2014. Note that the tagger must be used together with a lexicon (for example Hill &amp; Garrett 2017a). One must use one&#39;s own script to tag all words with all tags in the lexicon and then apply the tagger to remove incorrect tags.</p> <p>On the associated corpus of 318,230 words (Hill &amp; Garrett 2017b) the lexical tagger (i.e. simply applying all available tags to all words) tags 141,911 words with the correct unique tag, achieves as accuracy of 1.000 (by definition getting the right tag among others for each word) with an ambiguity of 2.73111. In contrast, the Rule Tagger tags 241,256 words with the correct unique tag, achieves an accuracy of 0.99893 and an ambiguity of 1.38577.</p> <p>Because this tagger does not achieve ambiguity 1.000 it is not suitable for tagging large scale corpora, but instead is useful for the creation of gold standard training data.</p> <p>N.B. In some rare cases the tagger removes all POS-tags.</p>

opencc-by-4.0May 2017View details →
zenodo40/100

Bacterial training dataset for Galaxy training network tutorials on Genome assembly

<p>This training dataset is from an imaginary <em>Staphylococcus aureus</em> bacterium with a miniature genome. There is a reference genome in various formats as well as some fastq reads of a closely related but also imaginary mutant strain.</p> <p>It is a useful dataset for demonstrating:</p> <ul> <li>de novo genome assembly</li> <li>read mapping and variant calling</li> <li>genome annotation</li> </ul> <p>The files included are:</p> <ul> <li><strong>wildtype.fna</strong>: the reference genome sequence of the wildtype strain in fasta format (a header line, then the nucleotide sequence of the genome.)</li> <li><strong>wildtype.gff</strong>: the reference genome sequence of the wildtype strain in general feature format (a list of features - one feature per line, then the nucleotide sequence of the genome.)</li> <li><strong>wildtype.gbk</strong>: the reference genome sequence in genbank format.</li> <li><strong>mutant_R1.fastq</strong> and <strong>mutant_R2.fastq</strong>: Fastq sequence reads of a closely related mutant strain. <ul> <li>The reads are paired-end.</li> <li>Each read is 150 bases long.</li> <li>The number of bases sequenced is equivalent to 19x the genome sequence of the wildtype strain. (Read coverage 19x - rather low!).</li> </ul> </li> </ul>

opencc-by-4.0May 2017View details →
zenodo40/100

Training material for small RNA-seq data analysis (Galaxy Training Network tutorial)

<p>The data provided here are part of a Galaxy Training Network tutorial that analyzes small RNA-seq (sRNA-seq) data from a study published by Harrington et al. (DOI:10.1186/s12864-017-3692-8) to detect differential abundance of various classes of endogenous short interfering RNAs (esiRNAs). The goal of this study was to investigate "connections between differential retroTn and hp-derived esiRNA processing and cellular location, and to investigate the potential link between mRNA 3’ end cleavage and esiRNA biogenesis." To this end, sRNA-seq libraries were constructed from triplicate <em>Drosophila</em> tissue culture samples under conditions of either control RNAi or RNAi knockdown of a factor involved in mRNA 3’ end processing, <em>Symplekin</em>. This dataset (GEO Accession: GSE82128) consists of single-end, size-selected, non-rRNA-depleted sRNA-seq libraries. Because of the long processing time for the large original files, we have downsampled the original raw data files to include only reads that align to a subset of interesting transcript features including: (1) transposable elements, (2) <em>Drosophila</em> piRNA clusters, (3) <em>Symplekin</em>, and (4) genes encoding mass spectrometry-defined protein binding partners of <em>Symplekin</em> from Additional File 2 in the indicated paper by Harrington et al. More details on features 1 and 2 can be found here: https://github.com/bowhan/piPipes/blob/master/common/dm3/genomic_features (piRNA_Cluster, Trn). All features are from the <em>Drosophila</em> genome Apr. 2006 (BDGP R5/<em>dm3</em>) release.</p>

opencc-by-4.0Jul 2017View details →
zenodo40/100

Datasets for Training and Inference of DeepRLI

<p>This repository contains datasets used in the development of the <a href="https://github.com/fairydance/DeepRLI" target="_blank" rel="noopener">DeepRLI</a> model for protein&ndash;ligand interaction prediction, which includes the training dataset for the model and data related to the PLK1 kinase involved in the case study.</p>

opencc-by-4.0May 2024View details →
zenodo40/100

Dataset for Review and Agenda of Digital Forensics Education and Training

<p>This repository has four .bib files encompassing 49 primary study entries, and one .CSV file with data extracted from such studies.</p> <p><strong>The authors gratefully acknowledge the support of the Technology in Forensic Sciences project (Instituto Nacional de Ci&ecirc;ncia e Tecnologia em Ci&ecirc;ncias Forenses - \textbf{INCT Forense}, Grant \#465450/2014-8) for funding this work.&nbsp;</strong></p> <p><strong>We also thank the \textbf{Arauc&aacute;ria Funding Agency of Paran&aacute;} for their financial and institutional support, as well as \textbf{NAPI - Public Security and Forensic Science} for research funding (Grant \#22.632.926-9). &nbsp;</strong></p> <p><strong>This work is also supported by CAPES Pro-Defesa (Grant \# V3084362P).</strong></p> <p><strong>Avelino Zorzo thanks \textbf{CNPq/Brazil} Grant \#306250/2021-7.&nbsp;</strong></p> <p><strong>Edson OliveiraJr thanks \textbf{CNPq/Brazil} Grant \#311503/2022-5.&nbsp;</strong></p>

opencc-by-4.0Jul 2024View details →
zenodo40/100

Training data for MaxQuant and Msstats label-free analysis in Galaxy

<p>The files serve as input and intermediate results for a MaxQuant and Msstats training on skin cancer tissues (<a href="https://doi.org/10.1016/j.matbio.2017.11.004">https://doi.org/10.1016/j.matbio.2017.11.004</a>) in the Galaxy training network (https://training.galaxyproject.org).</p> <p>Input files: human FASTA database for Maxquant. Annotation file and comparison matrix file for Msstats.</p> <p>Intermediate result files: MaxQuant protein groups, evidence and PTXQC.</p>

opencc-by-4.0Feb 2021View details →
zenodo40/100

HYPERCOG – training material dataset

<p>The Hypercog training pack is aimed at teaching cyberphysical systems to two different types of profiles :&nbsp;</p><p>1) the operators at the 3 industrial sites impacted by the implementation of the HyperCOG solution (Solvay, Sidenor, Çimsa);</p><p>&nbsp;2) Master students who would like to get familiar with cyberphysical systems.</p><p>The main objective for operators training was to understand the human environment where the solutions are implemented. For students it was to identify the skills needed to become familiar with cyber-physical systems. The ultimate goal was to train operators to overcome the skill gap, so they can use HyperCOG solutions in their daily work, and to produce training materials on cyber-physical systems based on the Hypercog experience, aimed at Master students .</p>

opencc-by-4.0Sep 2023View details →
zenodo40/100

Xception trained model for classifying large ornithopod dinosaur footprints

<p>PLOS ONE: Classification of large ornithopod dinosaur footprints using Xception transfer learning</p><p>The trained model using Xception transfer learning, provided in https://github.com/CNUGeophysics/Xception_ornithopod.git</p>

opencc-by-4.0Oct 2023View details →
zenodo40/100

Survey on the implementation of digital technology architecture in teacher training

<p>This file contains the results of a detailed survey exploring the implementation of digital technology architecture in teacher education. The survey was designed to assess how educational institutions are integrating advanced digital tools into their professional development programmes and the impact this has on teaching and learning. The data collected provides valuable insight into the effectiveness of these technologies in the educational environment, areas of success and challenges that still need to be addressed to optimise educator training in the digital age.</p>

opencc-by-4.0Nov 2023View details →
zenodo40/100

SURFBIO Training: "Analytical methods for the study of microbial cell-Surface and Surface-colloid interactions" (2021)

<p>2021. SURFBIO project training within WP1.</p><p>Content:</p><ul><li><strong>Webinar on Vertical scanning interferometry: a microscopic technique to analyze surface reactivity,&nbsp;</strong>by Dr. Cornelius Fischer (HZDR, Germany)&nbsp;</li><li><strong>Webinar on An introduction to radiolabelling as a versatile tool in colloid tracing, </strong>by Stefan Schymura (HZDR, Germany).</li><li><strong>Webinar on Development and construction of biocarriers and aggregates for potential industrial applications</strong>, by Dr. Andre Skirtach and Dr. Bogdan Parakhonskiy (GHENT University).&nbsp;</li><li><strong>Materials and fluidic design to study artificially functionalized microorganisms</strong>, by Dr. Andre Skirtach and Dr. Bogdan Parakhonskiy (GHENT University).&nbsp;</li></ul>

opencc-by-4.0Nov 2021View details →
zenodo40/100

Thai Word Embeddings (word2vec) Trained on Oscar Corpus

<p>A large Thai word2vec model trained on Oscar corpus and tokenized and normalized with&nbsp;PyThaiNLP. The model can be loaded using gensim, it is saved in binary format.&nbsp;</p>

opencc-by-4.0Apr 2023View details →
zenodo40/100

Training and test data, plus saved models for the upcoming paper `Top-down perceptual inference shaping the activity of early visual cortex'

<p>Each .pkl&nbsp;file contains a training or test dataset&nbsp;in the form of a Python dictionary (generated with Python 3.8.5) with the following fields:</p><ul><li>'train_images': 640,000 float32 images&nbsp;used&nbsp;for model training. These are 40px images that contain 1600 pixel intensities each.</li><li>'train_labels': float32 labels for each image in&nbsp;'train_images'. All natural images are&nbsp;labeled&nbsp;with 0.0. Texture images are labeled with 0.0, 1,0, 2.0, 3.0, or 4.0,&nbsp;according to their texture family.</li><li>'test_images': 64,000 float32 images&nbsp;used&nbsp;for model testing.&nbsp;These are 40px images that contain 1600 pixel intensities each.</li><li>'test_labels': float32 labels for each image in&nbsp;'test_images'. All natural images are&nbsp;labeled&nbsp;with 0.0. Texture images are labeled with 0.0, 1,0, 2.0, 3.0, or 4.0,&nbsp;according to their texture family.</li></ul><p>The .zip file contains a saved model snapshot and various intermediate evaluative data.&nbsp;Details on these are coming soon.</p>

opencc-by-4.0Nov 2023View details →
zenodo40/100

Dataset: Behavior of Participants in Hands-on Cybersecurity Training Suitable for Process Mining

<p>This repository contains supplementary materials for the following journal paper:</p> <p>Radek O&scaron;lej&scaron;ek, Martin Mac&aacute;k, Karol&iacute;na Dočkalov&aacute; Bursk&aacute;.<br><em>Hands-on cybersecurity training behavior data for process mining.</em><br>In Elsevier Data in Brief. 2023.<br>Available as open-access article on&nbsp;<a href="https://doi.org/10.1016/j.dib.2023.109956">https://doi.org/10.1016/j.dib.2023.109956</a></p> <p><strong>Contents</strong></p> <p>Datasets store event logs of trainees participating in hands-on cybersecurity exercises organized in the&nbsp;<a href="https://www.kypo.cz">KYPO Cyber Range</a>. The data includes training scenarios (expected behavior), raw event logs in the JSON format, and aggregated behavioral data suitable for process mining analysis.</p> <ol> <li><strong>Data1:</strong> A dataset of 52 trainees participating in the <a href="https://gitlab.ics.muni.cz/muni-kypo-trainings/games/locust-3302">Locust 3302</a> exercise adapted an insider attack scenario. No time restrictions were posed on playtime. The data file is structured as follows: <ul> <li>training_definition.json: The exercise content &ndash; cybersecurity tasks and hints. The training is based on the <a href="https://gitlab.ics.muni.cz/muni-kypo-trainings/games/locust-3302">Locust 3302</a> game adapted to an insider attack scenario.</li> <li>training_events: Recorded progress of trainees within the exercise, i.e., the status of completing tasks.</li> <li>command_histories: Recorded commands executed on network hosts.</li> <li>process_mining.csv: Complete PM-ready dataset suitable for process discovery or conformance analysis.</li> <li>process_mining_simplified.csv : Reduced PM-ready dataset with semantically identical events being removed.</li> </ul> </li> <li><strong>Data2:</strong> A dataset of 48 trainees participating in the original <a href="https://gitlab.ics.muni.cz/muni-kypo-trainings/games/locust-3302">Locust 3302</a> exercise. Three supervised training sessions were restricted to two hours of playtime. The structure follows the structure of Data1.</li> <li><strong>Tool:</strong> A Java application used to aggregate raw JSON data and transform them into a CSV format suitable for process mining techniques.</li> </ol> <p><strong>How to cite</strong></p> <p>If you use or build upon the materials, please use the BibTeX entry below to cite the original work.</p> <pre><code>@article{Oslejsek2023dataset, &nbsp; &nbsp; author = {Radek O\v{s}lej\v{s}ek and Martin Mac\'{a}k and Karol\'{i}na {Do\v{c}kalov\'{a} Bursk\'{a}}}, &nbsp; &nbsp; title = {Hands-on cybersecurity training behavior data for process mining}, &nbsp; &nbsp; journal = {{Data in Brief}}, publisher = {Elsevier}, issn = {2352-3409}, year = {2023}, volume = {52}, doi = {10.1016/j.dib.2023.109956}, url = {https://www.sciencedirect.com/science/article/pii/S2352340923009873} }</code></pre>

opencc-by-4.0Nov 2023View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record