Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
431
datasets available to search
ShareScore release 0.7.1
Dataset results
431 results for “Training Datasets”
Dataset for training the Surrogate Model of microlaser neurons on the reduced MNIST classification task
<p>This dataset was used to train a surrogate multilayer perceptron surrogate model of microlaser neurons.</p> <p>It is in csv format. It was generated using the Yamada Model as found in </p> <p><span>Selmi F, Braive R, Beaudoin G, Sagnes I, Kuszelewicz R and Barbay S 2014 Relative Refractory Period in an Excitable Semiconductor Laser <em>Phys. Rev. Lett.</em> <strong>112</strong> 183902</span>.</p>
A dataset recorded during development of an affective brain-computer music interface: training sessions
Open the record for dataset details and reuse information.
Smartbay Marine Species Object Detection Training dataset
<h1>Training dataset</h1> <p>The SmartBay Observatory in Galway Bay is an important contribution by Ireland to the growing global network of real-time data capture systems deployed within the ocean – technology giving us new insights into the ocean which we have not had before.</p> <p>The observatory was installed on the seafloor 1.5km off the coast of Spiddal, County Galway, Ireland . The observatory uses cameras, probes and sensors to permit continuous and remote live underwater monitoring. This observatory equipment allows ocean researchers unique real-time access to monitor ongoing changes in the marine environment. Data relating to the marine environment at the site is transferred in real-time from the SmartBay Observatory through a fibre optic telecommunications cable to the Marine Institute headquarters and onwards onto the internet. The data includes a live video stream, the depth of the observatory node, the sea temperature and salinity, and estimates of the chlorophyll and turbidity levels in the water which give an indication of the volume of phytoplankton and other particles, such as sediment, in the water.</p> <p>The Smartbay Marine Species Object Detection training Dataset is an initial Bounding Box Annotated image dataset used in attempting to Train a YOLOv8 Object Detection Model to classify the Marine Fauna observed in the Smartbay Observatory Video footage using species names.</p> <p>The imagery used in this training dataset consists of image frame captures from the <a href="https://smartbay.marine.ie">Smartbay</a> video Archive files, CC-BY imagery from the <a href="https://www.minka-sdg.org">www.minka-sdg.org</a> website and images taken by Eva Cullen in the "<a href="https://nationalaquarium.ie/">Galway Atlantaquaria</a>" Aquarium in Galway, Ireland.</p> <p>The imagery were annotated using CVAT, collated on <a href="https://www.roboflow.com/">Roboflow</a> and exported in YOLOv8 training dataset format. </p>
Smartbay Marine Types Object Detection Training dataset
<h1>Training Dataset</h1> <p>The SmartBay Observatory in Galway Bay is an important contribution by Ireland to the growing global network of real-time data capture systems deployed within the ocean – technology giving us new insights into the ocean which we have not had before.</p> <p>The observatory was installed on the seafloor 1.5km off the coast of Spiddal, County Galway, Ireland . The observatory uses cameras, probes and sensors to permit continuous and remote live underwater monitoring. This observatory equipment allows ocean researchers unique real-time access to monitor ongoing changes in the marine environment. Data relating to the marine environment at the site is transferred in real-time from the SmartBay Observatory through a fibre optic telecommunications cable to the Marine Institute headquarters and onwards onto the internet. The data includes a live video stream, the depth of the observatory node, the sea temperature and salinity, and estimates of the chlorophyll and turbidity levels in the water which give an indication of the volume of phytoplankton and other particles, such as sediment, in the water.</p> <p>The Smartbay Marine Types Object Detection training Dataset is an initial Bounding Box Annotated image dataset used in attempting to Train a YOLOv8 Object Detection Model to classify the Marine Fauna observed in the Smartbay Observatory Video footage using broad "Marine Type" classes.</p> <p>The imagery used in this training dataset consists of image frame captures from the <a href="https://smartbay.marine.ie">Smartbay</a> video Archive files, CC-BY imagery from the <a href="https://www.minka-sdg.org">www.minka-sdg.org</a> website and images taken by Eva Cullen in the "<a href="https://nationalaquarium.ie/">Galway Atlantaquaria</a>" Aquarium in Galway, Ireland.</p> <p>The imagery were annotated using CVAT, collated on <a href="https://www.roboflow.com/">Roboflow</a> and exported in YOLOv8 trainign dataset format. </p>
Training and test dataset of STED images of microtubules in fixed cells
<p>Training and test dataset of microtubule used in the manuscript "Denoising diffusion models for high-resolution microscopy image restoration".</p>
Super-resolving ocean dynamics from space with computer vision algorithms: training datasets
<p>We provide here the datasets used for the development of the dilated Adaptive Residual Network for the super-resolution of ocean Absolute Dynamic Topography described in <em>Buongiorno Nardelli et al.</em> (2022). The model is designed to combine satellite altimetry and thermal observations and provides super-resolved dynamic topography. The training/test datasets have been built starting from the data originally prepared for an Observing System Simulation Experiment carried out in the framework of the European Space Agency CIRCOL project [<em>Ciani et al.</em>, 2021]. They consist of one year of synthetic daily Absolute Dynamic Topography (ADT), surface geostrophic currents and sea surface temperature data obtained from Copernicus Marine Service Mediterranean Forecasting System (MFS) (Product ID: MEDSEA-ANALYSIS- FORECAST-PHY-006-013) [<em>Clementi et al. 2021</em>]. Synthetic Altimeter-derived ADT maps were obtained by first sampling the model output along the actual tracks of a synthetic constellation composed of 4 Radar Altimeters: Jason-3, Sentinel-3A, SARAL/Altika, and Cryosat-2 missions (this step is achieved by running the SWOT simulator software [<em>Gaultier et al.</em>, 2016]) and successively applying the DUACS (<em>Data Unification and Altimeter Combination System)</em> mapping method. The original input images cover the entire Mediterranean domain at 1/24° spatial resolution, leading to an individual image size of 380x1000 pixels. Here, we have randomly chosen 40 dates (~11% of the total) to be kept aside as fully independent test data, and successively re-sampled the original images extracting much smaller tiles (76x100), which are used as input to the network training. The tiles are extracted by going through a double loop on latitude and longitude, imposing a spatial overlap of 50%. Full details on data pre-processing (e.g.normalization strategies) are given in the paper:</p> <ul> <li>Buongiorno Nardelli, B.; Cavaliere, D.; Charles, E.; Ciani, D. Super-Resolving Ocean Dynamics from Space with Computer Vision Algorithms. <em>Remote Sens.</em>, <strong>2022</strong>, 14, 1159. https://doi.org/10.3390/rs14051159</li> </ul>
Training Datasets for Epilepsy Analysis: Preprocessing and Feature Extraction from EEG Time Series
<h2>The files include the 20 training datasets, in csv format, from 20 epileptic patients. Each set of data is described by 1080 features extracted using the sliding window technique.</h2>
Dataset for Training Material - Galaxy Workflow - Analyse unaligned ncRNAs
<p>Input dataset for Galaxy Training Material for the Analyze unaligned ncRNAs workflow.</p> <p>See https://github.com/galaxyproject/training-material for more information.</p>
Nephrops (Nephrops norvegicus) Burrow object detection simple training dataset from Irish Underwater TV surveys
<div> <div> <div> <div> <h1>Training dataset</h1> <p>Norway prawns (<em>Nephrops norvegicus</em>), also known as the Dublin Bay prawn, are common around the Irish coast. They are found in distinct sandy/muddy areas where the sediment is suitable for them to construct their burrows. <em>Nephrops </em>spend a great deal of time in their burrows and their emergence from these is related to time of year, light intensity and tidal strength. The Irish <em>Nephrops </em>fishery is extremely valuable with landings recently worth around €55m at first sale, supporting an important Irish fishing industry. </p> <p><em>Nephrops</em> are managed in Functional Units (FUs). The Marine Institute has conducted under water television surveys since 2002 to independently estimate abundance, distribution and stock sizes of <em>Nephrops</em> <em>norvegicus </em>for:</p> <ul> <li>Irish Sea <em>Nephrops</em> Grounds (FU 14 and 15) in collaboration with <a title="Link to 'Fisheries and Aquatic Ecosystems' work in AFBI Northern Ireland" href="https://www.afbini.gov.uk/area-of-expertise/fisheries-and-aquatic-ecosystems">AFBI</a> an <a title="Link to Cefas (the Centre for Environment, Fisheries, and Aquaculture Science) in the UK" href="https://www.cefas.co.uk/">CEFAS</a>.</li> <li>Porcupine Bank <em>Nephrops</em> Grounds (FU16)</li> <li>Aran, Galway Bay and Slyne Head <em>Nephrops</em> Grounds (FU17)</li> <li>South and South west Ireland <em>Nephrops</em> Grounds (FU19)</li> <li>Labadie, Jones and Cockburn <em>Nephrops</em> Grounds (FU20 and 21)</li> <li>“Smalls” <em>Nephrops</em> Grounds (FU22)</li> </ul> <p>Each year during the summer months, on average 300 stations are surveyed each year, in three survey legs, covering all the FUs in depths from 20 to 650 metres.</p> <p>A high definition camera system is towed over the sea bed for 10 minutes travelling approx. 200m at 0.8 knots on a purpose built sledge. The UWTV survey follows survey protocols available <a title="Link to survey protocols" href="https://doi.org/10.17895/ices.pub.8014">here</a> agreed by International Council for the Exploration of the Sea (ICES) Working Group on <em>Nephrops </em>surveys (WGNEPS). </p> <p>As part of the iMagine project a selection of images from the Underwater TV survey Functional Units were annotated with bounding boxes and labels in YOLOv8 format to train an YOLOv8 Object Detection Models. The training dataset is saved in YOLOv8 format. It is intended to train a YOLOv8 Nephrrops burrow object detection model to assess the utility of an Object Detection model is assisting Prawn Survey work in the semi automated annotation of prawn burrow imagery.</p> </div> </div> </div> </div>
AMFinder training dataset
<p>Soil fungi establish mutualistic interactions with the roots of most vascular land plants. Arbuscular mycorrhizal (AM) fungi are among the most extensively characterised mycobionts to date.</p> <p>The software <a href="https://github.com/SchornacklabSLCU/amfinder.git">AMFinder</a> allows for automatic computer vision-based identification and quantification of AM fungal colonisation and intraradical hyphal structures on ink-stained root images using convolutional neural networks.</p> <p><strong>This dataset contains ink-stained root images used for AMFinder training.</strong></p>
UDP Synthetic Dataset for training ML time series models
<p>The dataset available has been produced by the "Next-Generation IoT solutions for the universal supply chain" (iNGENIOUS) project’s consortium under EC grant agreement 957216, made publicly available as part of the Horizon 2020 Open Research Data Pilot (<a href="https://www.openaire.eu/what-is-the-open-research-data-pilot">ORD pilot</a>).<br> The European Commission is not liable for any use that may be made of the information contained herein.</p> <p>The available dataset is in csv format and contains synthetic data of UDP packets received and sent by a single User Plane Function (UPF) covering a span of 6 weeks. The format of the datafile is:</p> <ul> <li>index</li> <li>timestamp </li> <li>UDP packets_rcvd - Total number of UDP packets received</li> <li>UDP packets sent - Total number of UDP packets sent</li> </ul> <p>The simulation was performed based on behavior of UPF and 5GC Network functions inferred from stress tests performed in the iNGENIOUS project's Automated Robots with Heterogeneous Networks Use Case, as well as patterns in urban mobility taken from available UE datasets [NCS+19].</p> <p>More information on the iNGENIOUS project can be found on the project’s website: <a href="https://ingenious-iot.eu/">https://ingenious-iot.eu/</a></p> <p>[NCS+19] Noussan M, Carioni G, Sanvito FD, Colombo E. Urban Mobility Demand Profiles:<br> Time Series for Cars and Bike-Sharing Use as a Resource for Transport and Energy<br> Modeling. Data. 2019; 4(3):108. https://doi.org/10.3390/data4030108</p>
Pre-training with simulated ultrasound images for breast mass segmentation and classification - dataset
<p>Dataset assosiated with the MICCAI Workshop on Data Engineering in Medical Imaging paper: "Pre-training with Simulated Ultrasound Images for Breast Mass Segmentation and Classification"</p>
Dataset: Reinforcing Cybersecurity Hands-on Training With Adaptive Learning
<p>This repository contains supplementary materials for the following conference paper:<br> <br> Pavel Seda, Jan Vykopal, Valdemar Švábenský, Pavel Čeleda.<em><br> Reinforcing Cybersecurity Hands-on Training With Adaptive Learning. </em><br> In Proceedings of the 51st IEEE Frontiers in Education Conference (FIE 2021).<br> <a href="https://doi.org/10.1109/FIE49875.2021.9637252">https://doi.org/10.1109/FIE49875.2021.9637252</a><br> <br> Preprint available at: <a href="https://arxiv.org/abs/2201.01574">https://arxiv.org/abs/2201.01574</a></p> <p><strong>How to cite</strong></p> <p>If you use or build upon the materials, please use the BibTeX entry below to cite the original paper (not only this web link).</p> <p>Some of the linked repositories have their separate citation entry; please use that one as well, if possible.</p> <pre><code>@inproceedings{Seda2021reinforcing, author = {Seda, Pavel and Vykopal, Jan and \v{S}v\'{a}bensk\'{y}, Valdemar and \v{C}eleda, Pavel}, title = {{Reinforcing Cybersecurity Hands-on Training With Adaptive Learning}}, booktitle = {Proceedings of the 51st IEEE Frontiers in Education Conference}, series = {FIE '21}, location = {Lincoln, NE, USA}, publisher = {IEEE}, address = {New York, NY, USA}, month = {10}, year = {2021}, pages = {1--9}, numpages = {9}, isbn = {978-1-6654-3851-3}, url = {https://doi.org/10.1109/FIE49875.2021.9637252}, doi = {10.1109/FIE49875.2021.9637252}, }</code></pre> <p> </p>
Cu dataset – A copper ore labeled images dataset for segmentation training and testing
<p>This dataset is composed of 121 pairs of correlated images. Each pair contains one image of a copper ore sample acquired through reflected light microscopy (RGB, 24-bit), and the corresponding binary reference image (8-bit), in which the pixels are labeled as belonging to one of two classes: ore (0) or embedding resin (255).</p> <p>The sample came from a copper ore from Yauri Cusco (Peru) with a complex mineralogy, mainly composed of sulfides, oxides, silicates, and native copper. It was classified by size. The fraction +74-100 μm was cold mounted with epoxy resin and subsequently ground and polished.</p> <p>Correlative microscopy was employed for image acquisition. Thus, 121 fields were imaged on a reflected light microscope with a 20× (NA 0.40) objective lens and on a scanning electron microscope (SEM). In sequence, they were registered, resulting in images of 1017×753 pixels with a resolution of 0.53 µm/pixel. As matter of fact, some images (the images No. 2, 3, 24, 25, 46, 47, 69, 91, and 113) have slightly smaller sizes because they were cropped during the registration procedure to correct co-localization errors of the order of a few pixels. Finally, the images from SEM were thresholded to generate the reference images.</p> <p>Further description of this sample and its imaging procedure can be found in the work by Gomes and Paciornik (2012).</p> <p>This dataset was created for developing and testing deep learning models on semantic segmentation tasks. The paper of Filippo et al. (2021) presented a variant of the DeepLabv3+ model (Chen et al., 2018) that reached mean values of 90.56% and 92.12% for overall accuracy and F1 score, respectively, for 5 rounds of experiments (training and testing), each with a different, random initialization of network weights.</p> <p>For further questions and suggestions, please do not hesitate to contact us.</p> <p> </p> <p><strong>Contact email</strong>: ogomes@gmail.com</p> <p> </p> <p>If you use this dataset in your own work, please cite this DOI: 10.5281/zenodo.5020566</p> <p> </p> <p>Please also cite this paper, which provides additional details about the dataset:</p> <p>Michel Pedro Filippo, Otávio da Fonseca Martins Gomes, Gilson Alexandre Ostwald Pedro da Costa, Guilherme Lucio Abelha Mota. <em>Deep learning semantic segmentation of opaque and non-opaque minerals from epoxy resin in reflected light microscopy images</em>. <strong>Minerals Engineering</strong>, Volume 170, 2021, 107007, https://doi.org/10.1016/j.mineng.2021.107007.</p>
FoldingDiff CATH S40 training dataset
<p>Dataset used to develop and train FoldingDiff, a generative model for protein backbone structures. </p>
Training and test datasets for the PredictONCO tool
<p>This dataset was used for training and validating the <a href="https://loschmidt.chemi.muni.cz/predictonco/">PredictONCO </a>web tool, supporting decision-making in precision oncology by extending the bioinformatics predictions with advanced computing and machine learning. The dataset consists of 1073 single-point mutants of 42 proteins, whose effect was classified as Oncogenic (509 data points) and Benign (564 data points). All mutations were annotated with a clinically verified effect and were compiled from the ClinVar and OncoKB databases. The dataset was manually curated based on the available information in other precision oncology databases (The Clinical Knowledgebase by The Jackson Laboratory, Personalized Cancer Therapy Knowledge Base by MD Anderson Cancer Center, cBioPortal, DoCM database) or in the primary literature. To create the dataset, we also removed any possible overlaps with the data points used in the PredictSNP consensus predictor and its constituents. This was implemented to avoid any test set data leakage due to using the PredictSNP score as one of the features (see below).</p> <p>The entire dataset (<strong>SEQ</strong>) was further annotated by the pipeline of PredictONCO. Briefly, the following six features were calculated regardless of the structural information available: essentiality of the mutated residue (yes/no), the conservation of the position (the conservation grade and score), the domain where the mutation is located (cytoplasmic, extracellular, transmembrane, other), the PredictSNP score, and the number of essential residues in the protein. For approximately half of the data (<strong>STR</strong>: 377 and 76 oncogenic and benign data points, respectively), the structural information was available, and six more features were calculated: FoldX and Rosetta ddg_monomer scores, whether the residue is in the catalytic pocket (identification of residues forming the ligand-binding pocket was obtained from P2Rank), and the pKa changes (the minimum and maximum changes as well as the number of essential residues whose pKa was changed – all values obtained from PROPKA3). For both <strong>STR </strong>and <strong>SEQ </strong>datasets, 20% of the data was held out for testing. The data split was implemented at the position level to ensure that no position from the test data subset appears in the training data subset. </p> <p>For more details about the tool, please visit the <a href="https://loschmidt.chemi.muni.cz/predictonco/help">help page</a> or <a href="https://loschmidt.chemi.muni.cz/peg/contact/">get in touch with us</a>.</p> <p>14-Dec-2023 update: the file with features<em> PredictONCO-features.txt</em> now includes UniProt IDs, transcripts, PDB codes, and mutations.</p>
Dataset: The effects of class balance on the training energy consumption of logistic regression models
<p>Two synthetic datasets for binary classification, generated with the Random Radial Basis Function generator from WEKA. They are the same shape and size (104.952 instances, 185 attributes), but the "balanced" dataset has 52,13% of its instances belonging to class c0, while the "unbalanced" one only has 4,04% of its instances belonging to class c0. Therefore, this set of datasets is primarily meant to study how class balance influences the behaviour of a machine learning model.</p>
XR training and gameplay 6DoF mobility dataset
<p>User mobility in extended reality (XR) can have a major impact on millimeter-wave (mmWave) links and may require dedicated mitigation strategies to ensure reliable connections and avoid service outages. The available prior art has predominantly focused on XR applications with constrained user mobility and limited impact on mmWave channels.</p> <p>We have performed dedicated experiments to extend the characterisation of relevant future XR use cases featuring a high degree of user mobility. To this end, we have carried out a tailor-made XR mobility measurement campaign, capturing the movement of the head, hands, and body in 6DoF. </p> <p>For a usage example, see the provided Jupyter Notebook in cacerumd-usage-example.zip.</p> <p>A description of the measurement campaign and a characterisation of the recorded mobility can be found in the corresponding <a href="https://ieeexplore.ieee.org/abstract/document/10634047">IEEE Magazine paper</a> or a longer version, with more details about the experiment, on <a href="https://arxiv.org/abs/2407.02636">Arxiv</a>.</p>
Data-driven physics-based modeling of pedestrian dynamics - dataset: Pedestrian trajectories at Eindhoven train station
<p>Pedestrian trajectories measured at train station Eindhoven Centraal (the Netherlands) on platform 2 with acces to tracks 3 and 4.</p> <p>The dataset is partitioned in files containing 10 consecutive days each, recording 4 data fields:</p> <ul> <li><strong>time_ms:</strong> Passed time since start of the measurements. Unit: milliseconds.</li> <li><strong>object_identifier:</strong> unique id identifying an object.</li> <li><strong>x_position_mm: </strong>coordinates of the object along the x-axis at the given time. Unit: millimeters.</li> <li><strong>y_position_mm:</strong> coordinates of the object along the y-axis at the given time. Unit: millimeters.</li> </ul> <p>Each object resembles a pedestrian on the train platform recorded with 10 frames per second. We deliberately removed exact date and time information for privacy reasons (see additional note). The data set consists of 60 consecutive days starting at an unkown time between 00:00 AM and 01:00 AM of a random date between April 1st and May 1st 2022. An overhead image of the platform is included showing train track 3 in the bottom and train track 4 in the top of the image.</p> <p>The data set is supplemented to the paper <a title="Data-driven physics-based modeling of pedestrian dynamics" href="https://doi.org/10.48550/arXiv.2407.20794" target="_blank" rel="noopener">Data-driven physics-based modeling of pedestrian dynamics</a> and can be processed by the associated <a title="Software: Data-driven physics-based modeling of pedestrian dynamics" href="https://github.com/c-pouw/physics-based-pedestrian-modeling" target="_blank" rel="noopener">Python implementation</a> to create pedestrian models. </p>
Concentrating solar power (CSP) plants AI-training dataset for flux density measurements.
<p>In this dataset, the tools required for the training of a neural net in the context of flux density measurements in concentrating solar power (CSP) plants are included. An Excel file with 931 meteorological conditions and the positions of the power plant and the receiver is included, as well as 15928 pairs of images resulting from ray-tracing in Solarturm Juelich (STJ) each of these conditions with 17 different combinations of heliostats. <br> <br>This dataset is part of the WP1 of TOPCSP european project (funded by HORIZON MSCA Doctoral Network, Project number 101072537).</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.