Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

11,687

datasets available to search

ShareScore release 0.7.1

Reset

Dataset results

11,687 results for “training”

Learn how ShareScore rates datasets ↗
zenodo44/100

NewsEye / READ AS training dataset from French Newspapers (19th, early 20th C.)

<p>The dataset comprises French newspaper pages from 19th and early 20th century with annotated text. The page images were provided by the&nbsp;<a href="https://www.bnf.fr/en">French National Library</a> and comprise 183 pages (training set). The data are formed according to the PAGE format (cf.&nbsp;Cf.&nbsp;<a href="https://github.com/PRImA-Research-Lab/PAGE-XML/">https://github.com/PRImA-Research-Lab/PAGE-XML/</a>) and were produced with the <a href="http://read.transkribus.eu/">Transkribus </a>platform with support of the <a href="http://newseye.eu/">NewsEye</a>&nbsp;and the&nbsp;<a href="http://read.transkribus.eu/">READ </a>project. The guidelines with which the AS GT was created are uploaded here as well.</p>

opencc-by-4.0Mar 2021View details →
zenodo44/100

Dataset of behavioral and neurophysiological data of a virtual sailing task published in: "Providing task instructions during motor training enhances performance and modulates attentional brain networks"

<p>Dataset belonging to the behavioral and neurophysiological data of the publication: &quot;Providing task instructions during motor training enhances performance and modulates attentional brain networks&quot;. The two uploaded Zip files contain kinematic and electroencephalographic data of 36 participants for the Obstacle and HorizonTask.</p>

opencc-by-4.0Jul 2021View details →
zenodo44/100

The DR-Train dataset: dynamic responses, GPS positions and environmental conditions of two light rail vehicles in Pittsburgh

<p><strong>Note: Downloading the large data file could have a timeout issue. If you cannot directly download it here, please use the following link as a complementary method for getting the data.&nbsp;</strong></p> <p><a href="https://drive.google.com/drive/folders/1oKn7IN7zznQuhwjDCDdjq8r9wHJYBEhj?usp=sharing">https://drive.google.com/drive/folders/1oKn7IN7zznQuhwjDCDdjq8r9wHJYBEhj?usp=sharing</a></p> <p>&nbsp;</p> <p>This dataset contains the dynamic responses (acceleration records) of two passenger&nbsp;trains with corresponding GPS positions, environmental conditions and track maintenance&nbsp;schedules for a light rail network in the city of Pittsburgh, Pennsylvania in the United&nbsp;States of America.</p> <p>In particular, two light rail vehicles were instrumented (identified as LRV4306 and&nbsp;LRV4313):&nbsp;<br> LRV 4306 has 5 acceleration channels, corresponding to the two uni-axial accelerometers&nbsp;inside the train and the three channels of the tri-axial accelerometer on the wheel truck.</p> <p><em>- The last digit of each acceleration file: 1, 2, 3, 4, 5<br> - Corresponding sensor channels: tri-axial x, tri-axial y, tri-axial z, front cabinet uni-axial, back cabinet uni-axial</em></p> <p><br> LRV 4313 has 8 acceleration channels, corresponding to the two uni-axial accelerometer&nbsp;and the two tri-axial accelerometers inside the train.</p> <p><em>- The last digit of each acceleration file: 1, 2, 3, 4, 5, 6, 7, 8<br> - Corresponding sensor channels: front cabinet uni-axial, back cabinet uni-axial, front tri-axial x, front tri-axial y, front tri-axial z, back tri-axial x, back tri-axial y, back tri-axial z.<br> - x longitudinal (vehicle moving direction); y-axis, transverse; z-axis, vertical.</em></p> <p>The dataset contained in this repository is a condensed version of the original raw data.&nbsp;While the accelerometers on the train were sampled continuously, this dataset contains&nbsp;only those measurements for when the train was actually moving along the track (i.e. not idling at a terminal).</p> <p>The data is stored in binary MAT-files (a MATLAB/Octave data format). These files contain&nbsp;MATLAB objects of the class &quot;pass&quot;, which is defined in the file pass.m that can be&nbsp;found in the &quot;code&quot; folder. Specifically, two MAT-files named &quot;obj_dic.mat&quot;, and found in&nbsp;the &quot;LRV4306&quot; and &quot;LRV4313&quot; folders, contain the &quot;pass&quot; objects of the two trains,&nbsp;respectively.</p> <p>Each category is described in detail. For more detail on the regions of the track, refer to the &#39;region.fig&#39; file in this folder. The track was divided into distinct regions so&nbsp;that the data over specific sections of track could be compared. These regions were&nbsp;chosen for two reasons:&nbsp;<br> (1) within a region, the train always followed the same track and&nbsp;<br> (2) there are no tunnels in them so the GPS data is relatively consistent.&nbsp;</p> <p>To get started, using MATLAB or Octave try running &quot;main_script.m&quot; in the &quot;code&quot; folder.</p> <p>A data descriptor paper with details of the data collection process was published.</p> <p>Please cite as</p> <p><strong>Liu, J., Chen, S., Lederman, G., Kramer, D. B., Noh, H. Y., Bielak, J., Garrett, J. H., Kovačević, J., &amp; Berges, M.&nbsp;Dynamic responses, GPS positions and environmental conditions of two light rail vehicles in Pittsburgh. Scientific Data, 6, 146. <a href="https://doi.org/10.1038/s41597-019-0148-9">https://doi.org/10.1038/s41597-019-0148-9</a>(2019)</strong></p> <p><strong>Liu, J., Chen, S., Lederman, G., Kramer, D. B., Noh, H. Y., Bielak, J., Garrett, J. H., Kovačević, J., &amp; Berges, M. The DR-Train dataset: dynamic responses, GPS positions and environmental conditions of two light rail vehicles in Pittsburgh.&nbsp;Zenodo,&nbsp;<a href="https://doi.org/10.5281/zenodo.1432702">https://doi.org/10.5281/zenodo.1432702</a>(2018).</strong></p> <p>For questions or suggestions please e-mail Jingxiao Liu &lt;liujx@stanford.edu&gt;</p>

opencc-by-4.0Oct 2018View details →
zenodo44/100

Ortoimages for CNN trainning

<p>Preprocessed ortoimages from Castilla and Leon to use in the training of a neural network to find ruins of ancient roman camps among the fields.</p>

opencc-by-4.0Jan 2022View details →
zenodo44/100

Generated Data for the Manuscript "Nonideality-Aware Training for Accurate and Robust Low-Power Memristive Neural Networks"

<p>The file contains&nbsp;data generated and referred to in the text and the figures of the manuscript.</p>

opencc-by-4.0Dec 2021View details →
zenodo44/100

[DCASE2022 Task 3] Synthetic SELD mixtures for baseline training

<p><strong>DESCRIPTION:</strong><br> <br> This audio dataset serves serves as supplementary material for the&nbsp;<a href="http://Sound Event Localization and Detection Evaluated in Real Spatial Sound Scenes">DCASE2022 Challenge Task 3:&nbsp;Sound Event Localization and Detection Evaluated in Real Spatial Sound Scenes</a>. The dataset consists of synthetic spatial audio mixtures of sound events spatialized for two different spatial formats using real measured room impulse responses (RIRs) measured in various spaces of Tampere University (TAU). The mixtures are generated using the same process as the one used to generate the recordings of the <a href="https://zenodo.org/record/5476980">TAU-NIGENS Spatial Sound Scenes 2021</a>&nbsp;dataset for the&nbsp;<a href="https://dcase.community/challenge2021/task-sound-event-localization-and-detection-results">DCASE2021 Challenge Task 3</a>.&nbsp;</p> <p>The SELD task setup in DCASE2022 is based on spatial recordings of real scenes, captured in the <a href="https://zenodo.org/record/6387880">STARS22</a> dataset. Since the task setup allows use of external data, these synthetic mixtures serve as additional training material for the&nbsp;<a href="https://github.com/sharathadavanne/seld-dcase2022">baseline model</a>, and they are shared for reasons of reproducibility. For more details on the task setup, please refer&nbsp;to the <a href="http://Sound Event Localization and Detection Evaluated in Real Spatial Sound Scenes">task description</a>.</p> <p>Note that the generator code and the collection of room responses used to spatialize sound samples will be also be made available soon. For more details on the recording of RIRs, spatialization, and generation, see:</p> <ul> <li>Archontis Politis, Sharath Adavanne, Daniel Krause, Antoine Deleforge, Prerak Srivastava, Tuomas Virtanen (2021).&nbsp;A Dataset of Dynamic Reverberant Sound Scenes with Directional Interferers for Sound Event Localization and Detection.&nbsp;In&nbsp;<em>Proceedings of the Detection and Classification of Acoustic Scenes and Events 2020 Workshop (DCASE2021)</em>, Barcelona, Spain.</li> </ul> <p>available&nbsp;<a href="https://dcase.community/documents/workshop2021/proceedings/DCASE2021Workshop_Politis_43.pdf">here</a>.</p> <p><strong>SPECIFICATIONS:</strong></p> <ul> <li><strong>13 target sound classes</strong> (see task description for details)</li> <li>The sound event samples are sources from the&nbsp;<strong><a href="https://zenodo.org/record/4060432">FSD50K</a></strong>&nbsp;dataset, based on affinity of the labels in that dataset to the target classes. The selection on distinguishing which labels in FSD50K corresponded to the target ones, then selecting samples that were tagged with only those labels, and additionally that they had annotator rating of Present and Predominant (see FSD50K for more details). The list of the selected files is included here.</li> <li><strong>1200</strong> 1-minute long spatial recordings</li> <li>Sampling rate of<strong> 24kHz</strong></li> <li>Two 4-channel recording formats, first-order Ambisonics (<strong>FOA</strong>) and tetrahedral microphone array (<strong>MIC</strong>)</li> <li>Spatial events spatialized in <strong>9 unique rooms</strong>, using measured RIRs for the two formats</li> <li>Maximum <strong>polyphony of 2</strong> (with possible same-class events overlapping)</li> <li>Even though the whole set is used for training of the baseline without distinction between the mixtures, we have included a <strong>separation into a training and testing split</strong>, in case on one needs to&nbsp;test&nbsp;the performance purely on those&nbsp;synthetic conditions (for example for comparisons with training on mixed synthetic-real data, fine-tuning on real data, or training on real data only).</li> <li>The training split is indicated as <strong>fold1</strong>&nbsp;in the dataset, contains 900 recordings spatialized on 6 rooms (150 recordings/room) and it is based on samples from the development set of FSD50K.</li> <li>The testing split is indicated as <strong>fold2</strong>&nbsp;in the dataset, contains 300 recordings spatialized on 3 rooms (100 recordings/room) and it is based on samples from the evaluation set of FSD50K.</li> <li>Common metadata files for both formats are provided. For the file naming and the metadata format, refer to the task setup.</li> </ul> <p><strong>FSD50K SELECTION:</strong></p> <p>The list of selected sound event recordings is included along the recordings and metadata, as <strong>FSD50K_selected.txt</strong>. Each line in the text&nbsp;file has the following structure:</p> <pre><code>[target_label]/[train/test]/[FSD50K_label]/filename.wav</code></pre> <p>with an example:</p> <pre><code>domesticSounds/train/Boiling/16584.wav</code></pre> <p>meaning that the file 16584.wav from FSD50K, with the <em>Boiling</em> label of FSD50K, is included in the samples for the training split of those synthetic recordings, and it is mapped to the target class of <em>domestic sounds. </em>Note that there can be multiple FSD50K labels mapped the same target class. Also note that if these are downloaded from FSD50K, and a folder structure is created that replicates the structure in the list, the resulting folder can be used out-of-the-box with the scene generator to generate new mixtures with the same or different parameters.</p> <p>Note that no sounds form FSD50K have been selected for the <em>Music</em>&nbsp;target class. Background and pop music tracks from the public domain have been cropped and used instead.</p> <p><strong>DOWNLOAD INSTRUCTIONS:</strong></p> <p>Download the zip files and use your preferred compression tool to unzip these split zip files. To extract a split zip archive (named as zip, z01, z02, ...), you could use, for example, the following syntax in Linux or OSX terminal:</p> <ol> <li>Combine the split archive to a single archive: <pre>zip -s 0 split.zip --out single.zip</pre> </li> <li>Extract the single archive using unzip: <pre>unzip single.zip</pre> </li> </ol>

opencc-by-4.0Mar 2022View details →
zenodo44/100

Training data for 'Preparing genomic data for phylogeny reconstruction' (Galaxy Training Material)

<p>This data is used for Galaxy Training Network training &#39;Preparing genomic data for phylogeny reconstruction&#39;. There are four nucleotide sequences from chromosome 5 of four strains of S. cerevisiae. The GenBank annotated sequenced were produced using &#39;funannotate predict annotation&#39; (Galaxy Version 1.8.9+galaxy2) on the nucleotide sequences sequences. References: DOI: 10.1126/science.274.5287.546; DOI: 10.1126/science.1189015; DOI: 10.1016/j.cell.2016.08.020</p>

opencc-by-4.0May 2022View details →
zenodo44/100

PhasAGE Training School 2 - Phase separations and transitions by viral proteins: from viral factories to interference with host cell functions- LECTURE

<p>The Training School 2 &ldquo;Biomolecular condensates in cell function, aging and disease&rdquo; is the <strong>second</strong> edition of a series of PhasAGE training activities.</p> <p>&nbsp;</p> <p>The main goal of this training school is to raise awareness and provide expertise on fundamental aspects of phase separation and formation of <strong>biomolecular condensates</strong>, specifically covering the importance of this process to cellular biology and its contribution to the aging process and age-related diseases.</p>

opencc-by-4.0Jun 2022View details →
zenodo44/100

PhasAGE Training School 2 - Condensation through liquid-liquid separation-LECTURE

<p>PhasAGE Training School 2 &ldquo;Biomolecular condensates in cell function, aging and disease&rdquo; is the<strong> second</strong> edition of a series of PhasAGE training activities.</p> <p>The main goal of this training school is to raise awareness and provide expertise on fundamental aspects of phase separation and formation of <strong>biomolecular condensates</strong>, specifically covering the importance of this process to cellular biology and its contribution to the aging process and age-related diseases.</p>

opencc-by-4.0Jun 2022View details →
zenodo44/100

gmap - qgis training material: Ingenii Basin (moon)

<p>This dataset part of the Geology and Planetary Mapping Winter School 2022 featuring Ingenii Basin as a study area.<br> Ingenii Basin is located on the lunar farside centred at 33.7&deg;S 163.5&deg;E within the South Pole-Aitken basin. The floor of Ingenii Basin is filled with mare materials with the basin having a diameter of 282 km.<br> We compiled a beginners &ndash; intermediate level training package for the area. The package includes the Lunar Reconnaissance Orbiter Camera (LROC) Wide Angle Camera (WAC) global mosaic&nbsp;(Speyerer et al., 2011) as a basemap, the Lunar Orbiter Laser Altimeter (LOLA) and SELenological and Engineering Explorer (SELENE) Kaguya merged lunar digital elevation model (DEM) (Barker et al., 2016) and spectral data in the form of a clementine Ultraviolet/Visible (UVVIS)&nbsp;warped color ratio mosaic (Lucey et al., 2000). The data is cut to the area of interest and a training project is set up for QGIS.&nbsp;</p> <p>The training package is designed as a group exercise with four adjacent tiles covering the entirety of Ingenii basin. For beginners the aim is to create a low scale map of the area where the basin rim is distinguished from the basin floor and mare unit as well as detecting smaller craters that exist in the area. These units should then be put in a stratigraphic relationship based on superposition, degradation state and embayment. For intermediate mappers this task can be extended to include the swirl features and finding potential areas for crater size frequency distribution measurement to determine absolute ages for a more detailed stratigraphy.</p> <p><br> &nbsp;</p> <p>&nbsp;</p>

opencc-by-4.0Jun 2022View details →
zenodo44/100

POC detection training data and weights

<p>The data and labels used to train the POC detection <a href="https://github.com/climate-processes/poc-detection">algorithm </a>used in support of this publication: https://doi.org/10.1029/2020GL092213. The associated model weights were also saved after training to aid reproducibility.</p>

opencc-by-4.0Jul 2022View details →
zenodo44/100

TIME4CS WP4 Mapping of citizen science training resources

<p>This dataset was compiled as part of the TIME4CS project, WP4, and lists identified citizen science training resources, as of July 2022.</p> <p>The <a href="https://eu-citizen.science/">EU-citizen.science</a> platform provided the basis for mapping CS training in Europe, as the team behind the platform has put considerable effort into compiling, and encouraging the CS community to contribute, CS training resources. Additionally, training courses were identified based on the case studies in WP1, as most universities do not list their courses on the EU-citizen.science platform.</p>

opencc-by-4.0Jul 2022View details →
zenodo44/100

Phase Object Reconstruction for 4D-STEM using Deep Learning, (4D-STEM Training Data)

<p><strong>Overview </strong></p> <p>This repository contains 742,688 samples of simulated Convergent Beam Electron Diffraction patterns (CBEDs); the training data for the paper <a href="https://arxiv.org/abs/2202.12611">&quot;Phase Object Reconstruction for 4D-STEM using Deep Learning&quot;</a>. The folder contains multiple hdf5 datasets. Each dataset has a corresponding Excel-sheet containing detailed information and simulation parameters for every datapoint, as well as a summary-report containing the parameter distributions, hdf5-infos and random number generator settings. This makes every dataset reproducible, using the simulation codes provided in <a href="https://github.com/ThFriedrich/ap_data_generation">https://github.com/ThFriedrich/ap_data_generation</a>.</p> <p><strong>Technical details</strong></p> <p>Every Datapoint consists of a 3x3 set of adjacent Convergent Beam Electron Diffraction pattern (CBEDs), the coherent exit wave phase and amplitude in real and reciprocal space, and the probe functions phase and amplitude in real space. All patterns are 64x64 pixel in 16 bit unsigned integer data format.</p> <p>Every hdf5 file has the following structure:</p> <table> <tbody> <tr> <td>Attributes</td> <td>&#39;Seed&#39;:&nbsp; 6108236<br> &#39;State&#39;:&nbsp; 251786606 ...<br> &#39;Type&#39;:&nbsp; &#39;twister&#39;<br> &nbsp;&#39;arch&#39;:&nbsp; &#39;glnxa64&#39;<br> &#39;gpu&#39;:&nbsp; &#39;NVIDIA GeForce RTX 3080&#39;<br> &#39;matlab_ver&#39;:&nbsp; &#39;2021a&#39;</td> </tr> <tr> <td>Dataset &#39;features&#39;</td> <td> <p>Size: 64x64x9x5000<br> Datatype: H5T_STD_U16LE (uint16)</p> </td> </tr> <tr> <td>Dataset &#39;labels_k&#39;</td> <td> <p>Size: 64x64x2x5000<br> Datatype: H5T_STD_U16LE (uint16)</p> </td> </tr> <tr> <td>Dataset &#39;labels_r&#39;</td> <td> <p>Size: 64x64x2x5000<br> Datatype: H5T_STD_U16LE (uint16)</p> </td> </tr> <tr> <td>Dataset &#39;probe_r&#39;</td> <td> <p>Size: 64x64x2x5000<br> Datatype: H5T_STD_U16LE (uint16)</p> </td> </tr> <tr> <td>Dataset &#39;meta&#39;</td> <td> <p>Size: 19x5000<br> Datatype: H5T_IEEE_F32LE (single)</p> </td> </tr> </tbody> </table> <p>The data was written to hdf5 in matlab. When reading from these files consider possibly different storage conventions (Row major vs. column major format). Data may need to be transposed accordingly. The integer arrays were scaled to use the full range of the uint16 datatype. The scaling values are stored under &quot;meta&quot;. To restore the original values in floating point numbers, convert the arrays like this:</p> <p>Matlab:</p> <pre><code>hdf_file = ['db_h5_b_5_Training.h5']; n = 128; % load `n` k-space exit waves x = single(h5read(hdf_file, '/labels_k', [1,1,1,1], [64,64,2,n])); % `meta` contains parameters and scaling factors for a given datapoint in following order: [E_0(keV), cond_lens_outer_aper_ang(mrad), collection angle(rA), step_size(A), scale_cbed_1 ... scale_cbed_9, scale_phase_k, scale_amp_k, scale_phase_r, scale_amp_r, scale_probe_phase_r, scale_probe_amp_r] s = h5read(hdf_file, '/meta', [14,1], [2,n]); amplitude = zeros(64,64,n); phase = zeros(64,64,n); for ix = 1:n phase(:,:,n) = (x(:,:,1,n)*s(1,ix) / 65536) - pi; amplitude(:,:,n) = (x(:,:,2,n)*s(2,ix)) / 65536; end % The 9 CBEDs correspond to a 3x3 kernel of patterns. The order in [x,y] is: %[[3, 6, 9]; % [2, 5, 8]; % [1, 4, 7]] </code></pre>

opencc-by-4.0Aug 2022View details →
zenodo44/100

StarDist Adipocyte Segmentation Training data, Training Notebook and Model

<p>Data from H&amp;E human bone marrow whole slide scanner images used in the paper: &quot;MarrowQuant 2.0: a digital pathology workflow assisting bone marrow evaluation in clinical and experimental hematology&quot; (https://doi.org/10.21203/rs.3.rs-1860140/v1)</p> <p>&nbsp;</p> <p>292 image patches</p> <p>Ground truth were manually annotated using QuPath and split into 263 images for training and 29 for validation.</p> <p>Training in StarDist was done on a Windows 10 PC with an RTX 2080 GPU. The requirements file for installing a Python 3.7 environment to run the attached notebooks is provided (<strong>stardist-val.txt</strong>).</p> <p>The StarDist model configuration can be found in the Jupyter Notebook :</p> <pre><code>Adipocyte Training.ipynb</code></pre> <p>Model validation and metrics can be performed by running the notebook after finishing the <strong>Adipocyte Training</strong> notebook.</p> <pre><code>Quality Control.ipynb</code></pre> <p>&nbsp;</p>

opencc-by-4.0Aug 2022View details →
zenodo44/100

Training Deep Learning Models to Estimate Permeability using Geophysical Datasets

<p>This folder contains the dataset for training deep learning models to estimate permeability using hydro-geophysics simulations</p>

opencc-by-4.0Aug 2022View details →
zenodo44/100

Unsupervised New Physics detection at 40 MHz: Training Dataset

<p>Unsupervised New Physics detection at 40 MHz data challenge</p> <p>Training dataset, consisting of a cocktail of Standard Model collision events (simulation of LHC 13 TeV proton-proton collisions) pre-filtered by a requirement of a muon or electron with 23 GeV transverse momentum. Data format description available on the data challenge web page:&nbsp;https://mpp-hep.github.io/ADC2021/</p>

opencc-by-4.0Jun 2021View details →
zenodo44/100

Effects of audio-motor training on spatial representations in long-term late blindness

<p>Datasets for behavioural data:</p> <p>-<em>Auditory horizontal localization task</em></p> <p>-&nbsp;<em>Auditory vertical localization task</em></p> <p>-&nbsp;<em>Position matching task</em></p> <p>-&nbsp;<em>Proprioceptive midline task</em></p> <p>Dataset for EEG data:</p> <p>- Spatial bisection task: mean ERP amplitude for 50-90ms timew window for each trial, separately for condition, session and roi</p> <p>&nbsp;</p>

opencc-by-4.0Oct 2022View details →
zenodo44/100

The PANORAMA Challenge: Public Training and Development Dataset (3)

<p>This dataset represents the <strong><a href="https://panorama.grand-challenge.org/" target="_blank" rel="noopener">PANORAMA</a>: Public Training and Development Dataset</strong>. It contains 2238 anonymized contrast-enhanced CT (CECT) scans acquired at two centers (Radboud University Medical Center, University Medical Center Groningen) based in The Netherlands. Additionally, it contains 194 cases from the&nbsp;<strong><a href="http://medicaldecathlon.com/" target="_blank" rel="noopener">Medical Segmentation Decathlon</a>&nbsp;</strong>dataset and 80 cases from<strong>&nbsp;<a href="https://www.cancerimagingarchive.net/collection/pancreas-ct/" target="_blank" rel="noopener">National Institutes of Health</a></strong>. For all updates/fixes regarding this dataset, please join the challenge and check out our&nbsp;<a href="https://grand-challenge.org/forums/forum/panorama-pancreatic-cancer-diagnosis-radiologists-meet-ai-711/topic/public-training-and-development-dataset-updates-and-fixes-2213/" target="_blank" rel="noopener">dedicated forum post</a> on this topic. The corresponding labels of the PANORAMA dataset can be found <a href="https://github.com/DIAGNijmegen/panorama_labels">here</a>.&nbsp;</p> <p>The PANORAMA challenge is an all-new grand challenge that aims to validate the diagnostic performance of artificial intelligence and radiologists at pancreatic ductal adenocarcinoma (PDAC) detection/diagnosis in CECT, with histopathology and follow-up (&ge; 3 years) as the reference standard, in a retrospective setting in the hidden testing dataset. The study hypothesizes that state-of-the-art AI algorithms are non-inferior to radiologists reading CECT.</p> <p>Key aspects of the PANORAMA study design have been established in conjunction with an international scientific advisory board of 13 experts in AI and pancreas radiology as well as a patient representative &mdash;to unify and standardize present-day guidelines, and to ensure meaningful validation of pancreas AI towards clinical translation (<strong><a href="https://www.sciencedirect.com/science/article/pii/S2405456921001607">Reinke et al., 2021</a></strong>).</p> <p><em>This PANORAMA dataset contains: batch&nbsp;<strong>3</strong> <strong>out of 4</strong></em></p>

opencc-by-nc-4.0Apr 2024View details →
zenodo44/100

The PANORAMA Challenge: Public Training and Development Dataset (4)

<p>This dataset represents the <strong><a href="https://panorama.grand-challenge.org/" target="_blank" rel="noopener">PANORAMA</a>: Public Training and Development Dataset</strong>. It contains 2238 anonymized contrast-enhanced CT (CECT) scans acquired at two centers (Radboud University Medical Center, University Medical Center Groningen) based in The Netherlands. Additionally, it contains 194 cases from the&nbsp;<strong><a href="http://medicaldecathlon.com/" target="_blank" rel="noopener">Medical Segmentation Decathlon</a>&nbsp;</strong>dataset and 80 cases from<strong>&nbsp;<a href="https://www.cancerimagingarchive.net/collection/pancreas-ct/" target="_blank" rel="noopener">National Institutes of Health</a></strong>. For all updates/fixes regarding this dataset, please join the challenge and check out our&nbsp;<a href="https://grand-challenge.org/forums/forum/panorama-pancreatic-cancer-diagnosis-radiologists-meet-ai-711/topic/public-training-and-development-dataset-updates-and-fixes-2213/" target="_blank" rel="noopener">dedicated forum post</a> on this topic. The corresponding labels of the PANORAMA dataset can be found <a href="https://github.com/DIAGNijmegen/panorama_labels">here</a>.&nbsp;</p> <p>The PANORAMA challenge is an all-new grand challenge that aims to validate the diagnostic performance of artificial intelligence and radiologists at pancreatic ductal adenocarcinoma (PDAC) detection/diagnosis in CECT, with histopathology and follow-up (&ge; 3 years) as the reference standard, in a retrospective setting in the hidden testing dataset. The study hypothesizes that state-of-the-art AI algorithms are non-inferior to radiologists reading CECT.</p> <p>Key aspects of the PANORAMA study design have been established in conjunction with an international scientific advisory board of 13 experts in AI and pancreas radiology as well as a patient representative &mdash;to unify and standardize present-day guidelines, and to ensure meaningful validation of pancreas AI towards clinical translation (<strong><a href="https://www.sciencedirect.com/science/article/pii/S2405456921001607">Reinke et al., 2021</a></strong>).</p> <p><em>This PANORAMA dataset contains: batch&nbsp;<strong>4</strong> <strong>out of 4</strong></em></p>

opencc-by-nc-4.0Apr 2024View details →
zenodo44/100

Trackerless 3D Freehand Ultrasound Reconstruction Challenge 2024 - Train Dataset (Part 2)

<blockquote> <p><strong>This Challenge will be an open-ended challenge, and we welcome your submission. Please register your team via this ⁠<a title="https://forms.office.com/e/dPg47ktV7M" href="https://forms.office.com/e/dPg47ktV7M" target="_blank" rel="noopener">form</a>. You can submit the algorithm via this <a title="https://forms.office.com/e/QChhNkLYiu" href="https://forms.office.com/e/QChhNkLYiu" target="_blank" rel="noopener noreferrer">form</a> for TUS-REC2024 Challenge, and we will test your submitted docker on the test set.</strong></p> <p><strong>We are organising TUS-REC2025 at MICCAI2025. More information is available on the <a href="https://github-pages.ucl.ac.uk/tus-rec-challenge/" target="_blank" rel="noopener">TUS-REC2025 challenge website</a> and <a href="https://github.com/QiLi111/TUS-REC2025-Challenge_baseline" target="_blank" rel="noopener">Baseline code repo</a>.</strong></p> </blockquote> <p><strong>This is the second part of the Challenge dataset.&nbsp;<a href="../doi/10.5281/zenodo.11178509" target="_blank" rel="noopener">Link</a> to first part; <a href="../doi/10.5281/zenodo.11355500" target="_blank" rel="noopener">Link</a> to third part. <a href="../doi/10.5281/zenodo.12979481" target="_blank" rel="noopener">Link</a> to validation dataset.</strong></p> <p>Acquisition devices and config: The 2D US images were acquired using an Ultrasonix machine (BK, Europe) with a curvilinear probe (4DC7-3/40). The associated position information of each frame was recorded by an optical tracker (NDI Polaris Vicra, Northern Digital Inc., Canada). The acquired US frames were recorded at 20 fps, with an image size of 480&times;640, without speckle reduction. The frequency was set at 6MHz with a dynamic range of 83 dB, an overall gain of 48% and a depth of 9 cm.&nbsp;</p> <div> <p>Scanning protocol: Both left and right forearms of volunteers were scanned. For each forearm, the US probe moves in three different trajectories (straight line shape, "C" shape, and "S" shape), in a distal-to-proximal direction followed by a proximal-to-distal direction, with the US plane perpendicular of and parallel to the scanning direction. The train dataset contains 1200 scans in total, 24 scans associated with each subject.</p> <div> <div> <p>For detailed information please refer to the <a href="https://github-pages.ucl.ac.uk/tus-rec-challenge/TUS-REC2024/" target="_blank" rel="noopener">Challenge website</a>. Baseline code is also provided, which can be found at this <a href="https://github.com/QiLi111/tus-rec-challenge_baseline" target="_blank" rel="noopener">repo</a>.</p> <p>Dataset structure:&nbsp;</p> </div> <div> <ul> <li> <p>The dataset contains 50 folders (one subject per folder), each with 24 scans. Each .h5 file corresponds to one scan, storing image and transformation of each frame within this scan. Key-value pairs in each .h5 file are explained below.</p> <ul> <li> <p>&ldquo;frames&rdquo;&nbsp; - All frames in the scan; with a shape of [N,H,W], where N refers to the number of frames in the scan, H and W denote the height and width of a frame.&nbsp;</p> </li> <li> <p>&ldquo;tforms&rdquo; - All transformations in the scan; with a shape of [N,4,4], where N is the number of frames in the scan, and the transformation matrix denotes the transformation from tracker tool space to camera space.&nbsp;</p> </li> <li> <p>Notations in the name of each .h5 file: &ldquo;RH&rdquo;: right arm; &ldquo;LH&rdquo;: left arm; &ldquo;Per&rdquo;: perpendicular; &ldquo;Par&rdquo;: parallel; &ldquo;L&rdquo;: straight line shape; &ldquo;C&rdquo;: C shape; &ldquo;S&rdquo;: S shape; &ldquo;DtP&rdquo;: distal-to-proximal direction; &ldquo;PtD&rdquo;: proximal-to-distal direction; For example, &ldquo;RH_Per_L_DtP.h5&rdquo; denotes a scan on the right forearm, with ultrasound probe perpendicular of the forearm sweeping along straight line, in distal-to-proximal direction.</p> </li> </ul> </li> <li> <p>Calibration matrix: The calibration matrix was obtained using a pinhead-based method. The "scaling_from_pixel_to_mm" and "spatial_calibration_from_image_coordinate_system_to_tracking_tool_coordinate_system" are provided in the &ldquo;calib_matrix.csv&rdquo;.&nbsp;</p> </li> </ul> <div> <p><strong>Data Usage Policy:</strong></p> <ul> <li>The training and validation data provided may be utilized within the research scope of this challenge and in subsequent research-related publications. However, commercial use of the training and validation data is prohibited. In cases where the intended use is ambiguous, participants accessing the data are requested to abstain from further distribution or use outside the scope of this challenge.</li> <li>If you use our dataset in your publication, please cite the challenge paper and some of the following optional articles:&nbsp; <ul> <li>Challenge paper: <ul> <li><strong>Qi Li et al. "TUS-REC2024: A Challenge to Reconstruct 3D Freehand Ultrasound Without External Tracker." <em>arXiv preprint arXiv:<a title="https://arxiv.org/abs/2506.21765" href="https://doi.org/10.48550/arXiv.2506.21765" target="_blank" rel="noopener">2506.21765</a></em>&nbsp;(2025).</strong></li> </ul> </li> <li>Optional articles: <ul> <li>Qi Li, Ziyi Shen, Qianye Yang, Dean C. Barratt, Matthew J. Clarkson, Tom Vercauteren, and Yipeng Hu. "Nonrigid Reconstruction of Freehand Ultrasound without a Tracker." In&nbsp;<em>International Conference on Medical Image Computing and Computer-Assisted Intervention</em>, pp. 689-699. Cham: Springer Nature Switzerland, 2024. doi: <a href="https://doi.org/10.1007/978-3-031-72083-3_64" target="_blank" rel="noopener">10.1007/978-3-031-72083-3_64.</a></li> <li>Qi Li, Ziyi Shen, Qian Li, Dean C. Barratt, Thomas Dowrick, Matthew J. Clarkson, Tom Vercauteren, and Yipeng Hu. "Long-term Dependency for 3D Reconstruction of Freehand Ultrasound Without External Tracker." IEEE Transactions on Biomedical Engineering, vol. 71, no. 3, pp. 1033-1042, 2024. doi:&nbsp;<a href="https://ieeexplore.ieee.org/abstract/document/10288201" target="_blank" rel="noopener">10.1109/TBME.2023.3325551</a>.</li> <li>Qi Li, Ziyi Shen, Qian Li, Dean C. Barratt, Thomas Dowrick, Matthew J. Clarkson, Tom Vercauteren, and Yipeng Hu. "Trackerless freehand ultrasound with sequence modelling and auxiliary transformation over past and future frames." In 2023 IEEE 20th International Symposium on Biomedical Imaging (ISBI), pp. 1-5. IEEE, 2023. doi: <a href="https://doi.org/10.1109/ISBI53787.2023.10230773" target="_blank" rel="noopener">10.1109/ISBI53787.2023.10230773.</a></li> <li>Qi Li, Ziyi Shen, Qian Li, Dean C. Barratt, Thomas Dowrick, Matthew J. Clarkson, Tom Vercauteren, and Yipeng Hu. "Privileged Anatomical and Protocol Discrimination in Trackerless 3D Ultrasound Reconstruction." In International Workshop on Advances in Simplifying Medical Ultrasound, pp. 142-151. Cham: Springer Nature Switzerland, 2023. doi: <a href="https://doi.org/10.1007/978-3-031-44521-7_14" target="_blank" rel="noopener">https://doi.org/10.1007/978-3-031-44521-7_14.</a></li> </ul> </li> </ul> </li> </ul> </div> </div> </div> </div>

opencc-by-nc-sa-4.0May 2024View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record