Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
250
datasets available to search
ShareScore release 0.9.0
Dataset results
250 results for “synthetic dataset”
Dataset from "Synthetic Training Data for Semantic Segmentation of the Environment from UAV Perspective"
<p>This dataset contains the images and ground truth label masks for semantic segmentation created and described in "Hinniger, C.; Rüter, J. Synthetic Training Data for Semantic Segmentation of the Environment from UAV Perspective. Aerospace 2023, 10, 604. https://doi.org/10.3390/aerospace10070604".</p>
Synthetic Airborne Intruder Dataset: A dataset based on High-Resolution Inpainting for Safety Critical Detect and Avoid
<p>Modern machine learning techniques have shown tremendous potential, especially for object detection on camera images. For this reason, they are also used to enable safety-critical automated processes such as autonomous drone flights. We present a study on object detection for Detect and Avoid, a safety critical function for drones that detects air traffic during automated flights for safety reasons. An ill-posed problem is the generation of good and especially large data sets, since detection itself is the corner case. Most models suffer from limited ground truth in raw data, e.g. recorded air traffic or frontal flight with a small aircraft. It often leads to poor and critical detection rates. We overcome this problem by using inpainting methods to bootstrap the dataset such that it explicitly contains the corner cases of the raw data. We provide an overview of inpainting methods and generative models and present an example pipeline given a small annotated dataset. We validate our method by generating a high-resolution dataset and present it to an independent object detector that was fully trained on real data.</p> <p>This dataset is represented in the following repository. The dataset is structured as follows:</p> <p># Synthetic Airborne Intruder Dataset</p> <p>This dataset was syntheticaly generated using an adapted [Pix2Pix](https://github.com/junyanz/pytorch-CycleGAN-and-pix2pix) with different background images and object segementations. Each image contains one object instance.</p> <p>The annotations are in the COCO annotation format.</p> <p>## Data Structure<br> Synthetic Dataset Root:<br> --train<br> |--images<br> |--instances.json<br> --val<br> |--images<br> |--instances.json<br> --test<br> |--images<br> |--instances.json<br> --Background_Sources<br> |--sources_train.csv<br> |--sources_val.csv<br> |--sources_test.cs<br> --README.md</p> <p>## Categories</p> <p>| Id | Name | Instances over all splits |<br> | ---| --- | --- |<br> | 0 | large airplane | 1695 | <br> | 1 | small airplane | 1255 | <br> | 2 | very small airplane | 46 | <br> | 3 | helicopter | 2201 | <br> | 4 | drone | 961 |<br> | 5 | hot air balloon | 315 |<br> | 6 | paraglider | 565 |<br> | 7 | airship | 42 |<br> | 8 | UFO | 0 |</p> <p>### Note:<br> UFO is a placeholder for future expansion of the dataset.</p> <p>## Splits<br> The dataset consists of 3 splits: train 5900 images, val 590 images, test 590 images.<br> The Number of instances per class and per split can be seen in the table below:</p> <p>Class | train | val | test<br> -------|-------|-----|--------<br> large airplane | 1416 | 142 | 137<br> small airplane | 1046 | 96 | 113<br> very small airplane | 38 | 2 | 6<br> helicopter | 1812 | 206 | 183<br> drone | 800 | 86 | 75<br> hot air balloon | 268 | 21 | 26<br> paragliders | 492 | 32 | 41<br> airship | 28 | 5 | 9<br> UFO | 0 | 0 | 0</p> <p>## Sources<br> The sources of the background images can be found in the files [here](./Background_Sources/).</p>
6DOF pose estimation - synthetically generated dataset using BlenderProc
Open the record for dataset details and reuse information.
Synthetic Datasets for mpEDM
<p>Generated datasets are synthesized for benchmarking mpEDM. The datasets were generated with <em>numpy</em>.</p> <p>File Name: Generated_TimeSeries_{No. of time steps}_{No. of time series}.h5</p>
Afro-MNIST: Synthetic generation of MNIST-style datasets for low-resource languages
<p>We present Afro-MNIST, a set of synthetic MNIST-style datasets for four orthographies used in Afro-Asiatic and Niger-Congo languages: Ge`ez (Ethiopic), Vai, Osmanya, and N'Ko.<br> These datasets serve as ``drop-in'' replacements for MNIST. We hope that MNIST-style datasets will be developed for other numeral systems, and that these datasets vitalize machine learning education in underrepresented nations in the research community.</p>
Synthetic student dataset on levels of use of digital tools and frequency of personal activities with ICTs
<p>Dataset artificially generated through the Generative Adversarial Networks of students in postgraduate degrees at the University of La Laguna on levels of use of digital tools and frequency of personal activities with ICT.</p>
Synth-Yard-MCMOT-1 – A synthetic multi camera multi object dataset in the context of yard logistics
<p>The dataset "Synth-Yard-MCMOT-1" is a novel image dataset for multi-camera multi-object tracking (MCMOT). It is the first of its kind to be generated in a virtual environment with the main focus on the tracking of trucks in yard logistics environments. The dataset consists of a total of 12,008 images generated by eight different cameras. The images contain 44,232 bounding boxes and segmentation masks and 52 individual tracks. Additionally, we provide a ninth camera, which is used to generate unified ground-truth information for the whole scene from an orthographic, top-down perspective comparable to a bird’s eye or map-view. The purpose of this dataset is to provide yard management systems with relevant data, which can be employed when aiming to determine the exact position of a truck and specifically identifying which gateway or designated parking spot it is located in. <br>Additionally in this dataset we added the annotations in the COCO-style format as well as the raw image-data.<br>Initial benchmarks for single-camera tracking demonstrate a mean IDF1 score of 0.96 and a mean MOTA score of 0.94, laying the baseline for computing world coordinates via MCMOT.<br><br><span>This work is part of the project “Silicon Economy Logistics </span><span>Ecosystem” which is funded by the German Federal Ministry </span><span>of Transport and Digital Infrastructure. <br></span><span>This research has partly been funded by the Federal Ministry </span><span>of Education and Research of Germany and the state of North-</span><span>Rhine Westphalia as part of the Lamarr-Institute for Machine </span><span>Learning and Artificial Intelligence. <br></span><span>We would like to thank our colleagues Julian Hinxlage, </span><span>Oleg Belov, Habib Kilic, Nicole Wagner-Hanl, Leona Cerimi, </span><span>Antonia Ponikarov and Nico Lindemann for their support.</span></p> <p><span>The code to create this dataset can be found here: </span><span>https://git.openlogisticsfoundation.org/silicon-economy/base/ml-toolbox/</span><span>synth-data-mcmot</span></p> <div> <div> </div> <div> </div> </div> <p> </p>
Synthetic Speech Dataset
<p>This contains the speech data synthesized for the following paper:</p> <p>"Using Speech Synthesis to Train End-to-End Spoken Language Understanding Models" by Loren Lugosch, Brett Meyer, Derek Nowrouzezahrai, and Mirco Ravanelli</p>
A large synthetic dataset for machine learning applications in power transmission grids
<p>With the ongoing energy transition, power grids are evolving fast. They operate more and more often close to their technical limit, under more and more volatile conditions. Fast, essentially real-time computational approaches to evaluate their operational safety, stability and reliability are therefore highly desirable. Machine Learning methods have been advocated to solve this challenge, however they are heavy consumers of training and testing data, while historical operational data for real-world power grids are hard if not impossible to access. </p> <p>This dataset contains long time series for production, consumption, and line flows, amounting to 20 years of data with a time resolution of one hour, for several thousands of loads and several hundreds of generators of various types representing the ultra-high-voltage transmission grid of continental Europe. The synthetic time series have been statistically validated agains real-world data.</p> <h2>Data generation algorithm</h2> <p>The algorithm is described in a <a href="https://doi.org/10.1038/s41597-025-04479-x">Nature Scientific Data paper</a>. It relies on <a href="https://zenodo.org/records/2642175" target="_blank" rel="noopener">the PanTaGruEl model of the European transmission network</a> -- the admittance of its lines as well as the location, type and capacity of its power generators -- and aggregated data gathered from <a href="https://transparency.entsoe.eu/" target="_blank" rel="noopener">the ENTSO-E transparency platform</a>, such as power consumption aggregated at the national level.</p> <h2>Network</h2> <p>The network information is encoded in the file <a href="https://zenodo.org/records/13378476/files/europe_network.json">europe_network.json</a>. It is given in <a href="https://lanl-ansi.github.io/PowerModels.jl/stable/" target="_blank" rel="noopener">PowerModels format</a>, which it itself derived from <a href="https://matpower.org/" target="_blank" rel="noopener">MatPower</a> and compatible with <a href="https://www.pandapower.org/" target="_blank" rel="noopener">PandaPower</a>. The network features 7822 power lines and 553 transformers connecting 4097 buses, to which are attached 815 generators of various types.</p> <h2>Time series</h2> <p>The time series forming the core of this dataset are given in CSV format. Each CSV file is a table with 8736 rows, one for each hourly time step of a 364-day year. All years are truncated to exactly 52 weeks of 7 days, and start on a Monday (the load profiles are typically different during weekdays and weekends). The number of columns depends on the type of table: there are 4097 columns in load files, 815 for generators, and 8375 for lines (including transformers). Each column is described by a header corresponding to the element identifier in the network file. All values are given in per-unit, both in the model file and in the tables, i.e. they are multiples of a base unit taken to be 100 MW.</p> <p>There are 20 tables of each type, labeled with a reference year (2016 to 2020) and an index (1 to 4), zipped into archive files arranged by year. This amount to a total of 20 years of synthetic data. When using loads, generators, and lines profiles together, it is important to use the same label: for instance, the files <em>loads_2020_1.csv</em>, <em>gens_2020_1.csv</em>, and <em>lines_2020_1.csv</em> represent a same year of the dataset, whereas <em>gens_2020_2.csv</em> is unrelated (it actually shares some features, such as nuclear profiles, but it is based on a dispatch with distinct loads).</p> <h2>Usage</h2> <p>The time series can be used without a reference to the network file, simply using all or a selection of columns of the CSV files, depending on the needs. We show below how to select series from a particular country, or how to aggregate hourly time steps into days or weeks. These examples use Python and the data analyis library <em>pandas</em>, but other frameworks can be used as well (Matlab, Julia). Since all the yearly time series are periodic, it is always possible to define a coherent time window modulo the length of the series.</p> <h3>Selecting a particular country</h3> <p>This example illustrates how to select generation data for Switzerland in Python. This can be done without parsing the network file, but using instead <a href="https://zenodo.org/records/13378476/files/gens_by_country.csv">gens_by_country.csv</a>, which contains a list of all generators for any country in the network. We start by importing the <em>pandas</em> library, and read the column of the file corresponding to Switzerland (country code CH):</p> <pre><code>import pandas as pd CH_gens = pd.read_csv('gens_by_country.csv', usecols=['CH'], dtype=str)</code></pre> <p>The object created in this way is Dataframe with some null values (not all countries have the same number of generators). It can be turned into a list with:</p> <pre><code>CH_gens_list = CH_gens.dropna().squeeze().to_list()</code></pre> <p>Finally, we can import all the time series of Swiss generators from a given data table with</p> <pre><code>pd.read_csv('gens_2016_1.csv', usecols=CH_gens_list)</code></pre> <p>The same procedure can be applied to loads using the list contained in the file <a href="https://zenodo.org/records/13378476/files/loads_by_country.csv">loads_by_country.csv</a>.</p> <h3>Averaging over time</h3> <p>This second example shows how to change the time resolution of the series. Suppose that we are interested in all the loads from a given table, which are given by default with a one-hour resolution:</p> <pre><code>hourly_loads = pd.read_csv('loads_2018_3.csv')</code></pre> <p>To get a daily average of the loads, we can use: </p> <pre><code>daily_loads = hourly_loads.groupby([t // 24 for t in range(24 * 364)]).mean()</code></pre> <p>This results in series of length 364. To average further over entire weeks and get series of length 52, we use: </p> <pre><code>weekly_loads = hourly_loads.groupby([t // (24 * 7) for t in range(24 * 364)]).mean()</code></pre> <h2>Source code</h2> <p>The code used to generate the dataset is freely available at <a href="https://github.com/GeeeHesso/PowerData" target="_blank" rel="noopener">https://github.com/GeeeHesso/PowerData</a>. It consists in two packages and several documentation notebooks. The first package, written in Python, provides functions to handle the data and to generate synthetic series based on historical data. The second package, written in Julia, is used to perform the optimal power flow. The documentation in the form of Jupyter notebooks contains numerous examples on how to use both packages. The entire workflow used to create this dataset is also provided, starting from raw ENTSO-E data files and ending with the synthetic dataset given in the repository.</p> <h2>Funding</h2> <p>This work was supported by the <a href="https://www.cydcampus.admin.ch">Cyber-Defence Campus of armasuisse</a> and by an internal research grant of the Engineering and Architecture domain of <a href="https://www.hes-so.ch">HES-SO</a>.</p>
EmotionCaps: A Synthetic Emotion-Enriched Audio Captioning Dataset
<p>Version 1.0, October 2024</p> <h2>Created by</h2> <p>Mithun Manivannan (1), Vignesh Nethrapalli (1), Mark Cartwright (1)</p> <ol> <li>Sound Interaction and Computer Lab, New Jersey Institute of Technology</li> </ol> <h2>Publication</h2> <p>If using this data in an academic work, please reference the DOI and version, as well as cite the following paper, which presented the data collection procedure and the first version of the dataset:</p> <p>Manivannan, M., Nethrapalli, V., Cartwright, M. EmotionCaps: Enhancing Audio Captioning Through Emotion-Augmented Data Generation. arXiv preprint arXiv:2410.12028, 2024.</p> <h2>Description</h2> <p>EmotionCaps is a ChatGPT-assisted, weakly-labeled audio captioning dataset developed to bridge the gap between soundscape emotion recognition (SER) and automated audio captioning (AAC). Created through a three-stage pipeline, the dataset leverages ground-truth annotations from AudioSet SL, which are enhanced by ChatGPT using tailored prompts and emotions assigned via a soundscape emotion recognition model trained on Emo-Soundscapes Dataset. It comprises four subsets of captions for 120,071 audio clips, each reflecting a different prompt variation: WavCaps-like, Scene-Focused, Emotion Addon, and Emotion Rewrite. The average word counts for these subsets are: WavCaps-like (12.61), Scene-Focused (14.04), Emotion Addon (18.35), and Emotion Rewrite (18.65). The increase in word count for the emotion prompts illustrates the difference in sentence length when integrating emotion information into the captions.</p> <h2>Audio Data</h2> <p>The audio data is from AudioSet SL, the strongly-labled subset of 120,071 audio clips from the larger AudioSet dataset.</p> <h2>Synthetic Captions</h2> <p>The synthetic captions were generated using a three-stage pipeline, beginning with training a soundscape emotion recognition model. This model assesses the valence and arousal of each audio clip, mapping the resulting vector to an emotion identifier. Next, we leveraged the ground-truth annotations from AudioSet SL, and extracted the list of sound events. Using these sound events, we employed ChatGPT to create different variations of captions by applying distinct prompts.</p> <p>We first used the WavCaps prompt for AudioSet SL as a base, the output of which we call WavCaps-like. Building on this, we created three new prompt variations (1) <strong>scene-focused</strong> which is a modified WavCaps prompt that describes the scene, (2) <strong>emotion addon</strong> which is an extension of the scene-Focused prompt, where an emotion is appended to the list of sound events to guide the caption generation, and (3) <strong>emotion rewrite</strong> which consists of two-step prompt where ChatGPT first generates the scene-focused caption, then is instructed to rewrite it with a specific emotion in mind.</p> <p>Using these four prompt styles — WavCaps, Scene-Focused, Emotion Addon, and Emotion Rewrite — along with the AudioSet SL sound events and predicted emotions, we employed ChatGPT-3.5 Turbo to generate four corresponding caption variations for the dataset.</p> <p>Each caption variation has been organized into separate CSV files for clarity and accessibility. All files correspond to the same set of audio clips from AudioSet SL, with the key distinction being the caption variation associated with each clip. The different subsets are designed to be used independently, as they each fulfill specific roles in understanding the impact of emotion in audio captions.</p> <ul> <li> <p><strong>wavcaps-like.csv</strong>: Contains captions generated using the WavCaps prompt, serving as the baseline before emotion is introduced.</p> </li> <li> <p><strong>scene-focused.csv</strong>: Provides captions focused on describing the scene or environment of the audio clip, without emotion integration.</p> </li> <li> <p><strong>emotion-addon.csv</strong>: Captions where emotion data is appended to the scene-focused base caption.</p> </li> <li> <p><strong>emotion-rewrite.csv</strong>: Captions that are completely rewritten based on the scene-focused base caption and the assigned emotion.</p> </li> </ul> <p>This structure allows users to explore how emotional content influences captioning models by comparing the variations both with and without emotional enrichment.</p> <h2>Columns in CSV files</h2> <p><em><strong>segment_id</strong></em> : The ID of the audio recording in AudioSet SL. These are in the form <em><YouTube ID>_<start time in ms>_<end time in ms></em></p> <p><em><strong>caption</strong></em> : The caption generated for each audio clip, corresponding to the specific subset (e.g., WavCaps, Scene-Focused, Emotion Addon, or Emotion Rewrite) as indicated by the file name.</p> <h2>Conditions of use</h2> <p>Dataset created by Mithun Manivannan, Vignesh Nethrapalli, Mark Cartwright</p> <p>The EmotionCaps dataset is offered free of charge under the terms of the Creative Commons Attribution 4.0 International (CC BY 4.0) license:</p> <p><a href="https://creativecommons.org/licenses/by/4.0/">https://creativecommons.org/licenses/by/4.0/</a></p> <p>The dataset and its contents are made available on an “as is” basis and without warranties of any kind, including without limitation satisfactory quality and conformity, merchantability, fitness for a particular purpose, accuracy or completeness, or absence of errors. Subject to any liability that may not be excluded or limited by law, New Jersey Institute of Technology is not liable for, and expressly excludes all liability for, loss or damage however and whenever caused to anyone by any use of the EmotionCaps dataset or any part of it.</p> <h2>Feedback</h2> <p>Please help us improve EmotionCaps by sending your feedback to:</p> <ul> <li>Mithun Manivannan: <a href="mailto:mithun.mani01@gmail.com">mithun.mani01@gmail.com</a></li> <li>Mark Cartwright: <a href="mailto:mcartwright@gmail.com">mcartwright@gmail.com</a></li> </ul> <p>In case of a problem, please include as many details as possible.</p> <h2>Acknowledgments</h2> <p>This work was partially supported by the New Jersey Institute of Technology Honors Summer Research Institute (HSRI).</p>
RAFT synthetic tropical cyclones dataset for Balaguru et al. 2022 - Science Advances
<p>This is the RAFT synthetic tropical cyclone (TC) dataset generated for the paper "Increased US coastal hurricane risk under climate change" submitted to the journal Science Advances in 2022.<br> Each file contains 50,000 synthetic TCs from RAFT either for the historical period (1980-2014) or the future period (2066-2100) under “SSP585”, and from a CMIP6 global climate model.<br> intensity_model_output_corrVMPI_11vars_alltcs_cutoff15_CMIP6_{PERIOD} _{MODEL}.mat<br> To read a .mat file in Python, one can use “scipy.io.loadmat”.<br> There are several variables included in each file, and all have the same dimension [number of storms, number of timesteps]. Here are a list of variable names and what they represent:<br> ‘lat’: Storm latitude;<br> ‘lon’: Storm longitude;<br> ‘year’: year;<br> ‘jday_syn’: Julian day in the year;<br> ‘vs0_syn’: maximum surface wind (knot).<br> Please note that this version of synthetic TC dataset is only intended for assessing the large-scale change of hurricane risk under climate change (through statistical-dynamical downscaling of CMIP6 GCMs), which is addressed in the above mentioned paper. Due to the model biases in CMIP6 and the low temporal resolution (monthly) used for RAFT inputs, the synthetic TCs' life-time maximum intensity is underestimated. Therefore, the synthetic TCs here should not be treated directly as "example TCs of current or future climate" without bias correction on the TC intensity. The authors plan to release a separate version of RAFT simulated synthetic TCs with proper bias correction for localized TC impact assessment. Please email authors if you have questions.</p>
Synthetic dataset for dual-perspective self-supervised learning
<p>The synthetic Ca datasets for training and testing, including training dataset with bidirectional collinear scan (N<sub>y</sub> = 2N<sub>x</sub>) for MP-SSL, training dataset with normal scan (N<sub>y</sub> = N<sub>x</sub>) for TP-SSL, and testing data (N<sub>y</sub> = N<sub>x</sub>)</p> <p>If you use these data simulated using our modified <a href="https://doi.org/10.1016/j.jneumeth.2021.109173">NAOMi</a> model, please cite the corresponding work:</p> <p><a href="https://doi.org/10.1186/s43074-023-00117-0"><strong><span>B. Shen</span></strong><span>, C. Luo, W. Pang, Y. Jiang, W. Wu, R. Hu, J. Qu, B. Gu, L. Liu. Surmounting photon limits and motion artifacts for biological dynamics imaging via dual-perspective self-supervised learning. PhotoniX 5, 1 (2024). </span></a></p>
Synthetic Training Dataset for Real-World Terminal Strip Object Detection
<p>This dataset provides synthetic training data for the real-world industrial application of terminal strip object detection to investigate the sim-to-real generalization performance of modern object detectors based on state-of-the-art image synthesis methods. It consists of 30.000 randomly generated synthetic images of terminal strips covering 36 different terminal blocks in five colors and additional accessories such as plug-in bridges, test adapters, end covers and markings. Except from the markings and the DIN rail all objects of the terminal strips are labeled with a bounding box and the respective object class for supervised learning. Additionally, 300 real images of terminal strips were taken and manually labeled for the real-world test.</p> <p>If you use this datset for your research, please consider citing this: <a href="https://arxiv.org/abs/2403.04809">Investigation of the Impact of Synthetic Training Data in the Industrial Application of Terminal Strip Object Detection</a></p>
Synthetic Datasets from the Article titled Privacy-preserving Ground-truth Data for Evaluating Additive Feature Attribution in Regression Models with Additive CBR and CQV
<p>Synthetic datasets were generated as benchmarks capturing the intrinsic characteristics of original data to investigate the performance of additive feature attribution methods for regression tasks. The synthetic datasets were generated based on 2, 6 and 8 clusters formed with the original data. The 6-cluster dataset was used for primary analysis and the other two were used for sensitivity analysis.</p><p>The synthetic dataset was generated from the original data acquired from <a href="https://www.eurocontrol.int/dashboard/rnd-data-archive">Aviation Data for Research Repository</a>, which was collected and processed by <a href="https://www.eurocontrol.int/">EUROCONTROL</a> from the Enhanced Tactical Flow Management System (ETFMS) flight data messages containing all flights in Europe throughout the year 2019, from May to October. The original dataset consisted of fundamental details of the flights, flight status, preceding flight legs, ATFM regulations, weather conditions, calendar information, etc. </p><p>A brief description of the columns in the synthetic data files is presented in the file 'data_description.pdf' and a more detailed discussion on features can be found in the works of Koolen and Coliban [1] and Dalmau et al. [2].</p><p> </p><p><strong>References</strong><br>[1] H. Koolen and I. Coliban, <a href="https://www.eurocontrol.int/sites/default/files/2020-06/flight-progress-msg-update-230620.pdf">Flight Progress Messages Document</a>, EUROCONTROL, Brussels, Belgium, Tech. Rep., 2020.<br>[2] R. Dalmau, F. Ballerini, H. Naessens, S. Belkoura, and S. Wangnick, <a href="https://www.sciencedirect.com/science/article/pii/S0969699721000739">An Explainable Machine Learning Approach to Improve Take-off Time Predictions</a>, Journal of Air Transport Management, vol. 95, p. 102 090, Aug. 2021. doi: 10.1016/j.jairtraman.2021.102090.</p><p><br> </p>
Small spots DeepMIB project, synthetic dataset for testing 2D semantic segmentation
<p>A complete DeepMIB project with a synthetic dataset generated for quick tests of semantic segmentation approaches.<br>The dataset includes a trained U-net network for detection of small random spots of 2 colors on a black background.</p><p>The network can be opened by loading "2D_SmallSpots_3cl_Unet.mibCfg" file by</p><ul><li><i>MIB->Menu->Tools->Deep learning segmentation->Options tab->Config files->Load </i></li><li>Drag and drop of the config file into DeepMIB window</li></ul><p>Microscopy Image Browser: <a href="https://mib.helsinki.fi">https://mib.helsinki.fi</a></p>
Large spots DeepMIB project, synthetic dataset for testing 2D patch-wise segmentation
<p>A complete DeepMIB project with a synthetic dataset generated for quick tests of the patch-wise segmentation (classification) approaches<br>The dataset includes a trained Resnet18 network for detection of patches that belong to large white spots on black background.<br>The network can be opened by loading "2D_LargeSpots_Patchwise_Resnet18.mibCfg" by</p><ul><li><i>MIB->Menu->Tools->Deep learning segmentation->Options tab->Config files->Load </i></li><li>Drag and drop of the config file into DeepMIB window</li></ul><p>Microscopy Image Browser: <a href="https://mib.helsinki.fi">https://mib.helsinki.fi</a></p>
Large Spots DeepMIB project, synthetic dataset for testing 2.5D semantic segmentation
<p>A complete DeepMIB project with a synthetic dataset generated for quick tests of 2.5D semantic segmentation approaches.<br>The dataset includes trained</p><ul><li>2.5D DeepLabV3-Resnet18 depth2color (Spots_25D_DLv3RN18_Z2C_xy200z5)</li><li>2.5D U-net depth2color (Spots_25D_Unet_Z2C_xy200z5)</li></ul><p>networkd for detection of large white 3D spots on a black background. The spots that are present on a single slice only are considered as background.</p><p>The network can be opened by loading the config files "*.mibCfg" by</p><ul><li><i>MIB->Menu->Tools->Deep learning segmentation->Options tab->Config files->Load </i></li><li>Drag and drop of the config file into DeepMIB window</li></ul><p>Microscopy Image Browser: <a href="https://mib.helsinki.fi/">https://mib.helsinki.fi</a></p>
Synthetic GW dataset with latent diffusion models
<p>SyntheticHTR: Handwritten Text Image Synthesis based on Latent Diffusion Models</p>
Synthetic basement depth, gravity anomalies, density and observation points training dataset to train deep learning model
<p>Contains 200000 training data for our DNN model.</p>
SynthRAD2023 Grand Challenge test dataset: synthetizing computed tomography for radiotherapy
<div> <p><strong>DATASET STRUCTURE</strong></p> <p>The dataset can be downloaded from <a href="https://doi.org/10.5281/zenodo.7260705">https://doi.org/10.5281/zenodo.10514185</a> and a detailed description is offered at "synthRAD2023_dataset_description.pdf".</p> <p>The<strong> test datasets</strong> for Task1 is in Task1.tar.zst, while for Task2 in Task2.tar.zst. After unzipping, each Task is organized according to the following folder structure:</p> <p>Task1<br>└── mr<br>│ ├── brain<br>│ │ ├── 1BA003<br>│ │ │ ├── mask.nii.gz<br>│ │ │ └── mr.nii.gz<br>│ │ ├── 1BA029<br>│ │ │ ├── mask.nii.gz<br>│ │ │ └── mr.nii.gz<br>│ │ ├── ...<br>│ └── pelvis<br>│ ├── 1PA003<br>│ │ ├── mask.nii.gz<br>│ │ └── mr.nii.gz<br>│ ├── 1PA006<br>│ │ ├── mask.nii.gz<br>│ │ └── mr.nii.gz<br>│ ├── ...<br>├── ct<br>│ ├── 1BA003.nii.gz<br>│ ├── 1BA029.nii.gz<br>│ ├── 1BA063.nii.gz<br>│ ├── ...<br>├── doseplanning<br>│ ├── 1BA003<br>│ │ ├── 1BA003_RTplan_photons.mat<br>│ │ └── 1BA003_RTplan_protons.mat<br>│ ├── 1BA029<br>│ │ ├── 1BA029_RTplan_photons.mat<br>│ │ ├── 1BA029_RTplan_protons.mat<br>│ ├── ...<br><br></p> <p>Task2<br>├── cbct<br>│ ├── brain<br>│ │ ├── 2BA019<br>│ │ │ ├── cbct.nii.gz<br>│ │ │ └── mask.nii.gz<br>│ │ ├── 2BA021<br>│ │ │ ├── cbct.nii.gz<br>│ │ │ └── mask.nii.gz<br>│ │ ├── ...<br>│ └── pelvis<br>│ ├── 2PA022<br>│ │ ├── cbct.nii.gz<br>│ │ └── mask.nii.gz<br>│ ├── 2PA023<br>│ │ ├── cbct.nii.gz<br>│ │ └── mask.nii.gz<br>│ ├── ... <br>├── ct<br>│ ├── 2BA019.nii.gz<br>│ ├── 2BA021.nii.gz<br>│ ├── 2BA022.nii.gz<br>│ ├── ....<br>├── doseplanning<br>│ ├── 2BA019<br>│ │ ├── 2BA019_RTplan_photons.mat<br>│ │ └── 2BA019_RTplan_protons.mat<br>│ ├── 2BA021<br>│ │ ├── 2BA021_RTplan_photons.mat<br>│ │ └── 2BA021_RTplan_protons.mat<br>│ ├── ...<br><br></p> <p>Each patient folder has a unique name that contains information about the task, anatomy, center and a patient ID. The naming follows the convention below:</p> <table> <tbody> <tr> <td>[Task]</td> <td>[Anatomy]</td> <td>[Center]</td> <td>[PatientID]</td> </tr> <tr> <td>1</td> <td>B</td> <td>A</td> <td>001</td> </tr> </tbody> </table> <p>The zip contains the following structure: </p> <ul> <li> <p>ct/<patient code>.nii.gz: The gold-standard CT image.</p> </li> <li> <p>doseplanning/<patient code>/<patient code>_RTplan_[photons/protons].mat: The matRad photon and proton doseplans for this particular patient. </p> </li> <li> <p>[mr/cbct]/[brain/pelvis]/<patient code>/[mr/cbct].nii.gz: The corresponding CBCT or MR image.</p> </li> <li> <p>[mr/cbct]/[brain/pelvis]/<patient code>/mask.nii.gz: image containing a binary mask of the dilated patient outline.</p> </li> </ul> <p><strong>DATASET DESCRIPTION</strong></p> <p>This challenge dataset contains imaging data of patients who underwent radiotherapy in the brain or pelvis region. Overall, the population is predominantly adult and no gender restrictions were considered during data collection. For Task 1, the inclusion criteria were the acquisition of a CT and MRI during treatment planning while for task 2, acquisitions of a CT and CBCT, used for patient positioning, were required. Datasets for task 1 and 2 do not necessarily contain the same patients, given the different image acquisitions for the different tasks.</p> <p>Data was collected at 3 Dutch university medical centers:</p> <ul> <li> <p>Radboud University Medical Center</p> </li> <li> <p>University Medical Center Utrecht</p> </li> <li> <p>University Medical Center Groningen</p> </li> </ul> <p>For anonymization purposes, from here on, institution names are substituted with A, B and C, without specifying which institute each letter refers to.</p> <p>The following number of patients is available in the training set.</p> <p><strong>Training</strong></p> <table> <tbody> <tr> <td> </td> <td> <p><strong>Brain</strong></p> </td> <td> <p><strong>Pelvis</strong></p> </td> </tr> <tr> <td> </td> <td> <p><strong>Center A</strong></p> </td> <td> <p><strong>Center B</strong></p> </td> <td> <p><strong>Center C</strong></p> </td> <td> <p><strong>Total</strong></p> </td> <td> <p><strong>Center A</strong></p> </td> <td> <p><strong>Center B</strong></p> </td> <td> <p><strong>Center C</strong></p> </td> <td> <p><strong>Tota</strong>l</p> </td> </tr> <tr> <td> <p><strong>Task 1</strong></p> </td> <td> <p>60</p> </td> <td> <p>60</p> </td> <td> <p>60</p> </td> <td> <p>180</p> </td> <td> <p>120</p> </td> <td> <p>0</p> </td> <td> <p>60</p> </td> <td> <p>180</p> </td> </tr> <tr> <td> <p><strong>Task 2</strong></p> </td> <td> <p>60</p> </td> <td> <p>60</p> </td> <td> <p>60</p> </td> <td> <p>180</p> </td> <td> <p>60</p> </td> <td> <p>60</p> </td> <td> <p>60</p> </td> <td> <p>180</p> </td> </tr> </tbody> </table> <p>Each subset generally contains equal amounts of patients from each center, except for task 1 brain, where center B had no MR scans available. To compensate for this, center A provided twice the number of patients than in other subsets.</p> <p><strong>Validation</strong></p> <table> <tbody> <tr> <td> </td> <td> <p><strong>Brain</strong></p> </td> <td> <p><strong>Pelvis</strong></p> </td> </tr> <tr> <td> </td> <td> <p><strong>Center A</strong></p> </td> <td> <p><strong>Center B</strong></p> </td> <td> <p><strong>Center C</strong></p> </td> <td> <p><strong>Total</strong></p> </td> <td> <p><strong>Center A</strong></p> </td> <td> <p><strong>Center B</strong></p> </td> <td> <p><strong>Center C</strong></p> </td> <td> <p><strong>Tota</strong>l</p> </td> </tr> <tr> <td> <p><strong>Task 1</strong></p> </td> <td> <p>10</p> </td> <td> <p>10</p> </td> <td> <p>10</p> </td> <td> <p>30</p> </td> <td> <p>20</p> </td> <td> <p>0</p> </td> <td> <p>10</p> </td> <td> <p>30</p> </td> </tr> <tr> <td> <p><strong>Task 2</strong></p> </td> <td> <p>10</p> </td> <td> <p>10</p> </td> <td> <p>10</p> </td> <td> <p>30</p> </td> <td> <p>10</p> </td> <td> <p>10</p> </td> <td> <p>10</p> </td> <td> <p>30</p> </td> </tr> </tbody> </table> <p><strong>Testing</strong></p> <table> <tbody> <tr> <td> </td> <td> <p><strong>Brain</strong></p> </td> <td> <p><strong>Pelvis</strong></p> </td> </tr> <tr> <td> </td> <td> <p><strong>Center A</strong></p> </td> <td> <p><strong>Center B</strong></p> </td> <td> <p><strong>Center C</strong></p> </td> <td> <p><strong>Total</strong></p> </td> <td> <p><strong>Center A</strong></p> </td> <td> <p><strong>Center B</strong></p> </td> <td> <p><strong>Center C</strong></p> </td> <td> <p><strong>Total</strong></p> </td> </tr> <tr> <td> <p><strong>Task 1</strong></p> </td> <td> <p>20</p> </td> <td> <p>20</p> </td> <td> <p>20</p> </td> <td> <p>60</p> </td> <td> <p>40</p> </td> <td> <p>0</p> </td> <td> <p>20</p> </td> <td> <p>60</p> </td> </tr> <tr> <td> <p><strong>Task 2</strong></p> </td> <td> <p>20</p> </td> <td> <p>20</p> </td> <td> <p>20</p> </td> <td> <p>60</p> </td> <td> <p>20</p> </td> <td> <p>20</p> </td> <td> <p>20</p> </td> <td> <p>60</p> </td> </tr> </tbody> </table> <p>In total, for all tasks and anatomies combined, 1080 image pairs (720 training, 120 validation, 240 testing) are available in this dataset. <strong>This repository only contains the test data.</strong></p> <p>All images were acquired with the clinically used scanners and imaging protocols of the respective centers and reflect typical images found in clinical routine. As a result, imaging protocols and scanner can vary between patients. A detailed description of the imaging protocol for each image, can be found in spreadsheets that are part of the dataset release (see dataset structure).</p> <p>Data was acquired with the following scanners:</p> <ul> <li> <p>Center A:</p> <ul> <li> <p>MRI: Philips Ingenia 1.5T/3.0T</p> </li> <li> <p>CT: Philips Brilliance Big Bore or Siemens Biograph20 PET-CT</p> </li> <li> <p>CBCT: Elekta XVI</p> </li> </ul> </li> <li> <p>Center B:</p> <ul> <li> <p>MRI: Siemens MAGNETOM Aera 1.5T or MAGNETOM Avanto_fit 1.5T</p> </li> <li> <p>CT: Siemens SOMATOM Definition AS</p> </li> <li> <p>CBCT: IBA Proteus+ or Elekta XVI</p> </li> </ul> </li> <li> <p>Center C:</p> <ul> <li> <p>MRI: Siemens Avanto fit 1.5T or Siemens MAGNETOM Vida fit 3.0T</p> </li> <li> <p>CT: Philips Brilliance Big Bore</p> </li> <li> <p>CBCT: Elekta XVI</p> </li> </ul> </li> </ul> <p>For task 1, MRIs were acquired with a T1-weighted gradient echo or an inversion prepared - turbo field echo (TFE) sequence and collected along with the corresponding planning CTs for all subjects. The exact acquisition parameters vary between patients and centers. For centers B and C, selected MRIs were acquired with Gadolinium contrast, while the selected MRIs of center A were acquired without contrast.</p> <p>For task 2, the CBCTs used for image-guided radiotherapy ensuring accurate patient position were selected for all subjects along with the corresponding planning CT.</p> <p>The following pre-processing steps were performed on the data:</p> <ul> <li> <p>Conversion from dicom to compressed nifti (nii.gz)</p> </li> <li> <p>Rigid registration between CT and MR/CBCT</p> </li> <li> <p>Anonymization (face removal, only for brain patients)</p> </li> <li> <p>Patient outline segmentation (provided as a binary mask)</p> </li> <li> <p>Crop MR/CBCT, CT and mask to remove background and reduce file sizes</p> </li> </ul> <p>The code used to preprocess the images can be found at: <a href="https://github.com/SynthRAD2023/">https://github.com/SynthRAD2023/</a>. Detailed information about the dataset are provided in SynthRAD2023_dataset_description.pdf published here along with the data and will also be submitted to Medical Physics.</p> <p><strong>ETHICAL APPROVAL</strong></p> <p>Each institution received ethical approval from their internal review board/Medical Ethical committee:</p> <ul> <li> <p>UMC Utrecht approved not-WMO on 4/03/2022 with number 22/474 entitled: “Synthetizing computed tomography for radiotherapy Grand Challenge (SynthRAD)”.</p> </li> <li> <p>UMC Groningen approved not-WMO on 20/07/2022 with number 202200310 entitled: “Synthesizing computed tomography for radiotherapy - Grand Challenge”.</p> </li> <li> <p>Radboud UMC declared the study not-WMO on 17/10/2022 with number 2022-15950 entitled “Synthetizing computed tomography for radiotherapy Grand Challenge”.</p> </li> </ul> <p><strong>CHALLENGE DESIGN</strong></p> <p>The overall challenge design can be found at <a href="https://doi.org/10.5281/zenodo.7746020">https://doi.org/10.5281/zenodo.7746020</a>. </p> <h2>Notes</h2> <div>FUNDING BODIES: The challenge has been funded thanks to the support of the Seed Fund provided by the " EWUU Alliance TU/e, WUR, UU, UMCU" https://ewuu.nl/en/collaboration/seed-fund/.</div> </div>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.