Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

250

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

250 results for “Synthetic dataset”

Learn how ShareScore rates datasets ↗
zenodo40/100

TIMIT-TTS: a Text-to-Speech Dataset for Synthetic Speech Detection

<p>With the rapid development of deep learning techniques, the generation and counterfeiting of multimedia material are becoming increasingly straightforward to perform. At the same time, sharing fake content on the web has become so simple that malicious users can create unpleasant situations with minimal effort. Also, forged media are getting more and more complex, with manipulated videos (e.g., deepfakes where both the visual and audio contents can be counterfeited) that are taking the scene over still images.<br> The multimedia forensic community has addressed the possible threats that this situation could imply by developing detectors that verify the authenticity of multimedia objects. However, the vast majority of these tools only analyze one modality at a time.<br> This was not a problem as long as still images were considered the most widely edited media, but now, since manipulated videos are becoming customary, performing monomodal analyses could be reductive. Nonetheless, there is a lack in the literature regarding multimodal detectors (systems that consider both audio and video components). This is due to the difficulty of developing them but also to the scarsity of datasets containing forged multimodal data to train and test the designed algorithms.</p> <p>In this paper we focus on the generation of an audio-visual deepfake dataset.<br> First, we present a general pipeline for synthesizing speech deepfake content from a given real or fake video, facilitating the creation of counterfeit multimodal material. The proposed method uses Text-to-Speech (TTS) and Dynamic Time Warping (DTW) techniques to achieve realistic speech tracks. Then, we use the pipeline to generate and release TIMIT-TTS, a synthetic speech dataset containing the most cutting-edge methods in the TTS field. This can be used as a standalone audio dataset, or combined with DeepfakeTIMIT and VidTIMIT video datasets to perform multimodal research. Finally, we present numerous experiments to benchmark the proposed dataset in both monomodal (i.e., audio) and multimodal (i.e., audio and video) conditions.<br> This highlights the need for multimodal forensic detectors and more multimodal deepfake data.</p> <ul> <li>For the initial version of TIMIT-TTS&nbsp;<strong>v1.0</strong> <ul> <li>Arxiv: https://arxiv.org/abs/2209.08000</li> <li>TIMIT-TTS Database v1.0: https://zenodo.org/record/6560159</li> </ul> </li> </ul>

opencc-by-4.0Sep 2022View details →
zenodo40/100

TrueFace: a Dataset for the Detection of Synthetic Face Images from Social Networks

<p>TrueFace is a first dataset of social media processed real and synthetic faces, obtained by the successful StyleGAN generative models, and shared on Facebook, Twitter and Telegram.</p> <p>Images have historically been a universal and cross-cultural communication medium, capable of reaching people of any social background, status or education. Unsurprisingly though, their social impact has often been exploited for malicious purposes, like spreading misinformation and manipulating public opinion. With today&#39;s technologies, the possibility to generate highly realistic fakes is within everyone&#39;s reach. A major threat derives in particular from the use of synthetically generated faces, which are able to deceive even the most experienced observer. To contrast this fake news phenomenon, researchers have employed artificial intelligence to detect synthetic images by analysing patterns and artifacts introduced by the generative models. However, most online images are subject to repeated sharing operations by social media platforms. Said platforms process uploaded images by applying operations (like compression) that progressively degrade those useful forensic traces, compromising the effectiveness of the developed detectors. To solve the synthetic-vs-real problem &quot;in the wild&quot;, more realistic image databases, like TrueFace, are needed to train specialised detectors.</p>

opencc-by-4.0Oct 2022View details →
zenodo40/100

Synthetic single particle cryo-EM dataset of the SARS-CoV-2 spike protein

<p>PDBs were generated using molecular dynamics.<br> See DESRES_README.txt for more details on molecular dynamics simulation.<br> PDBs were converted to volumetric data using EMAN2.<br> The image stack contains 100 000 projection images each&nbsp;<br> of the 10 states (see PDBs), at an SNR of 1/10 in the following order:</p> <p>state00 (closed)<br> state01 (closed)<br> state02 (closed)<br> state10 (intermediate)<br> state11 (intermediate)<br> state12 (intermediate)<br> state13 (intermediate)<br> state20 (open)<br> state21 (open)<br> state22 (open)</p> <p>Projections were made using relion_project.&nbsp;<br> &nbsp;&nbsp;White gaussian noise with standard deviation 1.0<br> &nbsp;&nbsp;CTF multiplied signal<br> &nbsp;&nbsp;High signal-to-noise ratio<br> &nbsp;&nbsp;Image size 96x96x96<br> &nbsp;&nbsp;<br> MRC-files used for the projections not included, but can be generated using the PDB files.<br> Final RELION reconstruction resolution is 5.33334 Angstrom (Nyqvist is at 5.33334).</p> <p>Command line for RELION reconstruction:<br> relion_refine_mpi --o refine3d/run --auto_refine --split_random_halves --i rot_trans_ctf_noise/stack.star --ref pdb2mrc/state21.mrc --ini_high 20 --dont_combine_weights_via_disc --preread_images --pool 30 --pad 2 --ctf --particle_diameter 130 --flatten_solvent --zero_mask --oversampling 1 --healpix_order 2 --auto_local_healpix_order 4 --offset_range 5 --offset_step 2 --low_resol_join_halves 40 --norm --scale --j 2 --gpu --fristiter_cc --grad&nbsp;</p> <p>This dataset is generated as a testbed for cryo-EM heterogeneity analysis.</p>

opencc-by-4.0Oct 2022View details →
zenodo40/100

Synthetic tetrode recording dataset with spike-waveform drift

<p><strong>Introduction</strong></p> <p>This synthetic ground-truth dataset accurately models long-term, continuous extracellular tetrode recordings from the rodent brain over a time-period of 256 hours. Each "recording" comprises spiking of 8 distinct single-units with firing rates ranging from 0.1 - 6 Hz, superimposed on background multi-unit spiking activity at 20 Hz. The recording sampling rate is 30 kHz. Single-unit spike amplitudes drift over a range of 100 to 400 <span class="math-tex">\(\mu V\)</span> based on the drift we observe in our own long-term recordings from the rodent motor cortex and striatum. For more details, please see our paper "Automated long-term recording and analysis of neural activity in behaving animals" ( https://doi.org/10.1101/033266).</p> <p>These recordings can be used to test the accuracy of spike-sorting algorithms when clustering non-stationary spike waveform data, such as our own Fast Automated Spike Tracker (FAST) outlined in our paper and available at https://github.com/Olveczky-Lab/FAST.</p> <p> </p> <p><strong>Dataset</strong></p> <p>Due to size restrictions, we provide here 1 sample tetrode of the full 6 tetrode dataset. Please contact us (https://olveczkylab.oeb.harvard.edu/about) if you require access to the other 5 synthetic tetrode recordings.</p> <p> </p> <p><strong>Instructions</strong></p> <p>The dataset comprises spike times and spike waveform snippets extracted a continuous synthetic tetrode recording. Provided are...</p> <ul> <li>A <strong>SpikeTimes</strong> file with a list of sample numbers for detected events (spikes) at <em>uint64</em> precision.</li> <li>A <strong>Spikes</strong> file with the waveforms of the detected events in <em>int16</em> precision. Each event waveform comprises <em>4 channels X 64 samples</em> 16-bit words arranged in the order [Ch0-Sample0, Ch1-Sample0, Ch2-Sample0, Ch3-Sample0, Ch0-Sample1, etc.]. To convert to units of voltage, change type to double precision and multiply by 1.95e-7.</li> <li>A <strong>SnippeterSettings.xml</strong> file with snippeting parameters (this is auto-generated by the FAST snippeting algorithm).</li> <li>A <strong>dataset_params.mat</strong> MATLAB data file containing the simulation parameters. The most important variables in the mat file are <em>sp</em> which contains a list of true spike-times (in samples @ 30 kHz) for all single-units in the dataset, and <em>sp_u</em> which specifies which unit (1-8) each spike originates from. Spike-times are generated by a homogenous Poisson process with firing rate specified for each unit by the variable <em>uFRs</em> and an absolute refractory period of 2 ms. The variable <em>d_Amps</em> specifies the amplitude of each single-unit spikes. The basic spike-waveform shape of each unit is provided in the variable <em>uWVs</em>. The spike-times and identity of background (multi-unit) spikes are specified in <em>b_sp</em> and <em>b_sp_u</em>. </li> </ul>

opencc-by-4.0Sep 2017View details →
zenodo40/100

Synthetic dataset for end-to-end Relation Extraction of relationships between Organisms and Natural-Products with Mixtral-8x7B-Instruct-v0.1

<p>A new synthetic dataset (training/validation) for end-to-end Relation Extraction of relationships between Organisms and Natural-Products.</p> <p>The new dataset was generated using <a href="https://huggingface.co/mistralai/Mixtral-8x7B-Instruct-v0.1">Mixtral-8x7B-Instruct-v0.1</a>. Like the model, the produced synthetic data are also submitted to the License of the model used for generation (apache-2.0).</p> <p>The new dataset was created based on the top-1000 (per biological kingdom) LOTUS literature references extracted with the <a href="https://github.com/idiap/gme-sampler">GME-sampler</a>.</p> <p>The dataset contains 8,913 items in the training set and 344 items in the validation set.</p> <p>The dataset was generated using the same protocol as described in the article.</p>

opencc-by-4.0Mar 2024View details →
zenodo40/100

Synthetic Shape Dataset for Numerosity: Exploring Six Configurations of Complexity

<p>The Synthetic Shape Dataset for Numerosity is a meticulously crafted collection designed to advance the understanding and training of machine learning models in the realm of numerosity perception, mirroring human neurocognitive abilities. Comprising six distinct configurations ranging from fundamental to intricate, this dataset offers a comprehensive exploration of shape complexity.</p> <p>Each configuration presents a unique array of synthetic images, where shapes are dynamically generated and randomly distributed against contrasting backgrounds. Configuration 1 serves as the foundation, featuring only white circles against a black background, with each circle sharing a uniform size. As complexity escalates through the configurations, additional elements are introduced, including variations in shape type, size, orientation, and pixel intensity.</p> <p>One notable feature of this dataset is that no shapes overlap or touch each other, ensuring clarity and precision in each image. The total dataset comprises 73,686 images, with each configuration meticulously crafted to offer distinct challenges for numerosity perception tasks.</p> <p>The shape generation process is divided into two types of shape sizes:</p> <ul> <li> <p><strong>Bounded:</strong> In this category, there is no correlation between the pixel count for shapes contained in an image and the target numerosity count. Each of the six configurations contains 9,212 images, totaling 55,272 images.</p> </li> <li> <p><strong>Unbounded:</strong> Here, there is a correlation between the pixel count for shapes in an image and the target numerosity count. Each configuration consists of 3,069 images, totaling 18,414 images.</p> </li> </ul> <p>The configurations are as follows:</p> <ol> <li> <p><strong>Configuration 1:</strong> Features only white circles against a black background, with uniform circle size.</p> </li> <li> <p><strong>Configuration 2:</strong> Similar to Configuration 1, but circles vary in size.</p> </li> <li> <p><strong>Configuration 3:</strong> Presents a black background with full white circles, triangles, squares, and pentagons. Shapes have a uniform orientation but do not share a uniform size.</p> </li> <li> <p><strong>Configuration 4:</strong> Similar to Configuration 3, but shapes do not have a uniform orientation.</p> </li> <li> <p><strong>Configuration 5:</strong> Similar to Configuration 4, but shapes also vary in pixel intensity, exhibiting different shades of grey.</p> </li> <li> <p><strong>Configuration 6:</strong> Features a white background with full black circles, triangles, squares, and pentagons. Shapes do not have a uniform orientation.</p> </li> </ol> <p>The dataset is split into training and test sets:</p> <ul> <li> <p><strong>Training Set:</strong> Contains 61,440 images, with each image containing 1-8 shapes, evenly distributed across target numerosity counts.</p> </li> <li> <p><strong>Test Set:</strong> Comprises 12,246 images, including a variety of numerosity counts ranging from 0 to 12 shapes per image.</p> </li> </ul> <p>Additionally, each image is accompanied by structured label information provided in a CSV file:</p> <ul> <li><strong>id:</strong> A unique number assigned to each image.</li> <li><strong>config:</strong> Indicates the configuration of the image.</li> <li><strong>target:</strong> Denotes the label for the image, representing the number of shapes contained within it.</li> <li><strong>shape:</strong> Specifies whether the image contains bounded or unbounded shapes. Options include 'bounded' or 'unbounded'.</li> <li><strong>numerosity_id:</strong> Identifies the image in the creation process (not crucial for end-users).</li> <li><strong>shape_id:</strong> Identifies the image in the creation process (not crucial for end-users).</li> <li><strong>split:</strong> Labels each image as belonging to either the training or test dataset. Options are 'train' or 'test'.</li> <li><strong>path:</strong> Provides the path to the image file.</li> </ul> <p>This dataset serves as a valuable resource for researchers seeking to explore and enhance machine learning models' ability to comprehend numerosity in various contexts.</p>

opencc-by-4.0May 2024View details →
zenodo40/100

iHelp public synthetic dataset - model training

<p>This is the model training version of the iHelp public synthetic dataset that has been synthesized from primary and secondary data collected during the Medical University Plovdiv pilot implementing their cancer program. The difference from the parent dataset is that the mean pain in the third interval has been selected as the outcome to be predicted, and the dataset is split into input and output vectors for training, validation and testing. The dataset is described in the "iHelp public synthetic dataset - model training.docx" file.</p>

opencc-by-4.0Jun 2024View details →
zenodo40/100

Datasets for "Deep learning-based design of synthetic orthologs of SH3 signaling domains"

<p>Description of data for "Deep learning-based design of synthetic orthologs of SH3 signaling domains":<br><br>biochemistry_data.zip --&gt; contains the binding assay, melting temperature, and enthalpy measurements that reproduce table 1 in the main text.<br>sequence_data.zip --&gt; contains the sequences with relative enrichment measurements and other meta data information (e.g. latent embeddings, paralog labels, etc.).<br><br>sequence_data.zip &gt; SH3_Library_Natural.xlsx --&gt; contains the natural alleles with normalized relative enrichment scores, paralog labels, mmd latent coordinates, and among other meta data.<br>sequence_data.zip &gt; SH3_Library_Design.xlsx --&gt; contains the design alleles with normalized relative enrichment scores, mmd latent coordinates, and among other meta data.<br>sequence_data.zip &gt; paralog_mapping.xlsx --&gt; contains the mapping between paralog name and COG labels.<br><br>note: to access these spreadsheets for analysis, we recommend using pandas library in Python. To reproduce figures within the manuscript, please follow this github repo link: https://github.com/chemgeeklian/SH3_orthology_paper_analysis</p>

opencc-by-4.0Jul 2024View details →
zenodo40/100

X-ray diffraction dataset for PDB 9G3K LecB from PA01 in complex with synthetic beta-fucosylamide

<p>X-ray diffraction images collected on proxima 2 Soleil &nbsp;the 17th of november 2023 at SOLEIL synchrotron, Saint Aubin, France for PDB ID 9G3K using a DECTRIS EIGER X 9M detector. X-ray dataset and xdsme processing for the structure of LecB from Pseudomonas aeruginosa PA01 strain in complex with synthetic beta-fucosylamide-furan-phenyl derivative. Data were cut at 1.55 A.</p>

opencc-by-4.0Jul 2024View details →
zenodo40/100

Synthetic Dataset for Outlier Detection

<p>This synthetically generated dataset can be used to evaluate outlier detection algorithms. It has 10 attributes and 1000 observations, of which 100 are&nbsp;labeled as outliers. Two-dimensional combinations of attributes form differently shaped clusters.</p> <ul> <li>Attribute 0 &amp; Attribute&nbsp;1: Two circular clusters</li> <li>Attribute&nbsp;2 &amp; Attribute&nbsp;3: Two banana shaped clusters</li> <li>Attribute&nbsp;4 &amp; Attribute&nbsp;5: Three point clouds</li> <li>Attribute&nbsp;6 &amp; Attribute&nbsp;7: Two point clouds with variances</li> <li>Attribute&nbsp;8 &amp; Attribute&nbsp;9: Three anisotropic shaped clusters.&nbsp;</li> </ul> <p>The &quot;outlier&quot; column states whether an observation is an outlier or not. Additionally, the .zip file contains 10 stratified randomized train test splits (70% train, 30% test).</p>

opencc-by-4.0Feb 2018View details →
zenodo40/100

UDNN_Synthetic_Datasets

<p>These are the scripts in order to create the synthetic datasets used for training and testing of the UDNN network. In the Phantom3DLibrary.dat the corresponding phantoms are described as list of geometric shapes. The Create_Synth_Projections.py contains the function for the generantion of the training and testing datasets with noisy projections. For the execution of the function the TomoPhantom toolbox (https://doi.org/10.5281/zenodo.1232424) has to be installed. Apart from that, the noiseless training and testing datasets are also given.</p>

opencc-by-4.0Jun 2018View details →
zenodo40/100

BirdVox-scaper-10k: a synthetic dataset for multilabel species classification of flight calls from 10-second audio recordings

<p>BirdVox-scaper-10k: a synthetic dataset for multilabel species classification of flight calls from 10-second audio recordings<br> =============================================================================================<br> Version 1.0, September 2019.</p> <p>&nbsp;</p> <p>Created By<br> -------------</p> <p>Elizabeth Mendoza (1), Vincent Lostanlen (2, 3, 4), Justin Salamon (3, 4), Andrew Farnsworth (2), Steve Kelling (2), and Juan Pablo Bello (3, 4).</p> <p>&nbsp;</p> <p>(1): Forest Hills High School, New York, NY, USA<br> (2): Cornell Lab of Ornithology, Cornell University, Ithaca, NY, USA<br> (3): Center for Urban Science and Progress, New York University, New York, NY, USA<br> (4): Music and Audio Research Lab, New York University, New York, NY, USA</p> <p>https://wp.nyu.edu/birdvox</p> <p>&nbsp;</p> <p>Description<br> --------------</p> <p>The BirdVox-scaper-10k dataset contains 9983 artificial soundscapes. Each soundscape lasts exactly ten seconds and contains one or several avian flight calls from up to 30 different species of New World warblers (Parulidae). Alongside each audio file, we include an annotation file describing the start time and end time of each flight call in the corresponding soundscape, as well as the species of warbler it belongs to.</p> <p>In order to synthesize soundscapes in BirdVox-scaper-10k, we mixed natural sounds from various pre-recorded sources. First, we extracted isolated recordings of flight calls containing little or no background noise from the CLO-43SD dataset [1]. Secondly, we extracted 10-second &quot;empty&quot; acoustic scenes from the BirdVox-DCASE-20k dataset [2]. These acoustic scenes contain various sources of real-world background noise, including biophony (insects) and anthropophony (vehicles), yet are guaranteed to be devoid of any flight calls. Lastly, we &quot;fill&quot; each acoustic scene by mixing it with flight calls sampled at random.</p> <p>Although the BirdVox-scaper-10k does not consist of natural recordings, we have taken several measures to ensure the plausibility of each synthesized soundscape, both from qualitative and quantitative standpoints.<br> <br> The BirdVox-scaper-10k dataset can be used, among other things, for the research, development, and testing of bioacoustic classification models.</p> <p>For details on the hardware of ROBIN recording units, we refer the reader to [2].</p> <p>[1] J. Salamon, J. Bello. Fusing shallow and deep learning for bioacoustic bird species classification. Proc. IEEE ICASSP, 2017.</p> <p>[2] V. Lostanlen, J. Salamon, A. Farnsworth, S. Kelling, and J. Bello. BirdVox-full-night: a dataset and benchmark for avian flight call detection. Proc. IEEE ICASSP, 2018.</p> <p>[3] J. Salamon, J. P. Bello, A. Farnsworth, M. Robbins, S. Keen, H. Klinck, and S. Kelling. Towards the Automatic Classification of Avian Flight Calls for Bioacoustic Monitoring. PLoS One, 2016.</p> <p>&nbsp;</p> <p>&nbsp;</p> <p>&nbsp;</p> <p>@inproceedings{lostanlen2018icassp,<br> &nbsp; title = {BirdVox-full-night: a dataset and benchmark for avian flight call detection},<br> &nbsp; author = {Lostanlen, Vincent and Salamon, Justin and Farnsworth, Andrew and Kelling, Steve and Bello, Juan Pablo},<br> &nbsp; booktitle = {Proc. IEEE ICASSP},<br> &nbsp; year = {2018},<br> &nbsp; published = {IEEE},<br> &nbsp; venue = {Calgary, Canada},<br> &nbsp; month = {April},<br> }</p>

opencc-by-4.0Feb 2019View details →
zenodo40/100

Synthetic Lunar Terrain: A Multimodal Open Dataset for Training and Evaluating Neuromorphic Vision Algorithms

<p><strong>Synthetic Lunar Terrain (SLT) </strong>is a dataset based on a reconstruction of a typical <strong>cratered lunar surface landscape&nbsp;</strong>at the <a href="https://set.adelaide.edu.au/atcsr/space-research/exterres-laboratory" target="_blank" rel="noopener">EXTERRES Laboratory</a> at University of Adelaide, Roseworthy Campus. On a surface area of <strong>3.6m x 4.8m</strong>, multiple synthetic craters with different sizes and geometries were sculpted into<strong> lunar regolith simulant</strong>. A <strong>9kW metal-halide lamp</strong> illuminated the scene, providing high contrast drop-shadows from the rims of craters and similar surface features that are characteristic for the Earth's moon.</p> <p>The purpose of this dataset is to provide multimodal recordings of visual information to develop and test algorithms on a hardware analogue of the moon rather than relying on computer simulations. In particular, comparisons between <strong>neuromorphic vision sensors</strong> like <strong>event-based cameras</strong> and imaging with <strong>conventional monocular cameras</strong> are at the core of this work. For this purpose, an event-based camera (Gen4 Prophesee with Prophesee-Sony IMX636 sensor) was mounted downward-pointing next to a optical camera (Basler a2A1920-160ucPRO with Sony IMX392 sensor) on an extendable rod which was moved above the surface in a slow and continuous sweep. In total, SLT consists of camera recordings from 21 different positions/settings, with clockwise and anti-clockwise motions under varying, extreme lighting conditions.</p> <p>The event-stream and grayscale image data can be further referenced via a detailed <strong>3D point cloud</strong>&nbsp;obtained by a FARO Focus S70 3D Scanner. This 3D Scan was post-processed, realigned and resampled into a 3D point cloud of&nbsp;<strong>~6.25M points,&nbsp;</strong>with a surface density of <strong>1.862 p/mm&sup2;</strong>, providing a ground-truth for the crater geometries.</p> <p>In detail, SLT contains the following:</p> <ul> <li>eventbased.zip: <ul> <li><strong>42 camera orbits</strong> in the binary&nbsp;<strong>EVT 3.0</strong> format (<a href="https://docs.prophesee.ai/stable/data/encoding_formats/evt3.html" target="_blank" rel="noopener">Prophesee docs</a>)&nbsp;</li> <li>corresponding <strong>.mp4 </strong>event-frame video rendering for visualization purposes (33.333ms accumulation time at 30FPS)</li> <li>corresponding<strong> .bias</strong> file containing settings used during recording</li> </ul> </li> <li>code.zip: <ul> <li>Standalone C++ code of the <strong>metavision EVT3-to-RAW file decoder</strong>, allowing to convert the binary EVT 3.0 format into a plaintext <strong>.csv&nbsp;</strong>that includes <ul> <li>the coordinates of the event-pixel,</li> <li>the polarity change,</li> <li>and the time-stamp of the event.</li> </ul> </li> <li>This code is an unmodified redistribution from the <a href="https://www.prophesee.ai/metavision-intelligence/" target="_blank" rel="noopener">Metavision SDK</a>, version 4.6.0, released by Prophesee under Apache License 2.0.</li> </ul> </li> <li>&nbsp;optical.zip: <ul> <li><strong>42 image sequences</strong> in <strong>.tif</strong> format (LZW, 1920x1200px, 8bit, grayscale) <ul> <li>Length of image sequences varies between about 300 to 700 images per sequence</li> </ul> </li> </ul> </li> <li>3d_scan.zip: <ul> <li><strong>SLT3d_scan.ply:</strong> 3D point cloud of the scene Stanford Polygon File Format</li> <li><strong>SLT3d_scan.xyz:</strong> 3D point cloud with plaintext x y z coordinates, white-space separated</li> </ul> </li> <li>cratermap.png: <ul> <li>Annotations of <strong>130 different surface features</strong> that have been manually identified as crater-like with approximate x,y-coordinates.</li> </ul> </li> <li>positionmap.png: <ul> <li>Illustration of the different positions from which the rod was moved over the scene (not to scale).</li> </ul> </li> <li>sample.zip: <ul> <li>A sample containing 1 event-camera orbit with the corresponding image sequence (for convenience only, to test the dataset without the need to download it's entirety)</li> </ul> </li> </ul> <p>The global coordinate frame of this dataset puts the origin at the centre of the scene. The shorter side of the terrain is roughly aligned with the x-axis, the longer side with the y-axis. The z-axis represents height/depth (compare with <strong>cratermap.png</strong>). The different conditions (compare with <strong>positionmap.png</strong>) from which the data was taken are encoded as follows:</p> <ul> <li><strong>A1, ..., A9</strong> refer to the left side of the scene (negative x)</li> <li><strong>B1, ..., B9 </strong>refer to the right side of the scene (positive x)</li> <li><strong>S1, S2, S3</strong> and <strong>S4 </strong>describe special lighting conditions and/or parameter settings</li> <li><strong>CW </strong>refers to a "clockwise" sweeping of the camera-rod, relative to the position</li> <li><strong>ACW</strong> refers to an "anti-clockwise" sweeping of the camera-rod, relative to the position</li> </ul> <p>The light from the metal-halide lamp was directed through a small opening, shining along the positive y-axis. In some of the setups, an obstacle was placed between the surface and the opening, blocking out part of the light to create a light-dark separator on the surface, emulating the&nbsp;<strong>terminator</strong> on the Moon between it's day and night side, resulting in highly contrastive images.</p> <p>We encourage you to consult and cite our related publication, should you find SLT useful.</p> <ul> <li>M&auml;rtens, M., Farries, K., Culton, J. and Chin, TJ. "<strong>Synthetic Lunar Terrain: A Multimodal Open Dataset for Training and Evaluating Neuromorphic Vision Algorithms</strong>", Proceedings of&nbsp; "<em>International Symposium on Artificial Intelligence, Robotics and Automation in Space (I-SAIRAS), 2024</em>", pp. 609-614</li> </ul> <p>&nbsp;</p>

opencc-by-4.0Aug 2024View details →
zenodo40/100

Synthetic Aircraft Trajectory Dataset

<p>This dataset comprises synthetically generated aircraft trajectories for multiple specific airport pairs across Europe. Generated using advanced machine learning techniques, the dataset includes high-resolution spatial and temporal information for each trajectory. It features key flight parameters such as latitude, longitude, altitude, and time, along with synthetic identifiers for each flight. This dataset is ideal for air traffic management research, flight path analysis, and the development of predictive models in aviation.</p>

opencc-by-4.0Sep 2024View details →
zenodo40/100

SUSHI🍣: A Dataset of Synthetic Unichannel Signals Based on Heuristic Implementation (Tiny)

<div> <h2>Overview</h2> </div> <p>This dataset consists of pairs of time series signal data and corresponding natural language texts that describe the characteristics of these time series patterns. It has been designed with the objective of creating and evaluating a foundational model that facilitates natural language processing tasks, such as query-by-text retrieval for time series signals and captioning for these signals. All time series signals included in this dataset have been artificially generated from predefined combinations of multiple classes of functions. In addition, the paired natural language texts are primarily based on texts randomly selected from a pre-registered list corresponding to the functions, which have been manually refined through visual inspection.</p> <div> <h2>Specification</h2> </div> <p>We have two versions. This page provides "Tiny" under the Creative Commons Attribution 4.0 International license.</p> <table> <tbody> <tr> <th>Spec</th> <th>Tiny</th> <th>Base</th> </tr> </tbody> <tbody> <tr> <td><strong>Samples</strong></td> <td><strong>1.4K</strong></td> <td>140K</td> </tr> <tr> <td><strong>Time length</strong></td> <td><strong>2048 points</strong></td> <td>2048 points</td> </tr> <tr> <td><strong>Format</strong></td> <td><strong>CSV, NPY, PNG</strong></td> <td>CSV, NPY, PNG</td> </tr> </tbody> </table> <p>In addition, when &ldquo;SUSHI&rdquo; is referred to without the word &ldquo;Tiny&rdquo; in a paper or other document, it shall be taken to mean &ldquo;Base&rdquo;.</p> <h2>Detailed Information</h2> <p>The detailed infromation is availtable in the following documents:</p> <ul> <li>Arxiv preprint: Yohei Kawaguchi, Kota Dohi, and Aoi Ito, "SUSHI: A Dataset of Synthetic Unichannel Signals Based on Heuristic Implementation," in arXiv e-prints: zzzzzzz, 2024. <a href="http://,,,," target="_blank" rel="noopener">http://,,,,</a></li> <li>Githut repository: <a href="https://github.com/y-kawagu/SUSHI/" target="_blank" rel="noopener">https://github.com/y-kawagu/SUSHI/&nbsp;</a></li> </ul> <h2>Citation</h2> <p>If you use this dataset, please cite the following paper:</p> <ul> <li>Yohei Kawaguchi, Kota Dohi, and Aoi Ito, "SUSHI: A Dataset of Synthetic Unichannel Signals Based on Heuristic Implementation," in arXiv e-prints: zzzzzzz, 2024. [URL]</li> </ul>

opencc-by-4.0Sep 2024View details →
zenodo40/100

Efficient Semantic Diffusion Architectures for Model Training on Synthetic Echocardiograms Dataset

<p>This is the official data repository for the paper: "Efficient Semantic Diffusion Architectures for Model Training on Synthetic Echocardiograms", available at:<a href="https://www.arxiv.org/abs/2409.19371"> https://www.arxiv.org/abs/2409.19371</a>. The corresponding code is available at: <a href="https://github.com/david-stojanovski/EDMLX">https://github.com/david-stojanovski/echo_from_noise</a></p> <p>&nbsp;</p> <p>The synthetic data is produced using a variety of generative architectures, including the <strong>Elucidating Diffusion Model (EDM), Variance Exploding (VE), Variance Preserving (VP)</strong>, and our novel models, <strong>EDM-L64</strong> and <strong>EDM-L128</strong>, which employ <strong>latent diffusion</strong> strategies to significantly reduce computational cost. By incorporating&nbsp;<strong>spatially adaptive normalization (SPADE) blocks</strong> and <strong>&Gamma;-distribution-based Variational Autoencoders (&Gamma;-VAE)</strong>, these datasets ensure that the generated images preserve the essential semantic features required for training deep learning models.</p> <p>&nbsp;</p> <p>All pretrained classification and segmentation models can be found within the <strong>trained_models </strong>file.</p> <p>All generated images can be found within the&nbsp;<strong>generated_data&nbsp;</strong>file. Included is the <strong>CAMUS</strong> and original <strong>Semantic Diffusion Model (SDM)&nbsp;</strong>data, as well as a folder labelled&nbsp;<strong>easy_inference</strong> designed to contain all relevant labelmaps in a convenient folder for generating replicas of the dataset (detailed at codebase).</p>

opencc-by-4.0Oct 2024View details →
zenodo40/100

Synthetic LBW Dataset

<p>Dataset generated by Technical University of Wien and converted to compressed HDF5 format by AIMEN Technology Center.</p> <p>This dataset contains synthetic data extracted from LBW numerical simulations. Since the output files from these simulations are very heavy and in complex format, some data from these simulations was extracted to HDF5 files. Data extracted are phase and temperature in the top plate and some process features. These data are intended to study t he dynamics of the process and to find the theoretical image obtained by a perfect thermal imager (comparing these to real camera images could be used to find the function that transforms the image in real cameras.&nbsp;The dataset contains 12 files where each file corresponds to one experiment with a unique combination of process parameters.&nbsp;</p> <p>&copy; COPYRIGHT 2021 The CUSTODIAN Consortium. All rights reserved.</p>

opencc-by-4.0Aug 2021View details →
zenodo40/100

Synthetic dataset for mobile wireless networks with SUMO -- aggregated traces

<p>Aggregated dataset from a published wireless dataset generator base in SUMO mobility model.&nbsp;</p> <p>&nbsp;</p> <p>This work was supported by national funds through Funda&ccedil;&atilde;o para a Ci&ecirc;ncia e a Tecnologia (FCT) with reference UIDB/50021/2020 and SFRH/BD/132053/2017.</p>

opencc-by-4.0Nov 2021View details →
zenodo40/100

Synthetic noisy urban soundscapes: a dataset of synthetic soundscapes with real urban backgrounds

<p><strong>Publication</strong></p> <p>&nbsp;</p> <p>If you use this data in your work, please cite the following paper, which introduced this dataset:</p> <p>&nbsp;</p> <p>[1] Pishdadian, F., Wichern, G., &amp; Le Roux, J. (2020). Finding strength in weakness: Learning to separate sounds with weak supervision. IEEE/ACM Transactions on Audio, Speech, and Language Processing (TASLP). [<a href="https://arxiv.org/pdf/1911.02182">pdf</a>]</p> <p>[2] Cramer, A., Cartwright, M., Pishdadian, F., and Bello, J.P. Weakly Supervised Source-Specific Sound Level Estimation in Noisy Soundscapes. In Proceedings of the IEEE Workshop on Applications of Signal Processing to Audio and Acoustics (WASPAA), 2021. [<a href="https://arxiv.org/pdf/2105.02911">pdf</a>]</p> <p><br> <strong>Created by</strong></p> <p>Fatemeh Pishdadian (1), Gordon Wichern (2), Jonathan Le Roux (2), Aurora Cramer (3, 4), Mark Cartwright (5), and Juan Pablo Bello (3,4,6,7)</p> <p>&nbsp;&nbsp;&nbsp; 1. Interactive Audio Lab, Northwestern University<br> &nbsp;&nbsp;&nbsp; 2. Mitsubishi Electric Research Laboratory<br> &nbsp;&nbsp;&nbsp; 3. Music and Audio Research Lab, New York University<br> &nbsp;&nbsp;&nbsp; 4. Department of Electrical and Computer Engineering, New York University<br> &nbsp;&nbsp;&nbsp; 5. Department of Informatics, New Jersey Institute of Technology<br> &nbsp;&nbsp;&nbsp; 6. Center for Urban Science and Progress, New York University<br> &nbsp;&nbsp;&nbsp; 7. Department of Computer Science and Engineering, New York University</p> <p><br> <strong>Description</strong></p> <p>Synthetic noisy urban soundscapes (SNUSS) is a dataset of synthetic soundscapes with real urban background noise meant to mimic urban soundscapes. This dataset contains 30,000 10 second mixtures, their isolated components, and auto-generated annotations. This dataset was developed with the goal of synthesizing soundscapes with a diverse set of realistic sounding background activity, for use in developing and evaluating machine listening systems in urban settings.</p> <p><br> <strong>Mixture generation</strong></p> <p>We generate synthetic mixtures using a collection of isolated sound events with class annotations, as well as a collection of urban background noise.&nbsp; Audio mixtures are 4 seconds long (at 16kHz). We generate a foreground sub-mixture using a subset of clips from <a href="https://urbansounddataset.weebly.com/urbansound8k.html">UrbanSound8K</a> [3] from the <em>car horn</em>, <em>dog bark</em>, <em>gun shot</em>, <em>jackhammer</em> and <em>siren</em> classes. These clips range from 0.5 s to 4s. The number of events per mixture is sampled from a zero-truncated Poisson distribution with a rate parameter of 5. The class for each event is chosen uniformly at random from the five target classes. The particular sound event is chosen uniformly at random from the available clips for that class. The start time is chosen uniformly throughout the clip such that the entire clip is contained in the 4 second mixture. A brief fade-in and fade-out is applied to the clip to avoid discontinuities. Each clip is set to a sound level sampled uniformly at random in the range -30 to -20 dB LUFS.</p> <p>For the background audio, we use urban background recordings from the SONYC-Background dataset [2, 4], containing 441 10 second recordings of urban background noise in New York City. For more information, see the <a href="https://doi.org/10.5281/zenodo.5129078">SONYC-Backgrounds page</a>. For each mixture, a random background clip is chosen from which we extract a uniformly chosen 4 second segment.</p> <p>We create datasets using an foreground-to-background SNRs of -50, -20 -0 dB LUFS, (`n50dB`, `n20dB`, and `0dB` respectively), in addition to a noiseless dataset (`none`). The datasets are generated such that the only difference between them is the relative loudness between the foreground and background.</p> <p>For the training set, we generate 20,000 mixtures using folds 1-6 of UrbanSound8K and the training set of SONYC-Background. For the validation set, we generate 5,000 examples using folds 7-8 of UrbanSound8K and the validation set of SONYC-Background. For the test set, we generate 5,000 examples using folds 9-10 of UrbanSound8K and the test set of SONYC-Background.</p> <p>For additional details on the foreground mixture generation process, please refer to [1]. For additional details on generating the soundscapes with background, please refer to [2].</p> <p><br> <strong>Files</strong></p> <p>The dataset files are split into the following compressed archives:</p> <ul> <li>`synthetic-noisy-urban-soundscapes_mixtures-bkgr-none.tar.gz` - Noiseless mixtures</li> <li>`synthetic-noisy-urban-soundscapes_mixtures-bkgr-n50dB.tar.gz` - Mixtures with -50 dB LUFS SNR</li> <li>`synthetic-noisy-urban-soundscapes_mixtures-bkgr-n20dB.tar.gz` - Mixtures with -20 dB LUFS SNR</li> <li>`synthetic-noisy-urban-soundscapes_mixtures-bkgr-0dB.tar.gz` - Mixtures with 0 dB LUFS SNR</li> <li>`synthetic-noisy-urban-soundscapes_isolated_events.tar.gz` - Isolated sound events for each mixture</li> </ul> <p><br> <em>Preparing the files</em></p> <ol> <li>Download each of the tar.gz files to a new folder. You need at the `isolated_events` and one of the mixture datasets.</li> <li>Decompress all of the tar.gz files.</li> <li>Merge the contents of the extracted `isolated_events` folder into the extracted mixture folders. This makes sure the corresponding isolated events for each mixture are placed in its `XXXXX_events` folder.</li> </ol> <p><br> <em>File structure</em></p> <p>The mixture dataset folder for the desired background condition should have the format `synthetic-noisy-urban-soundscapes_mixtures-bkgr-&lt;condition&gt;/&lt;split&gt;`. Within each split folder are the mixture files, which are identified by an integer (padded with leading zeros up to 5 places). For a mixture `00001`, the mixture audio is `00001.wav` and the annotation file (in JAMS [5] format) is `00001.jams`. The isolated events can be found in the `00001_events` folder, where foreground events have the format `foreground&lt;fg-event-num&gt;_&lt;class-name&gt;.wav` and the background recording (if used) is called `background0_2017.wav`.</p> <p><br> <strong>Contact</strong></p> <p>If you have any questions, comments, or concerns, please direct correspondence to Aurora Cramer (aurora (dot) linh (dot) cramer (at) gmail (dot) com).</p> <p>&nbsp;</p> <p><strong>References</strong></p> <p>[1] Pishdadian, F., Wichern, G., &amp; Le Roux, J. (2020). Finding strength in weakness: Learning to separate sounds with weak supervision. IEEE/ACM Transactions on Audio, Speech, and Language Processing (TASLP).</p> <p>[2] Cramer, A., Cartwright, M., Pishdadian, F., and Bello, J. P. (2021). Weakly Supervised Source-Specific Sound Level Estimation in Noisy Soundscapes. In 2015 IEEE Workshop on Applications of Signal Processing to Audio and Acoustics (WASPAA).</p> <p>[3] Salamon, J., Jacoby, C., and Bello, J.P. (2014). A dataset and taxonomy for urban sound research. In 2014 ACM International Conference on Multimedia.</p> <p>[4] Cramer, A., Cartwright, M., Pishdadian, F., and Bello, J.P. (2021). SONYC-Backgrounds: a collection of urban background recordings from an acoustic sensor network (1.0.0). Zenodo. https://doi.org/10.5281/zenodo.5129078</p> <p>[5] Humphrey, E. J., Salamon, J., Nieto, O., Forsyth, J., Bittner, R. M., and Bello, J.P. (2014). JAMS: A JSON Annotated Music Specification for Reproducible MIR Research. In 2014 International Society for Music Information Retrieval Conference (ISMIR)</p> <p><br> <strong>Acknowledgements</strong></p> <p>This work is partially supported by National Science Foundation <a href="https://www.nsf.gov/awardsearch/showAward?AWD_ID=1633259">award 1633259</a> and <a href="https://www.nsf.gov/awardsearch/showAward?AWD_ID=1544753">award 1544753</a>.</p> <p>&nbsp;</p>

opencc-by-4.0Oct 2021View details →
zenodo40/100

Synthetic automotive LiDAR dataset with radial velocity additional feature - (x,y,z,v)

<p>The synthetic dataset was generated using KITTI-like specifications and annotations format. It is comprised by the KITTI&nbsp;standard&nbsp; folders: label_2, image_2 and calib. Furthermore, there is a velodyne file for each of the following use cases:</p> <ul> <li>Point cloud 1: (x,y,z, (Bool)Is_Object):&nbsp;In this point cloud, the best performance of the Deep Learning model is expected as ground truth information is provided as the additional feature of each point. <ul> <li>Point cloud 1A: (x,y,z, (Bool)Is_Car):&nbsp;the additional feature of each point that belongs to an object of the &rsquo;Car&rsquo; type has a Boolean 1.0 value; contrariwise, the 0.0 value was used. File:&nbsp;velodyne_1A_isCar;</li> <li>Point cloud 1B: (x,y,z, (Bool)Is_Ped):&nbsp;the additional feature of each point that belongs to an object of the &rsquo;Pedestrian&rsquo; type has a Boolean 1.0 value; contrariwise, the value 0.0 was used.&nbsp; File:&nbsp;velodyne_1A_isPed.</li> </ul> </li> <li>Point cloud 2: (x,y,z, (Float)Radial_Velocity): this point cloud has the relative radial velocity as an additional feature for each point. File:&nbsp;velodyne_2_radial_velocity;</li> <li>Point cloud 3: (x,y,z,(Float)Absolute_Speed): in this point cloud, every point has the absolute speed of the object as the additional feature.&nbsp;File:&nbsp;velodyne_3_abs_speed;</li> <li>Point cloud 4:&nbsp;(x,y,z,(Bool)Is_Moving):&nbsp;the additional feature of this point cloud is a Boolean value that is set to 1.0 if the object is moving; contrariwise, it is set to 0.0 for static objects. File:&nbsp;velodyne_4_is_moving;</li> <li>Point cloud 5:&nbsp;(x,y,z,0): no additional feature information. If desired, requires post-processing to convert to (x,y,z) or changing the toolbox point cloud configuration to not consider the additional feature.&nbsp;File:&nbsp;velodyne_5_xyz;</li> </ul> <p>Additionally, the label split for testing and training sets&nbsp;used can be found at file: Labels_split.</p> <p>This work was made as part of a master thesis. For further details, please check the dataset generation source code [1]. Any further questions please contact Leandro Alexandrino (l.alexandrino@ua.pt).</p> <p>&nbsp;</p> <p>[1] -&nbsp;Fork deepgtav-presil - leandro alexandrino, https://github.com/leandroalexandrino1995/DeepGTAVPreSIL.</p>

opencc-by-4.0Oct 2022View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record