Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
1,754
datasets available to search
ShareScore release 0.7.1
Dataset results
1,754 results for “synthetic”
HD-SIM-RBV: a synthetic dataset with model-based simulations of blood volume changes during hemodialysis
<p>The HD-SIM-RBV dataset is a synthetic (model-based) dataset generated to enable the study of blood volume (BV) or relative blood volume (RBV) changes during hemodialysis (HD).</p> <p>The dataset includes the profiles of BV changes during a standard 4-hour HD session simulated using a lumped-parameter, physiologically-based model of the cardiovascular system and the whole-body water and solute kinetics in 5,000 virtual patients with randomly adjusted values of 90 physiological parameters.</p> <p>For each of the 90 selected parameters, a random value was drawn from a normal distribution with the mean equal to the baseline value used originally in the model (with a few exceptions) and the standard deviation (SD) assumed at the level of 10%, 20%, or 40% of the baseline value, depending on the nature of the given parameter and the likelihood of its variation in the population (for some parameters, SD was set below 10% - see Parameters.xlsx). Only values within ±2SD from the mean were accepted. </p> <p>Ultrafiltration was set randomly within ±1 L from the assigned fluid overload. All other parameters as well as dialysis settings were kept constant for all virtual patients (at the levels used in our previous work - see the references below).</p> <p> </p> <p>When using the dataset, please cite the associated conference paper:</p> <p>Pstras L, Waniewski J. A Model-Based Dataset for In-Silico Exploration of the Patterns of Relative Blood Volume Changes During Hemodialysis. 2023 IEEE EMBS Special Topic Conference on Data Science and Engineering in Healthcare, Medicine and Biology, 149-150, 2023, doi: 10.1109/IEEECONF58974.2023.10404528.</p>
Synthetic Dataset of Citation Strings in 12 Styles
<p>This dataset was produced in the aim of testing different tools for citation string parsing, as part of the experiment reported in the paper:</p> <blockquote> <p>Iana Atanassova and Marc Bertin, 2024. "Breaking Boundaries in Citation Parsing: A Comparative Study of Generative LLMs and Traditional Out-of-the-box Citation Parsers", Bibliometric-enhanced Information Retrieval workshop (BIR), collocated with ECIR 2024, Glasgow, Scotland. </p> </blockquote> <h2><br>Data</h2> <p>The data that is provided here is organised as follows:</p> <ul> <li>the file <strong>citation-strings.zip</strong> contains raw citation strings that were generated for each of the 12 citation styles in txt format</li> <li>the file <strong>parsers-output.csv</strong> contains the output that was produced from the parsers: ChatGPT, Llama, and Neural ParsCit</li> </ul> <h2><br>To cite this work</h2> <p>To use this dataset and/or the results produced in the experiment, please cite the following article:</p> <blockquote> <p>@inproceedings{atanassova2024citparse,<br> title = {{Breaking Boundaries in Citation Parsing: A Comparative Study of Generative LLMs and Traditional Out-of-the-box Citation Parsers}}, <br> author = {Iana Atanassova and Marc Bertin},<br> year = {2024},<br> booktitle = {{International Workshop on Bibliometric-enhanced Information Retrieval (BIR 2024) co-located with the 46\textsuperscript{st} European Conference on Information Retrieval (ECIR 2024)}},<br> address = {Glasgow, Scotland}<br>}</p> </blockquote> <h3>Authors information</h3> <ul> <li>Iana Atanassova, ORCID https://orcid.org/0000-0003-3571-4006 URL https://iana-atanassova.github.io/</li> <li>Marc Bertin, ORCID https://orcid.org/0000-0003-1803-6952 URL https://elico-recherche.msh-lse.fr/membres/marc-bertin</li> </ul> <h3>Related github repository</h3> <p>https://github.com/iana-atanassova/citation-parsers-bir2024.git </p>
Dataset of "MoO3-xNiMoO4 nanorods synthetized using NiO nanoparticles for hydrogen evolution in anion exchange membrane water electrolysis"
<p>Novel method of Mo-Ni catalyst for hydrogen evolution reaction in anion exchange membrane water electrolysis was used. Complete physico-chemical and electrochemical characterization was done. Prepared material showed enhanced performance when compared to the similar Ni based materials. Physico-chemical characterization showed, that final material is formed by NiMoO4 nanorods coverd on the surface by the layer of the MoO3-x.</p>
Dataset - Generating reliable estimates of tropical cyclone induced coastal hazards along the Bay of Bengal for current and future climates using synthetic tracks
<p>This data is complementary to the paper by Leijnse et al. 2022 "Generating reliable estimates of tropical cyclone induced coastal hazards along the Bay of Bengal for current and future climates using synthetic tracks" <br> https://doi.org/10.5194/nhess-2021-181</p> <p>This data is made available in the hope that it will be useful, but WITHOUT ANY WARRANTY; without even the implied warranty of MERCHANTABILITY or FITNESS FOR A PARTICULAR PURPOSE</p> <p>For questions about the data ask: tim.leijnse@deltares.nl</p> <p>For more information about the tool to generate the used synthetic tracks TCWiSE see: <a href="https://www.deltares.nl/en/software/tcwise/">https://www.deltares.nl/en/software/tcwise/</a></p> <p> </p>
AraucanaXAI - HEPAR synthetic datasets
<p>Supporting datasets (iid and ood) used in the evaluation experiments of the paper "Why did AI get this one wrong? - tree-based explanations of machine learning model predictions" by Parimbelli, Buonocore, Nicora, Michalowski, Wilk and Bellazzi.</p>
The extrAIM dataset: A merged satellite-based daily precipitation dataset for the Mediterranean region (including an ensemble of 20 synthetic realisations)
<p><strong>extrAIM </strong>dataset is a <strong>new merged daily precipitation product</strong> (extraim_merged_data.nc) for the Mediterranean region with the following characteristics:</p> <ul> <li><strong>Dataset format:</strong> NetCDF</li> <li><strong>Spatial resolution:</strong> 25 x 25 km</li> <li><strong>Temporal resolution:</strong> 1 day</li> <li><strong>Spatial coverage:</strong> Longitude: from -6.25 to 38.25, Latitude: 27.75 to 49</li> <li><strong>Temporal coverage: </strong>01-01-2007 to 30-09-2021</li> <li><strong>Merging approach: </strong>Two-step merging (classification and regression) <ul> <li><strong>Algorithm: </strong>Random Forest for both classification and regression</li> <li><strong>Training strategy:</strong> Full training strategy</li> </ul> </li> <li><strong>Merged precipitation products: </strong>SM2Rain-ASCAT and GPM Late Run</li> <li><strong>Reference precipitation product:</strong> EMO5</li> <li><strong>Static covariates: </strong>Longitude, Latitude and Elevation, in both classification and regression step <ul> <li><strong>Classification step:</strong> probability dry and probability dry of the 5 neighboring points around the target locations</li> <li><strong>Regression step:</strong> mean, standard deviation and skewness of daily precipitation, of the entire series and non-zero amounts, as well as mean precipitation of the 5 neighboring points around the target locations</li> </ul> </li> </ul> <p>In addition, an <strong>ensemble of 20 synthetic realizations</strong> (equiprobable and bias-adjusted) of the merged dataset is provided (files named: “extraim_realisation_XX.nc”). The synthetic realisations were produced using the extrAIM’s uncertainty-quantification approach and the associated conditional sampling method.</p>
Dynamic X-ray CT of Synthetic magma for Digital Volume Correlation analysis
<p>Dataset of synthetic magma subjected to compression, useful for Digital Volume Correlation analysis, ref [1,2]. The data has been acquired at the Diamond Light Source synchrotron, with a bespoke thermo-mechanical rig (“P2R”) on the I12 beamline, ref [3,4,5]. Dataset 0 has no applied compression, while dataset 1 has applied compression.</p> <p>The data was saved with numpy 1.21 with <a href="https://numpy.org/doc/1.21/reference/generated/numpy.lib.format.html#format-version-1-0">NumPy format version 1.0</a> as dataset_0.npy and dataset_1.npy, and NumPy can be used to read it back in. Both data files have a header specifying how the data is stored, and following the header comes the array data.</p> <p>In particular the header length is 128 bytes, and the data consists of a 3 dimensional matrix of size (1520, 1257, 1260) stored in unsigned integer 8 bit, Fortran order. The screenshot named import_imagej.png shows how to import the data in with <a href="https://imagej.nih.gov/ij/">ImageJ</a>.</p> <p> </p> <p>A <a href="https://github.com/Kitware/MetaIO">METAImage</a> header describing the data in text form for each dataset is also provided, i.e. dataset_0.mhd and dataset_1.mhd,</p>
Synthetic and real EEG datasets for closed-loop neuroscience
<p>The dataset is made primarily for the task of real-time low latency filtering of the EEG data in the closed loop neuroscience experiments and for EEG forecasting task. The dataset consists of a real data and 5 options of the synthetic data of varying difficulty.</p><p>The real dataset consists of 25 people involved into the P4 alpha neurofeedback training. Its total size is about 16.3 hours. A more detailed instruction for this file is provided in the file Real dataset instructions.txt.</p><p>Synthetic data is generated in 5 different ways: sine wave with white noise, sine wave with pink noise, narrow-band filtered pink noise sample with pink noise, state-space model with white noise and state-space model with pink noise. Each of these datasets has about 34.5 hours of data. It is generated similarly to (Wodeyar et al, 2021). A more detailed instruction for the synthetic dataset can be found in the file Synthetic datasets instructions.txt.<br> </p><p>In LowLatencyEEGFiltering.zip one can find a code for the models used in our paper for low-latency filtering with this data.</p><p>NOTE: Code is also published in the following GitHub repository: https://github.com/ivsemenkov/LowLatencyEEGFiltering</p><p> </p><p>If you use our data or code please cite: https://www.doi.org/10.1088/1741-2552/acf7f3</p>
Synthetic images of cell nuclei in widefield microscopy
<p>The images were generated by <a href="http://www.cs.tut.fi/sgn/csb/simcep/tool.html">SIMCEP</a>, a widefield fluorescence microscopy biological images simulator.</p> <p>The dataset is used to demonstrate the execution of image analysis workflows with BIAFLOWS on a local machine from a jupyter notebook.</p>
Wind measurement data from the publication: "Development of a load model validation framework applied to synthetic turbulent wind field evaluation"
<h3>Dataset description:</h3> <p>This datasat represents supplementary material used in the contribution "Development of a load model validation framework applied to<br>synthetic turbulent wind field evaluation" by Meyer, Huhn and Gottschall.</p> <p>Wind measurements from the Testfeld BHV are made available. For installation details, see the mentioned reference.</p> <p> </p> <h3>File description:</h3> <ul> <li>Lidar_HWS.nc - Horizontal wind speed measurements (10 min averages) from a WindCube V2 vertical profiler for one day with a low-level jet occurrence ( <div> <div>2021-04-20)</div> </div> </li> <li>Cups_HWS.nc - Horizontal wind speed measurements (10 min averages) from cup anemometer installed on a met mast for the same day</li> <li>Ensemble_averaged_Spectra.nc - Ensemble averaged spectra for neutral and near neutral situations from a Gill Windmaster at 110m above ground level, used to fit the Mann and KSEC model parameters</li> </ul> <h3> </h3> <h3>Referencing:</h3> <p>When used, please cite like the following:</p> <p>Meyer, Paul J., Matthias L. Huhn, and Julia Gottschall. 2024. "Development of a Load Model Validation Framework Applied to Synthetic Turbulent Wind Field Evaluation" <em>Energies</em> 17, no. 4: 797. https://doi.org/10.3390/en17040797</p> <p> </p> <p> </p>
VoroCrack3d: An annotated data set of 3d CT concrete images with synthetic crack structures
<p>VoroCrack3d is an annotated data set of 3d CT images of concrete with synthetic crack structures. Its main purpose is the training and testing of machine learning models for 3d crack segmentation. The data set comprises 1344 images together with their corresponding ground truths. The concrete backgrounds are cropped out sections of size 400x400x400 voxels of CT images of concrete. To this end, several different concrete samples were scanned (normal concrete (NC), high-performance concrete (HPC), ultra-high-performance concrete (UHPC), air pore concrete; without and with reinforcements (straight steel fibers, crimped steel fibers, hooked-end steel fibers, polypropylene fibers, fibers made of glass fiber-reinforced polymer). The original concrete images have a resolution between 2.8 and 106 micrometers.</p> <p>The crack structures are modeled via minimum-weight surfaces in Voronoi diagrams according to the paper</p> <p>[1] C. Jung, C. Redenbach, Crack Modeling via Minimum-Weight Surfaces in 3d Voronoi Diagrams, Journal of Mathematics in Industry, 13, 10 (2023). https://doi.org/10.1186/s13362-023-00138-1.</p> <p>The surfaces are discretized, dilated and superimposed on the concrete backgrounds.</p> <p>The data set offers a high variety regarding concrete types, noise levels and crack widths, shapes, regularity and branching. This makes it suitable for studying the generalizability and robustness of 3d crack segmentation methods.</p> <p>______________________________________________________________________________________________</p> <p>The folder 'data' contains seven subfolders, each containing the data generated from a specific concrete type (NC, HPC, air pore concrete, polypropylene fiber-reinforced concrete, steel fiber-reinforced concrete (straight, crimped and hooked-end steel fibers)).</p> <p>Each subfolder again contains four subfolders according to the point process model that was used for generating the 3d Voronoi diagrams. The point processes and Voronoi diagrams are restricted to windows of size 400x150x400. </p> <p>- 'hc': Hard core point process with 60% volume density and intensity 0.000025 obtained from force-biased sphere packing.<br>- 'matclust': Matérn cluster process with parent intensity 0.0002/50, offspring intensity 50 and cluster radius 20.<br>- 'ppp': Poisson point process with intensity 0.0002.<br>- 'ppp-scaled': Poisson point process with intensity 0.0002 (but inside 200x150x200 window). The resulting Voronoi diagram is stretched in x- and z- direction by a factor of 2.</p> <p>Each of these contains five subfolders: one for the 3d input images, two for the corresponding labels (ground truths; one with and one without pores/fibers), one for the input and label previews (slice z=200 for each of the images) and a misc folder containing the concrete background without crack and, if applicable, the pore/fiber segmentation image.</p> <p>The data itself then contains 48 images:<br>1a-1d: crack with up to seven branches; fixed crack width (~1 voxel).<br>2a-2d: crack with up to four branches; fixed crack width (~1 voxel).<br>3a-3d: crack with up to one branch; fixed crack width (~1 voxel).<br>4a-4d: crack with no branches; fixed crack width (~1 voxel).<br>5a-5d: crack with no branches; fixed crack width (~3 voxels).<br>6a-6d: crack with no branches; fixed crack width (~5 voxels).<br>7a-7d: crack with no branches; fixed crack width (~7 voxels).<br>8a-8d: crack with up to seven branches; multiscale crack (bernoulli parameter 0.01);<br>9a-9d: crack with up to seven branches; multiscale crack (bernoulli parameter 0.02);<br>10a-10d: crack with up to seven branches; multiscale crack (bernoulli parameter 0.05);<br>11a-11d: crack with up to seven branches; multiscale crack (bernoulli parameter 0.1);<br>12a-12d: crack with up to seven branches; multiscale crack (bernoulli parameter 0.2);</p> <p>The names 'a'-'d' indicate level of added noise added to the image:<br>a: None.<br>b: Uniformly on [-sigma,sigma] <br>c: Uniformly on [-2*sigma,2*sigma] <br>d: Uniformly on [-4*sigma,4*sigma] <br>Negative values are mapped to 0. <br>For inputs of type int, noise values are rounded to the nearest integer.<br>(sigma = standard deviation of voxel greyvalues in image)</p> <p>Note that the grey values in the ground truths correspond to the local crack width. They can be thresholded to obtain binary masks.</p> <p>For more details, we refer to [1].</p>
High-Voltage Disconnector State Identification: synthetic and real images of substation disconnectors
<p>This dataset contains the training and test images used in the work detailed in the article: Barpp Gomes, V., Marchesi, B., Gruber, Y.A. <em>et al.</em> Exploring Synthetic Data for Training Deep Learning Models for High-Voltage Disconnector State Identification. <em>J Control Autom Electr Syst</em> (2025). <a href="https://doi.org/10.1007/s40313-025-01204-2">https://doi.org/10.1007/s40313-025-01204-2</a></p> <p>Contais about 940,000 synthetic (CGI-rendered) and 60,000 real (camera-captured) samples of four types of substation disconnectors, on both open and closed states:</p> <ul> <li>230 kV center break (type 1, as indicated in the article);</li> <li>230 kV center break (type 2);</li> <li>230 kV double side break</li> <li>525 kV horizontal semi-pantograph</li> </ul> <p>Each zip file contains images of one type of substation disconnector. Images are sized 320x128 and are organized in folders, as follows:</p> <ul> <li>00_train_synth: Synthetic training images.</li> <li>01_train_real: A small set of real training images, as indicated in the article.</li> <li>02_test_real_normal1: One set of real test images.</li> <li>03_test_real_normal2: Another set of real test images, from a different time period.</li> <li>04_test_real_maneuvers: A special set of real test images in which the switches have been operated (are in different states).</li> </ul>
Synthetic Dataset of Emergency Healthcare Services
<p>Synthetic dataset of emergency services comprised of several CSV files that we have generated using a simulation software. This dataset is open for public use; please cite our work if used in research or applications.</p> <p>## File Overview</p> <ol> <li>**CheckBloodPressure.csv** - (9 KB): Contains blood pressure Server records of patients.</li> <li>**CheckPatientType.csv** - (19 KB): Identifies the type of each patient (e.g., 1 or 3).</li> <li>**Fill_Information.csv** - (2 KB): Fill information records for new patients.</li> <li>**MedicalRecord1.csv** - (10 KB): Medical record dataset for patient type 1.</li> <li>**MedicalRecord2.csv** - (4 KB): Medical record dataset for patient type 2.</li> <li>**MedicalRecord3.csv** - (2 KB): Medical record dataset for patient type 3.</li> <li>**MedicalRecord4.csv** - (13 KB): Medical record dataset for patient type 4.</li> <li>**OutPatientDepartment.csv** - (18 KB): Data related to the satisfaction and length of stay of an given patient.</li> <li>**Triage.csv** - (13 KB): Data related to the triage process.</li> <li>**README.txt** - (4 KB): Documentation of the dataset, including structure, metadata, and usage.</li> </ol> <p> </p> <p>## Common Fields Across Files</p> <ol> <li>**Patient ID** *(Integer)*: Unique identifier for each patient.</li> <li>**Patient Type** *(Integer)*: Classification of patient (e.g., 1, 4).</li> <li>**Medical Records Arrival Time** *(DateTime)*: Timestamp of the patient's first arrival in the medical record department.</li> <li>**Exiting Time** *(DateTime)*: Timestamp when the patient exits a Server.</li> <li>**Waiting Time (min)** *(Real)*: Total waiting time before being attended to.</li> <li>**Resource Used** *(String)*: Resource (e.g., Operator) allocated to the patient.</li> <li>**Utilization %** *(Real)*: Utilization rate of the resource as a percentage.</li> <li>**Queue Count Before Processing** *(Integer)*: Number of patients in the queue before processing begins.</li> <li>**Queue Count After Processing** *(Integer)*: Number of patients in the queue after processing ends.</li> <li>**Queue Difference** *(Integer)*: Difference between the before and after queue counts.</li> <li>**Length of Stay (min)** *(Real)*: Total time spent in the simulation by the patient.</li> <li>**LOS without Queues (min)** *(Real)*: Length of stay excluding any queuing time.</li> <li>**Satisfaction %** *(Real)*: Patient satisfaction rating based on their experience.</li> <li>**New Patient?** *(String)*: Indicates if this is a new patient or a returning one.</li> </ol> <p>Ferreira, M. (2024). Synthetic Dataset of Emergency Healthcare Services [Data set]. Zenodo. https://doi.org/10.5281/zenodo.14212812</p> <p> </p>
Synthetic cryo electron subtomograms containing biomolecular complexes with continuous conformational variability, used for validating TomoFlow method
<p>Two datasets used for validating TomoFlow method, an optical-flow based approach for analyzing continuous conformational variability of biomolecular complexes in cryo electron subtomograms. The TomoFlow method and the methods used to synthesize the two test datasets have been fully described in the following article: "M. Harastani, M. Eltsov, A. Leforestier, S. Jonic, TomoFlow: Analysis of continuous conformational variability of macromolecules in cryogenic subtomograms based on 3D dense optical flow, Journal of Molecular Biology (2021), doi: https://doi.org/10.1016/j.jmb.2021.167381". Additionally, this article describes a test of TomoFlow using one experimental cryo electron tomography dataset (available in EMPIAR and EMDB databases under the accession codes EMPIAR-10679 and EMD-12699). </p>
Synthetic time series data generation for edge analytics
<p>In this research, we create synthetic data with features that are like data from IoT devices. We use an existing air quality dataset that includes temperature and gas sensor measurements. This real-time dataset includes component values for the Air Quality Index (AQI) and ppm concentrations for various polluting gas concentrations. We build a JavaScript Object Notation (JSON) model to capture the distribution of variables and structure of this real dataset to generate the synthetic data. Based on the synthetic dataset and original dataset, we create a comparative predictive model. Analysis of synthetic dataset predictive model shows that it can be successfully used for edge analytics purposes, replacing real-world datasets. There is no significant difference between the real-world dataset compared the synthetic dataset. The generated synthetic data requires no modification to suit the edge computing requirements. The framework can generate correct synthetic datasets based on JSON schema attributes. The accuracy, precision, and recall values for the real and synthetic datasets indicate that the logistic regression model is capable of successfully classifying data</p>
DESED_synthetic
<p>Link to the associated github repository: <a href="https://github.com/turpaultn/Desed">https://github.com/turpaultn/Desed</a></p> <p>Link to the papers: <a href="https://hal.inria.fr/hal-02160855"><em>https://hal.inria.fr/hal-02160855</em></a>, <a href="https://hal.inria.fr/hal-02355573v1">https://hal.inria.fr/hal-02355573v1</a></p> <p>Domestic Environment Sound Event Detection (DESED).</p> <p><strong>Description</strong><br> This dataset is the synthetic part of the DESED dataset. It allows creating mixtures of isolated sounds and backgrounds.</p> <p>There is the material to:</p> <ul> <li>Reproduce the DCASE 2019 task 4 synthetic dataset</li> <li>Reproduce the DCASE 2020 task 4 synthetic train dataset</li> <li>Creating new mixtures from isolated foreground sounds and background sounds.</li> </ul> <p><strong>Files:</strong></p> <p><strong>If you want to generate new audio mixtures yourself from the original files.</strong></p> <ol> <li><strong>DESED_synth_soundbank.tar.gz</strong> : Raw data used to generate mixtures.</li> <li><strong>DESED_synth_dcase2019jams.tar.gz</strong>: JAMS files, metadata describing how to recreate the dcase2019 synthetic dataset<strong> </strong></li> <li><strong>DESED_synth_dcase20_train_val_jams.tar: </strong>JAMS files, metadata describing how to recreate the dcase2020 synthetic train and valid dataset.</li> <li><strong>DESED_synth_dcase20_eval_jams.tar: </strong>JAMS files, metadata describing how to recreate the dcase2020 synthetic eval dataset (only the basic one, variants of it have been made but not presented here).</li> <li><strong>dcase_synth.zip </strong>synthetic part of the DCASE training set from 2021 and for 2022.</li> </ol> <p><strong>If you simply want the evaluation synthetic dataset used in DCASE 2019 task 4.</strong></p> <ol> <li><strong>DESED_synth_eval_dcase2019.tar.gz</strong><strong> </strong>:<strong> </strong>Evaluation audio and metadata files used in dcase 2019 task 4.</li> </ol> <p> </p> <p>The mixtures are generated using Scaper (https://github.com/justinsalamon/scaper) [1].</p> <p>* Background files are extracted from SINS [2], MUSAN [3] or Youtube and have been selected because they contain a very low amount of our sound event classes.<br> * Foreground files are extracted from Freesound [4][5] and manually verified to check the quality and segmented to remove silences.</p> <p><strong>References</strong><br> [1] J. Salamon, D. MacConnell, M. Cartwright, P. Li, and J. P. Bello. Scaper: A library for soundscape synthesis and augmentation<br> In IEEE Workshop on Applications of Signal Processing to Audio and Acoustics (WASPAA), New Paltz, NY, USA, Oct. 2017.</p> <p>[2] Gert Dekkers, Steven Lauwereins, Bart Thoen, Mulu Weldegebreal Adhana, Henk Brouckxon, Toon van Waterschoot, Bart Vanrumste, Marian Verhelst, and Peter Karsmakers.<br> The SINS database for detection of daily activities in a home environment using an acoustic sensor network.<br> In Proceedings of the Detection and Classification of Acoustic Scenes and Events 2017 Workshop (DCASE2017), 32–36. November 2017.</p> <p>[3] David Snyder and Guoguo Chen and Daniel Povey.<br> MUSAN: A Music, Speech, and Noise Corpus.<br> arXiv, 1510.08484, 2015.</p> <p>[4] F. Font, G. Roma & X. Serra. Freesound technical demo. In Proceedings of the 21st ACM international conference on Multimedia. ACM, 2013.<br> <br> [5] E. Fonseca, J. Pons, X. Favory, F. Font, D. Bogdanov, A. Ferraro, S. Oramas, A. Porter & X. Serra. Freesound Datasets: A Platform for the Creation of Open Audio Datasets.<br> In Proceedings of the 18th International Society for Music Information Retrieval Conference, Suzhou, China, 2017.</p> <p> </p>
Synthetic Simulations Of Extracellular Recordings (SSOER) Dataset
<p>This dataset contains synthetic data from simulations (for a total duration of 10 minutes) including the activity of one multi-unit and two single-units for different firing rates and signal-to-noise ratio levels. It is intended to be used as a standardized dataset to evaluate spike sorting algorithms.</p> <p>Recordings were taken using a sampling rate of 24 kHz, and are comprised of spikes from a database with 594 different average spike shapes, taken from real recordings from monkey neocortex and basal ganglia.</p> <p>This dataset is comprised of two files: <em>data.npy</em> and <em>labels.csv</em>.</p> <ul> <li><em>data.npy</em> contains 14,400,000 sampled voltage values, from a single channel, taken at a sampling rate of 24 kHz. </li> <li><em>labels.csv</em> contains the timestep, spike class, amplitude (SNR), and firing rate associated with each spiking event.</li> </ul> <p>The original samples used to construct this dataset where previously constructed and made available in [1]. This dataset is an amalgamation of simulation files, which were previously publicly accessible at: <a href="http://www2.le.ac.uk/departments/engineering/research/bioengineering/neuroengineering-lab/software">http://www2.le.ac.uk/departments/engineering/research/bioengineering/neuroengineering-lab/software</a>. Consequently, when using or making modifications to this dataset, in addition to acknowledging this record, [1] must also be acknowledged, as per the original author's request.</p> <p>[1] J. Martinez, C. Pedreira, M. J. Ison, and R. Quian Quiroga, “Realistic simulation of extracellular recordings,” Journal of Neuroscience Methods, vol. 184, no. 2, pp. 285–293, Nov. 2009, doi: 10.1016/j.jneumeth.2009.08.017.</p>
Synthetic datasets reflecting the shRNA-seq knockdown ENCODE data for HepG2 and K562 with coresponding GRN
<p>Synthetic data correspond to the ENCODE data for cell lines HepG2 (https://www.encodeproject.org/biosamples/ENCBS282XVK/) and K562 (https://www.encodeproject.org/biosamples/ENCBS023XVB/). The data and networks were generated using GeneSPIDER (publicly available at https://bitbucket.org/sonnhammergrni/genespider/).</p> <p> </p> <p><strong>Table.1 </strong>Description of the files</p> <table> <tbody> <tr> <td>data_HepG2like_SNR_L=0.0054699_diff=1.6188e-05.txt</td> <td>Synthetic gene expression knockdown (shRNA-seq) data immitating the ENCODE data for HepG2 cell line. Data size: 232 RBPs vs 464 experiments (2 replicates). SNR_L is the value of signal to noise ratio. Difference (diff) value tells the difference between replicate correlation coefficients of real and synthetic ENCODE data. Columns represent experiments, rows represent genes.</td> </tr> <tr> <td>data_K562like_SNR_L=0.0028692_diff=0.00017339.txt</td> <td>Synthetic gene expression knockdown (shRNA-seq) data immitating the ENCODE data for K562 cell line. Data size: 232 RBPs vs 464 experiments (2 replicates). SNR_L is the value of signal to noise ratio. Difference (diff) value tells the difference between replicate correlation coefficients of real and synthetic ENCODE data. Columns represent experiments, rows represent genes.</td> </tr> <tr> <td>network_HEPG2like_sparsity4.txt</td> <td>Synthetic scale-free gene regulatory network compatibile with data_HepG2like_SNR_L=0.0054699_diff=1.6188e-05.txt. Sparsity (average node degree) is 4 including selfloops. Direction should be read from columns to rows.</td> </tr> <tr> <td>network_K562like_sparsity4.txt</td> <td>Synthetic scale-free gene regulatory network compatibile with data_K562like_SNR_L=0.0028692_diff=0.00017339.txt. Sparsity (average node degree) is 4 including selfloops. Direction should be read from columns to rows.</td> </tr> <tr> <td>perturbations_HepG2&K562_2replicates.txt</td> <td>Perturbation matrix including information about knockeddown RBPs. Data size: 232 RBPs vs 464 experiments (2 replicates).</td> </tr> </tbody> </table> <p> </p> <p>Created by Garbulowski et al. (2024) as a part of the work entitled "Comprehensive analysis of the RBP regulome reveals functional modules and drug candidates in liver cancer"</p>
ncrncornell/ced2ar-synlbd-codebook: DDI Codebook for the Synthetic LBD
<p>Codebook for the Synthetic LBD, a Census Bureau data product, see <a href="https://www.census.gov/ces/dataproducts/synlbd/">https://www.census.gov/ces/dataproducts/synlbd/</a>.</p> <p>The SynLBD usage model relies on a Synthetic Data Server, maintained (as of 2018) by Cornell University, see <a href="https://www2.vrdc.cornell.edu/news/synthetic-data-server/">https://www2.vrdc.cornell.edu/news/synthetic-data-server/</a>.</p> <p>Live version of the DDI codebook at <a href="https://www2.ncrn.cornell.edu/ced2ar-web/codebooks/synlbd/">https://www2.ncrn.cornell.edu/ced2ar-web/codebooks/synlbd/</a></p>
S45 | SYNTHCANNAB | Synthetic Cannabinoids from CompTox
<p>This is the collection associated with list S45 SYNTHCANNAB on the NORMAN Suspect List Exchange.</p> <p><a href="https://www.norman-network.com/nds/SLE/">https://www.norman-network.com/nds/SLE/</a></p> <p>S45</p> <p>SYNTHCANNAB</p> <p><strong>Synthetic Cannabinoids </strong></p> <p>UPDATED 17/06/2019</p> <p>Now a list of synthetic cannabinoids aassembled from public resources, from CompTox. </p> <p>The psychoactive compounds are now in a new list with these substances, due to changes CompTox side, see DOI <a href="https://doi.org/10.5281/zenodo.2656740">10.5281/zenodo.2656740</a></p> <p> </p> <p> </p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.