Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
252
datasets available to search
ShareScore release 0.9.0
Dataset results
252 results for “Synthetic data”
pGAN Synthetic Dataset: A Deep Learning Approach to Private Data Sharing of Medical Images Using Conditional GANs
<p>Synthetic dataset for <strong>A Deep Learning Approach to Private Data Sharing of Medical Images Using Conditional GANs</strong></p> <p><strong> Dataset specification:</strong></p> <ul> <li>MRI images of Vertebral Units labelled based on region</li> <li>Dataset is comprised of 10000 pairs of images and labels</li> <li>Image and label pair number k can be selected by: synthetic_dataset['images'][k] and synthetic_dataset['regions'][k]</li> <li>Images are 3D of size (9, 64, 64)</li> <li>Regions are stored as an integer. Mapping is 0: cervical, 1: thoracic, 2: lumbar</li> </ul> <p>Arxiv paper: <a href="https://arxiv.org/abs/2106.13199">https://arxiv.org/abs/2106.13199</a><br> Github code: <a href="https://github.com/tcoroller/pGAN/">https://github.com/tcoroller/pGAN/</a></p> <p>Abstract:</p> <p>Sharing data from clinical studies can facilitate innovative data-driven research and ultimately lead to better public health. However, sharing biomedical data can put sensitive personal information at risk. This is usually solved by anonymization, which is a slow and expensive process. An alternative to anonymization is sharing a synthetic dataset that bears a behaviour similar to the real data but preserves privacy. As part of the collaboration between Novartis and the Oxford Big Data Institute, we generate a synthetic dataset based on COSENTYX Ankylosing Spondylitis (AS) clinical study. We apply an Auxiliary Classifier GAN (ac-GAN) to generate synthetic magnetic resonance images (MRIs) of vertebral units (VUs). The images are conditioned on the VU location (cervical, thoracic and lumbar). In this paper, we present a method for generating a synthetic dataset and conduct an in-depth analysis on its properties of along three key metrics: image fidelity, sample diversity and dataset privacy.</p>
3-D synthetic near surface data set with frequency-domain electromagnetic induction data
<p>Realistic three-dimensional exhaustive data set that mimics a near surface mining landfill deposit of waste fine-shaly sands. The data set is composed by petrophysical properties and frequency domain electromagnetic induction (FDEM) data and was created with the purpose of testing algorithms for near-surface modeling and characterization using electromagnetic data.</p> <p>The set of petrophysical properties include porosity, water saturation, particle density and density. Each property corresponds to a single geostatistical realization. The three-dimensional model has a dimension of 150 by 200 by 4 meters (i.e., length, width, depth) with a cell size of 0.5 m by 0.5 m by 0.1 m, respectively (grid size of 300 x 400 x 40). The model grid has 4.8 million cells.<br> Porosity and particle density were modelled based on samples of fine-shaly sands collected at a mine tailing in Portugal for which we investigated porosity, specific weight and particle density. The results of these investigations were used to generate three-dimensional models of subsurface rock properties with unconditional stochastic sequential simulation (Deutsch & Journel, 1998).<br> Porosity was modelled with an omnidirectional spherical variogram model in the horizontal direction. The variogram model has a horizontal range of 10 m, a vertical range of 1 m and a nugget effect of 0.2 % of the total variance of the data. This variogram model describes the expected spatial distribution of this property in the mine tailing.</p> <p>To ensure plausibility between rock properties, particle density and water saturation models were generated with stochastic sequential co-simulation (Deutsch & Journel, 1998) conditioned to the porosity model. For particle density we imposed an omnidirectional spherical variogram model in the horizontal direction with a range of 10 m, a vertical range of 1 m and a nugget effect of 0.2 % of the total variance of the data, and the correlation between porosity and particle density from the lab measurements. For water saturation we imposed an omnidirectional spherical variogram model in the horizontal direction with a range of 16 m, a vertical range of 2 m and a nugget effect of 0.1 (%). For the co-simulation we imposed a correlation between porosity and water content, borrowed from Bhanbhro et al. (2013) and Dumont et al. (2016).</p> <p>The pore fluid was defined as consisting in 80% of water and 20% of leachate, having a density of 0.99114 g/cm3 at a temperature of 30ºC (Souza et al., 2014). The density was mathematically calculated from porosity and particle density models and the density of the pore fluid by using a simple volumetric average of the geological material densities and its relationship to porosity (Mavko et al., 2009), <em>d</em><sub><em>b</em> </sub>= (1 - Ø) <em>d<sub>0</sub></em> Ø <em>d<sub>fl</sub></em> , where <em>d<sub>0</sub></em> is the density of the mineral grains, <em>d<sub>fl</sub></em> is the density of the pore fluids, and Ø is porosity.</p> <p>The electrical conductivity (EC) was created based on the well-known empirical relationship of Archie’s law (Archie, 1942). We first calculate electrical conductivity using the following equation, <em>R<sub>t</sub></em> = <em>a</em> <em>S<sub>w</sub><sup>-n</sup></em> Ø<sup><em>-m</em></sup> <em>R<sub>w</sub></em> , where <em>a</em> is the tortuosity constant, assumed as 0.88, <em>S<sub>w</sub></em> is the water saturation, <em>n</em> is the saturation exponent, assumed as 2, Ø is the porosity, <em>m</em> is the cementation exponent, assumed as 1.37, and <em>R<sub>w</sub></em> is the electrical resistivity of the pore fluid, assumed as 0.25. From the lithology and range of porosity values of the mining landfill model, the values of <em>a</em>, <em>n</em> and <em>m</em> were defined from Keller (1987). The electrical resistivity of the pore fluid was defined based on its composition and density (Keller, 1987). The EC was calculated based on Archie´s second law (Archie, 1942), where conductivity of the partially saturated rock (<em>c<sub>t</sub></em>) is the inverse of its resistivity (<em>R<sub>t</sub></em>), <em>c<sub>t</sub></em> = 1 / <em>R<sub>t </sub></em> (Mavko et al., 2009).</p> <p>Since the relationship between magnetic minerals and the magnetic properties of the rocks depends primarily of the composition and grain size of them (Butler, 2005), the magnetic susceptibility (MS) was modelled using the common range of magnetic susceptibility for unconsolidated sediments (Hudson et al., 1999) with unconditional stochastic sequential simulation (Deutsch & Journel, 1998), imposing an omnidirectional spherical variogram model in the horizontal direction with a range of 20 m, a vertical range of 4 m and a nugget effect of 0.1 % of the total variance.</p> <p>From the resulting three-dimensional models of EC and MS, we retrieved nine equally spaced boreholes along the same yz profile. These borehole data might be used as experimental data for modelling workflows, including geophysical inversion.</p> <p>FDEM data, both the in-phase (IP) and quadrature-phase (QP), were calculated using a 1-D forward model (Hanssens et al., 2019). The acquisition configuration replicates one of the most common sensors for FDEM near-surface surveys, namely the DUALEM-421S (DUALEM Inc., Milton, Canada). It considers two loop-loop coil orientations, a horizontal coplanar (HCP) and a perpendicular one (PRP), with the normal 3 offsets per coil orientation for this equipment, 1, 2 and 4 meters for HCP, and 1.1, 2.1 and 4.1 meters for PRP, plus an extra offset per coil orientation, 10 meters for HCP and 10.1 meters for PRP, ensuring a theoretical larger depth of investigation. The FDEM data were calculated defining the operating frequency of the sensor as 9000 Hz, with an elevation to the surface of 0.15 m.</p>
Pre-Processed Cancer Multi-Omic Data from TCGA and Synthetic Data
<p><strong>ABSTRACT </strong></p> <p>It contains the data of four omic profiles (CNV, mRNA, miRNA, and protein) obtained for BRCA, LGG, and LUAD obtained from the TCGA project. </p> <p>In addition, we provide synthetic data for a mixture of isotropic distributions.</p> <p><strong>Instructions: </strong></p> <p>Cancer data are identified by cancer type (LGG: low-grade glioma, BRCA: breast cancer, and LUAD: lung cancer). The data are scaled by using the minima and maxima of each column so that the values are between 0 and 1. In these files, the columns are the features and the rows correspond to the patients.</p> <p>The summary data contains only the numerical values. The columns are the features and the rows are the observations.</p> <p><strong>Inspiration:</strong></p> <p>This dataset uploaded to U-BRITE for "AI against CANCER DATA SCIENCE HACKATHON"</p> <p>https://cancer.ubrite.org/hackathon-2021/</p> <p><strong>Acknowledgements</strong></p> <p>Diego Salazar, June 20, 2021, "Pre-processed Cancer multi-omic data from TCGA and synthetic data", IEEE Dataport, doi: https://dx.doi.org/10.21227/pjb8-d090.</p> <p>https://ieee-dataport.org/documents/pre-processed-cancer-multi-omic-data-tcga-and-synthetic-data</p> <p><strong>U-BRITE last update date:</strong> 07/21/2021</p>
2D Synthetic Training Data For SyMBac
<p>Synthetic training datasets, used to train models to segment</p> <ul> <li><em>B. subtilis </em>growing in mother machine (100x oil, phase contrast)</li> <li><em>E. coli </em>growing on agar pads (100x oil, phase contrast)</li> <li><em>E. coli </em>streaked onto agar pads (60x air, fluorescence)</li> <li><em>E. coli </em>growing in a microfluidic turbidostat (100x oil, phase contrast)</li> </ul>
Taxonomy of Knowledge Types for Synthetic Data Generation
<p>The full taxonomy of knowledge types for synthetic data generation in production.</p> <p>For more information, see <a href="https://doi.org/10.54941/ahfe1002915">IHSI 2023 conference paper</a>.</p>
Synthetic fraud data
<p>This repository contains a preprocessed version of the synthetic fraud dataset published in: </p> <p>Padhi, I., Schiff, Y., Melnyk, I., Rigotti, M., Mroueh, Y., Dognin, P., ... & Altman, E. (2021, June). Tabular transformers for modeling multivariate time series. In <em>ICASSP 2021-2021 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)</em> (pp. 3565-3569). IEEE.</p> <p>The original dataset can be accessed <a href="https://ibm.box.com/v/tabformer-data">here</a></p> <p>The subsets are arranged in a rolling window. Each training subset contains 300 days of transaction data. The test sets contain 60 days of data. These subsets are intended for use with SHINE (<a href="https://GitHub.com/rafaelvanbelle/SHINE">https://github.com/rafaelvanbelle/SHINE</a>/)</p> <div> </div>
Engineering cellular communication between light-activated synthetic cells and bacteria (Source data)
<p>Source data files for supplementary figures for published version of "Engineering cellular communication between light-activated synthetic cells and bacteria" https://www.biorxiv.org/content/10.1101/2022.07.22.500923v1</p>
Synthetic flat ntuples in the ROOT data format
<p>These two datasets contain flat ntuples in the ROOT data format, synthetically produced using a ROOT macro. ROOT is the typical data format for storing High Energy Physics (HEP) data. In the ROOT context, a flat ntuple stores data not in the form of C++ object but each branch contains simple numbers (e.g. floating point numbers, integers) or vectors whose dimension might vary on an event-by-event basis.</p> <p>The two datasets have similar content and we distinguish them by adding the suffix <em>bkg</em> and <em>sgn</em> which stands for background and signal. There is no specific physics meaning behind these data. The two datasets contain both 5M of events and 50 branches, of which 10 contain integers, 10 floating point numbers, 10 vectors of integers, and 20 vectors of floating point numbers. The distributions used to fill the datasets are: gaussian, uniform, and exponential. The two datasets differ only on the content of two branches filled with vectors of floating point numbers, where a different mean and standard deviation have been set.</p> <p>These datasets have been produced to test the performance of the MLaaS4HEP framework. More info on the MLaaS4HEP framework in <a href="https://doi.org/10.1007/s41781-021-00061-3">https://doi.org/10.1007/s41781-021-00061-3</a></p>
Digital Twin of a Multi-Arm Robot Platform based on Isaac Sim for Synthetic Data Generation
<p>This data set is required by the following repository<br> https://github.com/AISciencePlatform/icra2023_synthetic_data_pretraining_for_robotics</p>
Wide range of Brachyceran fly taxa attracted to synthetic and semi-synthetic generic noctuid lures and the description of new attractants for Sciomyzidae and Heleomyzidae families - RAW Data
<p>Wide range of Brachyceran fly taxa attracted to synthetic and semi-synthetic generic noctuid lures and the description of new attractants for Sciomyzidae and Heleomyzidae families - RAW Data </p>
Raw data for Broaden the application of Yarrowia Lipolytica synthetic biology tools to explore the potential of Yarrowia clade biodiversity
<p>Raw data for growth curves and fluorescence/OD ratio for article "Broaden the application of Yarrowia Lipolytica synthetic biology tools to explore the potential of Yarrowia clade biodiversity"</p>
TP53 synthetic genomics data for benchmarking variant callers
<p>This is a synthetic genomics dataset generated with <a href="https://github.com/ncsa/NEAT">NEAT </a> for the gene TP53 for the use case of benchmarking somatic variant callers. The reports for all bam files where created using <a href="https://github.com/genome/bam-readcount">bam-readcount</a>.</p> <p>To find out more about our pipeline please visit <a href="https://github.com/BiodataAnalysisGroup/synth4bench">the Biodata Analysis Group GitHub</a> and also our <a href="https://biodataanalysisgroup.github.io/">GitHub page</a> :)</p>
NGS data from: Deploying synthetic coevolution and machine learning to engineer protein-protein interactions
<p>Fine-tuning of protein-protein interactions occurs naturally through coevolution, but this process is difficult to recapitulate in the laboratory. We describe a synthetic platform for protein-protein coevolution that can isolate matched pairs of interacting muteins from complex libraries. This large dataset of coevolved complexes<span class="Apple-converted-space"> </span>drove a systems-level analysis of molecular recognition between Z domain-affibody pairs spanning a wide range of structures, affinities, cross-reactivities, and orthogonalities, and captured a broad spectrum of coevolutionary networks. Furthermore, we harnessed pre-trained protein language models to expand, <em>in silico</em>, the amino acid diversity of our coevolution screen, predicting remodeled interfaces beyond the reach of the experimental library. The integration of these approaches provides a means of generating protein complexes with diverse molecular recognition properties as tools for biotechnology and synthetic biology.</p>
Synthetic along-track altimetry data over 1993-2018 from a NEMO-based simulation of the IMHOTEP project
<p>"Synthetic observations" of along-track SSH have been extracted online during the production of the global, NEMO-based experiment ** IMHOTEP-GAIc**, at every single time and locations where a true SLA observation exists in the AVISO database for the along-track altimetry from the TOPEX, Jason-1, Jason-2 and Jason-3 satellite continuous series over the period 1993-2018. This global ocean/sea-ice/iceberg simulation uses the NEMO model, and has a horizontal resolution of 1/4°. The atmospheric forcing applied at the surface is based on the JRA reanalysis (Kobayashi et al., 2015) and varies over the full range of time-scales from 6 hours to multi-decadal. The freshwater runoff forcing applied to the experiment is fully-variable (daily to multi-decadal) based on the ISBA hydrographic reanalysis for rivers (Decharme et al., 2019) and from altimeter data and regional GCM simulations for the liquid and solid discharges from the Greenland ice-sheet (Mouginot et al 2019). These runoffs are only climatological around Antarctica.<br>This synthetic along-track SSH dataset from the model is available over the altimetry period (1993-2018). It is provided there along with a time-mean model SSH (gridded model field) over the same period that can be used as a proxy for mean dynamic topography ("MDT").</p><p>See the README file for more information. And online documentation is also available here: https://doc-imhotep.readthedocs.io/en/latest/6-Synthetic-Obs.html</p>
Data from: Genome duplication effects on functional traits and fitness are genetic context and species dependent: studies of synthetic polyploid Fragaria
Open the record for dataset details and reuse information.
Data from: Nitroalkanes as ketone synthetic equivalent in C-N and C-S bond formation reaction
Open the record for dataset details and reuse information.
Generation of synthetic whole-slide image tiles of tumours from RNA-sequencing data via cascaded diffusion models
Open the record for dataset details and reuse information.
NGS data from: Deploying synthetic coevolution and machine learning to engineer protein-protein interactions
Open the record for dataset details and reuse information.
Data from: A synthetic biology and green bioprocess approach to recreate agarwood sesquiterpenoid mixtures
Open the record for dataset details and reuse information.
Data from: Genome-wide CRISPR synthetic lethality screen identifies a role for the ADP-ribosyltransferase PARP14 in replication fork stability controlled by ATR
Open the record for dataset details and reuse information.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.