Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

252

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

252 results for “Synthetic data”

Learn how ShareScore rates datasets ↗
zenodo32/100

the data of "High-pressure behavior of synthetic quenchable hydroxylbastnasite-(Sm): Implications for rare earth elements, carbon, and water transmission into the deep earth"

<p>This is the package of the Raman data of synthesized Sm(CO<sub>3</sub>)OH at high-pressure and ambient temperature &nbsp;and synchrotron radiation X-Ray&nbsp;data&nbsp;of synthesized Sm(CO<sub>3</sub>)OH&nbsp;&nbsp;at high-pressure and ambient temperature, they are&nbsp;in *.asc and *.chi respectively.</p>

opencc-by-4.0May 2022View details →
zenodo32/100

Tutorial video to use the 'sm-dtw' assessment tool with simulated synthetic phyllotaxis data

<p>This tutorial explain how to use &#39;sm-dtw&#39;, an assessment tool which has been designed to evaluate how good a phyllotaxis measure is from a plant phenotyping experiment. To get data to play with and explore all possible case scenarios, we also designed a program to generate phyllotaxis data and simulate typical errors produced by a phenotyping experiment.</p> <p>This tutorial explains:</p> <p>1) the context in which such a tool is useful (what is a phyllotaxis measure ? What kind of phenotyping experiment ? Why do need to evaluate your results ? What are typical errors you want to detect ?)</p> <p>2) how to download and use the two programs (&#39;sm-dtw&#39; and the generator of phyllotaxis data)</p> <p>3) how to play with the programs thanks to pedagogical demonstrator notebooks.</p> <p>In brief, there are 3 notebooks that are meant to be run as three successive steps:</p> <ul> <li>step1 / Notebook 1: it allows anybody to simulate phyllotaxis data (pair of sequences consisting of ground truth sequences and their error-prone measures),</li> <li>step2 / Notebook 2: assess the measure performance with our new program &lsquo;sm-dtw&rsquo; (detect errors and quantify precision)</li> <li>step3/Notebook 3: control that sm-dtw program correctly interprets the differences between the measure and its ground truth reference.</li> </ul>

opencc-by-4.0Jul 2022View details →
zenodo32/100

The material properties of a bacterial-derived biomolecular condensate tune biological function in natural and synthetic systems. Source data

<p>Supplementary information for manuscript titled &quot;The material properties of a bacterial-derived biomolecular condensate tune biological function in natural and synthetic systems&quot;</p>

openSep 2022View details →
zenodo32/100

Synthetic data 0.1

<p>Synthetic data generated with stable diffution. Consists of 6,390 images. Real dataset used for generating:&nbsp;<a href="../records/10203721" rel="nofollow">https://zenodo.org/records/10203721</a>. stable diffusiton model (img2img) used for generating:&nbsp;<a href="https://github.com/AUTOMATIC1111/stable-diffusion-webui">https://github.com/AUTOMATIC1111/stable-diffusion-webui</a>.&nbsp; denoising strength 0.1</p> <p>Project (practical wotk for Bachelor's paper) where data is used for model training: https://github.com/rkalvitis/Bakalaurs.&nbsp;</p>

opencc-by-4.0May 2024View details →
zenodo32/100

Synthetic data 0.15

<p>Synthetic data generated with stable diffution. Consists of 6,390 images. Real dataset used for generating:&nbsp;<a href="../records/10203721" rel="nofollow">https://zenodo.org/records/10203721</a>. stable diffusiton model (img2img) used for generating:&nbsp;<a href="https://github.com/AUTOMATIC1111/stable-diffusion-webui">https://github.com/AUTOMATIC1111/stable-diffusion-webui</a>.&nbsp; denoising strength 0.15</p> <p>Project (practical wotk for Bachelor's paper) where data is used for model training: https://github.com/rkalvitis/Bakalaurs.&nbsp;</p>

opencc-by-4.0May 2024View details →
zenodo32/100

Synthetic data RELION project archive for ccpem-pipeliner cryodrgn jobs slow tests

<p>Synthetic data reconstructed via RELION4.0 used for testing CryoDRGN jobs in ccpem-pipeliner (https://gitlab.com/ccpem/ccpem-pipeliner). Micrographs simulated via Roodmus (<span>https://doi.org/10.1101/2024.04.29.590932</span>) using molcular conformations originating from DE Shaw simulation DESRES-ANTON-11021571 (D. E. Shaw Research, "Molecular Dynamics Simulations Related to SARS-CoV-2," D. E. Shaw Research Technical Data, 2020. https://www.deshawresearch.com/downloads/download_trajectory_sarscov2.cgi/)</p>

opencc-by-4.0Jul 2024View details →
zenodo32/100

Synthetic images of fluorescent spots and ground truth data

<p>Synthetical images of fluorescent spots and ground truth data created with the simcep software.</p>

opencc-by-4.0Jul 2018View details →
zenodo32/100

Faraday Synthetic Data Outputs

<p>This dataset contains 10 million synthetic load profiles of trained on over 300M smart meter readings from 20K Octopus Energy UK households sampled between 2021 and 2022, and is conditioned on labels such as the:</p> <ol> <li> <p>Property types: house, flat, terraced, detached, semi-detached etc</p> </li> <li> <p>Energy performance certificate (EPC) rating: A/B/C, D/E, F/G etc</p> </li> <li> <p>Low Carbon Technology (LCT) ownership: heat pumps, electric vehicles, solar PVs etc</p> </li> <li> <p>Seasonality: days of the week and month of the year</p> </li> </ol> <p><strong>&nbsp;</strong></p> <p>For more information about Faraday, please refer to the <a href="https://www.climatechange.ai/papers/iclr2024/43">workshop paper</a> that Centre for Net Zero presented at ICLR 2024.&nbsp; For more information about OpenSynth, please visit our Github repository&nbsp;<a href="https://github.com/OpenSynth-energy/OpenSynth">https://github.com/OpenSynth-energy/OpenSynth</a>. For more news and updates on OpenSynth, please subscribe to our mailing list&nbsp;<a href="https://lists.lfenergy.org/g/opensynth-discussion">here</a>.</p>

openAug 2024View details →
zenodo32/100

Synthetic Datasets for "Binary Classification Optimisation with AI-Generated Data"

<p>Images of melanomas and Basal Cell Carcinoma generated with a stylegan2. Dataset corresponding to the article "Binary Classification Optimisation with&nbsp;AI-Generated Data"</p>

opencc-by-4.0Sep 2024View details →
zenodo32/100

Analysing detection thresholds of lithological complexity and karst overprinting in outcrop-scale seismic models (Synthetic seismic models - supplementary data)

<p>This dataset includes pre-loaded seismic models based on digital outcrop models (DOMs) representing 11 different case scenarios with varying signal processing setting (frequency, illumination angle, noise) and different geological models for the DOMs. The case scenarios are set up to investigate detection tresholds in seismic models of thin beds with a high degree of lithological variation.&nbsp;</p> <p>The dataset includes:</p> <ul> <li>A Petrel 2022 project with 11 pre-loaded seismic models based on two different geomodels presented in the paper.</li> <li>The input SGY files for each seismic section and PSF used.</li> </ul> <p>This dataset is related to a manuscript currently in review for Marine and Petroleum Geology.</p>

opencc-by-4.0Oct 2024View details →
zenodo32/100

Data for "Mechanism of Solid-State 1H Photochemically Induced Dynamic Nuclear Polarization in a Synthetic Donor−Chromophore−Acceptor at 0.3 T"

<p>This data is associated with the publication "Mechanism of Solid-State 1H Photochemically Induced Dynamic Nuclear Polarization in a Synthetic Donor&minus;Chromophore&minus;Acceptor at 0.3 T" and contains the NMR and photo-CIDNP data discussed in the publication.</p>

opencc-by-4.0Oct 2024View details →
zenodo32/100

Data archive for paper "Copula-based synthetic data augmentation for machine-learning emulators"

<p><strong>Overview</strong></p> <p>This is the data archive for paper &quot;<a href="https://doi.org/10.5194/gmd-14-5205-2021">Copula-based synthetic data augmentation for machine-learning emulators</a>&quot;. It contains the paper&rsquo;s data archive with model outputs (see <code>results</code> folder) and the Singularity image for (optionally) re-running experiments.</p> <p>For the Python tool used to generate synthetic data, please refer to <a href="https://github.com/dmey/synthia">Synthia</a>.</p> <p><strong>Requirements</strong></p> <ul> <li><a href="https://sylabs.io/singularity/">Singularity</a> &gt;= 3</li> <li><a href="https://en.wikipedia.org/wiki/Portable_Batch_System">Portable Batch System</a> (PBS) job scheduler*</li> <li>Today&#39;s high-performance computer (e.g. ~ 32 CPUs @ 2 500 MHz with 64 GB of RAM )</li> </ul> <p>*Although PBS in not a strict requirement, it is required to run all helper scripts as included in this repository. Please note that depending on your specific system settings and resource availability, you may need to modify PBS parameters at the top of submit scripts stored in the <code>hpc</code> directory (e.g. <code>#PBS -lwalltime=72:00:00</code>).</p> <p><strong>Usage</strong></p> <p>To reproduce the results from the experiments described in the paper, first fit all copula models to the reduced NWP-SAF dataset with:</p> <pre><code>qsub hpc/fit.sh</code></pre> <p>then, to generate synthetic data, run all machine learning model configurations, and compute the relevant statistics use:</p> <pre><code>qsub hpc/stats.sh qsub hpc/ml_control.sh qsub hpc/ml_synth.sh</code></pre> <p>Finally, to plot all artifacts included in the paper use:</p> <pre><code>qsub hpc/plot.sh</code></pre> <p><strong>Licence</strong></p> <p>Code released under <a href="./LICENSE.txt">MIT license</a>. Data from the reduced NWP-SAF dataset released under <a href="./data/LICENSE.txt">CC BY 4.0</a>.</p>

openother-atDec 2020View details →
zenodo32/100

Use of synthetic aperture radar data for the determination of vegetation indices

<p>Database for the article &quot;Use of synthetic aperture radar data for the determination of vegetation indices&quot;</p>

opencc-by-4.0Sep 2021View details →
zenodo32/100

Trusted World of Corona (TWOC) - WHO COVID-19 Case Report Form (CRF) Synthetic Data - FAIR

<p>The synthetic dataset that models COVID-19 real world observations from WHO COVID-19 RAPID Version CRFs of hospitalized patients for the hypothesis under study, originally created by the TWOC project.</p>

openother-pdSep 2021View details →
zenodo32/100

Synthetic data set to evaluate and benchmark the performance of multiple linear regression algorithms in Scikit-Learn and SANElib

<p>The datasets respresent different numbers of columns and rows to measure the scalability of linear regression algorihms in terms of columns and rows.</p>

opencc-by-4.0Aug 2021View details →
zenodo32/100

Input data for RemoTeC synthetic measurements and retrieval

<p>RemoTeC is a retrieval algorithm developed for the retrieval of trace gas column-averaged dry air mole fractions from measured level 1b radiance spectra in the near-infrared (NIR) and shortwave-infrared (SWIR) bands. It is open access software developed by The Netherlands Institute for Space Research (SRON) and Karlsruhe Institute for Technology (KIT). &nbsp;</p> <p>The dataset available here is the input data for the RemoTeC synthetic measurement and retrieval code. The source code for these algorithms can be found at:</p> <p><a href="https://bitbucket.org/sron_earth/remotec_synthetic_measurements/src/master/">https://bitbucket.org/sron_earth/remotec_multi_purpose/src/main/</a></p> <p>Here we provide the data for both of these algorithms. The retrieval code should be used in conjunction with the synthetic measurement generator, but the synthetic measurement generator can be used as stand alone code. Any publications from using either of these codes should reference DOI for this dataset.</p>

opencc-by-4.0Nov 2023View details →
zenodo32/100

ULTRA-SR: Synthetic Data

<p><strong>ULTRA-SR: Ultrasound Localisation and TRacking Algorithms for Super Resolution</strong></p> <p>The ULTRA-SR Challenge was held at the 2022 IEEE International Ultrasound Symposium in Venice, Italy.</p> <p>The aim of the ULTRA-SR challenge is to extensively evaluate the performance of localization and tracking algorithms for Ultrasound Localization Microscopy.</p> <p>Two realistic simulated datasets of microvascular flow generated, using flow physics, acoustic field simulation with low and high clinical ultrasound frequencies, integrated with non-linear bubble dynamics.</p> <p>The ULTRA-SR Challenge dataset is publicly available for researchers to benchmark new algorithms and software.</p> <p>Additional information&nbsp;is available at:&nbsp;<a href="https://ultra-sr.com/">https://ultra-sr.com/</a></p> <p>Pre-print paper to cite which describes the data generation:&nbsp;<a href="https://doi.org/10.48550/arXiv.2211.00754">https://doi.org/10.48550/arXiv.2211.00754</a></p>

opencc-by-4.0Oct 2022View details →
zenodo32/100

Dataset of CCFs, dispersion data, synthetic ambient noise data, and rupture simulation data in the Anninghe study

<p>This is the main dataset supporting the Anninghe paper.</p> <p>sta.loc represents the station list file.<br> 3DVelocityModel folder contains the final velocity model file Anninghe_Vs3D.dat.<br> CCFs_Original folder contains the original CCFs obtained from the Anninghe array.<br> CCFs_Separated folder contains the CCFs after model separation. In CCFs_Separated folder, the merge folder is the final average CCFs used for later tomography.<br> DisperData folder contains the seimiautopicked dispersion data from the separated CCFs.<br> RuptureSimulation folder contains the rupture simulation data in the paper.<br> SynNoiseData folder contains the CCFs calculated from synthetic ambient noise data.<br> &nbsp;</p>

opencc-by-4.0Aug 2022View details →
zenodo32/100

AIRRSHIP: Example synthetic B cell receptor repertoire data

<p>Example repertoire data generated by AIRRSHIP (https://github.com/Cowanlab/airrship). Four repertoires are available (two with SHM, two without), each of which contains 100,000 sequences produced using the default AIRRSHIP parameters. Sequence data is contained in the FASTA files, TSV files give details of each step in the generation process, summary file shows the command given to AIRRSHIP and the locus file contains the alleles used in the repertoire. See https://airrship.readthedocs.io/en/latest/output/ for more information on file format.</p> <p>Repertoires were created using version 0.1.2 of AIRRSHIP.</p>

opencc-by-4.0Jan 2023View details →
zenodo32/100

Synthetic data for Aligning Distant Sequences to Graphs using Long Seed Sketches

<p>Each directory inside the folder correponds to the number of levels used to generate the dataset (for more details, see the description written in the publication). Inside each directory, the files &quot;reference_X&quot; and &quot;mutated_X&quot; correpond to the sequences reference and mutated at rate X, respectively.&nbsp;</p>

opencc-by-4.0Mar 2023View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record