Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

15

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

15 results for “synthetic samples”

Learn how ShareScore rates datasets ↗
zenodo44/100

A Synthetic Global Spatiotemporal Sampled River Discharge Database for Different Satellite Altimetry Mission Orbits

<p><strong>Corresponding peer-reviewed publication</strong></p> <p>This dataset corresponds to all the RRR input and output files that were used in the study reported in:</p> <ul> <li> <p>Sikder, Md. S., Bonnema, M., Emery, C. M., David, C. H., Lin, P., Pan, M., et al. (2021). A Synthetic Data Set Inspired by Satellite Altimetry and Impacts of Sampling on Global Spaceborne Discharge Characterization. <em>Water Resources Research</em>, <em>57</em>(2), e2020WR029035. <a href="https://doi.org/10.1029/2020WR029035">https://doi.org/10.1029/2020WR029035</a></p> </li> </ul> <p>When making use of any of the files in this dataset, please cite both the aforementioned article and the dataset herein.&nbsp;</p> <p>Note that this dataset makes extensive use of the river network and RAPID simulations that were produced in the following study, and the paper is gratefully acknowledged here:</p> <ul> <li> <p>Lin, P., Pan, M., Beck, H. E., Yang, Y., Yamazaki, D., Frasson, R., et al. (2019). Global Reconstruction of Naturalized River Flows at 2.94 Million Reaches. <em>Water Resources Research</em>, <em>55</em>(8), 6499&ndash;6516. <a href="https://doi.org/10.1029/2019WR025287">https://doi.org/10.1029/2019WR025287</a></p> </li> </ul> <p><strong>Version of record and details of this version</strong></p> <p>The version of record for this dataset (i.e. the one used in the aforementioned paper) is version V1.1 available at <a href="https://doi.org/10.5281/zenodo.4064188">https://doi.org/10.5281/zenodo.4064188</a>. This version V2.1 was produced to facilitate testing of the RRR software (<a href="https://github.com/c-h-david/rrr">https://github.com/c-h-david/rrr</a>). Notable details regarding this version compared to V2.0 are as follows:</p> <ul> <li>The temporal sequence files (seq_TIM*.csv) of observations for regular temporal sampling now all have a sampling mean time of 0 second for every river reach instead of the previous value which corresponded to the cycle of observations (e.g. 259,200 seconds for a three-day regular temporal sampling). This allows to start sampling at the onset of each simulation instead of at the end of the first cycle. This change does impact the findings of the study.</li> <li>The sampled discharge files (Qout*.nc) where produced with an updated version of rrr_anl_spl_mod.py which now selects the time step at which a sample is retained using a slightly different approach. The update only impacts sampling results when the sampling time matches the river model output time step exactly, and is more accurate now. This change does impact the findings of the study.</li> </ul>

opencc-by-4.0Oct 2020View details →
zenodo44/100

Synthetic Escherichia coli mixture samples with variable coverage

<p>This dataset contains the synthetic mixture samples and reference sequences - as well as the appropriate metadata - that were originally used in the 2021 revision of the mSWEEP manuscript.<br> <br> There are 87 samples in total, each containing 100bp paired-end Illumina sequencing reads from 10 different&nbsp;<em>Escherichia coli&nbsp;</em>strains from 10 different lineages. The number of reads is set so that the sequencing coverage of the individual strains varies between 50x and 0.10x and sums up to 100x.</p>

opencc-by-4.0Sep 2021View details →
zenodo44/100

UDAPDR Synthetic Queries Datasets Sample

<p>Sample of synthetic queries datasets for&nbsp;<a href="https://arxiv.org/abs/2303.00807">UDAPDR: Unsupervised Domain Adaptation via LLM Prompting and Distillation of Rerankers</a>.</p>

opencc-by-4.0Mar 2023View details →
zenodo44/100

Synthetic data set "Synth1" for the paper "Adaptive Sampling of 3D Spatial Correlations for Focus+Context Visualization"

<p>Synthetic data set "Synth1" for the paper "Adaptive Sampling of 3D Spatial Correlations for Focus+Context Visualization".<br>Preprint of the paper available at: <a href="https://arxiv.org/abs/2309.03308">https://arxiv.org/abs/2309.03308</a></p>

opencc-by-4.0Oct 2023View details →
zenodo40/100

A synthetic sample of short-cadence solar-like oscillators for TESS

<p>These are the simulated lightcurves for a synthetic sample of short-cadence solar-like oscillators as they might be observed by the Transiting Exoplanet Survey Satellite (TESS), as presented by Ball et al. (2018),&nbsp;&quot;A synthetic sample of short-cadence solar-like oscillators for TESS&quot;.&nbsp;</p> <p>Each zip file contains the FITS lightcurves and mode data&nbsp;for stars observed in one sector, starting in the southern ecliptic&nbsp;hemisphere. The CSV file contains a selection of data contained in the FITS headers.</p> <p>For more detail, read the paper on <a href="https://arxiv.org/abs/1809.09108">arXiv</a>.</p> <p>Note that the FITS headers incorrectly report the units of the white noise level as <span class="math-tex">\(\mathrm{ppm}\)</span>&nbsp;rather than&nbsp;<span class="math-tex">\(\mathrm{ppm}\cdot\mathrm{hr}^{1/2}\)</span>. The white noise level excludes any systematic component, which should be added in quadrature if desired.</p>

opencc-by-sa-4.0Oct 2018View details →
zenodo36/100

Synthetic viral samples with DVGs

<p>Two synthetic&nbsp;groups of fastq files with a known population of Defective Viral Genomes (DVGs) and the full genome virus.&nbsp;</p> <p>- &#39;sars216&#39; dataset was&nbsp;generated from SARS-CoV-2 genome</p> <p>- &#39;tumv72&#39;&nbsp;dataset was generated from Turnip Mosaic Virus.</p> <p>The composition of each file is defined in the composition file uploaded in the GitHub repository:&nbsp;https://github.com/MJmaolu/SyntheticSamplesWithDVGs.</p> <p>Simulated samples used to evaluate the performance of the DVGfinder tool (https://github.com/MJmaolu/DVGfinder).</p>

opencc-by-4.0Apr 2022View details →
zenodo32/100

mGEMS synthetic mixed samples (Supplementary Table)

<p>Supplementary Table from the mGEMS publication, which contains the information about the synthetic mixed samples in the manuscript. The table contains the accession numbers of the isolate sequencing reads assigned to each mixed sample, their lineage assignments, as well as assembly statistics (total length, number of contigs, N50, L50) for both the isolate sequencing data (assembled with shovill v0.9.0) and the synthetic mixed samples processed with the mGEMS pipeline (mGEMS binner v0.1.1, Themisto v0.1.1, mSWEEP v1.3.2, and shovill v0.9.0).</p>

opencc-by-4.0Mar 2020View details →
zenodo32/100

Sample Synthetic Training Dataset: Part1

<p>A sample dataset that we sub-sampled in order to train the yes/no lizard and species identification machine learning model displayed in the nature methods brief communication paper.&nbsp;</p> <p>This Part Contains the Data For E. e. croceater and &quot;Blank Backgrounds&quot;.</p> <p>&nbsp;</p>

opencc-by-4.0May 2022View details →
zenodo32/100

Creating Site Specific Synthetic ML Training Datasets for Conservation: Sample Synthetic Training Dataset [Resized]

<p>These files support the paper &quot;Creating Site Specific Synthetic Machine Learning Training Datasets for Conservation&quot; by providing a sample of the synthetic data generated using the described methodology. Additionally, a sample of this dataset could and has been used to train&nbsp;different image classifier machine learning models.&nbsp;</p> <p>The provided ZIP file contains four folders of synthetic training data JPEG images, separated by species, then background-type within the subsequent sub-folders. All of these folders and sub-folders are clearly labeled. The folder marked &quot;Blank Backgrounds&quot; is the only folder with images that DO NOT contain any herpetofauna. Instead, this folder consists of only background images, which were used to train an&nbsp;image classifier model to recognize the presence of (or lack there-of) a herpetofauna specimen within an image.&nbsp;&nbsp;</p> <p><strong>*NOTE - These images were resized for the purposes of efficient upload to Github//Zenodo. There may be same quality/detail loss compared to the original images...*</strong></p>

opencc-by-4.0May 2022View details →
zenodo28/100

Synthetic brain tumor MRI samples - 2022-09

<p>Synthetic brain tumor MRI samples - 2022-09</p>

opencc-by-4.0Sep 2022View details →
zenodo28/100

Synthetic brain tumor MRI samples - 2022-val

<p>Synthetic brain tumor MRI samples - 2022-val</p>

opencc-by-4.0Nov 2022View details →
zenodo28/100

Synthetic brain tumor MRI samples - 2022-11

<p>Synthetic brain tumor MRI samples - 2022-11</p>

opencc-by-4.0Nov 2022View details →
zenodo28/100

Synthetic brain tumor MRI samples - 2022-12

<p>Synthetic brain tumor MRI samples - 2022-12</p>

opencc-by-4.0Nov 2022View details →
zenodo28/100

Synthetic brain tumor MRI samples - 2022-test

<p>Synthetic brain tumor MRI samples - 2022-test</p>

opencc-by-4.0Dec 2022View details →
nasa16/100

DXC'09 Synthetic Track Sample Data

Sample data for the DXC'09 Synthetic Track.

restrictednotspecifiedMar 2025View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record