Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

37

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

37 results for “subsampling”

Learn how ShareScore rates datasets ↗
zenodo52/100

Sparse observations induce large biases in estimates of the global ocean CO2 sink: an ocean model subsampling experiment

<p>Dataset underlying the analysis in Hauck et al., 2023: Sparse observations induce large biases in estimates of the global ocean CO<sub>2</sub> sink - an ocean model subsampling experiment, Philosophical Transactions A</p> <p>Surface ocean partial pressure of CO<sub>2 </sub>(pCO<sub>2</sub>) and air-sea CO<sub>2</sub> flux reconstructions, using two mapping methods (MPI-SOM-FFN, CarboScope) three different sampling masks: SOCAT, SOCAT+SOCCOM, IDEAL (based on bgcArgo, Roemmich et al., 2019).</p> <p>Also, all FESOM-REcoM output fields that were used in the reconstructions are provided.</p> <p>We further provide the three masks that were used for subsampling: SOCAT, SOCAT+SOCCOM, IDEAL (bgcArgo).</p> <p>&nbsp;</p>

opencc-by-4.0May 2023View details →
zenodo48/100

Subsample of the maximum Water Area Extent of Telangana Rainwater Harvesting System from Pléiades DEM

<p>Small Reservoirs maximum water area extent polygones composing the Rainwater Harvesting System over the Telangana state, South-India. Maximum Water Area Extent are elevation isolines selected by hand from Very High Resolution Digital Elevation Model (VHR DEM at 2 meters resolution) derived from pairs of stereoscopic Pl&eacute;iades images at 50cm resolution (10.5281/zenodo.10403040). The selection is made to find the area that contain both MWAE derived from Sentinel-2 (10.5281/zenodo.10402199) and Landsat archives (Global Surface Water, <span><a href="https://doi.org/10.1038/nature20584" target="_blank" rel="noopener">10.1038/nature20584</a></span>) curated from rivers and big dams (10.5281/zenodo.10402096).</p>

opencc-by-4.0Dec 2023View details →
zenodo48/100

Subsample of the maximum Water Area Extent of Telangana Rainwater Harvesting System from Sentinel-2

<p>Small Reservoirs Maximum Water Area Extent polygones (MWAE) composing the Rainwater Harvesting System (RHS) derieved from Sentinel-2 Multispectral data in the Telangana state, South-India. MWAE is extracted from Sentinel-2 cloud free images time serie collected from 2016 to 2021 (last access in 2021) over the area covered by stereoscopic images acquired from Pl&eacute;iades satellites (DEM available 10.5281/zenodo.10403040). A random forest classification is used with a set of training and validation samples. These samples are Sentinel pixel locations (10 x 10 meters) corresponding to permanent water pixels extracted from Global Surface Water datasets (doi:10.1038/nature20584) and never flooded pixels derived from Height Above Nearest Drainage data-set (10.1016/j.jhydrol.2011.03.051).</p>

opencc-by-4.0Dec 2023View details →
edi44/100

NEON Biorepository Soil Microbe Collection (Bulk Subsamples) (repackaging of occurrences published by the NEON Biorepository Data Portal)

This collection contains samples collected during periodic soil sampling and frozen at ultra-low temperatures in order to provide material for microbial sequencing or other microbial analyses (NEON sample classes: sls_soilCoreCollection_in.geneticArchiveSample1ID, sls_soilCoreCollection_in.geneticArchiveSample2ID, sls_soilCoreCollection_in.geneticArchiveSample3ID, sls_soilCoreCollection_in.geneticArchiveSample4ID, sls_soilCoreCollection_in.geneticArchiveSample5ID,sls_metagenomicsPooling_in.compositeSampleID). Archive samples are collected during each soil sampling bout and are promptly frozen. Three unique locations are sampled per plot, with ten plots per site. Bouts occur three times per year in order to capture the prevailing conditions at the site during different seasons, except in Alaska where only 1 bout is possible. Soil sampling is conducted to a maximum depth of 30 &plusmn; 1 cm, and when organic (O) and mineral (M) horizons are present within a single profile, they are separated prior to analysis and archiving. However, other sub-horizons are not separated. During the majority of bouts, only the top horizon (O if present, else M) is collected and archived. Soils are homogenized and non-soil material is removed by hand in the field, then subsamples are immediately frozen on dry ice. They are maintained in ultra-low temperature freezers until shipment to the Biorepository. See links below for NEON data products that provide physical, chemical, and biological measurements for these same soils (soil pH and moisture are always measured; chemical properties as well as microbial community composition and biomass are determined only for a subset of collection bouts). In addition, a more detailed characterization of the dominant soil types at each site, including taxonomy, texture, bulk density, and geochemical properties, occurred during the construction period of NEON through two projects. These data are available in NEON data products Soil physical and chemical

openCustomFeb 2023View details →
zenodo40/100

Subsampled fastq from GSM7890929 and GSM7890951

<p>This record contains:<br>- 4 fastqs which are subsets of fastqs corresponding to mESC and 168h gastruloids (the whole fastqs are available at SRA).<br>- All command lines which have been used to generate these fastqs:<br>&nbsp; &nbsp;- The main script is pipeline.sh<br>&nbsp; &nbsp;- Other files are temporary files or accessory scripts<br><br>The subsetting is totally biased and promote cells from different clusters so the proportion are absolutely artifical!</p><p>The goal of this subsetting is to be able to use this data in training and workflow testing.</p>

opencc-by-4.0Nov 2023View details →
zenodo40/100

dataset for "basic setting", "+ binary semantic loss", "+ class weights", "+ height weights", "+ region weights", "+ elastic distortion and subsampling", "+ TreeMix" in paper Automated forest inventory: analysis of high-density airborne LiDAR point clouds with 3D deep learning

<p>dataset for "basic setting", "+ binary semantic loss", "+ class weights", "+ height weights", "+ region weights", "+ elastic distortion and subsampling", "+ TreeMix" in paper Automated forest inventory: analysis of high-density airborne LiDAR point clouds with 3D deep learning</p>

opencc-by-4.0Mar 2024View details →
zenodo40/100

Course 27255 - Subsampled Sequencing Results

<p>This dataset contains a subset of the&nbsp;raw fastq data of the microbial isolates collected during the course <strong>27255 - Advanced Experimental Prokaryotic Molecular Biology and Ecology (Spring 2022) </strong>at DTU.</p>

opencc-by-4.0Apr 2022View details →
zenodo40/100

STEAD subsample 4 CDiffSD

<h1>STEAD Subsample Dataset for CDiffSD Training</h1> <h2>Overview</h2> <p>This dataset is a subsampled version of the STEAD dataset, specifically tailored for training our CDiffSD model (Cold Diffusion for Seismic Denoising). It consists of four HDF5 files, each saved in a format that requires Python's `h5py` method for opening.</p> <h2>Dataset Files</h2> <p>The dataset includes the following files:</p> <ul> <li>train: Used for both training and validation phases (with validation train split). Contains earthquake ground truth traces.</li> <li>noise_train: Used for both training and validation phases. Contains noise used to contaminate the traces.</li> <li>test: Used for the testing phase, structured similarly to train.</li> <li>noise_test: Used for the testing phase, contains noise data for testing.</li> </ul> <p>Each file is structured to support the training and evaluation of seismic denoising models.</p> <h2>Data</h2> <p>The HDF5 files named noise contain two main datasets:</p> <ul> <li>traces: This dataset includes N number of events, with each event being 6000 in size, representing the length of the traces. Each trace is organized into three channels in the following order: E (East-West), N (North-South), Z (Vertical).</li> <li>metadata: This dataset contains the names of the traces for each event.</li> </ul> <p>Similarly, the train and test files, which contain earthquake data, include the same traces and metadata datasets, but also feature two additional datasets:</p> <ul> <li>p_arrival: Contains the arrival indices of P-waves, expressed in counts.</li> <li>s_arrival: Contains the arrival indices of S-waves, also expressed in counts.</li> </ul> <h2><br>Usage</h2> <p>To load these files in a Python environment, use the following approach:</p> <p><code>```python</code></p> <p><code>import h5py</code><br><code>import numpy as np</code></p> <p><code># Open the HDF5 file in read mode</code><br><code>with h5py.File('train_noise.hdf5', 'r') as file:</code><br><code>&nbsp; &nbsp; # Print all the main keys in the file</code><br><code>&nbsp; &nbsp; print("Keys in the HDF5 file:", list(file.keys()))</code></p> <p><code>&nbsp; &nbsp; if 'traces' in file:</code><br><code>&nbsp; &nbsp; &nbsp; &nbsp; # Access the dataset</code><br><code>&nbsp; &nbsp; &nbsp; &nbsp; data = file['traces'][:10] &nbsp;# Load the first 10 traces</code></p> <p><code>&nbsp; &nbsp; if 'metadata' in file:</code><br><code>&nbsp; &nbsp; &nbsp; &nbsp; # Access the dataset</code><br><code>&nbsp; &nbsp; &nbsp; &nbsp; trace_name = file['metadata'][:10] &nbsp;# Load the first 10 metadata entries```</code></p> <p>Ensure that the path to the file is correctly specified relative to your Python script.</p> <h2>Requirements</h2> <p>To use this dataset, ensure you have Python installed along with the Pandas library, which can be installed via pip if not already available:</p> <p><code>```bash</code><br><code>pip install numpy</code><br><code>pip install h5py</code><br><code>```</code></p>

opencc-by-4.0Apr 2024View details →
zenodo40/100

Updated MCMC chains and subsamples for Hector calibration paper

<p>These csvs contain MCMC chains and sampled subsets for calibration of the Hector simple climate model (<a href="https://github.com/JGCRI/hector">https://github.com/JGCRI/hector</a>, DOI:10.5194/gmd-8-939-2015). This series of five calibrations with different observational constraints and free parameters is described in Vega-Westhoff et al. (2019, DOI:10.1029/2018EF001082;&nbsp;see their&nbsp;table 1).</p> <p>The calibrations use a version of Hector that includes the BRICK sea-level module (<a href="https://github.com/scrim-network/BRICK">https://github.com/scrim-network/BRICK</a>, DOI:10.5194/gmd-10-2741-2017). Hector with BRICK is available on my fork of the Hector model (https://github.com/bvegawe/hector/tree/dev_slr). The calibration process is also adapted from BRICK. The code used to produce these chains can be found at&nbsp;https://github.com/bvegawe/hector_probabilistic, DOI:10.5281/zenodo.3236411.</p> <p>These five sets of&nbsp;MCMC chains were produced using hector_calib_driver.R. Inputs used to create each calibration are specified below:&nbsp;</p> <p>T.csv:&nbsp; Rscript hector_calib_driver.R -f *output folder* -n 200000 --forcing TRUE --np 10 --model_set onlyT_model --obs_set onlyT_obs</p> <p>TOHC.csv: Rscript hector_calib_driver.R -f *output folder* --forcing TRUE --np 10 --model_set doeclim_model --obs_set doeclim_obs</p> <p>TOHCGICGISAIS.csv: Rscript hector_calib_driver.R -f *output folder* --forcing TRUE --np 10 --model_set noTE_model --obs_set noTE_obs</p> <p>TTE.csv: Rscript hector_calib_driver.R -f *output folder* --forcing TRUE --np 10 --model_set onlyTE_model --obs_set doeclimTE_obs</p> <p>TTEGICGISAIS.csv: Rscript hector_calib_driver.R -f *output folder* --forcing TRUE --np 10 --model_set all_model --obs_set noOcheat_obs</p> <p>&nbsp;</p> <p>V1.1 - Updated with new calibrations. Previously only included two calibration files, <a href="https://zenodo.org/api/files/5f877c29-3990-4d12-92f0-1a7c6f7d271c/wSLR_calib.csv?versionId=b9c5ea10-6f52-4b91-b30a-be68cb58fba6">wSLR_calib.csv</a>, which was&nbsp;a &#39;full&#39; model calibration, including both ocean heat and thermosteric sea level historical constraints, and woSLR_calib.csv, which included only ocean heat and temperature constraints. After review, we redid the calibration experiment, including additional&nbsp;combinations of constraints, and removing the &#39;full&#39; calibration, which treated the ocean heat and thermosteric sea level constraints as independent.</p>

opencc-by-4.0Jun 2019View details →
zenodo36/100

Subsampled 20x of sequence reads from study PRJNA315192

<p>subsampled reads from study PRJNA315192 for study purposes.</p> <p>Nanopore version</p> <p>&nbsp;</p> <p>Original works:</p> <p>Greig, D.R., Do Nascimento, V., Gally, D.L.&nbsp;<em>et al.</em>&nbsp;Re-analysis of an outbreak of Shiga toxin-producing&nbsp;<em>Escherichia coli</em>&nbsp;O157:H7 associated with raw drinking milk using Nanopore sequencing.&nbsp;<em>Sci Rep</em>&nbsp;<strong>14</strong>, 5821 (2024). https://doi.org/10.1038/s41598-024-54662-0</p>

opencc-by-4.0Oct 2024View details →
zenodo36/100

Parhyale hawaiensis regenerating leg (subsampled dataset)

<p>This is a demo dataset for the ELEPHANT tracking software stored in tif format. The dataset is a short subset of a live imaging of a regenerating leg of the crustacean&nbsp;<em>Parhyale hawaiensis</em> acquired on Zeiss LSM 800 confocal microscope. Please see the paper for details.</p> <p>Original data:</p> <p><a href="https://doi.org/10.5281/zenodo.4630932">10.5281/zenodo.4630932</a></p> <p><a href="https://zenodo.org/records/4630932">https://zenodo.org/records/4630932</a></p> <p>License:</p> <p>Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 International (CC BY-NC-ND 4.0)</p> <p>https://creativecommons.org/licenses/by-nc-nd/4.0/</p>

opencc-by-nc-nd-4.0Oct 2024View details →
zenodo32/100

Multi-label Pathway Prediction based on Active Dataset Subsampling

<p>We include samples of various data types used in the work &quot;Multi-label Pathway Prediction based on Active Dataset Subsampling&quot; (under-review)</p> <p>More information about the software package and instructions are provided in&nbsp;<a href="https://github.com/hallamlab/leADS">hallamlab/leADS</a></p>

opencc-by-4.0Jul 2020View details →
dryad32/100

Data from: Subsampling reveals that unbalanced sampling affects STRUCTURE results in a multi-species dataset

Studying the genetic population structure of species can reveal important insights into several key evolutionary, historical, demographic, and anthropogenic processes. One of the most important statistical tools for inferring genetic clusters is the program STRUCTURE. Recently, several papers have pointed out that STRUCTURE may show a bias when the sampling design is unbalanced, resulting in spurious joining of underrepresented populations and spurious separation of overrepresented populations. Suggestions to overcome this bias include subsampling and changing the ancestry model, but the performance of these two methods has not yet been tested on actual data. Here, I use a dataset of twelve high-alpine plant species to test whether unbalanced sampling affects the STRUCTURE inference of population differentiation between the European Alps and the Carpathians. For four of the twelve species, subsampling of the Alpine populations –to match the sample size between the Alps and the Carpathians– resulted in a drastically different clustering than the full dataset. On the other hand, STRUCTURE results with the alternative ancestry model were indistinguishable from the results with the default model. Based on these results, the subsampling strategy seems a more viable approach to overcome the bias than the alternative ancestry model. However, subsampling is only possible when there is an a priori expectation of what constitute the main clusters. Though these results do not mean that the use of STRUCTURE should be discarded, it does indicate that users of the software should be cautious about the interpretation of the results when sampling is unbalanced.

opencc-zeroDec 2017View details →
zenodo32/100

McKinney Livox Avai -- 50 pcnt subsample

A 50% subsampling of the lidar points for the original McKinney House lidar scan. This should make for a faster model loading experience. Source: Objaverse 1.0 / Sketchfab

opencc-by-nc-sa-2.0Oct 2021View details →
zenodo32/100

A test dataset for LiverCT (Lu2022_subsampled)

<p>A random subset of Lu2022, including 10000 cells.</p> <p>Lu, Y., Yang, A., Quan, C.&nbsp;<em>et al.</em>&nbsp;A single-cell atlas of the multicellular ecosystem of primary and metastatic hepatocellular carcinoma.&nbsp;<em>Nat Commun</em>&nbsp;<strong>13</strong>, 4594 (2022). https://doi.org/10.1038/s41467-022-32283-3</p>

opencc-by-4.0Aug 2024View details →
zenodo32/100

Assessment of Subsampling Schemes for Compressive Nano-FTIR Imaging: Underlying Dataset

<p>The dataset in this publication is related to the following Publication:</p> <p>Metzner, S., K&auml;stner, B., Marschall, M., Wubbeler, G., Wundrack, S., Bakin, A., Hoehl, A., Ruhl, E., &amp; Elster, C. (2022).<br> Assessment of Subsampling Schemes for Compressive Nano-FTIR Imaging.<br> <em>IEEE Transactions on Instrumentation and Measurement</em>, <em>71</em>, 1&ndash;8.<br> https://doi.org/10.1109/TIM.2022.3204072</p> <p>It contains the underlying code and the data for generating the figures.</p>

opencc-by-4.0Mar 2023View details →
dryad32/100

Data from: Subsampling reveals that unbalanced sampling affects STRUCTURE results in a multi-species dataset

Open the record for dataset details and reuse information.

publicJul 2018View details →
dryad32/100

Data from: Phylogenomic subsampling and the search for phylogenetically reliable loci

Open the record for dataset details and reuse information.

publicJun 2021View details →
dryad28/100

Ordered phylogenomic subsampling enables diagnosis of systematic errors in the placement of the enigmatic arachnid order Palpigradi

<p><span><span><span><span><span><span><span><span><span><span><span>The miniaturized arachnid order Palpigradi has ambiguous phylogenetic affinities, due to its odd combination of plesiomorphic and derived morphological traits. This lineage has never been sampled in phylogenomic datasets because of its small body size and fragility of most species, a sampling gap of immediate concern to recent disputes over arachnid monophyly. To redress this gap, we sampled a population of the cave-inhabiting species <i>Eukoenenia spelaea</i> from Slovakia and inferred its placement in the phylogeny of Chelicerata using dense phylogenomic matrices of up to 1450 loci, drawn from high-quality transcriptomic libraries and complete genomes. The complete matrix included exemplars of all extant orders of Chelicerata. Analyses of the complete matrix recovered palpigrades as the sister group of the long-branch order Parasitiformes (ticks) with high support. However, sequential deletion of long-branch taxa revealed that the position of palpigrades is prone to topological instability. Phylogenomic subsampling approaches that maximized taxon or dataset completeness recovered palpigrades as the sister group of camel spiders (Solifugae), with modest support. While this relationship is congruent with the location and architecture of the coxal glands, a long-forgotten character system that opens in the pedipalpal segments only in palpigrades and solifuges, we show that nodal support values in concatenated supermatrices can mask high levels of underlying topological conflict in the placement of the enigmatic Palpigradi. </span></span></span></span></span></span></span></span></span></span></span></p>

opencc-zeroNov 2019View details →
dryad28/100

Data from: Taxon-rich phylogenomic analyses resolve the eukaryotic tree of life and reveal the power of subsampling by sites

Most eukaryotic lineages are microbial, and many have only recently been sampled for phylogenetic studies or remain in the 'dark area' of the tree of life where there are no molecular data. To assess relationships among eukaryotic lineages, we perform a taxon-rich phylogenomic analysis including 232 eukaryotes selected to maximize taxonomic diversity and up to 1554 genes chosen as vertically inherited based on their broad distribution among eukaryotes. We also include sequences from 486 bacteria and 84 archaea to assess the impact of endosymbiotic gene transfer (EGT) from plastids and to detect contamination. Overall, our analyses are consistent with other less taxon-rich estimates of the eukaryotic tree of life and we recover strong support for five major clades: Amoebozoa, Excavata (without the genus Malawimonas), Opisthokonta, Archaeplastida and SAR (Stramenopila, Alveolata and Rhizaria). Our analyses also highlight the existence of 'orphan' lineages, lineages that lack robust placement in the eukaryotic tree of life and indicate the possibility of as yet undiscovered diversity. In analyses including bacteria and archaea, we find that ~10% of the 1554 genes, which we choose because they are found in four or five of the five major eukaryotic clades and hence may be more likely to be inherited vertically, appear to have been acquired from cyanobacteria through EGT in photosynthetic lineages. Removing these EGT genes places the green algae as sister to the glaucophytes instead of the red algae, suggesting that unknowingly including of genes of plastid origin, and combining them with genes of nuclear origin, may mislead phylogenetic estimates. Finally, the large size of our dataset allows comparative analyses of subsets of data; alignments built from randomly sampled sites provide greater support, particularly for deep relationships, than do equivalent sized datasets built from randomly sampled genes.

opencc-zeroDec 2013View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record