Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

23

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

23 results for “Sampling coverage”

Learn how ShareScore rates datasets ↗
edi48/100

GIS30 GIS Coverages Defining Sample Locations for Abiotic Datasets on Konza Prairie (1972-present)

These data show sample locations for various abiotic data collected on Konza Prairie (rain gauges, soil moisture, and stream data). Included in these data are the locations for 12 rain gauges (GIS300) on Konza Prairie. The Konza headquarters weather station formerly consisted of two gauges which were operated year-round. The Konza headquarters weather station currently consists of one Otto-Pluvio2 gauge which is operated year-round. The remaining Konza-operated gauges run from April 1 to November 1. These data are to be used in conjunction with the APT01 (precipitation) dataset. GIS305 defines the locations where measurements of soil moisture (%volume) are taken on Konza Prairie. These data are to be used in conjunction with the ASM01 (soil moisture) dataset. GIS309 defines the locations within watershed N4D of soil sampler nests. In Jan 2020, we separated the original GIS310 file 'Wells in N4D' into GIS310 'Wells in N4D' and GIS309 'Soil Sampler Nests'. Prior to then, soil sampler nests and wells were combined in GIS310. GIS310 defines the locations within watershed N4D where samples are taken for analyzing the belowground water chemistry of the watershed. These data are to be used in conjunction with the AGW01 dataset. GIS311 defines the locations of 14 wells at two sites along Kings Creek. Depth and nutrient content of groundwater is measured at these sites. These data are to be used in conjunction with the AGW02 dataset. GIS315 defines the locations of stream sampling stations within multiple Konza watersheds. These data are to be used in conjunction with the NWC, ASS, ASD, and ASW datasets. GIS320 defines the locations of the rainfall collectors used to collect the samples analyzed as a part of the National Atmospheric Deposition Program. These data are to be used in conjunction with the ANA01 dataset. These data are available to download as zipped shapefiles (.zip), compressed Google Earth KML layers (.kmz).

openCC0Jan 2023View details →
edi48/100

GIS35 GIS Coverages Defining Sample Locations for Belowground Datasets on Konza Prairie (1982-present)

These data show the locations of research conducted at the below ground plots near Konza Headquarters. Record type 1 (GIS350) describes the 64 belowground plots receiving a variety of nutrient, burn, and mowing treatments. Data for BMS01, BMS02, and BNS01 are collected on these plots. Record type 6 (GIS355) describes the locations of the Micro-Rhizotrons. Two spatial datasets lie on the belowground plots, but are classified separately. These are the Lysimeters on belowground plots (GIS455) and Aboveground biomass on belowground plots (GIS505) datasets. GIS505 may be used alongside the BGPVC dataset, because it shares sample locations with PBB01. These data are available to download as zipped shapefiles (.zip), compressed Google Earth KML layers (.kmz), and associated EML metadata (.xml).

openCC0Jan 2023View details →
edi48/100

GIS40 GIS Coverages Defining the Sample Locations of Konza Consumer Data (1982-present)

These data show the sampling locations for the consumer datasets at Konza Prairie. GIS400 defines the starting points for sweep samples of grasshoppers across Konza Prairie. These data may be used in conjunction with the sweep sample datasets (CGR02). GIS401 defines the starting points for sweep samples of grasshoppers across Konza Prairie, focusing on grazing impact. These data may be used in conjunction with the sweep sample datasets (CGR02Z). GIS405 defines the trap locations for small mammal sampling across Konza Prairie. These data may be used in conjunction with CSM0X. GIS 406 defines the locations of small mammal host parasite sampling at Konza Prairie. These data may be used in conjunction with CSM08. GIS410 defines the stream stretches for fish sampling across Konza Prairie. These data may be used in conjunction with CFC01. These data are available to download as zipped shapefiles (.zip), compressed Google Earth KML layers (.kmz), and associated EML metadata (.xml).

openCC0Jan 2023View details →
edi48/100

GIS45 GIS Coverages Defining the Konza Nutrient Data Sample Locations (1982-present)

These data show the sample locations for soil bulk density and chemical characteristics along LTER vegetation plots. This dataset contains the transect lines (GIS450) and sample locations(GIS451) at which the soil cores are sampled. These data may be used in conjunction with the Soil Chemistry and Bulk Density (NSC01) datasets. GIS455 contains the locations of the lysimeters used to measure soil water chemistry on the belowground plots. These data may be used in conjunction with the NBS01 dataset. GIS460 contains the locations of the bulk precipitation collectors on Konza Prairie. These data may be used in conjunction with the NBP01 dataset. These data are available to download as zipped shapefiles (.zip), compressed Google Earth KML layers (.kmz), and associated EML metadata (.xml).

openCC0Jan 2023View details →
edi48/100

GIS50 Coverages Defining the Konza Producer Data Sample Locations (1982-present)

These data show the sample locations for datasets pertaining to primary production at Konza Prairie. These data reference various treatments across Konza including varying burn frequencies, belowground plots, patch burn, exclosures, etc.Record type one (GIS500) contains sample locations for estimated standing crop biomass in various burning-grazing treatments (PABXX). Record type six (GIS505) contains sample locations for peak foliage biomass measured at the belowground plot experiments (PBBXX).Record type 11 and 12 contain the transect (GIS510) and plot (GIS511) locations for plots in the patch-burn experiments (PBGXX). Record type 16 (GIS515) contains the locations of exclosures used to sample primary productivity in bison grazed watersheds (PEB01). Record type 21 (GIS520) contains the locations of exclosures used to sample primary productivity in cattle grazed watersheds (PEB01X). Record type 26 (GIS525) contains the locations of sample sites for litterfall (PGLXX) collectors in the gallery forest. Record type 31 (GIS530) contains species composition transects, and (GIS531) provides locations for species transect plots in the patch-burn experiments for Konza Prairie. These data may be used in conjunction species composition (PVC01 and PVC02), primary production in grazing exclosures (PEB01, PEB01_X), soil chemistry and bulk density (NSC01) and primary production (PAB01). These data are available to download as zipped shapefiles (.zip), compressed Google Earth KML layers (.kmz).

openCC0Jan 2023View details →
edi48/100

GIS60 GIS Coverages Defining Other Konza Sample and Research Areas (1982-present)

These data show locations of samples and research areas at Konza that do not fit under our standard classifications. GIS 600 contains the locations of the Hulbert plots on Konza Prairie. GIS605 contains locations for rainfall shelters, ramps, experimental streams, restoration plots, the weather station, grasshopper cages, the climate extremes project. Currently no associated LTER datasets exist for these locations. GIS 610 provides a record of the historic Konza gridded location system. Older datasets may reference these locations with a column letter and row number. GIS615 contains the location for the Clean Air Status and Trends Network (CASTNET) site on Konza Prairie. For more information, visit the following link: http://www.epa.gov/castnet/javaweb/site_pages/KNZ184.html. GIS620 contains the location for the USGS gauging station. These data may be used in conjunction with the Stream Discharge for Kings Creek Measured at USGS Gauging Station (ASD01) dataset. For more information, visit the following link: http://waterdata.usgs.gov/nwis/nwisman/?site_no=06879650. GIS630) and GIS635 contain the location and treatment information for two bison grant grazing experiments. Currently, no associated LTER datasets exist for these data. These data are available to download as zipped shapefiles (.zip), compressed Google Earth KML layers (.kmz).

openCC0Jan 2023View details →
zenodo44/100

Data for: "Comprehensive sampling of coverage effects in catalysis by leveraging generalization in neural network models"

<p>This repository contains the raw data to reproduce the paper: "Comprehensive sampling of coverage effects in catalysis by leveraging generalization in neural network models". Within the .tar.gz file, you will find the directory structure described above.</p> <h2>Directory Structure</h2> <h3>`data`</h3> <p>Contains the data to reproduce all figures in the manuscript. Used primarily by the Jupyter Notebooks that plot the data from the paper.</p> <h3>`eval`</h3> <p>Contains the predicted energies according to a MACE model for the following systems and facets:<br>- covsplit (100, 111, 211, 331, 410, 711): The NN model is trained on low-coverage structures and tested on high-coverage structures for a single facet<br>- evencov (100, 111, 211, 331, 410, 711): The NN is trained on even coverages and tested on odd coverages for a single facet<br>- facet (100, 111, 211, 331, 410, 711): the NN is trained on the facet indicated by the folder name (e.g., facet-100 means that the model was trained on Cu(100)) and tested on all of the other facets.<br>- full: the model was trained on all facets and all coverages<br>- slopes (various versions and configurations): the models were trained with different body-order correlation (v) for the Cu(711) facet and tested only on the Cu(711) facet<br>- Rh111: Energies for the Rh(111) + CHOH + CO systems.</p> <h3>`mcmc`</h3> <p>Contains the data for MCMC (Markov Chain Monte Carlo) evaluations for two systems: Cu and Rh<br>- copper-mcmc-public.tar.gz<br>- rhodium-mcmc-public.tar.gz</p> <h3>`models`</h3> <p>Contains the weights and parameters of the best-performing MACE models trained in this work, as selected by the validation loss:</p> <p>File formats: `.model` and `_swa.model` relate to the first-stage of training and the second-stage of training.</p> <h3>`pyscripts`</h3> <p>Python scripts to perform the MCMC sampling given the custom configuration file `sample_cfg.json`.</p> <h3>`scripts`</h3> <p>Shell scripts for evaluation and training the MACE models, along with the hyperparameters used in doing so.</p> <p>- Evaluation scripts (eval-*.sh)<br>- Training scripts (train-*.sh)</p> <h3>`train`</h3> <p>Training, validation, and testing data for all Cu and Rh facets in this work, according to the naming scheme described above.</p> <p>- Rh111<br>- covsplit<br>- evencov<br>- facet<br>- full<br>- slopes</p>

opencc-by-4.0Sep 2024View details →
zenodo44/100

Synthetic Escherichia coli mixture samples with variable coverage

<p>This dataset contains the synthetic mixture samples and reference sequences - as well as the appropriate metadata - that were originally used in the 2021 revision of the mSWEEP manuscript.<br> <br> There are 87 samples in total, each containing 100bp paired-end Illumina sequencing reads from 10 different&nbsp;<em>Escherichia coli&nbsp;</em>strains from 10 different lineages. The number of reads is set so that the sequencing coverage of the individual strains varies between 50x and 0.10x and sums up to 100x.</p>

opencc-by-4.0Sep 2021View details →
zenodo44/100

Variation of and associations with the depth and evenness of sequencing coverage in a sample of archived plastid genomes

<p>Depth and evenness of sequencing coverage are considered potential indicators of genome assembly quality. In plastid genomics, where new data generation has outpaced the development of suitable assembly quality indicators, these coverage metrics could offer insights into the quality of plastomes of different sizes, structures, or taxonomic origins. However, the typical variation of sequencing depth and evenness among archived plastid genomes, their variability between plastome partitions, and any association with methodological factors have yet to be evaluated. This study explores the variation of sequencing depth and evenness across a sample of publicly accessible plastid genomes and their potential associations with plastome structure, assembly accuracy, and the methodological provenance of the genome data using statistical tests. Our results indicate significant differences in sequencing depth across the four structural partitions as well as between the coding and non-coding sections of the genomes, a significant correlation between sequencing evenness and the number of ambiguous nucleotides, and a significant difference in sequencing evenness between several DNA sequencing platforms. These findings highlight that many publicly accessible plastid genomes are based on sequence data with highly variable sequencing depth and evenness and that this variation is influenced, at least partially, by genome structure and methodological factors.</p>

opencc-by-4.0May 2024View details →
zenodo40/100

Global Pasture Watch - Grassland sampling design derived by Feature Space Coverage Sampling (FSCS) at 1-km spatial resolution

<p>Sampling design used in the production of the <strong>global maps of grassland dynamics 2000&ndash;2022 at 30 m spatial resolution</strong> in the scope of the Global Pasture Wath initiative. The sampling desing was based in Feature Space Coverage Sampling and resulted in 10,000 sample tiles (1x1 km) distributed across the World, which were visual interpreted in Very-High Resolution imagery thorugh the QGIS plugin&nbsp;<a href="https://plugins.qgis.org/plugins/qgis-fgi-plugin/">QGIS Fast Grid Inspection</a>.</p> <p>FSCS steps include:</p> <ul> <li>Short vegetation mask that includes all pixels mapped as mosaic, shrubland, grassland, and sparse vegetation in at least one year from 1993 to 2021 according to <a href="https://www.esa-landcover-cci.org/">ESA/CCI global land cover</a> (<code>gpw_short.veg.mask_esacci.lc_p_1km_s_19920101_20201231_go_epsg.3857_v1.tif</code>),</li> <li>87 input raster layers (including vegetation indices, terrain, land temperature, climate and water variable),</li> <li>Principal Components Analysis (PCA) using all input layers,</li> <li>Selection of the 10 first components (explaining 75% of variance),</li> <li>&nbsp;K-Means with 10,000 clusters (targeted number of samples - &nbsp; &nbsp;&nbsp;<br><code>gpw_grassland_fscs.kmeans.cluster_c_1km_20000101_20221231_go_epsg.3857_v1.tif</code>)</li> <li>Calculation of euclidean distance (in the principal component space) of all 1-km pixels to the centre of each cluster,</li> <li>Selection of the pixel with the shortest distance for each cluster,</li> <li>Conversion of the selected pixels into sample tiles ()</li> </ul> <p>The file&nbsp;<code>gpw_grassland_fscs_tile.samples_1km_20000101_20221231_go_epsg.3857_v1.gpkg</code> provides the sample tiles and include the follow collumns:</p> <ul> <li><strong>X</strong>: Latitude in Web Mercator projection (EPSG:3857),</li> <li><strong>Y</strong>: Longitude in Web Mercator projection (EPSG:3857),</li> <li><strong>cluster_id</strong>: K-Means output ranging from 0&mdash;9999,</li> <li><strong>cluster_distance</strong>: Distance from the selected sample to the centre of the cluster,</li> <li><strong>cluster_size</strong>: Number o 1-km pixels inside the K-Means cluster, estimated using Web Mercator projection (<a href="https://epsg.io/3857">EPSG:3857</a>)</li> <li><strong>cluster_size_equal_area</strong>: Number o 1-km pixels inside the K-Means cluster, estimated using Goode Homolosine Land projection (<a href="https://epsg.io/54052">ESRI:54052</a>)</li> <li><strong>cluster_size_corr</strong>: Correction factor to adjust the area distortion due to Web Mercator projection, estimated by the difference in normalized propotional values of cluster_size and cluster_size_equal_area.</li> <li><strong>rf_n_pred</strong>: Number of pixels predicted by a RF model trained to estimate probability to select the pixel closer to the centre of the KMeans cluster. The RF models were trained individually per each cluster using the 10 first components derived by PCA (<code>gpw_comps_fscs.pca_m_1km_20000101_20221231_go_epsg.3857_v1.tar.gz</code>).</li> <li><strong>rf_samp_prob</strong>: Sampling probability based on RF model (<em>rf_n_pred / cluster_size</em>)</li> <li><strong>rf_samp_wei</strong>: Sampling weight estimated in Web Mercator projection.</li> <li><strong>rf_samp_wei_coor</strong>: Corrected sampling weight estimated in Goode Homolosine Land projection.</li> </ul> <h3>Related resources</h3> <ul> <li><strong>Maps of dominant grassland:</strong><br><a href="https://zenodo.org/records/13890400">2000-2002</a> <a href="https://zenodo.org/records/13890402">2003-2005</a> <a href="https://zenodo.org/records/13890404">2006-2008</a> <a href="https://zenodo.org/records/13890408">2009-2011</a> <a href="https://zenodo.org/records/13890410">2012-2014</a> <a href="https://zenodo.org/records/13890412">2015-2017</a> <a href="https://zenodo.org/records/13890414">2018-2020</a> <a href="https://zenodo.org/records/13890416">2021-2022</a></li> <li><strong>Probability maps of cultivated grassland:</strong><br><a href="https://zenodo.org/records/13890401/files/ggc-30m.csv?download=1">2000-2022 (All URLs)</a></li> <li><strong>Probability maps of natural/semi-natural grassland:</strong><br><a href="https://zenodo.org/records/13890401/files/ggc-30m.csv?download=1">2000-2022 (All URLs)</a></li> <li><strong>Grassland reference samples based on VHR imagery (2000&ndash;2022):</strong><br><a href="https://doi.org/10.5281/zenodo.11281157">GeoPackage files</a></li> <li><strong>Global machine learning models (Random Forest):</strong><br><a href="https://doi.org/10.5281/zenodo.13952806">Parquet and joblib python files</a></li> <li><strong>Reference sampling design derived by FSCV:</strong><br><a href="https://doi.org/10.5281/zenodo.11391517">GeoPackage and raster files</a></li> <li><strong>Harmonized reference samples based on existing LULC dataset:</strong><br><a href="https://doi.org/10.5281/zenodo.13951976">GeoPackage and raster files</a></li> <li><strong>Source code for reproducibility:<br></strong><a href="https://doi.org/10.5281/zenodo.13952867">GitHub release</a><strong><br></strong></li> <li><strong>Mapping feedback tool:</strong><br><a href="https://geo-wiki.org">GeoWiki</a></li> <li><strong>Data catalogues:</strong><br><a href="https://stac.openlandmap.org/gpw_ggc-30m/collection.json?.language=en">OpenLandMap STAC</a> <a href="https://global-pasture-watch.projects.earthengine.app/view/ggc-30m">Google Earth Engine</a></li> </ul> <h3>Support</h3> <p>For questions of bugs/inconsistencies related to the dataset raise a GitHub issue in&nbsp;<a href="https://github.com/wri/global-pasture-watch">https://github.com/wri/global-pasture-watch</a></p>

opencc-by-4.0Nov 2024View details →
zenodo40/100

somalier files for thousand genomes high-coverage VCF for 2504 samples

<p>somalier files extracted from the thousand genomes VCF to be used for ancestry prediction with somalier.</p>

opencc-by-4.0Sep 2019View details →
zenodo40/100

Fig. 4. Goods coverage index for each sample using the OTUs. Retained 330,000 in Bacterial community of ticks (Acari: Ixodidae) and mammals from Arauca, Colombian Orinoquia

Fig. 4. Goods coverage index for each sample using the OTUs. Retained 330,000 (49.90%) features in 15 (93.75%) samples at the specified sampling depth (22,000).

opencc-by-4.0Aug 2024View details →
zenodo36/100

VaMoS 2022 Submission 12 - Continuous T-Wise Sampling: Increasing Coverage over Time

<p>Evaluation results for the VaMoS 2022 submission 12. For more details regarding the dataset please refer to the included readme.md file.</p>

opencc-by-4.0Nov 2021View details →
zenodo36/100

rarefaction and extrapolation of sample coverage based on sample-based abundance data

<p>#R code.txt is the R code for plotting figures and constructing Tables</p> <p>#bciabun1010.txt is the species_by_plot matrix that BCI forest Plot is divided by a plot with size 10m*<em>10m</em></p> <p><em>#bciabun2020.txt is the species_by_plot matrix that BCI forest Plot is divided by a plot with size 20m*</em>20m</p> <p><em>#bciabun5050.txt is the species_by_plot matrix that BCI forest Plot is divided by a plot with size 50m*5</em>0m</p> <p>#fus10.txt is the species_by_plot matrix that Fushan forest Plot is divided by a plot with size 10m*<em>10m</em></p> <p><em>#fus20.txt is the species_by_plot matrix that Fushan forest Plot is divided by a plot with size 20m*</em>20m</p> <p>#fus50.txt is the species_by_plot matrix that Fushan forest Plot is divided by a plot with size 50m*<em>50m</em></p> <p><em>#lhc10.txt is the species_by_plot matrix that Lianhuachi forest Plot is divided by a plot with size 10m*</em>10m</p> <p>#lhc20.txt is the species_by_plot matrix that Lianhuachi forest Plot is divided by a plot with size 20m*20m</p> <p>#lhc50.txt is the species_by_plot matrix that Lianhuachi forest Plot is divided by a plot with size 50m*50m</p>

opencc-by-4.0Aug 2022View details →
dryad36/100

Data from: Standardising fossil disparity metrics using sample coverage

Open the record for dataset details and reuse information.

publicOct 2024View details →
zenodo32/100

Code Coverage Dataset Sample

<p>This is a code coverage dataset sample for submission the The Web Conference 2023. It is submitted anonymously to preserve the double-blind review process.</p>

opencc-by-4.0Oct 2022View details →
dryad28/100

Data from: Practical low-coverage genomewide sequencing of hundreds of individually barcoded samples for population and evolutionary genomics in nonmodel species

Today most population genomic studies of nonmodel organisms either sequence a subset of the genome deeply in each individual or sequence pools of unlabelled individuals. With a step-by-step workflow, we illustrate how low-coverage whole-genome sequencing of hundreds of individually barcoded samples is now a practical alternative strategy for obtaining genomewide data on a population scale. We used a highly efficient protocol to generate high-quality libraries for ~6.5 USD from each of 876 Atlantic silversides (a teleost fish with a genome size ~730 Mb) that we sequenced to 1–4× genome coverage. In the absence of a reference genome, we developed a bioinformatic pipeline for mapping the genomic reads to a de novo assembled reference transcriptome. This provides an 'in silico' method for exome capture that avoids the complexities and expenses of using wet chemistry for target isolation. Using novel tools for analysis of low-coverage data, we extracted population allele frequencies, individual genotype likelihoods and polymorphism data for 2 504 335 SNPs across the exome for the 876 fish. To illustrate the use of the resulting data, we present a preliminary analysis of geographical patterns in the exome data and a comparison of complete mitochondrial genome sequences for each individual (constructed from the low-coverage data) that show population colonization patterns along the US east coast. With a total cost per sample of less than 50 USD (including sequencing) and ability to prepare 96 libraries in only 5 h, our approach adds a viable new option to the population genomics toolbox.

opencc-zeroDec 2015View details →
dryad28/100

Data from: Practical low-coverage genomewide sequencing of hundreds of individually barcoded samples for population and evolutionary genomics in nonmodel species

Open the record for dataset details and reuse information.

publicAug 2016View details →
dryad28/100

Data from: Sequencing degraded DNA from non-destructively sampled museum specimens for RAD-tagging and low-coverage shotgun phylogenetics

Open the record for dataset details and reuse information.

publicJul 2015View details →
geo24/100

Chromosomal microarray data for validation of copy-number variants detection from a low-coverage whole-genome sequencing approach in clinical samples

GEO Series GSE73191. Homo sapiens. 72 samples. Type: Genome variation profiling by array; Genome variation profiling by SNP array; SNP genotyping by SNP array.

openGEO-OpenOct 2015View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record