Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

88

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

88 results for “Sampling bias”

Learn how ShareScore rates datasets ↗
dryad32/100

Data from: Breaking down the lithification bias: the effect of preferential sampling of larger specimens on the estimate of species richness, evenness, and average specimen size

Open the record for dataset details and reuse information.

publicDec 2017View details →
dryad32/100

Data from: The scale-of-choice effect and how estimates of assortative mating in the wild can be biased due to heterogeneous samples

Open the record for dataset details and reuse information.

publicJun 2015View details →
dryad32/100

Sampling biases shape our view of the natural world

Open the record for dataset details and reuse information.

publicJul 2021View details →
dryad28/100

Data from: Sampling schemes and drift can bias admixture proportions inferred by STRUCTURE

<p><span>The interbreeding of individuals coming from genetically differentiated but incompletely isolated populations can lead to the formation of admixed populations, having important implications in ecology and evolution. In this simulation study, we evaluate how individual admixture proportions estimated by the software <span>structure</span> are quantitatively affected by different factors. Using various scenarios of admixture between two diverging populations, we found that unbalanced sampling from parental populations may seriously bias the inferred admixture proportions; moreover, proportionally large samples from the admixed population can also decrease the accuracy and precision of the inferences. As expected, weak differentiation between parental populations and drift after the admixture event strongly increase the biases caused by uneven sampling. We also show that admixture proportions are generally more biased when parental populations unequally contributed to the admixed population. Finally, with few exceptions, using a large number of markers reduces those biases, but using alternative priors for individual ancestry or the uncorrelated allele model only marginally affect the inference of admixture in most situations. We conclude that unbalanced sampling may cause important biases in the admixture proportions estimated by <span>structure</span>, especially when a small number of markers are used, and those biases can be worsened by the effect of drift and unequal genetic contribution of parental populations. Empirical studies should thus be careful with their sampling design and consider historical characteristics when using this software to estimate the ancestry of individuals from admixed populations.</span></p>

opencc-zeroJul 2020View details →
dryad28/100

Data from: How many dinosaur species were there? Fossil bias and true richness estimated using a Poisson sampling model

The fossil record is a rich source of information about biological diversity in the past. However, the fossil record is not only incomplete but has also inherent biases due to geological, physical, chemical and biological factors. Our knowledge of past life is also biased because of differences in academic and amateur interests and sampling efforts. As a result, not all individuals or species that lived in the past are equally likely to be discovered at any point in time or space. To reconstruct temporal dynamics of diversity using the fossil record, biased sampling must be explicitly taken into account. Here, we introduce an approach that uses the variation in the number of times each species is observed in the fossil record to estimate both sampling bias and true richness. We term our technique TRiPS (True Richness estimated using a Poisson Sampling model) and explore its robustness to violation of its assumptions via simulations. We then venture to estimate sampling bias and absolute species richness of dinosaurs in the geological stages of the Mesozoic. Using TRiPS, we estimate that 1936 (1543–2468) species of dinosaurs roamed the Earth during the Mesozoic. We also present improved estimates of species richness trajectories of the three major dinosaur clades: the sauropodomorphs, ornithischians and theropods, casting doubt on the Jurassic–Cretaceous extinction event and demonstrating that all dinosaur groups are subject to considerable sampling bias throughout the Mesozoic.

opencc-zeroDec 2015View details →
zenodo28/100

Supplementary material 3 from: Molloy SW, Davis RA, Dunlop JA, van Etten EJB (2017) Applying surrogate species presences to correct sample bias in species distribution models: a case study using the Pilbara population of the Northern Quoll. Nature Conservation 18: 27-46. https://doi.org/10.3897/natureconservation.18.12235

Weighted mean SDMs for individual algorithms and evaluation statistics (biomod2) :

opencc-by-4.0May 2017View details →
zenodo28/100

Supplementary material 2 from: Molloy SW, Davis RA, Dunlop JA, van Etten EJB (2017) Applying surrogate species presences to correct sample bias in species distribution models: a case study using the Pilbara population of the Northern Quoll. Nature Conservation 18: 27-46. https://doi.org/10.3897/natureconservation.18.12235

Full readout for the MaxEnt northern quoll SDM :

opencc-by-4.0May 2017View details →
zenodo28/100

Supplementary material 1 from: Molloy SW, Davis RA, Dunlop JA, van Etten EJB (2017) Applying surrogate species presences to correct sample bias in species distribution models: a case study using the Pilbara population of the Northern Quoll. Nature Conservation 18: 27-46. https://doi.org/10.3897/natureconservation.18.12235

GIS data sets used in variable assessments and map of Pilbara vegetation systems :

opencc-by-4.0May 2017View details →
dryad28/100

Data from: Origin of tensile strength of a woven sample cut in bias directions

Textile fabrics are highly anisotropic, so that their mechanical properties including strengths are a function of direction. An extreme case is when a woven fabric sample is cut in such a way where the bias angle and hence the tension loading direction is around 45° relative to the principal directions. Then, once loaded, no yarn in the sample is held at both ends, so the yarns have to build up their internal tension entirely via yarn–yarn friction at the interlacing points. The overall fabric strength in such a sample is a result of contributions from the yarns being pulled out and those broken during the process, and thus becomes a function of the bias direction angle θ, sample width W and length L, along with other factors known to affect fabric strength tested in principal directions. Furthermore, in such a bias sample when the major parameters, e.g. the sample width W, change, not only the resultant strengths differ, but also the strength generating mechanisms (or failure types) vary. This is an interesting problem and is analysed in this study. More specifically, the issues examined in this paper include the exact mechanisms and details of how each interlacing point imparts the frictional constraint for a yarn to acquire tension to the level of its strength when both yarn ends were not actively held by the testing grips; the theoretical expression of the critical yarn length for a yarn to be able to break rather than be pulled out, as a function of the related factors; and the general relations between the tensile strength of such a bias sample and its structural properties. At the end, theoretical predictions are compared with our experimental data.

opencc-zeroDec 2014View details →
dryad28/100

Data from: RADseq underestimates diversity and introduces genealogical biases due to nonrandom haplotype sampling

Reduced representation genome-sequencing approaches based on restriction digestion are enabling large-scale marker generation and facilitating genomic studies in a wide range of model and nonmodel systems. However, sampling chromosomes based on restriction digestion may introduce a bias in allele frequency estimation due to polymorphisms in restriction sites. To explore the effects of this nonrandom sampling and its sensitivity to different evolutionary parameters, we developed a coalescent-simulation framework to mimic the biased recovery of chromosomes in restriction-based short-read sequencing experiments (RADseq). We analysed simulated DNA sequence datasets and compared known values from simulations with those that would be estimated using a RADseq approach from the same samples. We compare these 'true' and 'estimated' values of commonly used summary statistics, π, θw, Tajima's D and FST. We show that loci with missing haplotypes have estimated summary statistic values that can deviate dramatically from true values and are also enriched for particular genealogical histories. These biases are sensitive to nonequilibrium demography, such as bottlenecks and population expansion. In silico digests with 102 completely sequenced Drosophila melanogaster genomes yielded results similar to our findings from coalescent simulations. Though the potential of RADseq for marker discovery and trait mapping in nonmodel systems remains undisputed, our results urge caution when applying this technique to make population genetic inferences.

opencc-zeroDec 2012View details →
dryad28/100

Current nest box designs may not be optimal for the larger forest dormice; pre-hibernation increase in body mass might lead to sampling bias in ecological data

<p>Biologists commonly use nest boxes to study small arboreal mammals, including forest dormouse (Dryomys nitedula). Hibernating dormouse species often experience pronounced seasonal variations in body mass, which might lead to sampling biases if it is not taken into account when designing nest boxes. In our study of forest dormouse, we noticed that the entrance hole of nest boxes had been gnawed on. We hypothesized that this behavior was exhibited by individual dormice who had higher body mass and, therefore, were unable to pass through the entrance holes. To test our hypothesis, we categorized individual dormice present inside nest boxes based on their body mass; then compared the seasonal body mass dynamics with the timing of the gnawing behavior. We also compared nest box occupancy by forest dormouse before and after the gnawing behavior. Interestingly, we found that the gnawing behavior was displayed exclusively when part of the dormouse population increased considerably in body mass, which supports our hypothesis. Additionally, nest box occupancy decreased significantly from 20% before to 4.6% after the gnawing behavior. We suggest that researchers use nest boxes with entrance holes larger than 4 cm in future studies of forest dormouse to prevent the possible exclusion of the conspecifics that have higher body mass before hibernation. This type of sampling bias can probably happen in studies of other species, such as fat dormouse, that similarly show pronounced seasonal variations in body mass. We recommend that biologists consider the seasonal body mass dynamics of the target species when designing nest boxes to minimize bias in ecological data and improve management actions.</p>

opencc-zeroDec 2020View details →
dryad28/100

Data from: Origin of tensile strength of a woven sample cut in bias directions

Open the record for dataset details and reuse information.

publicApr 2015View details →
dryad28/100

Data from: The fossil record of ichthyosaurs, completeness metrics and sampling biases

Open the record for dataset details and reuse information.

publicFeb 2016View details →
dryad28/100

Data from: How many dinosaur species were there? Fossil bias and true richness estimated using a Poisson sampling model

Open the record for dataset details and reuse information.

publicMar 2017View details →
dryad28/100

Data from: Sampling bias and the fossil record of planktonic foraminifera on land and in the deep sea

Open the record for dataset details and reuse information.

publicApr 2012View details →
dryad28/100

Data from: RADseq underestimates diversity and introduces genealogical biases due to nonrandom haplotype sampling

Open the record for dataset details and reuse information.

publicFeb 2013View details →
dryad28/100

Spatial sampling bias and model complexity in stream-based species distribution models: a case study of Paddlefish (Polyodon spathula) in the Arkansas River basin, U.S.A.

Open the record for dataset details and reuse information.

publicNov 2019View details →
dryad28/100

Data from: Sampling schemes and drift can bias admixture proportions inferred by STRUCTURE

Open the record for dataset details and reuse information.

publicJul 2020View details →
dryad28/100

Data from: Approach-induced biases in human information sampling

Open the record for dataset details and reuse information.

publicJan 2017View details →
dryad28/100

Current nest box designs may not be optimal for the larger forest dormice; pre-hibernation increase in body mass might lead to sampling bias in ecological data

Open the record for dataset details and reuse information.

publicDec 2022View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record