Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
88
datasets available to search
ShareScore release 0.9.0
Dataset results
88 results for “Sampling bias”
Data from: Breaking down the lithification bias: the effect of preferential sampling of larger specimens on the estimate of species richness, evenness, and average specimen size
Open the record for dataset details and reuse information.
Data from: The scale-of-choice effect and how estimates of assortative mating in the wild can be biased due to heterogeneous samples
Open the record for dataset details and reuse information.
Sampling biases shape our view of the natural world
Open the record for dataset details and reuse information.
Data from: Sampling schemes and drift can bias admixture proportions inferred by STRUCTURE
<p><span>The interbreeding of individuals coming from genetically differentiated but incompletely isolated populations can lead to the formation of admixed populations, having important implications in ecology and evolution. In this simulation study, we evaluate how individual admixture proportions estimated by the software <span>structure</span> are quantitatively affected by different factors. Using various scenarios of admixture between two diverging populations, we found that unbalanced sampling from parental populations may seriously bias the inferred admixture proportions; moreover, proportionally large samples from the admixed population can also decrease the accuracy and precision of the inferences. As expected, weak differentiation between parental populations and drift after the admixture event strongly increase the biases caused by uneven sampling. We also show that admixture proportions are generally more biased when parental populations unequally contributed to the admixed population. Finally, with few exceptions, using a large number of markers reduces those biases, but using alternative priors for individual ancestry or the uncorrelated allele model only marginally affect the inference of admixture in most situations. We conclude that unbalanced sampling may cause important biases in the admixture proportions estimated by <span>structure</span>, especially when a small number of markers are used, and those biases can be worsened by the effect of drift and unequal genetic contribution of parental populations. Empirical studies should thus be careful with their sampling design and consider historical characteristics when using this software to estimate the ancestry of individuals from admixed populations.</span></p>
Data from: How many dinosaur species were there? Fossil bias and true richness estimated using a Poisson sampling model
The fossil record is a rich source of information about biological diversity in the past. However, the fossil record is not only incomplete but has also inherent biases due to geological, physical, chemical and biological factors. Our knowledge of past life is also biased because of differences in academic and amateur interests and sampling efforts. As a result, not all individuals or species that lived in the past are equally likely to be discovered at any point in time or space. To reconstruct temporal dynamics of diversity using the fossil record, biased sampling must be explicitly taken into account. Here, we introduce an approach that uses the variation in the number of times each species is observed in the fossil record to estimate both sampling bias and true richness. We term our technique TRiPS (True Richness estimated using a Poisson Sampling model) and explore its robustness to violation of its assumptions via simulations. We then venture to estimate sampling bias and absolute species richness of dinosaurs in the geological stages of the Mesozoic. Using TRiPS, we estimate that 1936 (1543–2468) species of dinosaurs roamed the Earth during the Mesozoic. We also present improved estimates of species richness trajectories of the three major dinosaur clades: the sauropodomorphs, ornithischians and theropods, casting doubt on the Jurassic–Cretaceous extinction event and demonstrating that all dinosaur groups are subject to considerable sampling bias throughout the Mesozoic.
Supplementary material 3 from: Molloy SW, Davis RA, Dunlop JA, van Etten EJB (2017) Applying surrogate species presences to correct sample bias in species distribution models: a case study using the Pilbara population of the Northern Quoll. Nature Conservation 18: 27-46. https://doi.org/10.3897/natureconservation.18.12235
Weighted mean SDMs for individual algorithms and evaluation statistics (biomod2) :
Supplementary material 2 from: Molloy SW, Davis RA, Dunlop JA, van Etten EJB (2017) Applying surrogate species presences to correct sample bias in species distribution models: a case study using the Pilbara population of the Northern Quoll. Nature Conservation 18: 27-46. https://doi.org/10.3897/natureconservation.18.12235
Full readout for the MaxEnt northern quoll SDM :
Supplementary material 1 from: Molloy SW, Davis RA, Dunlop JA, van Etten EJB (2017) Applying surrogate species presences to correct sample bias in species distribution models: a case study using the Pilbara population of the Northern Quoll. Nature Conservation 18: 27-46. https://doi.org/10.3897/natureconservation.18.12235
GIS data sets used in variable assessments and map of Pilbara vegetation systems :
Data from: Origin of tensile strength of a woven sample cut in bias directions
Textile fabrics are highly anisotropic, so that their mechanical properties including strengths are a function of direction. An extreme case is when a woven fabric sample is cut in such a way where the bias angle and hence the tension loading direction is around 45° relative to the principal directions. Then, once loaded, no yarn in the sample is held at both ends, so the yarns have to build up their internal tension entirely via yarn–yarn friction at the interlacing points. The overall fabric strength in such a sample is a result of contributions from the yarns being pulled out and those broken during the process, and thus becomes a function of the bias direction angle θ, sample width W and length L, along with other factors known to affect fabric strength tested in principal directions. Furthermore, in such a bias sample when the major parameters, e.g. the sample width W, change, not only the resultant strengths differ, but also the strength generating mechanisms (or failure types) vary. This is an interesting problem and is analysed in this study. More specifically, the issues examined in this paper include the exact mechanisms and details of how each interlacing point imparts the frictional constraint for a yarn to acquire tension to the level of its strength when both yarn ends were not actively held by the testing grips; the theoretical expression of the critical yarn length for a yarn to be able to break rather than be pulled out, as a function of the related factors; and the general relations between the tensile strength of such a bias sample and its structural properties. At the end, theoretical predictions are compared with our experimental data.
Data from: RADseq underestimates diversity and introduces genealogical biases due to nonrandom haplotype sampling
Reduced representation genome-sequencing approaches based on restriction digestion are enabling large-scale marker generation and facilitating genomic studies in a wide range of model and nonmodel systems. However, sampling chromosomes based on restriction digestion may introduce a bias in allele frequency estimation due to polymorphisms in restriction sites. To explore the effects of this nonrandom sampling and its sensitivity to different evolutionary parameters, we developed a coalescent-simulation framework to mimic the biased recovery of chromosomes in restriction-based short-read sequencing experiments (RADseq). We analysed simulated DNA sequence datasets and compared known values from simulations with those that would be estimated using a RADseq approach from the same samples. We compare these 'true' and 'estimated' values of commonly used summary statistics, π, θw, Tajima's D and FST. We show that loci with missing haplotypes have estimated summary statistic values that can deviate dramatically from true values and are also enriched for particular genealogical histories. These biases are sensitive to nonequilibrium demography, such as bottlenecks and population expansion. In silico digests with 102 completely sequenced Drosophila melanogaster genomes yielded results similar to our findings from coalescent simulations. Though the potential of RADseq for marker discovery and trait mapping in nonmodel systems remains undisputed, our results urge caution when applying this technique to make population genetic inferences.
Current nest box designs may not be optimal for the larger forest dormice; pre-hibernation increase in body mass might lead to sampling bias in ecological data
<p>Biologists commonly use nest boxes to study small arboreal mammals, including forest dormouse (Dryomys nitedula). Hibernating dormouse species often experience pronounced seasonal variations in body mass, which might lead to sampling biases if it is not taken into account when designing nest boxes. In our study of forest dormouse, we noticed that the entrance hole of nest boxes had been gnawed on. We hypothesized that this behavior was exhibited by individual dormice who had higher body mass and, therefore, were unable to pass through the entrance holes. To test our hypothesis, we categorized individual dormice present inside nest boxes based on their body mass; then compared the seasonal body mass dynamics with the timing of the gnawing behavior. We also compared nest box occupancy by forest dormouse before and after the gnawing behavior. Interestingly, we found that the gnawing behavior was displayed exclusively when part of the dormouse population increased considerably in body mass, which supports our hypothesis. Additionally, nest box occupancy decreased significantly from 20% before to 4.6% after the gnawing behavior. We suggest that researchers use nest boxes with entrance holes larger than 4 cm in future studies of forest dormouse to prevent the possible exclusion of the conspecifics that have higher body mass before hibernation. This type of sampling bias can probably happen in studies of other species, such as fat dormouse, that similarly show pronounced seasonal variations in body mass. We recommend that biologists consider the seasonal body mass dynamics of the target species when designing nest boxes to minimize bias in ecological data and improve management actions.</p>
Data from: Origin of tensile strength of a woven sample cut in bias directions
Open the record for dataset details and reuse information.
Data from: The fossil record of ichthyosaurs, completeness metrics and sampling biases
Open the record for dataset details and reuse information.
Data from: How many dinosaur species were there? Fossil bias and true richness estimated using a Poisson sampling model
Open the record for dataset details and reuse information.
Data from: Sampling bias and the fossil record of planktonic foraminifera on land and in the deep sea
Open the record for dataset details and reuse information.
Data from: RADseq underestimates diversity and introduces genealogical biases due to nonrandom haplotype sampling
Open the record for dataset details and reuse information.
Spatial sampling bias and model complexity in stream-based species distribution models: a case study of Paddlefish (Polyodon spathula) in the Arkansas River basin, U.S.A.
Open the record for dataset details and reuse information.
Data from: Sampling schemes and drift can bias admixture proportions inferred by STRUCTURE
Open the record for dataset details and reuse information.
Data from: Approach-induced biases in human information sampling
Open the record for dataset details and reuse information.
Current nest box designs may not be optimal for the larger forest dormice; pre-hibernation increase in body mass might lead to sampling bias in ecological data
Open the record for dataset details and reuse information.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.