Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
88
datasets available to search
ShareScore release 0.9.0
Dataset results
88 results for “Sampling bias”
Sampling bias exaggerates a textbook example of a trophic cascade
Open the record for dataset details and reuse information.
Data from: Integrated species distribution models to account for sampling biases and improve range wide occurrence predictions
Open the record for dataset details and reuse information.
Piecewise continuous sampling: a method for minimizing bias and sampling effort for estimated metrics of animal behavior
Open the record for dataset details and reuse information.
LAGOS Lake nutrient, carbon and chlororphyll data to evaluate biases in lake water quality sampling practices in a 17-state region of the US
This dataset includes data for eight major limnological variables in LAGOS-NE_LIMNO v. 1.087.1 that were used to evaluate biases in lake water quality sampling and implications for macroscale research (Stanley et al. In Revision, Limnology and Oceanography). Most observations came from LAGOS-NE_LIMNO v. 1.087.1, an integrated database of lake ecosystems (Soranno et al. 2015, Soranno et al. 2017) but were supplemented with additional data from the State of New Hampshire. LAGOS-NE contains information on lakes great than or equal to 1 ha (originally derived from the U.S. Geological Survey's 2013 National Hydrography Dataset) for a 17-state region of the U.S., and a subset of the lakes has observational data on lake chemistry and productivity. Approximately 87 different sources of data were compiled for the LAGOS-NELIMNO v. 1.087.1 dataset and were mostly generated by government agencies (state, federal, tribal) and universities. In this analysis, we compiled data for eight major limnological variables (Secchi disk depth, chlorophyll, total phosphorus, total nitrogen, nitrate, ammonium, true water color, and dissolved organic carbon) and geographic characteristics of lakes (location, lake area, depth, perimeter, watershed area) to evaluate biases in different limnological properties over space and time.
The apparent exponential radiation of Phanerozoic land vertebrates is an artefact of spatial sampling biases
There is no consensus about how terrestrial biodiversity was assembled through deep time, and in particular whether it has risen exponentially over the Phanerozoic. Using a database of 38,711 fossil occurrences, we show that the spatial extent of the 'global' terrestrial tetrapod fossil record itself expands exponentially through the Phanerozoic, and that this spatial variation explains around 75% of the variation in known fossil species counts. Controlling for this bias, we find that regional-scale terrestrial tetrapod diversity was constrained over timespans of tens to hundreds of millions of years, and similar patterns are recovered for major subgroups, such as dinosaurs, mammals and squamates. The Cretaceous/Paleogene mass extinction, 66 million years ago, fundamentally disrupted terrestrial ecosystems, catalysing an abrupt increase in its aftermath. Nevertheless, this was followed by general stasis and recent diversity levels do not exceed those of the Paleogene. These findings parallel those recovered in analyses of local community-level richness, suggesting that tetrapod beta diversity has also not shown a general increase through time. Taken together, our findings strongly contradict past studies that suggested unbounded diversity increases over the last 100 million years.
Data for: The meta-analysis of the effects of spatial sampling bias correction on presence only species distribution models
<p>This dataset contains information extracted from 70 studies identified through a systematic review of the peer-reviewed literature (Web of Science and SCOPUS databases both searched on the 13/02/2023) to evaluate the effect of spatial sampling bias correction methods in presence-only species distribution models.</p>
Data used in "Chemically specific sampling bias: the ratio of PM2.5 to surface AOD on average and peak days in the U.S."
<p>This dataset contains all relevant data used in the manuscript "Chemically specific sampling bias: the ratio of PM2.5 to surface AOD on average and peak days in the U.S." published in Environmental Science: Atmospheres.</p>
A large-scale assessment of ant diversity across the Brazilian Amazon Basin: integrating geographic, ecological, and morphological drivers of sampling bias
<p>Tropical ecosystems are often biodiversity hotspots, and invertebrates represent the main underrepresented component of diversity in large-scale analyses. This problem is partly related to the scarcity of data widely available to conduct these studies and the lack of systematic organization of knowledge about invertebrates' distributions in biodiversity hotspots. Here, we introduce and analyze a comprehensive data compilation of Amazonian ant diversity. Using records from 1817 to 2020 from both published and unpublished sources, we describe the diversity and distribution of ant species in the Brazilian Amazon Basin. Further, using high-definition images and data from taxonomic publications, we build a comprehensive database of morphological traits for the ant species that occur in the region. In total, we recorded 1,067 nominal species in the Brazilian Amazon Basin, with sampling locations strongly biased by access routes, urban centers, research institutions, and major infrastructure projects. Large areas where ant sampling is non-existent represent about 52% of the basin and are concentrated mainly in the North, Southeastern, and Western Brazilian Amazon. We found that distance to roads is the main driver of ant sampling in the Amazon. Contrary to our expectations, morphological traits had lower predictive power in predicting sample bias than purely geographic variables. However, when geographic predictors were controlled, habitat stratum and traits contribute to explain the remaining variance. More species were recorded in better-sampled areas, but species richness estimation models suggest that areas in South Amazonian edge forests are associated with especially high species richness. Our results represent the first trait-based, large-scale study for insects in Amazonian forests and a starting point for macroecological studies focusing on insect diversity in the Amazon Basin.</p>
Meta-analysis of Antarctic phylogeography reveals strong sampling bias and critical knowledge gaps
<p>Much of Antarctica's highly endemic terrestrial biodiversity is found in small ice-free patches. Substantial genetic differentiation has been detected among populations across spatial scales. Sampling is, however, often restricted to commonly-accessed sites, and we therefore lack a comprehensive understanding of broad-scale biogeographic patterns, which could impede forecasts of the nature and impacts of future change. Here, we present a synthesis of published genetic studies across terrestrial Antarctica and the broader Antarctic region, aiming to identify current biogeographic patterns, environmental drivers of diversity, and future research priorities. A database of all published genetic research from terrestrial fauna and flora (excl. microbes) across the Antarctic region was constructed. This database was then filtered to focus on the most well-represented taxa and markers (mitochondrial COI for fauna, and nuclear ITS for flora). The final dataset comprised 7222 records, spanning 153 studies of 335 different species. There was strong taxonomic bias towards flowering plants (52% of all floral data sets) and springtails (54% of all faunal data sets), and geographic bias towards the Antarctic Peninsula and Victoria Land. Recent connectivity between the Antarctic continent and neighbouring landmasses, such as South America and the Southern Ocean Islands (SOIs), was inferred for some groups, but patterns observed for most taxa were strongly influenced by sampling biases. Above-ground wind speed and habitat heterogeneity were positively correlated with genetic diversity indices overall, though environment was a generally poor predictor of genetic diversity. The low resolution and variable coverage of data may also have reduced the power of our comparative inferences. In the future, higher-resolution data, such as genomic SNPs and environmental modelling, alongside targeting sampling of remote sites and under-sampled taxa, will address current knowledge gaps and greatly advance our understanding of evolutionary processes across the Antarctic region.</p>
Considering sampling bias in close-kin mark-recapture (CKMR) abundance estimates of Atlantic salmon
<p>Genetic methods for the estimation of population size can be powerful alternatives to conventional methods. Close-kin mark-recapture (CKMR) is based on the principles of conventional mark-recapture, but instead of being physically marked, individuals are marked through their close kin. The aim of this study was to evaluate the potential of CKMR for the estimation of spawner abundance in Atlantic salmon and how age, sex, spatial, and temporal sampling bias may affect CKMR estimates. Spawner abundance in a wild population was estimated from genetic samples of adults returning in 2018 and of their potential offspring collected in 2019. Adult samples were obtained in two ways. First, adults were sampled and released alive in the breeding habitat during spawning surveys. Second, genetic samples were collected from out-migrating smolts PIT tagged in 2017 and registered when returning as adults in 2018. CKMR estimates based on adult samples collected during spawning surveys were somewhat higher than conventional counts. Uncertainty was small (CV<0.15), due to the detection of a high number of parent-offspring-pairs. Sampling of adults was age- and size-biased and correction for those biases resulted in moderate changes in the CKMR estimate. Juvenile dispersal was limited, but spatially balanced sampling of adults rendered CKMR estimates robust to spatially biased sampling of juveniles. CKMR estimates based on returning PIT tagged adults were approximately twice as high as estimates based on samples collected during spawning surveys. We suggest that estimates based on PIT tagged fish reflect the total abundance of adults entering the river, while estimates based on samples collected during spawning surveys reflect the abundance of adults present in the breeding habitat at the time of spawning. Our study showed that CKMR can be used to estimate spawner abundance in Atlantic salmon, with a moderate sampling effort, but a carefully designed sampling regime is required.</p>
Pooling robustness in distance sampling: Avoiding bias when there is unmodelled heterogeneity
<p>Data from a two-visit line transect survey of four songbird species gathered in spring 2004. Study area size was 33.2 ha of woodland and parkland on the Montrave Estate near Leven in Fife, Scotland.</p>
Biases and distribution patterns in hard-bodied microscopic animals (Acari: Halacaridae): Size doesn't matter, but generalism and sampling effort do
<span>Aim</span> <p><span>The interplay between distribution ranges, species traits, and sampling and taxonomic biases remain elusive amongst microscopic animals. This ignorance obscures our understanding of the diversity patterns of a major component of biodiversity. Here, we used marine Halacaridae to explore whether differences between marine provinces can explain their distribution patterns or if differential sampling efforts across regions prevent any macroecological inference. Furthermore, we test if certain functional traits influence their distribution patterns.</span></p> <span>Location</span> <p><span>Europe.</span></p> <span>Results</span> <p><span>Whereas geographical variables provided a better explanation for differences in species composition, sampling effort and distance from marine biological stations accounted for the majority of differences in European Halacaridae richness. Species occurring in more habitats showed broader geographical ranges and accumulated more records. Species traits like body size affected the distribution of halacarid species.</span></p> <span>Main conclusions</span> <p><span>We propose that the sampling effort of halacarid mites in Europe might be explained by two different cognitive biases: the convenience of selecting certain sampling localities compared to others, and the tendency of zoologists to scrutinize habitats where their target organisms are more common.</span></p>
Toileting behaviours of the UK public: insights for reducing gender bias wastewater-based epidemiology sampling strategies
<p>Cross-sectional survey results from a toileting behaviour survey conducted between the 27th to the 28th of June 2022. Participants (<em>n</em> = 2109) were aged 18 years or older and were living in the UK. The survey consisted of 17 closed-ended questions, with 7 of the questions addressing specific demographic topics and 10 questions addressed toileting behaviour. The questionnaire was designed by a team consisting of environmental microbiologists, public health specialists, wastewater-based epidemiologists, and social scientists, based on the study objectives and incorporating information from previous studies on the same topic. First, self-report questions were asked on typical frequency of urination and defecation, followed by self-reports of frequency of urination and defecation at a variety of locations including at home, at work, educational buildings, transport hubs, and in public toilets. The comfort in urination/defecation at these locations for defecation and urination was also measured. Other questions about toileting behaviour and health monitoring were measured by statements with a 5-point Likert scale (e.g., strongly disagree to strongly agree). </p>
Dataset associated with "Effect of sampling bias on global estimates of ocean carbon export"
<p>Dataset and Matlab code for plotting the figures in the manuscript "Effect of sampling bias on global estimates of ocean carbon export", submitted to Geophysical Research Letters.</p>
Considering sampling bias in close-kin mark-recapture (CKMR) abundance estimates of Atlantic salmon
Open the record for dataset details and reuse information.
The apparent exponential radiation of Phanerozoic land vertebrates is an artefact of spatial sampling biases
Open the record for dataset details and reuse information.
Data for: The meta-analysis of the effects of spatial sampling bias correction on presence only species distribution models
Open the record for dataset details and reuse information.
Biases and distribution patterns in hard-bodied microscopic animals (Acari: Halacaridae): Size doesn’t matter, but generalism and sampling effort do
Open the record for dataset details and reuse information.
A large-scale assessment of ant diversity across the Brazilian Amazon Basin: integrating geographic, ecological, and morphological drivers of sampling bias
Open the record for dataset details and reuse information.
Data from: Dealing with assumptions and sampling bias in the estimation of effective population size: A case study in an amphibian population
Open the record for dataset details and reuse information.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.