Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

88

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

88 results for “Sampling bias”

Learn how ShareScore rates datasets ↗
dryad40/100

Sampling bias exaggerates a textbook example of a trophic cascade

Open the record for dataset details and reuse information.

publicNov 2023View details →
dryad40/100

Data from: Integrated species distribution models to account for sampling biases and improve range wide occurrence predictions

Open the record for dataset details and reuse information.

publicNov 2023View details →
dryad40/100

Piecewise continuous sampling: a method for minimizing bias and sampling effort for estimated metrics of animal behavior

Open the record for dataset details and reuse information.

publicApr 2024View details →
edi40/100

LAGOS Lake nutrient, carbon and chlororphyll data to evaluate biases in lake water quality sampling practices in a 17-state region of the US

This dataset includes data for eight major limnological variables in LAGOS-NE_LIMNO v. 1.087.1 that were used to evaluate biases in lake water quality sampling and implications for macroscale research (Stanley et al. In Revision, Limnology and Oceanography). Most observations came from LAGOS-NE_LIMNO v. 1.087.1, an integrated database of lake ecosystems (Soranno et al. 2015, Soranno et al. 2017) but were supplemented with additional data from the State of New Hampshire. LAGOS-NE contains information on lakes great than or equal to 1 ha (originally derived from the U.S. Geological Survey's 2013 National Hydrography Dataset) for a 17-state region of the U.S., and a subset of the lakes has observational data on lake chemistry and productivity. Approximately 87 different sources of data were compiled for the LAGOS-NELIMNO v. 1.087.1 dataset and were mostly generated by government agencies (state, federal, tribal) and universities. In this analysis, we compiled data for eight major limnological variables (Secchi disk depth, chlorophyll, total phosphorus, total nitrogen, nitrate, ammonium, true water color, and dissolved organic carbon) and geographic characteristics of lakes (location, lake area, depth, perimeter, watershed area) to evaluate biases in different limnological properties over space and time.

openCC (other)Jan 2019View details →
dryad36/100

The apparent exponential radiation of Phanerozoic land vertebrates is an artefact of spatial sampling biases

There is no consensus about how terrestrial biodiversity was assembled through deep time, and in particular whether it has risen exponentially over the Phanerozoic. Using a database of 38,711 fossil occurrences, we show that the spatial extent of the 'global' terrestrial tetrapod fossil record itself expands exponentially through the Phanerozoic, and that this spatial variation explains around 75% of the variation in known fossil species counts. Controlling for this bias, we find that regional-scale terrestrial tetrapod diversity was constrained over timespans of tens to hundreds of millions of years, and similar patterns are recovered for major subgroups, such as dinosaurs, mammals and squamates. The Cretaceous/Paleogene mass extinction, 66 million years ago, fundamentally disrupted terrestrial ecosystems, catalysing an abrupt increase in its aftermath. Nevertheless, this was followed by general stasis and recent diversity levels do not exceed those of the Paleogene. These findings parallel those recovered in analyses of local community-level richness, suggesting that tetrapod beta diversity has also not shown a general increase through time. Taken together, our findings strongly contradict past studies that suggested unbounded diversity increases over the last 100 million years.

opencc-zeroMar 2020View details →
dryad36/100

Data for: The meta-analysis of the effects of spatial sampling bias correction on presence only species distribution models

<p>This dataset contains information extracted from 70 studies identified through a systematic review of the peer-reviewed literature (Web of Science and SCOPUS databases both searched on the 13/02/2023) to evaluate the effect of spatial sampling bias correction methods in presence-only species distribution models.</p>

opencc-zeroDec 2023View details →
zenodo36/100

Data used in "Chemically specific sampling bias: the ratio of PM2.5 to surface AOD on average and peak days in the U.S."

<p>This dataset contains all relevant data used in the manuscript "Chemically specific sampling bias: the ratio of PM2.5 to surface AOD on average and peak days in the U.S." published in Environmental Science: Atmospheres.</p>

opencc-by-4.0Nov 2023View details →
dryad36/100

A large-scale assessment of ant diversity across the Brazilian Amazon Basin: integrating geographic, ecological, and morphological drivers of sampling bias

<p>Tropical ecosystems are often biodiversity hotspots, and invertebrates represent the main underrepresented component of diversity in large-scale analyses. This problem is partly related to the scarcity of data widely available to conduct these studies and the lack of systematic organization of knowledge about invertebrates' distributions in biodiversity hotspots. Here, we introduce and analyze a comprehensive data compilation of Amazonian ant diversity. Using records from 1817 to 2020 from both published and unpublished sources, we describe the diversity and distribution of ant species in the Brazilian Amazon Basin. Further, using high-definition images and data from taxonomic publications, we build a comprehensive database of morphological traits for the ant species that occur in the region. In total, we recorded 1,067 nominal species in the Brazilian Amazon Basin, with sampling locations strongly biased by access routes, urban centers, research institutions, and major infrastructure projects. Large areas where ant sampling is non-existent represent about 52% of the basin and are concentrated mainly in the North, Southeastern, and Western Brazilian Amazon. We found that distance to roads is the main driver of ant sampling in the Amazon. Contrary to our expectations, morphological traits had lower predictive power in predicting sample bias than purely geographic variables. However, when geographic predictors were controlled, habitat stratum and traits contribute to explain the remaining variance. More species were recorded in better-sampled areas, but species richness estimation models suggest that areas in South Amazonian edge forests are associated with especially high species richness. Our results represent the first trait-based, large-scale study for insects in Amazonian forests and a starting point for macroecological studies focusing on insect diversity in the Amazon Basin.</p>

opencc-zeroMay 2022View details →
dryad36/100

Meta-analysis of Antarctic phylogeography reveals strong sampling bias and critical knowledge gaps

<p>Much of Antarctica's highly endemic terrestrial biodiversity is found in small ice-free patches. Substantial genetic differentiation has been detected among populations across spatial scales. Sampling is, however, often restricted to commonly-accessed sites, and we therefore lack a comprehensive understanding of broad-scale biogeographic patterns, which could impede forecasts of the nature and impacts of future change. Here, we present a synthesis of published genetic studies across terrestrial Antarctica and the broader Antarctic region, aiming to identify current biogeographic patterns, environmental drivers of diversity, and future research priorities. A database of all published genetic research from terrestrial fauna and flora (excl. microbes) across the Antarctic region was constructed. This database was then filtered to focus on the most well-represented taxa and markers (mitochondrial COI for fauna, and nuclear ITS for flora). The final dataset comprised 7222 records, spanning 153 studies of 335 different species. There was strong taxonomic bias towards flowering plants (52% of all floral data sets) and springtails (54% of all faunal data sets), and geographic bias towards the Antarctic Peninsula and Victoria Land. Recent connectivity between the Antarctic continent and neighbouring landmasses, such as South America and the Southern Ocean Islands (SOIs), was inferred for some groups, but patterns observed for most taxa were strongly influenced by sampling biases. Above-ground wind speed and habitat heterogeneity were positively correlated with genetic diversity indices overall, though environment was a generally poor predictor of genetic diversity. The low resolution and variable coverage of data may also have reduced the power of our comparative inferences. In the future, higher-resolution data, such as genomic SNPs and environmental modelling, alongside targeting sampling of remote sites and under-sampled taxa, will address current knowledge gaps and greatly advance our understanding of evolutionary processes across the Antarctic region.</p>

opencc-zeroSep 2022View details →
dryad36/100

Considering sampling bias in close-kin mark-recapture (CKMR) abundance estimates of Atlantic salmon

<p>Genetic methods for the estimation of population size can be powerful alternatives to conventional methods. Close-kin mark-recapture (CKMR) is based on the principles of conventional mark-recapture, but instead of being physically marked, individuals are marked through their close kin. The aim of this study was to evaluate the potential of CKMR for the estimation of spawner abundance in Atlantic salmon and how age, sex, spatial, and temporal sampling bias may affect CKMR estimates. Spawner abundance in a wild population was estimated from genetic samples of adults returning in 2018 and of their potential offspring collected in 2019. Adult samples were obtained in two ways. First, adults were sampled and released alive in the breeding habitat during spawning surveys. Second, genetic samples were collected from out-migrating smolts PIT tagged in 2017 and registered when returning as adults in 2018. CKMR estimates based on adult samples collected during spawning surveys were somewhat higher than conventional counts. Uncertainty was small (CV&lt;0.15), due to the detection of a high number of parent-offspring-pairs. Sampling of adults was age- and size-biased and correction for those biases resulted in moderate changes in the CKMR estimate. Juvenile dispersal was limited, but spatially balanced sampling of adults rendered CKMR estimates robust to spatially biased sampling of juveniles. CKMR estimates based on returning PIT tagged adults were approximately twice as high as estimates based on samples collected during spawning surveys. We suggest that estimates based on PIT tagged fish reflect the total abundance of adults entering the river, while estimates based on samples collected during spawning surveys reflect the abundance of adults present in the breeding habitat at the time of spawning. Our study showed that CKMR can be used to estimate spawner abundance in Atlantic salmon, with a moderate sampling effort, but a carefully designed sampling regime is required.</p>

opencc-zeroJan 2022View details →
dryad36/100

Pooling robustness in distance sampling: Avoiding bias when there is unmodelled heterogeneity

<p>Data from a two-visit line transect survey of four songbird species gathered in spring 2004. Study area size was 33.2 ha of woodland and parkland on the Montrave Estate near Leven in Fife, Scotland.</p>

opencc-zeroNov 2022View details →
dryad36/100

Biases and distribution patterns in hard-bodied microscopic animals (Acari: Halacaridae): Size doesn't matter, but generalism and sampling effort do

<span>Aim</span> <p><span>The interplay between distribution ranges, species traits, and sampling and taxonomic biases remain elusive amongst microscopic animals. This ignorance obscures our understanding of the diversity patterns of a major component of biodiversity. Here, we used marine Halacaridae to explore whether differences between marine provinces can explain their distribution patterns or if differential sampling efforts across regions prevent any macroecological inference. Furthermore, we test if certain functional traits influence their distribution patterns.</span></p> <span>Location</span> <p><span>Europe.</span></p> <span>Results</span> <p><span>Whereas geographical variables provided a better explanation for differences in species composition, sampling effort and distance from marine biological stations accounted for the majority of differences in European Halacaridae richness. Species occurring in more habitats showed broader geographical ranges and accumulated more records. Species traits like body size affected the distribution of halacarid species.</span></p> <span>Main conclusions</span> <p><span>We propose that the sampling effort of halacarid mites in Europe might be explained by two different cognitive biases: the convenience of selecting certain sampling localities compared to others, and the tendency of zoologists to scrutinize habitats where their target organisms are more common.</span></p>

opencc-zeroJan 2023View details →
zenodo36/100

Toileting behaviours of the UK public: insights for reducing gender bias wastewater-based epidemiology sampling strategies

<p>Cross-sectional survey results from a toileting behaviour survey conducted&nbsp;between the 27th to the 28th of June 2022. Participants (<em>n</em>&nbsp;= 2109) were aged 18 years or older and were living in the UK.&nbsp;The survey consisted of 17 closed-ended questions, with 7 of the questions addressing specific demographic topics and 10 questions addressed toileting behaviour. The questionnaire was designed by a team consisting of environmental microbiologists, public health specialists, wastewater-based epidemiologists, and social scientists, based on the study objectives and incorporating information from previous studies on the same topic.&nbsp; First, self-report questions were asked on typical frequency of urination and defecation, followed by self-reports of frequency of urination and defecation at a variety of locations including at home, at work, educational buildings, transport hubs, and in public toilets. The comfort in urination/defecation at these locations for defecation and urination was also measured. Other questions about toileting behaviour and health monitoring were measured by statements with a 5-point Likert scale (e.g., strongly disagree to strongly agree).&nbsp;</p>

opencc-by-4.0Jul 2023View details →
zenodo36/100

Dataset associated with "Effect of sampling bias on global estimates of ocean carbon export"

<p>Dataset and Matlab code for plotting the figures in the manuscript &quot;Effect of sampling bias on global estimates of ocean carbon export&quot;, submitted to Geophysical Research Letters.</p>

opencc-by-4.0Aug 2023View details →
dryad36/100

Considering sampling bias in close-kin mark-recapture (CKMR) abundance estimates of Atlantic salmon

Open the record for dataset details and reuse information.

publicOct 2021View details →
dryad36/100

The apparent exponential radiation of Phanerozoic land vertebrates is an artefact of spatial sampling biases

Open the record for dataset details and reuse information.

publicMar 2020View details →
dryad36/100

Data for: The meta-analysis of the effects of spatial sampling bias correction on presence only species distribution models

Open the record for dataset details and reuse information.

publicDec 2023View details →
dryad36/100

Biases and distribution patterns in hard-bodied microscopic animals (Acari: Halacaridae): Size doesn’t matter, but generalism and sampling effort do

Open the record for dataset details and reuse information.

publicJan 2023View details →
dryad36/100

A large-scale assessment of ant diversity across the Brazilian Amazon Basin: integrating geographic, ecological, and morphological drivers of sampling bias

Open the record for dataset details and reuse information.

publicJun 2022View details →
dryad36/100

Data from: Dealing with assumptions and sampling bias in the estimation of effective population size: A case study in an amphibian population

Open the record for dataset details and reuse information.

publicSep 2024View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record