Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
88
datasets available to search
ShareScore release 0.9.0
Dataset results
88 results for “Sampling bias”
Pooling robustness in distance sampling: Avoiding bias when there is unmodelled heterogeneity
Open the record for dataset details and reuse information.
Meta-analysis of Antarctic phylogeography reveals strong sampling bias and critical knowledge gaps
Open the record for dataset details and reuse information.
Data from: Dispersal of stream salmonids from nests and stocking sites: patterns, variability, and sampling bias
Open the record for dataset details and reuse information.
Data from: Sexual dimorphism and sex-biased sampling influence analyses of trait evolution and diversity in a major songbird clade
Open the record for dataset details and reuse information.
Data from: Sampling beetle communities: trap design interacts with weather and species traits to bias capture rates
<p>Globally, many insect populations are declining, prompting calls for action. Yet these findings have also prompted discussion about sampling methods and interpretation of long-term datasets. As insect monitoring and research efforts increase, it is critical to quantify the effectiveness of sampling methods. This is especially true if sampling biases of different methods covary with climate, which is also changing over time. We assess the effectiveness of two types of flight intercept traps commonly used for beetles, a diverse insect group responsible for numerous ecosystem services, under different climatic conditions in Norwegian boreal forest. One of these trap designs includes a device to prevent rainwater from entering the collection vial, diluting preservatives and flushing out beetles. This design is compared to a standard trap. We ask how beetle capture rates vary between these traps, and how these differences vary based on precipitation levels and beetle body size, an important species trait. Bayesian mixed models reveal that the standard and modified traps differ in their beetle capture rates, but that the magnitude and direction of these differences change with precipitation levels and beetle body size. At low rainfall levels standard traps catch more beetles, but as precipitation increases the catch rates of modified traps overtake those of standard traps. This effect is most pronounced for large-bodied beetles. Sampling methods are known to differ in their effectiveness. Here, we present evidence for a less well-known but likely common phenomenon - an interaction between climate and sampling, such that relative effectiveness of trap types for beetle sampling differs depending on precipitation levels and species traits. This highlights a challenge for long-term monitoring programs, where both climate and insect populations are changing. Sampling methods should be sought that eliminate climate interactions, any biases should be quantified, and all insect datasets should include detailed methodological metadata.</p>
Data from: Breaking down the lithification bias: the effect of preferential sampling of larger specimens on the estimate of species richness, evenness, and average specimen size
Lithification, the transition of unconsolidated sediments to fully indurated rocks, can potentially bias estimates of species richness, evenness, and body size distribution derived from fossil assemblages. Fossil collections made from well-indurated rocks consistently exhibit lower species richness, lower evenness, and a specimen size distribution skewed towards larger specimens relative to collections made from unconsolidated sediments, even when collections are drawn from the same assemblage. This phenomenon is known as the lithification bias. While the bias itself has been demonstrated empirically, much less attention has been paid to its causes. Proposed causes include taphonomic processes (e.g., destruction of small specimens during early diagenesis) or methodological differences (e.g., sieving vs. counting specimens on outcrops, bedding surfaces, or mechanically split surfaces). Here we investigate the potential effects of preferential intersection that could also result in a methodologically related bias: the preferential sampling of larger specimens relative to smaller ones when fossils are counted on rock surfaces. We used an analog model to simulate preferential intersection (fossil collection via splitting fossiliferous rock) and compare the results to a random draw model that approximates the effects of sieving. The model was parameterized using nine different combinations of species abundance and species size distributions. The results show that, with rare exceptions, species richness is 5–23% lower, evenness 5-25% lower, and average specimen size 24–150% higher in preferential intersection than in random draw simulations. We conclude that preferential intersection can impose a significant bias independent of other mechanisms (e.g., preferential destruction of smaller specimens during diagenetic or sampling processes), that the magnitude of this bias is partially dependent on the species abundance and size distributions, and that this bias alone does not fully account for empirically observed lithification bias on species richness (i.e., other sources of bias are also at work).
Data from: The scale-of-choice effect and how estimates of assortative mating in the wild can be biased due to heterogeneous samples
The mode in which sexual organisms choose mates is a key evolutionary process, as it can have a profound impact on fitness and speciation. One way to study mate choice in the wild is by measuring trait correlation between mates. Positive assortative mating is inferred when individuals of a mating pair display traits that are more similar than those expected under random mating while negative assortative mating is the opposite. A recent review of 1134 trait correlations found that positive estimates of assortative mating were more frequent and larger in magnitude than negative estimates. Here we describe the scale-of-choice effect (SCE), which occurs when mate choice exists at a smaller scale than that of the investigator's sampling, while simultaneously the trait is heterogeneously distributed at the true scale-of-choice. We demonstrate the SCE by Monte Carlo simulations and estimate it in two organisms showing positive (Littorina saxatilis) and negative (L. fabalis) assortative mating. Our results show that both positive and negative estimates are biased by the SCE by different magnitudes, typically towards positive values. Therefore, the low frequency of negative assortative mating observed in the literature may be due to the SCE's impact on correlation estimates, which demands new experimental evaluation.
Data from: A survey of palaeontological sampling biases in fishes based on the phanerozoic record of Great Britain
Fishes represent more than half of all living vertebrate species, but patterns of fish diversity remain little explored in the fossil record. A compendium of fossil occurrences from Great Britain was assembled in order to address a series of questions concerning the palaeontological record of fishes. There are broad similarities between British richness trajectories and those compiled from global data, including an initial peak in the mid-Palaeozoic (Devonian or Carboniferous, depending on the compilation), with a late Palaeozoic trough followed by a sharp rise in diversity in the Late Cretaceous and Paleogene. The British dataset is too small to reveal any significant differences in richness between time bins using subsampling, but a modeling approach based on sampling and geological proxies consistently shows lower-than-predicted richness in the Silurian-Devonian and higher-than-predicted richness in the Late Cretaceous and Eocene. This positive excursion is robust to the exclusion of data from the early Eocene London Clay Lagerstätte. Chondrichthyans (sharks, rays, and ratfishes) and osteichthyans (ray-finned and lobe-finned fishes) show contrasting relationships with geological and sampling proxies, possibly reflecting different taphonomic profiles or idiosyncratic variation in the relative proportion of freshwater and marine deposits over the British Phanerozoic.
Polyester simulation (with GC bias) samples 1-4
<p>Simulated reads for samples 1-4 of an 8 sample (and 2 condition) dataset. </p>
Polyester simulation (with GC bias) samples 5-8
<p>Simulated reads for samples 5-8 of an 8 sample (and 2 condition) dataset. </p>
Dataset for paper: How Twitter Data Sampling Biases U.S. Voter Behavior Characterizations
<p>This repository contains the data and code for the paper "How Twitter Data Sampling Biases U.S. Voter Behavior Characterizations."</p>
Assessing and Correcting Neighborhood Socioeconomic Spatial Sampling Biases in Citizen Science Mosquito Data Collection
<p>Reporting data from the Mosquito Alert citizen science system, active catch basin surveillance, and mosquito trap surveillance used in "Assessing and Correcting Neighborhood Socioeconomic Spatial Sampling Biases in Citizen Science Mosquito Data Collection."</p> <p>The file named mosquito_alert_adult_bite_reports_Barcelona_2014_2023.Rds includes all adult mosquito and mosquito bite reports received from Barcelona Municipality from the start of the Mosqiuto Alert project in 2014 through the end of 2023. The file named mosquito_alert_validated_albopictus_reports_Barcelona_2014_23.Rds includes all expert-validated <em>Ae. albopictus </em>reports received from Barcelona Municipality during the same time period. The data is stored as RDS files and contain the following fields:</p> <ul> <li><strong>year </strong>- the year in which the report was made. Class = dbl.</li> <li><strong>date </strong>- the date om which the report was made. Class = date.</li> <li><strong>type </strong>- the report type, either adult mosquito ("adult") or mosquito breeding site ("site"). Class = chr.</li> <li><strong>lon</strong> - the longitude of the report location. Class = dbl.</li> <li><strong>lat</strong> - the latitude of the report location. Class = dbl.</li> <li><strong>validation_score</strong> - Entolab validation score. Either 1 (possible <em>Ae. albopictus</em>) or 2 (probable <em>Ae. albopictus</em>). This field is present only in the validated reports data. </li> </ul> <p>The file named active_catch_basin_drain_data.Rds includes information about all catch basin drains in Barcelona Municipality in which the Barcelona Public Health Agency (ASPB) detected mosquito activity as part of its continuous monitoring and control of mosquitoes from 2019 through 2023. The data is stored in an RDS file with the following fields:</p> <ul> <li><strong>any_reports </strong>- dummy variable indicating whether any Mosquito Alert adult mosquito or mosquito bite reports were sent through Mosquito Alert from within 200 m of the catch basin drain during the year in which the ASPB detected mosquito activity in hte catch basin drain. Class = lgl.</li> <li><strong>se_expected</strong> - sampling effort for the 0.025 degree lon/lat sampling cell in which the catch basin drain lies during the year in which the ASPB detected mosquito activity in the drain. This value is taken from the SE_expected variable in the sampling_effort_daily_cellres_025.csv.gz file available at https://zenodo.org/records/12602985. Sampling effort is estimated as the expected number of participants sending at least one report from the cell during the day in question given the the number of participants recorded in the cell that day and the amount of time elapsed since each one began participating in the project. Class = dbl.</li> <li><strong>p_singlehh</strong> - proportion of single-member households in the population of the census tract in which the catch basin drain is located. Class = dbl.</li> <li><strong>mean_age </strong>- mean age of the population of the census tract in which the catch basin drain is located. Class = dbl.</li> <li><strong>mean_rent_consumption_unit</strong> - mean income per consumption unit in the census tract in which the catch basin drain is located. Class = dbl.</li> <li><strong>popd</strong> - population density of the census tract in which the catch basin drain is located. Class = dbl.</li> <li><strong>id_item </strong>- unique identifier given to the catch basin drain. Drain itentifiers appear multiple times in the data when the ASPB detected activity in the drain in multiple years. Class = dbl.</li> </ul> <p>The file named trap_data.Rds includes information on the adult mosquito trap surveillance analyzed in this article. The data is stored in an RDS file with the following fields:</p> <ul> <li><strong>females </strong>- number of Ae. albopictus females found in the trap. Class = dbl.</li> <li><strong>trap_name</strong> - unique identifier for the trap. Class = chr.</li> <li><strong>trapping_effort</strong> - number of days from when the trap was set to when it was checked. Class = dbl.</li> <li><strong>date</strong> - date on which the trap was checked. Class = date.</li> <li><strong>mean_tm30</strong> - mean temperature for the 30 days leading up to the date on which the trap was checked. Class = dbl.</li> <li><strong>mean_rent_consumption_unit </strong>- mean income per consumption unit for the census tract in which the trap was located. Class = dbl.</li> </ul>
Data from: Across space and time: a review of sampling and analytical biases in fossil data across macroecological scales
<p>Quantitative studies of fossil data have proven critical to a number of major macroevolutionary and macroecological discoveries, such as the 'Big 5' mass extinctions of the Phanerozoic. The development and easy accessibility of major meta-data sources such as the Paleobiology Database and Geobiodiversity Database have also spurred the widespread application of these data to testing ecological hypotheses at finer spatiotemporal and phylogenetic scales. However, issues of preservational/taphonomic biases, sampling/collecting biases, taxonomic issues, and analytical choice can impact the degree of interpretative resolution possible, and even obscure biological 'signal' from error/bias-introduced 'noise'. The degree to which these factors can impact analytical interpretations is not well-documented in comparison to the scale of use of these data sources. Here, we review the many forms of systematic error that can creep into a paleoecological study, from the stage of data collection to the interpretation of analytical results, and provide two case studies based upon re-analysis of previously-published datasets to illustrate the varying impacts of such biases. The first case study focuses on the Cambrian Burgess Shale, and the second on the Belly River Group, with both representing highly-sampled, taphonomically characterized, and spatiotemporally-constrained datasets developed through multiple years of sustained field collecting. In the former, we illustrate the impacts of collecting bias through quantitative comparisons of collected vs. discarded specimens over multiple field seasons, illustrating the impact of this data loss on ecological reconstructions and analysis. In the latter case study, we review the impact of preservational biases, the approaches to their quantification and mitigation, where these approaches have led to misinterpretations in the past, and the differences in ecological resolution that result from occurrence vs abundance approaches in macroecological analysis. Lastly, we synthesize these case studies with our review of past approaches to propose a series of recommendations for future paleoecological and macroecological studies, emphasizing the continued importance of high-quality primary data and ongoing need for a first-principles approach to address existing issues of missing data.</p>
Data from: Using data from related species to overcome spatial sampling bias and associated limitations in ecological niche modeling
Open the record for dataset details and reuse information.
Data from: Who escapes detection? Quantifying the causes and consequences of sampling biases in a long-term field study
Open the record for dataset details and reuse information.
Data from: Across space and time: a review of sampling and analytical biases in fossil data across macroecological scales
Open the record for dataset details and reuse information.
Data from: Taxonomic structure of the fossil record is shaped by sampling bias
Open the record for dataset details and reuse information.
Data from: A survey of palaeontological sampling biases in fishes based on the phanerozoic record of Great Britain
Open the record for dataset details and reuse information.
Data from: Sampling beetle communities: trap design interacts with weather and species traits to bias capture rates
Open the record for dataset details and reuse information.
Data from: Lepidosaurian diversity in the Mesozoic–Paleogene: the potential roles of sampling biases and environmental drivers
Open the record for dataset details and reuse information.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.