Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
88
datasets available to search
ShareScore release 0.7.1
Dataset results
88 results for “sampling bias”
Widespread Sampling Biases in Herbaria Revealed from Large-Scale Digitization 1656-2016
Non-random collecting practices may bias conclusions drawn from analyses of herbarium records. Recent efforts to fully digitize and mobilize regional floras offer a timely opportunity to assess commonalities and differences in herbarium sampling biases. We determined spatial, temporal, trait, phylogenetic, and collector biases in ~5 million herbarium records, representing three of the most complete digitized floras of the world: Australia (AU), South Africa (SA), and New England, USA (NE) We identified numerous shared and unique biases among these regions. Shared biases included specimens i) collected close to roads and herbaria; ii) collected more frequently during spring; iii) of threatened species collected less frequently; and iv) of close relatives collected in similar numbers. Regional differences included i) over-representation of graminoids in SA and AU and of annuals in AU; and ii) peak collection during the 1910s in NE, 1980s in SA, and 1990s in AU. Finally, in all regions, a disproportionately large percentage of specimens were collected by a few individuals. These mega-collectors, and their associated preferences and idiosyncrasies, may have shaped patterns of collection bias via ‘founder effects’. Studies using herbarium collections should account for sampling biases and future collecting efforts should avoid compounding these biases.
Improving the Efficiency of Variationally Enhanced Sampling with Wavelet-Based Bias Potentials
<p>Archive with data supporting the paper "Improving the Efficiency of Variationally Enhanced Sampling with Wavelet-Based Bias Potentials" and the related PhD thesis by B. Pampel</p>
Data from: Integrated species distribution models to account for sampling biases and improve range wide occurrence predictions
<p><strong><span>Aim</span></strong></p> <p><span>Species distribution models (SDMs) that integrate presence-only and presence-absence data offer a promising avenue to improve information on species' geographic distributions. The use of such 'integrated SDMs' on a species range-wide extent has been constrained by the often-limited presence-absence data and by the heterogeneous sampling of the presence-only data. Here, we evaluate integrated SDMs for studying species ranges with a novel expert range map-based evaluation. We build a new understanding about how integrated SDMs address issues of estimation accuracy and data deficiency and thereby offer advantages over traditional SDMs.</span></p> <p><strong><span>Location</span></strong></p> <p><span>South and Central America.</span></p> <p><strong><span>Time period</span></strong></p> <p><span>1979-2017.</span></p> <p><strong><span>Major taxa studied</span></strong></p> <p><span>Hummingbirds.</span></p> <p><strong><span>Methods</span></strong></p> <p><span>We build integrated SDMs by linking two observation models – one for each data type – to the same underlying spatial process.</span> <span>We validate SDMs with two schemes: i) cross-validation with presence-absence data and ii) comparison with respect to the species' whole range as defined with IUCN range maps. We also compare models relative to the estimated response curves and compute the association between the benefit of the data integration and the number of presence records in each data set.</span></p> <p><strong><span>Results</span></strong></p> <p><span>The integrated SDM accounting for the spatially varying sampling intensity of the presence-only data was one of the top-performing models in both model validation schemes. Presence-only data alleviated overly large niche estimates, and data integration was beneficial compared to modelling solely presence-only data for species that had few presence points when predicting the species' whole range. On the community level, integrated models improved the species richness prediction.</span></p> <p><strong><span>Main conclusions</span></strong></p> <p><span>Integrated SDMs combining presence-only and presence-absence data are successfully able to borrow strengths from both data types and offer improved predictions of species' ranges. Integrated SDMs can potentially alleviate the impacts of taxonomically and geographically uneven sampling and to leverage the detailed sampling information in presence-absence data.</span></p>
Piecewise continuous sampling: a method for minimizing bias and sampling effort for estimated metrics of animal behavior
<p>Capturing qualitative features of animal behavior requires recording occurrences of behavior over time. Continuous sampling is best for capturing brief behaviors, but can be very time consuming. Instantaneous sampling can reduce the amount of labor required, but can miss short-duration behaviors. We therefore synthesized these techniques by continuously sampling during randomly scattered time intervals; a technique we call piecewise continuous sampling. To optimize and test the efficacy of this technique, we collected a continuous behavioral dataset of harvester ant workers, and then we developed a protocol to estimate the amount of sampling time necessary to reconstruct the proportion of time animals spend in different behavioral states. This protocol finds the sample size needed for the variance of the sample to converge on the variation of the population. We then divided this estimated time into equal-duration intervals that were randomly distributed across the entire continuous dataset. Finally, we calculated both time-dependent and time-independent error from this sample. We found that 4 to 16 sampling intervals minimize both types of error simultaneously. This finding was robust to differences in underlying behavior and was validated with simulations, implying that this method could be used for many types of organisms.</p>
Sampling bias exaggerates a textbook example of a trophic cascade
<p>Understanding trophic cascades in terrestrial wildlife communities is a major challenge because these systems are difficult to sample properly. We show how a tradition of nonrandom sampling has confounded this understanding in a textbook system (Yellowstone National Park) where carnivore [<em>Canis lupus</em> (wolf)] recovery is associated with a trophic cascade involving changes in herbivore [<em>Cervus canadensis</em> (elk)] behavior and density that promote plant regeneration. Long-term data indicate a practice of sampling only the tallest young plants overestimated regeneration of overstory aspen (<em>Populus tremuloides</em>) by a factor of 3-8 compared to random sampling because it favored plants taller than the preferred browsing height of elk and overlooked non-regenerating aspen stands. Random sampling described a trophic cascade, but it was weaker than the one that nonrandom sampling described. Our findings highlight the critical importance of basic sampling principles (e.g., randomization) for achieving an accurate understanding of trophic cascades in terrestrial wildlife systems.</p>
Understanding Sampling Bias in the Global Heat Flow Compilation
<p>Geothermal heat flow measurements, including calculated weights and geological, tectonic and topographic settings.</p> <p>Paper in review (July 14th, 2022)</p>
How Confidence in Prior Attitudes, Social Tag Popularity, and Source Credibility Shape Confirmation Bias Toward Antidepressants and Psychotherapy in a Representative German Sample: Randomized Controlled Web-Based Study
<p>ABSTRACT</p> <p>Background: In health-related, Web-based information search, people should select information in line with expert (vs nonexpert) information, independent of their prior attitudes and consequent confirmation bias.</p> <p>Objective: This study aimed to investigate confirmation bias in mental health–related information search, particularly (1) if high confidence worsens confirmation bias, (2) if social tags eliminate the influence of prior attitudes, and (3) if people successfully distinguish high and low source credibility.</p> <p>Methods: In total, 520 participants of a representative sample of the German Web-based population were recruited via a panel company. Among them, 48.1% (250/520) participants completed the fully automated study. Participants provided <em>prior attitudes</em> about antidepressants and psychotherapy. We manipulated (1) <em>confidence</em> in prior attitudes when participants searched for blog posts about the treatment of depression, (2) <em>tag popularity</em> —either psychotherapy or antidepressant tags were more popular, and (3) <em>source credibility</em> with banners indicating high or low expertise of the tagging community. We measured <em>tag</em> and <em>blog post</em> selection, and <em>treatment</em><em>efficacy ratings</em> after navigation.</p> <p>Results: Tag popularity predicted the proportion of selected antidepressant tags (beta=.44, SE 0.11; <em>P</em><.001) and blog posts (beta=.46, SE 0.11; <em>P</em><.001). When confidence was low (−1 SD), participants selected more blog posts consistent with prior attitudes (beta=−.26, SE 0.05; <em>P</em><.001). Moreover, when confidence was low (−1 SD) and source credibility was high (+1 SD), the efficacy ratings of attitude-consistent treatments increased (beta=.34, SE 0.13; <em>P</em>=.01).</p> <p>Conclusions: We found correlational support for defense motivation account underlying confirmation bias in the mental health–related search context. That is, participants tended to select information that supported their prior attitudes, which is not in line with the current scientific evidence. Implications for presenting persuasive Web-based information are also discussed.</p> <p>Trial Registration: ClinicalTrials.gov NCT03899168; https://clinicaltrials.gov/ct2/show/NCT03899168 (Archived by WebCite at http://www.webcitation.org/77Nyot3Do)</p> <p>J Med Internet Res 2019;21(4):e11081</p> <p>doi:10.2196/11081</p>
Fig. 6 in Using abundance data to assess the relative role of sampling biases and evolutionary radiations in Upper Muschelkalk ammonoids
Fig. 6. Percent similarity among bins averaged to 1 degree bins. A. om7 interval. B. om8 interval. C. om9 interval. The thicker the line, the greater the similarity between the two cells connected by the line.
Fig. 4 in Using abundance data to assess the relative role of sampling biases and evolutionary radiations in Upper Muschelkalk ammonoids
Fig. 4. Correlations between richness per map and number of localities. A. om7 interval. B. om8 interval. C. om9 interval. The gap in the distribution of points for the om8 interval highlights the discontinuity between a group of maps with few taxa at a few localities and other maps with a large number of localities and high richness.
Fig. 5 in Using abundance data to assess the relative role of sampling biases and evolutionary radiations in Upper Muschelkalk ammonoids
Fig. 5. Rarefaction curves for each interval, based on number of occurrences. The confidence envelope of the species richness for om9 departs significantly from those of om7 and om8 above 50 occurrences, but the significantly higher species−richness of om8 only becomes apparent at sample sizes of around 250 specimens, indicating that a few, rare taxa are boosting richness in the om8 interval.
Fig. 2 in Using abundance data to assess the relative role of sampling biases and evolutionary radiations in Upper Muschelkalk ammonoids
Fig. 2. Distribution of Muschelkalk ammonoid localities used in this study plotted on a map of modern Germany. The overall geographic spread of localities does not change greatly over time.
Fig. 3 in Using abundance data to assess the relative role of sampling biases and evolutionary radiations in Upper Muschelkalk ammonoids
Fig. 3. Correlations between richness per map and number of occurrences. A. om7 interval. B. om8 interval. C. om9 interval.
Fig. 1 in Using abundance data to assess the relative role of sampling biases and evolutionary radiations in Upper Muschelkalk ammonoids
Fig. 1. Chart of stratigraphic interval names and durations for the Muschelkalk of the Germanic Basin with ammonoid immigration events marked (simplified from Klug et al. 2005: fig. 1).
FIGURE 5 in A simulation-based examination of residual diversity estimates as a method of correcting for sampling bias
FIGURE 5. The performance of different implementations of the residual diversity estimate when a specific bias is forced to be the dominant influence. (5.1) Mean Spearman's rho values of four implementations of the RDE using the Smith and McGowan method. PFORM and PTAPH are set at 0.9 to minimise their influence, PLOC is variable. PMIST set at 0.1. (5.2) Mean Spearman's rho values of four implementations of the RDE using the Smith and McGowan method. LOC and PTAPH are set at 0.9 to minimise their influence, PFORM is variable. PMIST set at 0.1. The dashed red line indicates the critical value at p=0.05. Abbreviations as in Table 1.
FIGURE 7 in A simulation-based examination of residual diversity estimates as a method of correcting for sampling bias
FIGURE 7. Sample simulation comparing the results of the taxic, phylogenetic and residual diversity estimates to the true diversity. PFROM, PLOC and PTAPH set at 0.25. PMIST set at 0.1. Black box highlights instance where the Signor Lipps effect has been exaggerated by the PDE; the TDE and RDE both identify the rapid diversity decrease present in the true diversity. Abbreviations as in Table 1.
FIGURE 6 in A simulation-based examination of residual diversity estimates as a method of correcting for sampling bias
FIGURE 6. The performance of the phylogenetic diversity estimate when errors are introduced to the phylogeny. Mean Spearman's rho values of the PDE, TDE and the best performing implementation of the RDE. PLOC, PFORM and PTAPH set at 0.25. PMIST variable. The dashed red line indicates the critical value at p=0.05. Abbreviations as in Table 1.
FIGURE 4 in A simulation-based examination of residual diversity estimates as a method of correcting for sampling bias
FIGURE 4. The performance of different implementations of the residual diversity estimate examining faunas with varying degrees of homogeneity. PFORM, PLOC and PTAPH are set at 0.25. The rate of dispersal is increased relative to the rate of local extinction to increase the homogeneity of the faunas. The dashed red line indicates the critical value at p=0.05. Abbreviations as in Table 1.
FIGURE 1 in A simulation-based examination of residual diversity estimates as a method of correcting for sampling bias
FIGURE 1. An illustration of the taphonomic filter in the simulation, shown applied to a single taxon in a single time bin. The taxon is originally present in every locality in each region it occupies, but the taphonomic filter removes it from randomly selected localities
FIGURE 3 in A simulation-based examination of residual diversity estimates as a method of correcting for sampling bias
FIGURE 3. The performance of different implementations of the residual diversity estimate (RDE) under different sampling regimes. (3.1) Mean Spearman's rho values of four implementations of the RDE using Formations as a proxy, with values of PFORM, PLOC and PTAPH variable but equal. (3.2) Mean Spearman's rho values of four implementations of the RDE using Localities as a proxy. (3.3) Mean Spearman's rho values of four implementations of the RDE, all using the Smith and McGowan method. (3.4) Mean Spearman's rho values of the taxic and phylogenetic diversity estimate compared to those of the optimum implementation of the RDE. The dashed red line indicates the critical value at p=0.05. Abbreviations as in Table 1.
FIGURE 2 in A simulation-based examination of residual diversity estimates as a method of correcting for sampling bias
FIGURE 2. An illustration of how sampling proxies are generated in this simulation. This schematic illustrates which formations and localities in a single time bin contain fossils of at least one species of the simulated clade after application of the taphonomic filter. Formations and localities are removed at random, representing a lack of sampling. Note that the number of clade-bearing formations and localities does not necessarily equal the number of formations and localities sampled, allowing the generation of four sampling proxies.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.