Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

88

datasets available to search

ShareScore release 0.7.1

Reset

Dataset results

88 results for “sampling bias”

Learn how ShareScore rates datasets ↗
edi56/100

Widespread Sampling Biases in Herbaria Revealed from Large-Scale Digitization 1656-2016

Non-random collecting practices may bias conclusions drawn from analyses of herbarium records. Recent efforts to fully digitize and mobilize regional floras offer a timely opportunity to assess commonalities and differences in herbarium sampling biases. We determined spatial, temporal, trait, phylogenetic, and collector biases in ~5 million herbarium records, representing three of the most complete digitized floras of the world: Australia (AU), South Africa (SA), and New England, USA (NE) We identified numerous shared and unique biases among these regions. Shared biases included specimens i) collected close to roads and herbaria; ii) collected more frequently during spring; iii) of threatened species collected less frequently; and iv) of close relatives collected in similar numbers. Regional differences included i) over-representation of graminoids in SA and AU and of annuals in AU; and ii) peak collection during the 1910s in NE, 1980s in SA, and 1990s in AU. Finally, in all regions, a disproportionately large percentage of specimens were collected by a few individuals. These mega-collectors, and their associated preferences and idiosyncrasies, may have shaped patterns of collection bias via ‘founder effects’. Studies using herbarium collections should account for sampling biases and future collecting efforts should avoid compounding these biases.

openCC0Dec 2023View details →
zenodo44/100

Improving the Efficiency of Variationally Enhanced Sampling with Wavelet-Based Bias Potentials

<p>Archive with data supporting the paper &quot;Improving the Efficiency of Variationally Enhanced Sampling with Wavelet-Based Bias Potentials&quot; and the related PhD thesis by B. Pampel</p>

opencc-by-4.0Jan 2022View details →
dryad40/100

Data from: Integrated species distribution models to account for sampling biases and improve range wide occurrence predictions

<p><strong><span>Aim</span></strong></p> <p><span>Species distribution models (SDMs) that integrate presence-only and presence-absence data offer a promising avenue to improve information on species' geographic distributions. The use of such 'integrated SDMs' on a species range-wide extent has been constrained by the often-limited presence-absence data and by the heterogeneous sampling of the presence-only data. Here, we evaluate integrated SDMs for studying species ranges with a novel expert range map-based evaluation. We build a new understanding about how integrated SDMs address issues of estimation accuracy and data deficiency and thereby offer advantages over traditional SDMs.</span></p> <p><strong><span>Location</span></strong></p> <p><span>South and Central America.</span></p> <p><strong><span>Time period</span></strong></p> <p><span>1979-2017.</span></p> <p><strong><span>Major taxa studied</span></strong></p> <p><span>Hummingbirds.</span></p> <p><strong><span>Methods</span></strong></p> <p><span>We build integrated SDMs by linking two observation models – one for each data type – to the same underlying spatial process.</span> <span>We validate SDMs with two schemes: i) cross-validation with presence-absence data and ii) comparison with respect to the species' whole range as defined with IUCN range maps. We also compare models relative to the estimated response curves and compute the association between the benefit of the data integration and the number of presence records in each data set.</span></p> <p><strong><span>Results</span></strong></p> <p><span>The integrated SDM accounting for the spatially varying sampling intensity of the presence-only data was one of the top-performing models in both model validation schemes. Presence-only data alleviated overly large niche estimates, and data integration was beneficial compared to modelling solely presence-only data for species that had few presence points when predicting the species' whole range. On the community level, integrated models improved the species richness prediction.</span></p> <p><strong><span>Main conclusions</span></strong></p> <p><span>Integrated SDMs combining presence-only and presence-absence data are successfully able to borrow strengths from both data types and offer improved predictions of species' ranges. Integrated SDMs can potentially alleviate the impacts of taxonomically and geographically uneven sampling and to leverage the detailed sampling information in presence-absence data.</span></p>

opencc-zeroNov 2023View details →
dryad40/100

Piecewise continuous sampling: a method for minimizing bias and sampling effort for estimated metrics of animal behavior

<p>Capturing qualitative features of animal behavior requires recording occurrences of behavior over time. Continuous sampling is best for capturing brief behaviors, but can be very time consuming. Instantaneous sampling can reduce the amount of labor required, but can miss short-duration behaviors. We therefore synthesized these techniques by continuously sampling during randomly scattered time intervals; a technique we call piecewise continuous sampling. To optimize and test the efficacy of this technique, we collected a continuous behavioral dataset of harvester ant workers, and then we developed a protocol to estimate the amount of sampling time necessary to reconstruct the proportion of time animals spend in different behavioral states. This protocol finds the sample size needed for the variance of the sample to converge on the variation of the population. We then divided this estimated time into equal-duration intervals that were randomly distributed across the entire continuous dataset. Finally, we calculated both time-dependent and time-independent error from this sample. We found that 4 to 16 sampling intervals minimize both types of error simultaneously. This finding was robust to differences in underlying behavior and was validated with simulations, implying that this method could be used for many types of organisms.</p>

opencc-zeroApr 2024View details →
dryad40/100

Sampling bias exaggerates a textbook example of a trophic cascade

<p>Understanding trophic cascades in terrestrial wildlife communities is a major challenge because these systems are difficult to sample properly. We show how a tradition of nonrandom sampling has confounded this understanding in a textbook system (Yellowstone National Park) where carnivore [<em>Canis lupus</em> (wolf)] recovery is associated with a trophic cascade involving changes in herbivore [<em>Cervus canadensis</em> (elk)] behavior and density that promote plant regeneration. Long-term data indicate a practice of sampling only the tallest young plants overestimated regeneration of overstory aspen (<em>Populus tremuloides</em>) by a factor of 3-8 compared to random sampling because it favored plants taller than the preferred browsing height of elk and overlooked non-regenerating aspen stands. Random sampling described a trophic cascade, but it was weaker than the one that nonrandom sampling described. Our findings highlight the critical importance of basic sampling principles (e.g., randomization) for achieving an accurate understanding of trophic cascades in terrestrial wildlife systems.</p>

opencc-zeroNov 2021View details →
zenodo40/100

Understanding Sampling Bias in the Global Heat Flow Compilation

<p>Geothermal heat flow measurements, including calculated weights and geological, tectonic and topographic settings.</p> <p>Paper in review (July 14th, 2022)</p>

opencc-by-4.0Jun 2022View details →
zenodo40/100

How Confidence in Prior Attitudes, Social Tag Popularity, and Source Credibility Shape Confirmation Bias Toward Antidepressants and Psychotherapy in a Representative German Sample: Randomized Controlled Web-Based Study

<p>ABSTRACT</p> <p>Background: In health-related, Web-based information search, people should select information in line with expert (vs nonexpert) information, independent of their prior attitudes and consequent confirmation bias.</p> <p>Objective: This study aimed to investigate confirmation bias in mental health&ndash;related information search, particularly (1) if high confidence worsens confirmation bias, (2) if social tags eliminate the influence of prior attitudes, and (3) if people successfully distinguish high and low source credibility.</p> <p>Methods: In total, 520 participants of a representative sample of the German Web-based population were recruited via a panel company. Among them, 48.1% (250/520) participants completed the fully automated study. Participants provided <em>prior attitudes</em> about antidepressants and psychotherapy. We manipulated (1) <em>confidence</em> in prior attitudes when participants searched for blog posts about the treatment of depression, (2) <em>tag popularity</em> &mdash;either psychotherapy or antidepressant tags were more popular, and (3) <em>source credibility</em> with banners indicating high or low expertise of the tagging community. We measured <em>tag</em> and <em>blog post</em> selection, and <em>treatment</em><em>efficacy ratings</em> after navigation.</p> <p>Results: Tag popularity predicted the proportion of selected antidepressant tags (beta=.44, SE 0.11; <em>P</em>&lt;.001) and blog posts (beta=.46, SE 0.11; <em>P</em>&lt;.001). When confidence was low (&minus;1 SD), participants selected more blog posts consistent with prior attitudes (beta=&minus;.26, SE 0.05; <em>P</em>&lt;.001). Moreover, when confidence was low (&minus;1 SD) and source credibility was high (+1 SD), the efficacy ratings of attitude-consistent treatments increased (beta=.34, SE 0.13; <em>P</em>=.01).</p> <p>Conclusions: We found correlational support for defense motivation account underlying confirmation bias in the mental health&ndash;related search context. That is, participants tended to select information that supported their prior attitudes, which is not in line with the current scientific evidence. Implications for presenting persuasive Web-based information are also discussed.</p> <p>Trial Registration: ClinicalTrials.gov NCT03899168; https://clinicaltrials.gov/ct2/show/NCT03899168 (Archived by WebCite at http://www.webcitation.org/77Nyot3Do)</p> <p>J Med Internet Res 2019;21(4):e11081</p> <p>doi:10.2196/11081</p>

opencc-by-4.0Apr 2019View details →
zenodo40/100

Fig. 6 in Using abundance data to assess the relative role of sampling biases and evolutionary radiations in Upper Muschelkalk ammonoids

Fig. 6. Percent similarity among bins averaged to 1 degree bins. A. om7 interval. B. om8 interval. C. om9 interval. The thicker the line, the greater the similarity between the two cells connected by the line.

opencc-by-4.0Jan 2012View details →
zenodo40/100

Fig. 4 in Using abundance data to assess the relative role of sampling biases and evolutionary radiations in Upper Muschelkalk ammonoids

Fig. 4. Correlations between richness per map and number of localities. A. om7 interval. B. om8 interval. C. om9 interval. The gap in the distribution of points for the om8 interval highlights the discontinuity between a group of maps with few taxa at a few localities and other maps with a large number of localities and high richness.

opencc-by-4.0Jan 2012View details →
zenodo40/100

Fig. 5 in Using abundance data to assess the relative role of sampling biases and evolutionary radiations in Upper Muschelkalk ammonoids

Fig. 5. Rarefaction curves for each interval, based on number of occurrences. The confidence envelope of the species richness for om9 departs significantly from those of om7 and om8 above 50 occurrences, but the significantly higher species−richness of om8 only becomes apparent at sample sizes of around 250 specimens, indicating that a few, rare taxa are boosting richness in the om8 interval.

opencc-by-4.0Jan 2012View details →
zenodo40/100

Fig. 2 in Using abundance data to assess the relative role of sampling biases and evolutionary radiations in Upper Muschelkalk ammonoids

Fig. 2. Distribution of Muschelkalk ammonoid localities used in this study plotted on a map of modern Germany. The overall geographic spread of localities does not change greatly over time.

opencc-by-4.0Jan 2012View details →
zenodo40/100

Fig. 3 in Using abundance data to assess the relative role of sampling biases and evolutionary radiations in Upper Muschelkalk ammonoids

Fig. 3. Correlations between richness per map and number of occurrences. A. om7 interval. B. om8 interval. C. om9 interval.

opencc-by-4.0Jan 2012View details →
zenodo40/100

Fig. 1 in Using abundance data to assess the relative role of sampling biases and evolutionary radiations in Upper Muschelkalk ammonoids

Fig. 1. Chart of stratigraphic interval names and durations for the Muschelkalk of the Germanic Basin with ammonoid immigration events marked (simplified from Klug et al. 2005: fig. 1).

opencc-by-4.0Jan 2012View details →
zenodo40/100

FIGURE 5 in A simulation-based examination of residual diversity estimates as a method of correcting for sampling bias

FIGURE 5. The performance of different implementations of the residual diversity estimate when a specific bias is forced to be the dominant influence. (5.1) Mean Spearman's rho values of four implementations of the RDE using the Smith and McGowan method. PFORM and PTAPH are set at 0.9 to minimise their influence, PLOC is variable. PMIST set at 0.1. (5.2) Mean Spearman's rho values of four implementations of the RDE using the Smith and McGowan method. LOC and PTAPH are set at 0.9 to minimise their influence, PFORM is variable. PMIST set at 0.1. The dashed red line indicates the critical value at p=0.05. Abbreviations as in Table 1.

opencc-by-4.0Nov 2015View details →
zenodo40/100

FIGURE 7 in A simulation-based examination of residual diversity estimates as a method of correcting for sampling bias

FIGURE 7. Sample simulation comparing the results of the taxic, phylogenetic and residual diversity estimates to the true diversity. PFROM, PLOC and PTAPH set at 0.25. PMIST set at 0.1. Black box highlights instance where the Signor Lipps effect has been exaggerated by the PDE; the TDE and RDE both identify the rapid diversity decrease present in the true diversity. Abbreviations as in Table 1.

opencc-by-4.0Nov 2015View details →
zenodo40/100

FIGURE 6 in A simulation-based examination of residual diversity estimates as a method of correcting for sampling bias

FIGURE 6. The performance of the phylogenetic diversity estimate when errors are introduced to the phylogeny. Mean Spearman's rho values of the PDE, TDE and the best performing implementation of the RDE. PLOC, PFORM and PTAPH set at 0.25. PMIST variable. The dashed red line indicates the critical value at p=0.05. Abbreviations as in Table 1.

opencc-by-4.0Nov 2015View details →
zenodo40/100

FIGURE 4 in A simulation-based examination of residual diversity estimates as a method of correcting for sampling bias

FIGURE 4. The performance of different implementations of the residual diversity estimate examining faunas with varying degrees of homogeneity. PFORM, PLOC and PTAPH are set at 0.25. The rate of dispersal is increased relative to the rate of local extinction to increase the homogeneity of the faunas. The dashed red line indicates the critical value at p=0.05. Abbreviations as in Table 1.

opencc-by-4.0Nov 2015View details →
zenodo40/100

FIGURE 1 in A simulation-based examination of residual diversity estimates as a method of correcting for sampling bias

FIGURE 1. An illustration of the taphonomic filter in the simulation, shown applied to a single taxon in a single time bin. The taxon is originally present in every locality in each region it occupies, but the taphonomic filter removes it from randomly selected localities

opencc-by-4.0Nov 2015View details →
zenodo40/100

FIGURE 3 in A simulation-based examination of residual diversity estimates as a method of correcting for sampling bias

FIGURE 3. The performance of different implementations of the residual diversity estimate (RDE) under different sampling regimes. (3.1) Mean Spearman's rho values of four implementations of the RDE using Formations as a proxy, with values of PFORM, PLOC and PTAPH variable but equal. (3.2) Mean Spearman's rho values of four implementations of the RDE using Localities as a proxy. (3.3) Mean Spearman's rho values of four implementations of the RDE, all using the Smith and McGowan method. (3.4) Mean Spearman's rho values of the taxic and phylogenetic diversity estimate compared to those of the optimum implementation of the RDE. The dashed red line indicates the critical value at p=0.05. Abbreviations as in Table 1.

opencc-by-4.0Nov 2015View details →
zenodo40/100

FIGURE 2 in A simulation-based examination of residual diversity estimates as a method of correcting for sampling bias

FIGURE 2. An illustration of how sampling proxies are generated in this simulation. This schematic illustrates which formations and localities in a single time bin contain fossils of at least one species of the simulated clade after application of the taphonomic filter. Formations and localities are removed at random, representing a lack of sampling. Note that the number of clade-bearing formations and localities does not necessarily equal the number of formations and localities sampled, allowing the generation of four sampling proxies.

opencc-by-4.0Nov 2015View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record