Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
50
datasets available to search
ShareScore release 0.9.0
Dataset results
50 results for “integrated samples”
Water Quality Sampling - integrated measurements for the Virginia Coast, 1992-2025
This dataset contains information about the aquatic environment along two transects that run from inlet to the mainland shore on the southern part of the Delmarva Peninsula since 1992. It has a large number of columns (111) that integrate information on water column and benthic measurements. Frequency of sampling varies from monthly to quarterly, and not all variables are necessarily measured on the same dates. However, all the data from a given date and location appears on a single line of the dataset.
CCE LTER process cruise, in the California Current region, event log records including date, time, position and activity for use in post-cruise data integration based on co-sampling indexes. From 2006 to 2019 CCE LTER used a locally developed event logging system. During P2107, CCE LTER started to utilize the R2R Event Logger on UNOL ships, 2006 - 2024 (ongoing).
The event logger program developed and maintained by the California Cooperative Oceanic Fisheries Investigations, SIO, program is used aboard CCE LTER process cruises to create indexes with temporal, spatial and activity information for post-cruise data integration. The event log is configured aboard the ship for the recording of sampling events by both ship crew personnel on the bridge, and research personnel in the lab. The event log is processed post-cruise to correct for various errors.
Data from: Integrated species distribution models to account for sampling biases and improve range wide occurrence predictions
<p><strong><span>Aim</span></strong></p> <p><span>Species distribution models (SDMs) that integrate presence-only and presence-absence data offer a promising avenue to improve information on species' geographic distributions. The use of such 'integrated SDMs' on a species range-wide extent has been constrained by the often-limited presence-absence data and by the heterogeneous sampling of the presence-only data. Here, we evaluate integrated SDMs for studying species ranges with a novel expert range map-based evaluation. We build a new understanding about how integrated SDMs address issues of estimation accuracy and data deficiency and thereby offer advantages over traditional SDMs.</span></p> <p><strong><span>Location</span></strong></p> <p><span>South and Central America.</span></p> <p><strong><span>Time period</span></strong></p> <p><span>1979-2017.</span></p> <p><strong><span>Major taxa studied</span></strong></p> <p><span>Hummingbirds.</span></p> <p><strong><span>Methods</span></strong></p> <p><span>We build integrated SDMs by linking two observation models – one for each data type – to the same underlying spatial process.</span> <span>We validate SDMs with two schemes: i) cross-validation with presence-absence data and ii) comparison with respect to the species' whole range as defined with IUCN range maps. We also compare models relative to the estimated response curves and compute the association between the benefit of the data integration and the number of presence records in each data set.</span></p> <p><strong><span>Results</span></strong></p> <p><span>The integrated SDM accounting for the spatially varying sampling intensity of the presence-only data was one of the top-performing models in both model validation schemes. Presence-only data alleviated overly large niche estimates, and data integration was beneficial compared to modelling solely presence-only data for species that had few presence points when predicting the species' whole range. On the community level, integrated models improved the species richness prediction.</span></p> <p><strong><span>Main conclusions</span></strong></p> <p><span>Integrated SDMs combining presence-only and presence-absence data are successfully able to borrow strengths from both data types and offer improved predictions of species' ranges. Integrated SDMs can potentially alleviate the impacts of taxonomically and geographically uneven sampling and to leverage the detailed sampling information in presence-absence data.</span></p>
Figure 3. Data Integration Sample-A Proposed Data Driven Architecture for Cardiology Network Application
<p>Pentaho Data Integration has implemented a metadata-driven approach where you<br> only specify the data you want integrated, but you do not specify the way you want it done.<br> One of the most important advantages of Pentaho is that one can create complex<br> transformations and jobs in a graphical, drag-and-drop environment without having to create<br> proprietary custom code that will work only with some proprietary application.</p>
Data from: Integrated species distribution models to account for sampling biases and improve range wide occurrence predictions
Open the record for dataset details and reuse information.
A novel method to assess the integrity of frozen archival DNA samples: Alpha-diversity ratios of short and long-read 16S rRNA gene sequences
Open the record for dataset details and reuse information.
A large-scale assessment of ant diversity across the Brazilian Amazon Basin: integrating geographic, ecological, and morphological drivers of sampling bias
<p>Tropical ecosystems are often biodiversity hotspots, and invertebrates represent the main underrepresented component of diversity in large-scale analyses. This problem is partly related to the scarcity of data widely available to conduct these studies and the lack of systematic organization of knowledge about invertebrates' distributions in biodiversity hotspots. Here, we introduce and analyze a comprehensive data compilation of Amazonian ant diversity. Using records from 1817 to 2020 from both published and unpublished sources, we describe the diversity and distribution of ant species in the Brazilian Amazon Basin. Further, using high-definition images and data from taxonomic publications, we build a comprehensive database of morphological traits for the ant species that occur in the region. In total, we recorded 1,067 nominal species in the Brazilian Amazon Basin, with sampling locations strongly biased by access routes, urban centers, research institutions, and major infrastructure projects. Large areas where ant sampling is non-existent represent about 52% of the basin and are concentrated mainly in the North, Southeastern, and Western Brazilian Amazon. We found that distance to roads is the main driver of ant sampling in the Amazon. Contrary to our expectations, morphological traits had lower predictive power in predicting sample bias than purely geographic variables. However, when geographic predictors were controlled, habitat stratum and traits contribute to explain the remaining variance. More species were recorded in better-sampled areas, but species richness estimation models suggest that areas in South Amazonian edge forests are associated with especially high species richness. Our results represent the first trait-based, large-scale study for insects in Amazonian forests and a starting point for macroecological studies focusing on insect diversity in the Amazon Basin.</p>
A large-scale assessment of ant diversity across the Brazilian Amazon Basin: integrating geographic, ecological, and morphological drivers of sampling bias
Open the record for dataset details and reuse information.
Examination of sample size determination in integration studies based on the coefficient of variation (ICV)
<p>Although there are various indices available for calculating morphological integration, the integration coefficient of variation (ICV) is most suited for assessing magnitudes of integration within and between morphological variance/covariance (V/CV) matrices. However, it is currently not known what the effects of varying sample sizes are on the reliable estimation of distributions of ICV scores.<b> </b>In this regard, the effects of varying sample size on ICV was examined by simulating parameter V/CV matrices with varying underlying magnitudes of average trait correlation (r<sup>2</sup>). ICV distributions were generated using a trait resampling protocol for various sample sizes (11 through 150) within various parameter r<sup>2 </sup>values. Next, empirical r<sup>2</sup> values were calculated based on data from 22 skeletal elements of 40 <i>Macaca fascicularis</i> specimens to examine whether the results from the simulation corresponded to real biological data. Mean ICV scores of various sample sizes were compared using Mann-Whitney U tests to examine which minimum sample sizes are required to reliably calculate mean ICV.<b> </b>Mann-Whitney U test results based on the simulated data showed that a sample size of 51 may be sufficient even for relatively low r<sup>2 </sup>values of 0.05. The empirical macaque data showed that 30‒40 individuals may be sufficient to reliably calculate mean ICV scores across skeletal elements.<b> </b>Our results correspond closely with previous assessments by Cheverud and colleagues that argued that a sample size of 40 is necessary to accurately estimate the structure of V/CV matrices.</p>
Data from: How diverse is Mitopus morio? Integrative taxonomy detects cryptic species in a small-scale sample of a widespread harvestman
Mitopus morio is a widespread harvestman species occurring in most of Europe and in moderate and cold-moderate zones of Asia and North America. The species is characterized by extreme variability in body size and leg length. As leg length is correlated with habitat temperature, M. morio has been considered as an example of Allen's rule. Recently, observations for a single location in Tyrol, Austria, indicated the absence of mating between short- and long-legged individuals. This study examines for signs of putative cryptic species in M. morio using an integrative approach that combines mating trials, amplified fragment length polymorphism whole-genome scans, mitochondrial sequences and morphometrics. The mating trials did not corroborate the initial hypothesis of a reproductive barrier associated with leg size. Both types of genetic data revealed the existence of three distinct groups, in line with the mating results but largely unrelated to leg morphology and geographical origin of specimens. Morphometric characters supporting the findings of the other disciplines were identified using a supervised approach. We infer from all data together the existence of strongly diverged cryptic lineages among the analysed individuals, cautiously interpret them as three sympatric species and conclude that in these harvestmen Allen's rule applies at different levels. Due to the unexpected amount of differentiation found within a geographical scale very small compared with the distribution of M. morio, we suggest a thorough revision of the genus prior to formal taxonomic changes. Our case study underlines the general applicability of the integrative taxonomic protocol used and highlights the relevance of several rationales implemented in the protocol.
FIGURE 7. Sample sites and climatic suitability for all Alpinobombus species combined estimated using Maxent from climate variables for the year 2000 in The arctic and alpine bumblebees of the subgenus Alpinobombus revised from integrative assessment of species' gene coalescents and morphology (Hymenoptera, Apidae, Bombus)
FIGURE 7. Sample sites and climatic suitability for all Alpinobombus species combined estimated using Maxent from climate variables for the year 2000. Bright yellow and brown areas show where the logistic prediction of occurrence is p> 0.5 (the outlier records with exceptionally low probabilities of suitability were excluded in step-wise iterations); brown spots show all consistent Alpinobombus site records; '?' shows the five records that are in climatically outlying sites. Polar projection, North Pole (starred) at the centre of the map, international boundaries and the Arctic Circle shown as narrow grey lines.
APPENDIX. List of sequenced specimens of Triphosa, with identification, Sampling sites collecting data, Accession numbers, and process ID in BOLD database. Data taken from BOLD and generated by Axel Hausmann (1); Bernd Müller (2); Dirk Stadie (3); Iva Mihoci 4); Marco Infusino, Stefano Scalercio (5); Norbert Poell (6); Wanke et al. (7). in An integrative taxonomic revision of the genus Triphosa Stephens, 1829 (Geometridae: Larentiinae) in the Middle East and Central Asia, with description of two new species
APPENDIX. List of sequenced specimens of Triphosa, with identification, Sampling sites collecting data, Accession numbers, and process ID in BOLD database. Data taken from BOLD and generated by Axel Hausmann (1); Bernd Müller (2); Dirk Stadie (3); Iva Mihoci 4); Marco Infusino, Stefano Scalercio (5); Norbert Poell (6); Wanke et al. (7).
Data from: How diverse is Mitopus morio? Integrative taxonomy detects cryptic species in a small-scale sample of a widespread harvestman
Open the record for dataset details and reuse information.
Examination of sample size determination in integration studies based on the coefficient of variation (ICV)
Open the record for dataset details and reuse information.
Data from: How many more? Sample size determination in studies of morphological integration and evolvability
The variational properties of living organisms are an important component of current evolutionary theory. As a consequence, researchers working on the field of multivariate evolution have increasingly used integration and evolvability statistics as a way of capturing the potentially complex patterns of trait association and their effects over evolutionary trajectories. Little attention has been paid, however, to the cascading effects that inaccurate estimates of trait covariance have on these widely used evolutionary statistics. Here, we analyze the relationship between sampling effort and inaccuracy in evolvability and integration statistics calculated from 10-trait matrices with varying patterns of covariation and magnitudes of integration. We then extrapolate our initial approach to different numbers of traits and different magnitudes of integration and estimate general equations relating the inaccuracy of the statistics of interest to sampling effort. We validate our equations using a dataset of cranial traits, and use them to make sample size recommendations. Our results suggest that highly inaccurate estimates of evolvability and integration statistics resulting from small sample sizes are likely common in the literature, given the sampling effort necessary to properly estimate them. We also show that patterns of covariation have no effect on the sampling properties of these statistics, but overall magnitudes of integration interact with sample size and lead to varying degrees of bias, imprecision, and inaccuracy. Finally, we provide R functions that can be used to calculate recommended sample sizes or to simply estimate the level of inaccuracy that should be expected in these statistics, given a sampling design.
Data from: Integration of 168,000 samples reveals global patterns of the human gut microbiome
<p>The data and code in this archive was used to generate the analyses and figures from the revised version of "Integration of 168,000 samples reveals global patterns of the human gut microbiome." The first version is available as <a href="https://doi.org/10.1101/2023.10.11.560955">a bioRxiv preprint</a>. Code for the text-mining work is <a href="https://github.com/krishnanlab/microbiome-metadata-annotation">available on GitHub</a>.</p> <p> </p>
Data from: On the sampling design of spatially explicit integrated population models
Open the record for dataset details and reuse information.
Data from: How many more? Sample size determination in studies of morphological integration and evolvability
Open the record for dataset details and reuse information.
Data from: ANDe™ : a fully integrated environmental DNA sampling system
Open the record for dataset details and reuse information.
Integrative taxonomy and geographic sampling underlie successful species delimitation
Open the record for dataset details and reuse information.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.