Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

1,940

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

1,940 results for “data sample”

Learn how ShareScore rates datasets ↗
dryad32/100

Data from: The use of MSR (Minimum Sample Richness) for sample assemblage comparisons

Minimum Sample Richness (MSR) is defined as the smallest number of taxa that must be recorded in a sample to achieve a given level of inter-assemblage classification accuracy. MSR is calculated from known or estimated richness and taxonomic similarity. Here we test MSR for strengths and weaknesses by using 167 published mammalian local faunas from the Paleogene and early Neogene of the Query and Liane area (Massif Central, southwestern France), and then apply MSR to 84 Oligo-Miocene faunas from Riversleigh, northwestern Queensland, Australia. In many cases, MSR is able to detect the assemblages in the data set that are potentially too incomplete to be used in a similarity-based comparative taxonomic analysis. The results show that the use of MSR significantly improves the quality of the clustering of fossil assemblages. We conclude that this method can screen sample assemblages that are not representative of their underlying original living communities. Ultimately, it can be used to identify which assemblages require further sampling before being included in a comparative analysis.

opencc-zeroDec 2010View details →
dryad32/100

Data from: Deep COI sequencing of standardized benthic samples unveils overlooked diversity of Jordanian coral reefs in the Northern Red Sea

High-Throughput Sequencing (HTS) of DNA barcodes (metabarcoding), particularly when combined with standardized sampling protocols, is one of the most promising approaches for censusing overlooked cryptic invertebrate communities. We present biodiversity estimates based on sequencing of the cytochrome c oxidase subunit 1 (COI) gene for coral reefs of the Gulf of Aqaba, a semi-enclosed system in the Northern Red Sea. Samples were obtained from standardized sampling devices [Autonomous Reef Monitoring Structures (ARMS)] deployed for 18 months. DNA barcoding of non-sessile specimens >2mm revealed 83 OTUs in six phyla, of which only 25% matched a reference sequence in public databases. Metabarcoding of the 2mm-500μm and sessile bulk fractions revealed 1197 OTUs in 15 animal phyla, of which only 4.9% matched reference barcodes. These results highlight the scarcity of COI data for cryptobenthic organisms of the Red Sea. Compared with data obtained using similar methods, our results suggest that Gulf of Aqaba reefs are less diverse than two Pacific coral reefs but much more diverse than an Atlantic oyster reef at a similar latitude. The standardized approaches used here show promise for establishing baseline data on biodiversity, monitoring the impacts of environmental change, and quantifying patterns of diversity at regional and global scales.

opencc-zeroDec 2015View details →
dryad32/100

Data from: Genotyping-in-Thousands by sequencing (GT-seq) panel development and application to minimally-invasive DNA samples to support studies in molecular ecology

Minimally-invasive sampling (MIS) is widespread in wildlife studies; however, its utility for massively parallel DNA sequencing (MPS) is limited. Poor sample quality and contamination by exogenous DNA can make MIS challenging to use with modern genotyping-by-sequencing approaches, which have been traditionally developed for high-quality DNA sources. Given that MIS is often more appropriate in many contexts, there is a need to make such samples practical for harnessing MPS. Here, we test the ability for Genotyping-in-Thousands by sequencing (GT-seq), a multiplex amplicon sequencing approach, to effectively genotype minimally-invasive cloacal DNA samples collected from the Western Rattlesnake (Crotalus oreganus), a threatened species in British Columbia, Canada. As there was no previous genetic information for this species, an optimized panel of 362 SNPs was selected for use with GT-seq from a de novo restriction-site associated DNA sequencing (RADseq) assembly. Comparisons of genotypes generated within and among RADseq and GT-seq for the same individuals found low rates of genotyping error (GT-seq: 0.50%; RADseq: 0.80%) and discordance (2.57%), the latter likely due to the different genotype calling models employed. GT-seq mean genotype discordance between blood and cloacal swab samples collected from the same individuals was also minimal (1.37%). Estimates of population diversity parameters were similar across GT-seq and RADseq datasets, as were inferred patterns of population structure. Overall, GT-seq can be effectively applied to low quality DNA samples, minimizing the inefficiencies presented by exogenous DNA typically found in minimally-invasive samples and continuing the expansion of molecular ecology and conservation genetics in the genomics era.

opencc-zeroAug 2019View details →
dryad32/100

Data from: Give me a sample of air and I will tell which species are found from your region – molecular identification of fungi from airborne spore samples

Fungi are a megadiverse group of organisms, they play major roles in ecosystem functioning, and are important for human health, food production, and nature conservation. Our knowledge on fungal diversity and fungal ecology is however still very limited, in part because surveying and identifying fungi is time demanding and requires expert knowledge. We present a method that allows anyone to generate a list of fungal species likely to occur in a region of interest, with minimal effort and without requiring taxonomical expertise. The method consists of using a cyclone sampler to acquire fungal spores directly from the air to an Eppendorf tube, and applying DNA barcoding with probabilistic species identification to generate a list of species from the sample. We tested the feasibility of the method by acquiring replicate air samples from different geographical regions within Finland. Our results show that air sampling is adequate for regional-level surveys, with samples collected >100 km apart varying but samples collected <10 km apart not varying in their species composition. The data show marked phenology, and thus that obtaining a representative species list requires aerial sampling that covers the entire fruiting season. In sum, aerial sampling combined with probabilistic molecular species identification offers a highly effective method for generating a species list of airborne dispersed fungi. The method presented here has the potential to revolutionize fungal surveys, as it provides a highly cost-efficient way to include fungi as a part of large-scale biodiversity assessments and monitoring programs.

opencc-zeroDec 2017View details →
zenodo32/100

Sample data for CellPhoneDB

<p>Sample data for the example test of CellPhoneDB, downloaded from https://github.com/Teichlab/cellphonedb</p>

opencc-by-4.0Jun 2021View details →
dryad32/100

Data from: Effects of distance on detectability of Arctic waterfowl using double-observer sampling during helicopter surveys

Aerial survey is an important, widely employed approach for estimating free‐ranging wildlife over large or inaccessible study areas. We studied how a distance covariate influenced probability of double‐observer detections for birds counted during a helicopter survey in Canada's central Arctic. Two observers, one behind the other but visually obscured from each other, counted birds in an incompletely shared field of view to a distance of 200 m. Each observer assigned detections to one of five 40‐m distance bins, guided by semi‐transparent marks on aircraft windows. Detections were recorded with distance bin, taxonomic group, wing‐flapping behavior, and group size. We compared two general model‐based estimation approaches pertinent to sampling wildlife under such situations. One was based on double‐observer methods without distance information, that provide sampling analogous to that required for mark–recapture (MR) estimation of detection probability, urn:x-wiley:20457758:media:ece34824:ece34824-math-0001, and group abundance, urn:x-wiley:20457758:media:ece34824:ece34824-math-0002, along a fixed‐width strip transect. The other method incorporated double‐observer MR with a categorical distance covariate (MRD). A priori, we were concerned that estimators from MR models were compromised by heterogeneity in urn:x-wiley:20457758:media:ece34824:ece34824-math-0003 due to un‐modeled distance information; that is, more distant birds are less likely to be detected by both observers, with the predicted effect that urn:x-wiley:20457758:media:ece34824:ece34824-math-0004 would be biased high, and urn:x-wiley:20457758:media:ece34824:ece34824-math-0005 biased low. We found that, despite increased complexity, MRD models (ΔAICc range: 0–16) fit data far better than MR models (ΔAICc range: 204–258). However, contrary to expectation, the more naïve MR estimators of urn:x-wiley:20457758:media:ece34824:ece34824-math-0006 were biased low in all cases, but only by 2%–5% in most cases. We suspect that this apparently anomalous finding was the result of specific limitations to, and trade‐offs in, visibility by observers on the survey platform used. While MR models provided acceptable point estimates of group abundance, their far higher stranded errors (0%–40%) compared to MRD estimates would compromise ability to detect temporal or spatial differences in abundance. Given improved precision of MRD models relative to MR models, and the possibility of bias when using MR methods from other survey platforms, we recommend avian ecologists use MRD protocols and estimation procedures when surveying Arctic bird populations.

opencc-zeroDec 2018View details →
dryad32/100

Data from: A non-invasive method for sampling the body odour of mammals

1. Olfaction is a central aspect of mammalian communication, providing information about individual attributes such as identity, sex, group membership or genetic quality. Yet, the chemical underpinnings of olfactory cues remain little understood, one of the reasons being the difficulty in obtaining high quality samples for chemical analysis. 2. In the present study we adjusted and evaluated the use of thermal desorption (TD) tubes, commonly used in plant metabolomic and environmental studies, for non-invasive sampling of mammalian body odour. We obtained chemical profiles of meerkat (Suricata suricatta) body odour samples using TD tubes analysed with gas chromatography – mass spectrometry (GC-MS). 3. TD tubes captured a wide range of volatile and semi-volatile organic compounds including compounds likely originating from the target animals. Adjustment of sampling parameters (distance, volume, flow rate, interruption of sampling) to increase the feasibility for a non-invasive application yielded samples of adequate quality. However, to minimize the variability between samples, sampling parameters should be kept constant and samples should be collected when no conspecifics are close-by. 4. The method was sensitive enough to pick up population differences in the chemical profiles of two captive groups of meerkats, demonstrating its applicability to biological questions. With sufficiently habituated animals, the method is applicable non-invasively, allowing short- and long-term studies on a wide range of questions, including e.g. chemical signatures of kinship, diet, individual health or reproductive state.

opencc-zeroDec 2016View details →
dryad32/100

Data from: Evaluation of potential overwinter mortality of age-0 walleye and appropriate age-1 sampling gear

Potential recruitment of age-0 Walleye Sander vitreus to adults is often indexed by the relative abundance of age-0 individuals during their first summer or fall. However, relationships between age-0 and adult Walleye abundance are often weak or nonsignificant in many waters. Overwinter mortality during the first year of life has been hypothesized as an important limitation to Walleye recruitment in lakes, but limited evidence of such mortality exists, likely due to difficulties in sampling age-1 Walleye during spring. The objectives of this study were to: 1) compare results from nighttime electrofishing to index relative abundance of age-1 Walleyes with relative abundance indices of minifyke nets in four eastern South Dakota lakes; 2) determine whether size-selective mortality was occurring in those four lakes; and 3) if size-selective mortality was occurring in these lakes, determine whether that mortality was attributed to body condition. We sampled four natural lakes in eastern South Dakota 2 wk after ice-off in 2013 and 2014. Precision of nighttime electrofishing (coefficient of variation = 216.6) was greater than that estimated for minifyke nets (coefficient of variation = 338.5) across both years. We detected no differences in length-frequency distributions of collected spring age-1 Walleye between the two gears. Age-0 fall relative abundance indices from electrofishing were significantly greater (P &lt; 0.01) than spring age-1 nighttime electrofishing indices of relative abundance at three of the four study lakes, indicating that overwinter mortality may occur at a substantial rate during the first year of life for Walleye in these systems. Quantile–quantile regression plots showed evidence of size-selective mortality in three of four lakes sampled. However, body condition of age-0 Walleye appeared to have little to no influence on overwinter mortality. Instead, we suggest that smaller-sized walleye may be more vulnerable to overwinter predation. Collectively, these results provide evidence of previously hypothesized overwinter mortality within the first year for Walleye and indicate possibilities for indexing potential adult recruitment of Walleye just after this critical period.

opencc-zeroDec 2017View details →
dryad32/100

Data from: Noninvasive sampling reveals population genetic structure in the Royle's pika, Ochotona roylei, in the western Himalaya

Understanding population genetic structure of climate-sensitive herbivore species is important as it provides useful insights on how shifts in environmental conditions can alter their distribution and abundance. Herbivore responses to the environment can have a strong indirect cascading effect on community structure. This is particularly important for Royle's pika (Lagomorpha: Ochotona roylei), a herbivorous talus-dwelling species in alpine ecosystem, which forms a major prey base for many carnivores in the Himalayan arc. In this study, we used seven polymorphic microsatellite loci to detect evidence for recent changes in genetic diversity and population structure in Royle's pika across five locations sampled between 8 km to 160 km apart in the western Himalaya. Using four clustering approaches, we found the presence of significant contemporary genetic structure in Royle's pika populations. The detected genetic structure could be primarily attributed to the landscape features in alpine habitat (e.g. wide lowland valleys, rivers) that may act as semi-permeable barriers to gene flow and distribution of food plants, which are key determinants in spatial distribution of herbivores. Pika showed low inbreeding coefficients (FIS) and a high level of pairwise relatedness for individuals within 1km suggesting low dispersal abilities of talus-dwelling pikas. We have found evidence of a recent population bottleneck, possibly due to effects of environmental disturbances (e.g. snow melting patterns or thermal stress). Our results reveal significant evidence of isolation by distance in genetic differentiation (FST range = 0.04−0.19). This is the first population genetics study on Royle's pika, which helps to address evolutionary consequences of climate change which are expected to significantly affect the distribution and population dynamics in this talus dwelling species.

opencc-zeroDec 2017View details →
dryad32/100

Data from: Scrutinizing key steps for reliable metabarcoding of environmental samples

1. Metabarcoding of environmental samples has many challenges and limitations that require carefully considered laboratory and analysis pipelines to ensure reliable results. We explore how decisions regarding study design, laboratory work and bioinformatic processing affect the final results, and provide guidelines for reliable study of environmental samples. 2. We evaluate the performance of four primer sets targeting COI and 16S regions characterising arthropod diversity in bat faecal samples, and investigate how metabarcoding results are affected by parameters including: i) number of PCR replicates per sample, ii) sequencing depth, iii) PCR replicate processing strategy (i.e. either additively, by combining the sequences obtained from the PCR replicates, or restrictively, by only retaining sequences that occur in multiple PCR replicates for each sample), iv) minimum copy number for sequences to be retained, v) chimera removal, and vi) similarity thresholds for OTU clustering. Lastly, we measure within- and between-taxa dissimilarities when using sequences from public databases to determine the most appropriate thresholds for OTU clustering and taxonomy assignment. 3. Our results show that the use of multiple primer sets reduces taxonomic biases and increases taxonomic coverage. Taxonomic profiles resulting from each primer set are principally affected by how many PCR replicates are carried out per sample and how sequences are filtered across them, the sequence copy number threshold and the OTU clustering threshold. We also report considerable diversity differences between PCR replicates from each sample. Sequencing depth increases the dissimilarity between PCR replicates unless the bioinformatic strategies to remove allegedly artefactual sequences are adjusted according to the number of analysed sequences. Finally, we show that the appropriate identity thresholds for OTU clustering and taxonomy assignment differ between target markers. 4. Metabarcoding of complex environmental samples ideally requires i) investigation of whether more than one primer sets targeting the same taxonomic group is needed to offset the effect of primer biases, ii) more than one PCR replicate per sample, iii) bioinformatic processing approaches of sequences that balance diversity detection with removal of artificial sequences, and iv) empirical selection of OTU clustering and taxonomy assignment thresholds tailored to each genetic marker and the obtained taxa.

opencc-zeroDec 2016View details →
zenodo32/100

Sample data for sequencing reads alignment

<p>These data are used for learning sequencing reads alignment and cluster usage</p>

opencc-by-4.0Jun 2021View details →
zenodo32/100

Sample data for COMUNET

<p>Sample data for the tutorial of COMUNET, downloaded from https://github.com/ScialdoneLab/COMUNET</p>

opencc-by-4.0Jun 2021View details →
zenodo32/100

Sample data for SingleCellSignalR

<p>Sample data for the tutorial of SingleCellSignalR, downloaded from https://github.com/SCA-IRCM/Demo</p>

opencc-by-4.0Jun 2021View details →
zenodo32/100

Sample data for NATMI

<p>Sample data for the tutorial of NATMI, downloaded from <a href="https://github.com/forrest-lab/NATMI">https://github.com/forrest-lab/NATMI</a></p>

opencc-by-4.0Jun 2021View details →
zenodo32/100

Sample data for celltalker

<p>Sample data for the tutorial of celltalker, downloaded from https://www.ncbi.nlm.nih.gov/geo/query/acc.cgi?acc=GSE139324</p>

opencc-by-4.0Jul 2021View details →
dryad32/100

Image-based automated species identification: Can virtual data augmentation overcome problems of insufficient sampling?

<p></p><p>Automated species identification and delimitation is challenging, particularly in rare and thus often scarcely sampled species, which do not allow sufficient discrimination of infraspecific versus interspecific variation. Typical problems arising from either low or exaggerated interspecific morphological differentiation are best met by automated methods of machine learning that learn efficient and effective species identification from training samples. However, limited infraspecific sampling remains a key challenge also in machine learning.</p> <p>In this study, we assessed whether a data augmentation approach may help to overcome the problem of scarce training data in automated visual species identification. The stepwise augmentation of data comprised image rotation as well as visual and virtual augmentation. The visual data augmentation applies classic approaches of data augmentation and generation of artificial images using a Generative Adversarial Networks (GAN) approach. Descriptive feature vectors are derived from bottleneck features of a VGG-16 convolutional neural network (CNN) that are then stepwise reduced in dimensionality using Global Average Pooling and PCA to prevent overfitting. Finally, data augmentation employs synthetic additional sampling in feature space by an oversampling algorithm in vector space (SMOTE). Applied on four different image datasets, which include scarab beetle genitalia (Pleophylla, Schizonycha) as well as wing patterns of bees (Osmia) and cattleheart butterflies (Parides), our augmentation approach outperformed a deep learning baseline approach by means of resulting identification accuracy with non-augmented data as well as a traditional 2D morphometric approach (Procrustes analysis of scarab beetle genitalia).</p><p></p>

opencc-zeroJul 2021View details →
dryad32/100

Data from: High-throughput SNP genotyping of historical and modern samples of five bird species via sequence capture of ultraconserved elements

Sample availability limits population genetics research on many species, especially taxa from regions with high diversity. However, many such species are well represented in museum collections assembled before the molecular era. Development of techniques to recover genetic data from these invaluable specimens will benefit biodiversity science. Using a mixture of freshly preserved and historical tissue samples, and a sequence capture probe set targeting &gt;5000 loci, we produced high-confidence genotype calls on thousands of single nucleotide polymorphisms (SNPs) in each of five South-East Asian bird species and their close relatives (N = 27–43). On average, 66.2% of the reads mapped to the pseudo-reference genome of each species. Of these mapped reads, an average of 52.7% was identified as PCR or optical duplicates. We achieved deeper effective sequencing for historical samples (122.7×) compared to modern samples (23.5×). The number of nucleotide sites with at least 8× sequencing depth was high, with averages ranging from 0.89 × 106 bp (Arachnothera, modern samples) to 1.98 × 106 bp (Stachyris, modern samples). Linear regression revealed that the amount of sequence data obtained from each historical sample (represented by per cent of the pseudo-reference genome recovered with ≥8× sequencing depth) was positively and significantly (P ≤ 0.013) related to how recently the sample was collected. We observed characteristic post-mortem damage in the DNA of historical samples. However, we were able to reduce the error rate significantly by truncating ends of reads during read mapping (local alignment) and conducting stringent SNP and genotype filtering.

opencc-zeroDec 2015View details →
zenodo32/100

Machine Learning-Assisted Sampling of SERS Substrates Improves Data Collection Efficiency: raw data and code

<p>Raw datasets and media accompanying the manuscript:&nbsp;<strong>Machine Learning-Assisted Sampling of SERS Substrates Improves Data Collection Efficiency</strong>: data, published in <em>Applied Spectroscopy </em>in 2021</p>

opencc-zeroJul 2021View details →
dryad32/100

Data from: Rearing and sampling methods for estimating spruce budworm development rates at constant temperatures

<p>We describe an experimental protocol for measuring the response of spruce budworm post-diapause larval development to temperature. This protocol is specifically designed to include measurements of development near their upper and lower thermal thresholds. The application of this protocol to a laboratory colony allowed for the first experimental evidence that spruce budworm larval development occurs at temperatures as low as 5 ºC and as high as 35 ºC and provides data to estimate development rates at temperatures from 5–35 ºC in 5 ºC increments. Our protocol is also designed to minimize mortality near the thermal development thresholds thus allowing for multi-generational studies. We observed developmental plasticity in larvae reared at constant temperatures, particularly the occurrence of up to 42% of some individuals requiring only five instars to complete development, compared to the expected six instars. An occurrence that exhibited no clear relation to temperature. While this protocol is specifically designed for spruce budworm, it provides a template for the study of other species' developmental responses to temperature.</p>

opencc-zeroJul 2021View details →
zenodo32/100

Optimal Spectral Sampling Forward Model Test Data

<p>Binary and ASCII files for tests of the OSS forward model application.</p>

opencc-by-4.0Aug 2021View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record