Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
595
datasets available to search
ShareScore release 0.9.0
Dataset results
595 results for “High-throughput sequencing”
Data from: Digital fragment analysis of short tandem repeats by high-throughput amplicon sequencing
Open the record for dataset details and reuse information.
Data from: High-throughput sequencing of ancient plant and mammal DNA preserved in herbivore middens
Open the record for dataset details and reuse information.
Data from: Development, validation and high-throughput analysis of sequence markers in nonmodel species
Open the record for dataset details and reuse information.
Data from: High-throughput sequencing of transposable element insertions suggests adaptive evolution of the invasive Asian Tiger Mosquito towards temperate environments
Open the record for dataset details and reuse information.
Data from: Seasonal diversity and dynamics of haptophytes in the Skagerrak, Norway, explored by high-throughput sequencing
Open the record for dataset details and reuse information.
Data from: Diet assessment of two land planarian species using high-throughput sequencing data
Open the record for dataset details and reuse information.
A high-throughput skim-sequencing approach for genotyping, dosage estimation and identifying translocations
Open the record for dataset details and reuse information.
Data from: Who’s for dinner? High-throughput sequencing reveals bat diet differentiation in a biodiversity hotspot where prey taxonomy is largely undescribed
Open the record for dataset details and reuse information.
Data from: Use of genotyping-by-sequencing data to develop a high-throughput and multi-functional SNP panel for conservation applications in Pacific lamprey
Open the record for dataset details and reuse information.
Data from: High-throughput sequencing of Bacillus anthracis in France: investigating genome diversity and population structure using whole-genome SNP discovery
Open the record for dataset details and reuse information.
Data from: A phylogeny of birds based on over 1,500 loci collected by target enrichment and high-throughput sequencing
Open the record for dataset details and reuse information.
Data from: High-throughput amplicon sequencing of rRNA genes requires a copy number correction to accurately reflect the effects of management practices on soil nematode community structure
Open the record for dataset details and reuse information.
Data from: More affordable and effective noninvasive SNP genotyping using high-throughput amplicon sequencing
<p>Non-invasive genotyping methods have become key elements of wildlife research over the last two decades, but their widespread adoption is limited by high costs, low success rates, and high error rates. <span>The information lost when genotyping success is low may lead to decreased precision in animal population densities, which could misguide conservation and management actions.</span> <span>Single nucleotide polymorphisms (SNPs) provide a promising alternative to traditionally used microsatellites as SNPs allow amplification of shorter DNA fragments, are less prone to genotyping errors, and produce results that are easily shared among laboratories.</span> Here, we outline a detailed protocol for cost-effective and accurate noninvasive SNP genotyping using multiplexed amplicon sequencing optimized for degraded DNA. <span>We validated this method for individual identification by genotyping 216 scats, 18 hairs and 15 tissues from coyotes (<i>Canis latrans</i>) using 26 SNPs. </span><a name="_Hlk33181599">Our genotyping success rate for scat samples was 93%, and 100% for hair and tissue, representing a substantial increase compared to previous microsatellite-based studies while remaining at a low cost of under $5 per PCR replicate (excluding labor). </a>The accuracy of the genotypes was further corroborated in that genotypes from scats matching known, GPS-collared coyotes were always located within the territory of the known individual. We also show that different levels of multiplexing produced similar results, but that PCR product cleanup strategies can have substantial effects on genotyping success. By making noninvasive genotyping more affordable, accurate, and efficient, this research may allow for a substantial increase in the use of noninvasive methods to monitor and conserve free-ranging wildlife populations.</p>
Data from: Robust DNA isolation and high-throughput sequencing library construction for herbarium specimens
Herbaria are an invaluable source of plant material that can be used in a variety of biological studies. The use of herbarium specimens is associated with a number of challenges including sample preservation quality, degraded DNA, and destructive sampling of rare specimens. In order to more effectively use herbarium material in large sequencing projects, a dependable and scalable method of DNA isolation and library preparation is needed. This paper demonstrates a robust, beginning-to-end protocol for DNA isolation and high-throughput library construction from herbarium specimens that does not require modification for individual samples. This protocol is tailored for low quality dried plant material and takes advantage of existing methods by optimizing tissue grinding, modifying library size selection, and introducing an optional reamplification step for low yield libraries. Reamplification of low yield DNA libraries can rescue samples derived from irreplaceable and potentially valuable herbarium specimens, negating the need for additional destructive sampling and without introducing discernible sequencing bias for common phylogenetic applications. The protocol has been tested on hundreds of grass species, but is expected to be adaptable for use in other plant lineages after verification. This protocol can be limited by extremely degraded DNA, where fragments do not exist in the desired size range, and by secondary metabolites present in some plant material that inhibit clean DNA isolation. Overall, this protocol introduces a fast and comprehensive method that allows for DNA isolation and library preparation of 24 samples in less than 13 hours, with only 8 hours of active hands-on time with minimal modifications.
Data from: Prevention, diagnosis, and treatment of high-throughput sequencing data pathologies
High Throughput Sequencing (HTS) technologies generate millions of sequence reads from DNA/RNA molecules rapidly and cost-effectively, enabling single investigator laboratories to address a variety of "omics" questions in non-model organisms, fundamentally changing the way genomic approaches are used to advance biological research. One major challenge posed by HTS is the complexity and difficulty of data quality control (QC). While QC issues associated with sample isolation, library preparation, and sequencing are well known and protocols for their handling are widely available, the QC of the actual sequence reads generated by HTS is often overlooked. HTS-generated sequence reads can contain various errors, biases, and artefacts whose identification and amelioration can greatly impact subsequent data analysis. However, a systematic survey on QC procedures for HTS data is still lacking. In this review, we begin by presenting standard "health check-up" QC procedures recommended for HTS datasets and establishing what "healthy" HTS data look like. We next proceed by classifying errors, biases and artifacts present in HTS data into three major types of "pathologies", discussing their causes and symptoms, and illustrating with examples their diagnosis and impact on downstream analyses. We conclude this review by offering examples of successful "treatment" protocols and recommendations on standard practices and treatment options. Notwithstanding the speed with which HTS technologies–and consequently their pathologies–change, we argue that careful QC of HTS data is an important–yet often neglected–aspect of their application in molecular ecology, and lay the groundwork for developing a HTS data QC "best practices" guide.
Data from: Universal and blocking primer mismatches limit the use of high-throughput DNA sequencing for the quantitative metabarcoding of arthropods
The quantification of the biological diversity in environmental samples using high-throughput DNA sequencing is hindered by the PCR bias caused by variable primer–template mismatches of the individual species. In some dietary studies, there is the added problem that samples are enriched with predator DNA, so often a predator-specific blocking oligonucleotide is used to alleviate the problem. However, specific blocking oligonucleotides could coblock nontarget species to some degree. Here, we accurately estimate the extent of the PCR biases induced by universal and blocking primers on a mock community prepared with DNA of twelve species of terrestrial arthropods. We also compare universal and blocking primer biases with those induced by variable annealing temperature and number of PCR cycles. The results show that reads of all species were recovered after PCR enrichment at our control conditions (no blocking oligonucleotide, 45 °C annealing temperature and 40 cycles) and high-throughput sequencing. They also show that the four factors considered biased the final proportions of the species to some degree. Among these factors, the number of primer–template mismatches of each species had a disproportionate effect (up to five orders of magnitude) on the amplification efficiency. In particular, the number of primer–template mismatches explained most of the variation (~3/4) in the amplification efficiency of the species. The effect of blocking oligonucleotide concentration on nontarget species relative abundance was also significant, but less important (below one order of magnitude). Considering the results reported here, the quantitative potential of the technique is limited, and only qualitative results (the species list) are reliable, at least when targeting the barcoding COI region.
Data from: Defining the alloreactive T cell repertoire using high-throughput sequencing of mixed lymphocyte reaction culture
The cellular immune response is the most important mediator of allograft rejection and is a major barrier to transplant tolerance. Delineation of the depth and breadth of the alloreactive T cell repertoire and subsequent application of the technology to the clinic may improve patient outcomes. As a first step toward this, we have used MLR and high-throughput sequencing to characterize the alloreactive T cell repertoire in healthy adults at baseline and 3 months later. Our results demonstrate that thousands of T cell clones proliferate in MLR, and that the alloreactive repertoire is dominated by relatively high-abundance T cell clones. This clonal make up is consistently reproducible across replicates and across a span of three months. These results indicate that our technology is sensitive and that the alloreactive TCR repertoire is broad and stable over time. We anticipate that application of this approach to track donor-reactive clones may positively impact clinical management of transplant patients.
Data from: Sequence Capture using PCR-generated Probes (SCPP): a cost-effective method of targeted high-throughput sequencing for non-model organisms
Recent advances in high-throughput sequencing library preparation and subgenomic enrichment methods have opened new avenues for population genetics and phylogenetics of non-model organisms. To multiplex large numbers of indexed samples while sequencing predominantly orthologous, targeted regions of the genome, we propose modifications to an existing, in-solution capture that utilizes PCR products as target probes to enrich library pools for the genomic subset of interest. The sequence capture using PCR-generated probes (SCPP) protocol requires no specialized equipment, is highly flexible, and significantly reduces experimental costs for projects where a modest scale of genetic data is optimal (25-100 genomic loci). Our alterations enable application of this method across a wider phylogenetic range of taxa and result in higher capture efficiencies and coverage at each locus. Efficient and consistent capture over multiple SCPP experiments and at various phylogenetic distances is demonstrated, extending the utility of this method to both phylogeographic and phylogenomic studies.
Figure 5 from: Siddique AB, Khokon AM, Unterseher M (2017) What do we learn from cultures in the omics age? High-throughput sequencing and cultivation of leaf-inhabiting endophytes from beech (Fagus sylvatica L.) revealed complementary community composition but similar correlations with local habitat conditions. MycoKeys 20: 1-16. https://doi.org/10.3897/mycokeys.20.11265
Figure 5 - Relative abundance distribution of fungal leaf-inhabiting endophytes of beech among the five main trophic guilds as revealed by analysis with FUNGuild (Nguyen et al. 2016). A compares the two localities for each trophic guild on the basis of Illumina data B compares the two localities for each trophic guild on the basis of cultivation data.
Figure 4 from: Siddique AB, Khokon AM, Unterseher M (2017) What do we learn from cultures in the omics age? High-throughput sequencing and cultivation of leaf-inhabiting endophytes from beech (Fagus sylvatica L.) revealed complementary community composition but similar correlations with local habitat conditions. MycoKeys 20: 1-16. https://doi.org/10.3897/mycokeys.20.11265
Figure 4 - Relative abundance of fungal leaf-inhabiting endophytes of beech among the five main trophic guilds as revealed by analysis with FUNGuild (Nguyen et al. 2016). A compares the two methods for each trophic guild and unassigned data. B displays the trophic guilds and unassigned taxa for Illumina data, C for cultivation data. Abbreviations in [B and C]: U = Unassigned, P = Pathotrophs, PSa = Patho-Saprotrophs, PSy = Patho-Symbiotrophs, Sa = Saprotrophs, Sy = Symbiotrophs
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.