Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
484
datasets available to search
ShareScore release 0.9.0
Dataset results
484 results for “Next-Generation Sequencing”
North Carolina Genomic Evaluation by Next-generation Exome Sequencing, 2
ClinicalTrials.gov study NCT03548779. IPD Sharing: YES. Countries: 1. Publications: 61.
Supplemental data from: Next-generation sequencing base calls for mosaic mutations
Open the record for dataset details and reuse information.
Sequences of bacterial and fungal communities by Next-Generation Sequencing (NGS) associated to wall patinas
Open the record for dataset details and reuse information.
Development and application of Faba_bean_130K Targeted Next-Generation Sequencing SNP genotyping platform based on transcriptome sequencing
Open the record for dataset details and reuse information.
A next-generation sequencing study of arthropods in the diet of Laysan Teal (Anas laysanensis)
Open the record for dataset details and reuse information.
Data from: Evaluating culture-free targeted next-generation sequencing for diagnosing drug-resistant tuberculosis: A multicentre clinical study of two end-to-end commercial workflows
Open the record for dataset details and reuse information.
Utilizing next-generation sequencing to identify prey DNA in western North Atlantic grey seal (Halichoerus grypus) diet
<p>Increasing grey seal (<i>Halichoerus grypus</i>) abundance in coastal New England is leading to social, political, economic, and ecological controversies. We studied grey seal feeding habits through next-generation sequencing of prey DNA using 16S amplicons from seal scat (N = 74) collected from a breeding colony on Monomoy Island in Massachusetts, U.S. and report frequency of occurrence and relative read abundance. We also assigned seal sex to scat samples using a revised PCR assay. In contrast to current understanding of grey seal diet from hard parts and fatty acid analysis, we found no significant difference between male and female diet measured by alpha and beta diversity. Overall, we detected 24 prey groups, 18 of which resolved to species. Sand lance (<i>Ammodytes</i> spp.) was the most frequently consumed prey group, with a frequency of occurrence (FO) of 97.3%, consistent with previous studies, but Atlantic menhaden (<i>Brevoortia tyrannus</i>), the second most frequently consumed species (FO = 60.8%), has not been documented in U.S. grey seal diet previously. Our results suggest that a metabarcoding approach to seal food habits can yield important new ecological insights, but that traditional hard parts analysis does not underestimate consumption of Atlantic cod (<i>Gadus morhua; </i>FO =<i> </i>6.7% Gadidae spp.) and salmon (<i>Salmo salar; </i>FO = 0%), two particularly valuable species of concern.</p>
Next-generation sequencing of DNA from resting eggs: signatures of eutrophication in a lake's sediment
<p>Supplementary data</p>
Data from: Targeted next-generation sequencing panels in the diagnosis of Charcot Marie Tooth disease
Objective: To investigate the effectiveness of targeted NGS panels in achieving a molecular diagnosis in CMT and related disorders in a clinical setting Methods: We prospectively enrolled 220 patients from two tertiary referral centres, one in London, UK (n=120) and one in Iowa, US (n=100) in whom a targeted CMT NGS panel had been requested as a diagnostic test. PMP22 duplication/deletion was previously excluded in demyelinating cases. We reviewed the genetic and clinical data upon completion of the diagnostic process. Results: After targeted NGS sequencing a definite molecular diagnosis, defined as a pathogenic or likely pathogenic variant, was reached in 30% of cases (n=67). The diagnostic rate was similar in London (32%) and Iowa (29%). Variants of unknown significance were found in an additional 33% of cases. Mutations in GJB1, MFN2, MPZ accounted for 39% of cases who received genetic confirmation, while the remainder of positive cases had mutations in diverse genes, including SH3TC2, GDAP1, IGHMBP2, LRSAM1, FDG4, GARS and another 12 less common genes. Copy number changes in PMP22, MPZ, MFN2, SH3TC2 and FDG4 were also accurately detected. A definite genetic diagnosis was more likely in cases with an early onset, a positive family history of neuropathy or consanguinity and a demyelinating neuropathy. Conclusions: NGS panels are effective tools in the diagnosis of CMT leading to the genetic confirmation in one third cases negative for PMP22 duplication/deletion, thus highlighting how rarer and previously undiagnosed subtypes represent today a relevant part of the genetic landscape of CMT.
Data from: Parallel tagged next-generation sequencing on pooled samples – a new approach for population genetics in ecology and conservation
Next-generation sequencing (NGS) on pooled samples has already been broadly applied in human medical diagnostics and plant and animal breeding. However, thus far it has been only sparingly employed in ecology and conservation, where it may serve as a useful diagnostic tool for rapid assessment of species genetic diversity and structure at the population level. Here we undertake a comprehensive evaluation of the accuracy, practicality and limitations of parallel tagged amplicon NGS on pooled population samples for estimating species population diversity and structure. We obtained 16S and Cyt b data from 20 populations of Leiopelma hochstetteri, a frog species of conservation concern in New Zealand, using two approaches – parallel tagged NGS on pooled population samples and individual Sanger sequenced samples. Data from each approach were then used to estimate two standard population genetic parameters, nucleotide diversity (π) and population differentiation (FST), that enable population genetic inference in a species conservation context. We found a positive correlation between our two approaches for population genetic estimates, showing that the pooled population NGS approach is a reliable, rapid and appropriate method for population genetic inference in an ecological and conservation context. Our experimental design also allowed us to identify both the strengths and weaknesses of the pooled population NGS approach and outline some guidelines and suggestions that might be considered when planning future projects.
Data from: Biodiversity assessment using next-generation sequencing: comparison of phylogenetic and functional diversity between Nebraska grasslands
Global biodiversity is declining rapidly as a consequence of anthropogenic changes to the environment. Traditional diversity indices such as species richness have been used to assess biodiversity, but recent arguments call for a more comprehensive assessment that includes both phylogenetic and functional diversity (PD and FD, respectively). Many PD metrics have been developed, but few empirical studies have compared metrics across sites with the goal of understanding their application to characterizing biodiversity. In this study, 17 PD metrics, four traditional diversity indices, and one measure of FD were calculated and compared between two Nebraska grasslands. PD metrics were calculated from robust phylogenies estimated from next-generation sequencing data of 45 species. Traditional indices were calculated using species abundance data, and FD was quantified by measuring the phylogenetic signal, K, of specific leaf area (SLA). Results showed that PD metrics and traditional indices were not always correlated, and various PD metrics characterized biodiversity differently. In addition, phylogenies estimated from >80 genes were more robust than single- or dual-gene phylogenies resulting in more reliable PD metrics. K of SLA indicated random trait assembly in all sites. Results suggested that metrics that identify phylogenetic structure and relatedness can provide information to conservation planners about the ability of a community to persist in an unpredictable future. A combination of these results with those of future investigations applying PD and FD metrics to varying communities will support concrete recommendations to conservation planners about how to incorporate these metrics into the selection of priority regions.
Data from: "Genome-wide microsatellite marker development from next-generation sequencing of two non-model bat species impacted by wind turbine mortality: Lasiurus borealis and L. cinereus (Vespertilionidae)" in Genomic Resources Notes accepted 1 October 2013 to 30 November 2013
Tree-roosting bats in the genus Lasiurus are widespread, migratory species that have not been well characterized for population genetic diversity and structure due to a lack of genetic resources. Generating genetic resources in Lasiurus is made pressing by the need for conservation genetic assessments of demographic trends in this genus, which comprise a large percentage of bat mortalities at wind turbine sites across North America. We report on marker development from whole-genome Illumina sequencing of the red bat (Lasirus borealis) and the hoary bat (L. cinereus). We generated paired-end libraries for a single individual of each species, sequenced on the Illumina HiSeq platform. We mapped a total of 46.6 million reads to the Myotis lucifigus reference genome, and used bioinformatics searches to identify tends of thousands of simple sequence repeats (SSRs) distributed across the bat genome. We selected 48 candidate microsatellite loci to develop cross-species primer sequences for Lasiurus, assembled these into multiplex combinations, and tested for amplification and polymorphism levels in a sample of 23 individuals from each of L. borealis and L. cinereus. In total, we identified 42 highly polymorphic loci that could be robustly amplified and scored, the majority of which (39) were also combinable into highly multiplexed assays of 4-8 loci each. The combination of new genomic sequence assemblies, a large set of highly polymorphic microsatellite loci, and the ability to efficiently multiplex represents a significant contribution to the genetic resources available for population and comparative genetic studies of bats.
Data from: A pragmatic approach to the analysis of diets of generalist predators: the use of next-generation sequencing with no blocking probes
Predicting whether a predator is capable of affecting the dynamics of a prey species in the field implies the analysis of the complete diet of the predator, not simply rates of predation on a target taxon. Here, we employed the Ion Torrent next-generation sequencing technology to investigate the diet of a generalist arthropod predator. A complete dietary analysis requires the use of general primers, but these will also amplify the predator unless suppressed using a blocking probe. However, blocking probes can potentially block other species, particularly if they are phylogenetically close. Here, we aimed to demonstrate that enough prey sequence could be obtained without blocking probes. In communities with many predators, this approach obviates the need to design and test numerous blocking primers, thus making analysis of complex community food webs a viable proposition. We applied this approach to the analysis of predation by the linyphiid spider Oedothorax fuscus in an arable field. We obtained over two million raw reads. After discarding the low-quality and predator reads, the libraries still contained over 61 000 prey reads (3% of the raw reads; 6% of reads passing quality control). The libraries were rich in Collembola, Lepidoptera, Diptera and Nematoda. They also contained sequences derived from several spider species and from horticultural pests (aphids). Oedothorax fuscus is common in UK cereal fields, and the results showed that it is exploiting a wide range of prey. Next-generation sequencing using general primers but without blocking probes provided ample sequences for analysis of the prey range of this spider and proved to be a simple and inexpensive approach.
Data from: Discrimination of grasshopper (Orthoptera: Acrididae) diet and niche overlap using next-generation sequencing of gut contents
Species of grasshopper have been divided into three diet classifications based on mandible morphology: forbivorous (specialist on forbs), graminivorous (specialist on grasses), and mixed feeding (broad-scale generalists). For example, Melanoplus bivittatus and Dissosteira carolina are presumed to be broad-scale generalists, Chortophaga viridifasciata is a specialist on grasses, and Melanoplus femurrubrum is a specialist on forbs. These classifications, however, have not been verified in the wild. Multiple specimens of these four species were collected, and diet analysis was performed using DNA metabarcoding of the gut contents. The rbcLa gene region was amplified and sequenced using Illumina MiSeq sequencing. Levins' measure and the Shannon–Wiener measure of niche breadth were calculated using family-level identifications and Morisita's measure of niche overlap was calculated using operational taxonomic units (OTUs). Gut contents confirm both D. carolina and M. bivittatus as generalists and C. viridifasciata as a specialist on grasses. For M. femurrubrum, a high niche breadth was observed and species of grasses were identified in the gut as well as forbs. Niche overlap values did not follow predicted patterns, however, the low values suggest low competition between these species.
Data from: Developing nuclear DNA phylogenetic markers in the angiosperm genus Leucadendron (Proteaceae): a next-generation sequencing transcriptomic approach
Despite the recent advances in generating molecular data, reconstructing species-level phylogenies for non-models groups remains a challenge. The use of a number of independent genes is required to resolve phylogenetic relationships, especially for groups displaying low polymorphism. In such cases, low-copy nuclear exons and non-coding regions, such as 3′ untranslated regions (3′-UTRs) or introns, constitute a potentially interesting source of nuclear DNA variation. Here, we present a methodology meant to identify new nuclear orthologous markers using both public-nucleotide databases and transcriptomic data generated for the group of interest by using next generation sequencing technology. To identify PCR primers for a non-model group, the genus Leucadendron (Proteaceae), we adopted a framework aimed at minimizing the probability of paralogy and maximizing polymorphism. We anchored when possible the right-hand primer into the 3′-UTR and the left-hand primer into the coding region. Seven new nuclear markers emerged from this search strategy, three of those included 3′-UTRs. We further compared the phylogenetic potential between our new markers and the ribosomal internal transcribed spacer region (ITS). The sequenced 3′-UTRs yielded higher polymorphism rates than the ITS region did. We did not find strong incongruences with the phylogenetic signal contained in the ITS region and the seven new designed markers but they strongly improved the phylogeny of the genus Leucadendron. Overall, this methodology is efficient in isolating orthologous loci and is valid for any non-model group given the availability of transcriptomic data.
Data from: Two new phragmotic ant species from Africa: morphology and next-generation sequencing solve a caste association problem in the genus Carebara Westwood
Phragmotic or "door head" ants have evolved independently in several ant genera across the world, but in Africa only one case has been documented until now. Carebara elmenteitae (Patrizi) is known from only a single phragmotic major worker collected from sifted leaf-litter near Lake Elmenteita in Kenya, but here the worker castes of two species collected from Kakamega Forest, a small rainforest in Western Kenya, are studied. Phragmotic major workers were previously identified as Carebara elmenteitae and non-phragmotic major and minor workers were assigned to C. thoracica (Weber). Using evidence of both morphological and next-generation sequencing analysis, it is shown that phragmotic and non-phragmotic workers of the two different species are actually the same and that neither name – C. elmenteitae or C. thoracica – correctly applies to them. Instead, this and another closely related species from Ivory Coast are both morphologically different from C. elmenteitae, and thus they are described as the new species Carebara phragmotica sp. n. and Carebara lilith sp. n.
Data from: The transcriptomics of sympatric dwarf and normal lake whitefish (Coregonus clupeaformis spp., Salmonidae) divergence as revealed by next-generation sequencing
Gene expression divergence is one of the mechanisms thought to be involved in the emergence of incipient species. Next-generation sequencing has become an extremely valuable tool for the study of this process by allowing whole transcriptome sequencing, or RNA-Seq. We have conducted a 454 GS-FLX pyrosequencing experiment in order to refine our understanding of adaptive divergence between dwarf and normal lake whitefish species (Coregonus clupeaformis spp.). The objectives were to: (1) investigate transcriptomic divergence as measured by liver RNA-Seq; (2) test the correlation between divergence in expression and sequence polymorphism and (3) investigate the extent of allelic imbalance. We also compared the results of RNA-seq with those of a previous microarray study performed on the same fish. Following de novo assembly, results showed that normal whitefish over-expressed more contigs associated with protein synthesis while dwarf fish over-expressed more contigs related to energy metabolism, immunity and DNA replication and repair. Moreover, 63 SNPs showed significant allelic imbalance, and this phenomenon prevailed in the recently diverged dwarf whitefish. Results also showed an absence of correlation between gene expression divergence as measured by RNA-Seq and either polymorphism rate or sequence divergence between normal and dwarf whitefish. This study reiterates an important role for gene expression divergence, and provides evidence for allele-specific expression divergence as well as evolutionary decoupling of regulatory and coding sequences in the adaptive divergence of normal and dwarf whitefish. It also demonstrates how next-generation sequencing can lead to a more comprehensive understanding of transcriptomic divergence in a young species pair.
Data from: Exploring evolution and diversity of Chinese Dipterocarpaceae using next-generation sequencing
Tropical forests, a key-category of land ecosystems, are faced with the world's highest levels of habitat conversion and associated biodiversity loss. In tropical Asia, Dipterocarpaceae are one of the economically and ecologically most important tree families, but their genomic diversity and evolution remain understudied, hampered by a lack of available genetic resources. Southern China represents the northern limit for Dipterocarpaceae, and thus changes in habitat ecology, community composition and adaptability to climatic conditions are of particular interest in this group. Phylogenomics is a tool for exploring both biodiversity and evolutionary relationships through space and time using plastome, nuclear and mitochondrial genome. We generated full plastome and Nuclear Ribosomal Cistron (NRC) data for Chinese Dipterocarpaceae species as a first step to improve our understanding of their ecology and evolutionary relationships. We generated the plastome of Dipterocarpus turbinatus, the species with the widest distribution using it as a baseline for comparisons with other taxa. Results showed low level of genomic diversity among analysed range-edge species, and different evolutionary history of the incongruent NRC and plastome data. Genomic resources provided in this study will serve as a starting point for future studies on conservation and sustainable use of these dominant forest taxa, phylogenomics and evolutionary studies.
Data from: Carnivore diet analysis based on next-generation sequencing: application to the leopard cat (Prionailurus bengalensis) in Pakistan
Diet analysis is a prerequisite to fully understand the biology of a species and the functioning of ecosystems. For carnivores, traditional diet analyses mostly rely upon the morphological identification of undigested remains in the feces. Here, we developed a methodology for carnivore diet analyses based on next generation sequencing. We applied this approach to the analysis of the vertebrate component of leopard cat diet in two ecologically distinct regions in northern Pakistan. Despite being a relatively common species with a wide distribution in Asia, little is known about this elusive predator. We analyzed a total of 38 leopard cat feces. After a classical DNA extraction, the DNA extracts were amplified using primers for vertebrates targeting about 100 bp of the mitochondrial 12S rRNA gene, with and without a blocking oligonucleotide specific to the predator sequence. The amplification products were then sequenced on a next generation sequencer. We identified a total of 18 prey taxa, including eight mammals, eight birds, one amphibian, and one fish. In general, our results confirmed that the leopard cat has a very eclectic diet, and feeds mainly on rodents, and particularly on the Muridae family. The DNA-based approach we propose here represents a valuable complement to current conventional methods. It can be applied to other carnivore species with only a slight adjustment relating to the design of the blocking oligonucleotide. It is robust, simple to implement, and allows the possibility of very large-scale analyses.
Data from: Utilizing next-generation sequencing to resolve the backbone of the Core Goodeniaceae and inform future taxonomic and floral form studies
Though considerable progress has been made in inferring phylogenetic relationships of many plant lineages, deep unresolved nodes remain a common problem that can impact downstream efforts, including taxonomic decision-making and character reconstruction. The Core Goodeniaceae is a group affected by this issue: data from the plastid regions trnL-trnF and matK have been insufficient to generate adequate support at key nodes along the backbone of the phylogeny. We performed genome skimming for 24 taxa representing major clades within Core Goodeniaceae. The plastome coding regions (CDS) and nuclear ribosomal repeats (NRR) were assembled and complemented with additional accessions sequenced for nuclear G3PDH and plastid trnL-trnF and matk. The CDS, NRR, and G3PDH alignments were analyzed independently and topology tests were used to detect the alignments' ability to reject alternative topologies. The CDS, NRR, and G3PDH alignments independently supported a Brunonia (Scaevola s.l. (Coopernookia (Goodenia s.l.))) backbone topology, but within Goodenia s.l., the strongly-supported plastome topology (Goodenia A (Goodenia B (Velleia + Goodenia C))) contrasts with the poorly supported nuclear topology ((Goodenia A + Goodenia B) (Velleia + Goodenia C)). A fully resolved and maximally supported topology for Core Goodeniaceae was recovered from the plastome CDS, and there is excellent support for most of the major clades and relationships among them in all alignments. The composition of these seven major clades renders many of the current taxonomic divisions non-monophyletic, prompting us to suggest that Goodenia may be split into several segregate genera.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.