Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
71
datasets available to search
ShareScore release 0.9.0
Dataset results
71 results for “Next Generation DNA Sequencing”
Genome-wide SNP discovery in native American and Hungarian Robinia pseudoacacia genotypes using next-generation double-digest restriction-site-associated DNA sequencing (ddRAD-Seq)
<p>Initial filtered ddRADseq dataset with highly variable SNP markers from native American and Hungarian <em>Robinia pseudoacacia</em> L. individuals</p>
Next-generation Sequencing Data Associated with "Genome Editing Outcomes Reveal Mycobacterial NucS Participates in a Short-Patch Repair of DNA Mismatches"
Open the record for dataset details and reuse information.
Utilizing next-generation sequencing to identify prey DNA in western North Atlantic grey seal (Halichoerus grypus) diet
<p>Increasing grey seal (<i>Halichoerus grypus</i>) abundance in coastal New England is leading to social, political, economic, and ecological controversies. We studied grey seal feeding habits through next-generation sequencing of prey DNA using 16S amplicons from seal scat (N = 74) collected from a breeding colony on Monomoy Island in Massachusetts, U.S. and report frequency of occurrence and relative read abundance. We also assigned seal sex to scat samples using a revised PCR assay. In contrast to current understanding of grey seal diet from hard parts and fatty acid analysis, we found no significant difference between male and female diet measured by alpha and beta diversity. Overall, we detected 24 prey groups, 18 of which resolved to species. Sand lance (<i>Ammodytes</i> spp.) was the most frequently consumed prey group, with a frequency of occurrence (FO) of 97.3%, consistent with previous studies, but Atlantic menhaden (<i>Brevoortia tyrannus</i>), the second most frequently consumed species (FO = 60.8%), has not been documented in U.S. grey seal diet previously. Our results suggest that a metabarcoding approach to seal food habits can yield important new ecological insights, but that traditional hard parts analysis does not underestimate consumption of Atlantic cod (<i>Gadus morhua; </i>FO =<i> </i>6.7% Gadidae spp.) and salmon (<i>Salmo salar; </i>FO = 0%), two particularly valuable species of concern.</p>
Next-generation sequencing of DNA from resting eggs: signatures of eutrophication in a lake's sediment
<p>Supplementary data</p>
Data from: Developing nuclear DNA phylogenetic markers in the angiosperm genus Leucadendron (Proteaceae): a next-generation sequencing transcriptomic approach
Despite the recent advances in generating molecular data, reconstructing species-level phylogenies for non-models groups remains a challenge. The use of a number of independent genes is required to resolve phylogenetic relationships, especially for groups displaying low polymorphism. In such cases, low-copy nuclear exons and non-coding regions, such as 3′ untranslated regions (3′-UTRs) or introns, constitute a potentially interesting source of nuclear DNA variation. Here, we present a methodology meant to identify new nuclear orthologous markers using both public-nucleotide databases and transcriptomic data generated for the group of interest by using next generation sequencing technology. To identify PCR primers for a non-model group, the genus Leucadendron (Proteaceae), we adopted a framework aimed at minimizing the probability of paralogy and maximizing polymorphism. We anchored when possible the right-hand primer into the 3′-UTR and the left-hand primer into the coding region. Seven new nuclear markers emerged from this search strategy, three of those included 3′-UTRs. We further compared the phylogenetic potential between our new markers and the ribosomal internal transcribed spacer region (ITS). The sequenced 3′-UTRs yielded higher polymorphism rates than the ITS region did. We did not find strong incongruences with the phylogenetic signal contained in the ITS region and the seven new designed markers but they strongly improved the phylogeny of the genus Leucadendron. Overall, this methodology is efficient in isolating orthologous loci and is valid for any non-model group given the availability of transcriptomic data.
Supplementary material 1 from: Binh HT, Ngoc NV, Tagane S, Toyama H, Mase K, Mitsuyuki C, Strijk JS, Suyama Y, Yahara T (2018) A taxonomic study of Quercus langbianensis complex based on morphology, and DNA barcodes of classic and next generation sequences. PhytoKeys 95: 37-70. https://doi.org/10.3897/phytokeys.95.21126
Figure S1, S2 : Explanation note: Figure S1. Bayesian phylogeny of 29 samples of Quercus and one Trigonobalanus (outgroup) based on ITS sequences. Branches are labeled with posterior probabilites. Figure S2. Bayesian phylogeny of 29 samples of Quercus and one Trigonobalanus (outgroup) based on concatenated rbcL and matK sequences. Branches are labeled with posterior probabilities.
Utilizing next-generation sequencing to identify prey DNA in western North Atlantic grey seal (Halichoerus grypus) diet
Open the record for dataset details and reuse information.
Data from: Developing nuclear DNA phylogenetic markers in the angiosperm genus Leucadendron (Proteaceae): a next-generation sequencing transcriptomic approach
Open the record for dataset details and reuse information.
Raw sequence of: A comparative analysis of spider prey spectra analyzed through the next-generation sequencing of individual and mixed DNA samples
Open the record for dataset details and reuse information.
Utilizing field collected insects for next generation sequencing: effects of sampling, storage, and DNA extraction methods
Open the record for dataset details and reuse information.
Data from: DNA barcodes from century-old type specimens using next generation sequencing
Type specimens have high scientific importance because they provide the only certain connection between the application of a Linnean name and a physical specimen. Many other individuals may have been identified as a particular species, but their linkage to the taxon concept is inferential. Because type specimens are often more than a century old and have experienced conditions unfavorable for DNA preservation, success in sequence recovery has been uncertain. The present study addresses this challenge by employing next generation sequencing (NGS) to recover sequences for the barcode region of the cytochrome c oxidase 1 gene from small amounts of template DNA. DNA quality was first screened in more than 1800 century-old type specimens of Lepidoptera by attempting to recover 164bp and 94bp reads via Sanger sequencing. This analysis permitted the assignment of each specimen to one of three DNA quality categories – high (164bp sequence), medium (94bp sequence), or low (no sequence). Ten specimens from each category were subsequently analyzed via a PCR-based NGS protocol requiring very little template DNA. It recovered sequence information from all specimens with average read lengths ranging from 458bp to 610bp for the three DNA categories. By sequencing ten specimens in each NGS run, costs were similar to Sanger analysis. Future increases in the number of specimens processed in each run promise substantial reductions in cost, making it possible to anticipate a future where barcode sequences are available from most type specimens.
Data from: PCR-Free enrichment of mitochondrial DNA from human blood and cell lines for high quality next-generation DNA sequencing
Recent advances in sequencing technology allow for accurate detection of mitochondrial sequence variants, even those in low abundance at heteroplasmic sites. Considerable sequencing cost savings can be achieved by enriching samples for mitochondrial (relative to nuclear) DNA. Reduction in nuclear DNA (nDNA) content can also help to avoid false positive variants resulting from nuclear mitochondrial sequences (numts). We isolate intact mitochondrial organelles from both human cell lines and blood components using two separate methods: a magnetic bead binding protocol and differential centrifugation. DNA is extracted and further enriched for mitochondrial DNA (mtDNA) by an enzyme digest. Only 1 ng of the purified DNA is necessary for library preparation and next generation sequence (NGS) analysis. Enrichment methods are assessed and compared using mtDNA (versus nDNA) content as a metric, measured by using real-time quantitative PCR and NGS read analysis. Among the various strategies examined, the optimal is differential centrifugation isolation followed by exonuclease digest. This strategy yields >35% mtDNA reads in blood and cell lines, which corresponds to hundreds-fold enrichment over baseline. The strategy also avoids false variant calls that, as we show, can be induced by the long-range PCR approaches that are the current standard in enrichment procedures. This optimization procedure allows mtDNA enrichment for efficient and accurate massively parallel sequencing, enabling NGS from samples with small amounts of starting material. This will decrease costs by increasing the number of samples that may be multiplexed, ultimately facilitating efforts to better understand mitochondria-related diseases.
Data from: A long PCR based approach for DNA enrichment prior to next-generation sequencing for systematic studies
Premise of the study: We present an alternative approach for molecular systematic studies that combines long PCR and next-generation sequencing (NGS). Our approach can be used to generate templates from any DNA source for NGS. Here we test our approach by amplifying complete chloroplast genomes and we present a set of 58 potentially universal primers for angiosperms to do so. Additionally, this approach is likely to be particularly useful for nuclear regions. Methods and Results: Chloroplast genomes of 30 species across angiosperms were amplified to test our approach. Amplification success varied depending on whether PCR conditions were optimized for a given taxon. To further test our approach, some amplicons were sequenced on an Illumina HiSeq 2000. Conclusions: Although here we tested this approach by sequencing plastomes, long PCR amplicons could be generated using DNA from any genome, expanding the possibilities of this approach for molecular systematic studies.
Mullus surmuletus environmental DNA intraspecific metabarcoding Next-Generation Sequencing data
<p>Four 250-liter aquariums were bleached clean one day prior to be used (filled with seawater; fish transfer) in Montpellier (France). Seawater collected by the French Research Institute for Exploitation of the Sea at Palavas-les-Flots (France) was first stored in a 1,000 L tank for two weeks, under UV treatment to avoid any contamination. The aquariums were then filled with 120 L of this water. Each aquarium had a closed-circuit water circulation and was equipped with an air bubbles exhauster in a tube that brought up the water on a neutral synthetic foam filter. The aquariums were thus oxygenated and the coarsest suspended matter was filtered out. The remaining seawater in the tank was used as a negative control (Aquarium 1). Nine to eleven fish were added to each of the four aquariums (Fig. 1). The aquarium water was sampled six hours after introducing the fish into the aquariums using an Athena peristaltic pump (SPYGEN, Le Bourget-du-Lac, France) with a nominal flow of 1.0 L/min to filter 30 L, and VigiDNA 0.22 μm crossflow filtration capsules (SPYGEN) with disposable sterile tubing. After filtration, 80 mL of CL1 conservation buffer (SPYGEN) was added before storing the samples at ambient temperature.</p> <p> </p> <p>We reanalyzed here two eDNA samples of 30 L replicate each, collected in the Mediterranean Sea, at Banyuls (France, coordinates: 42.41568, 3.17110) and Calvi (France, coordinates: 42.62964, 8.89161) published in a previous metabarcoding analysis and known to contain <em>M. surmuletus</em> sequences (detected with the metabarcode teleo 12S) (Boulanger <em>et al.</em> 2021). These two Mediterranean eDNA samples were amplified and sequenced using the primers developed for this study and then analyzed using the best-performing pipeline as determined by our evaluation. These two samples were used as proof of concept of the possibility to estimate within site variability in real conditions.</p> <p> </p> <p> </p> <p>DNA extraction and amplification from eDNA samples were performed by the company SPYGEN (Le Bourget du Lac, France) in separate, dedicated rooms following the protocol described by Polanco Fernández <em>et al.</em> (2020). The amplification was performed in a final volume of 25 μL including 1 U of AmpliTaq Gold DNA Polymerase (Applied Biosystems, Foster City, CA, USA), 10 mM of Tris-HCl, 50 mM of KCl, 2.5 mM of MgCl2, 0.2 mM of each dNTP, 0.2 μM of each primer, 0.2 μg/μL of bovine serum albumin (Roche Diagnostics, Basel, Switzerland) and 3 μL of DNA template. The PCR mixture was denatured at 95°C for 10 min, followed by 50 cycles of 30 s at 95°C, 30 s at 47°C and 1 min at 72°C and a final elongation step at 72°C for 7 min. The primers were 5’-labelled with an eight-nucleotide tag unique to each DNA sample, allowing each sequence to be assigned to the corresponding sample during the sequence analysis. Twelve replicate PCRs were run per sample. Two libraries were prepared using the MetaFast protocol (Fasteris 2020, <a href="https://www.fasteris.com/dna/">https://www.fasteris.com/dna/</a>) and the sequencing was performed by Fasteris (Geneva, Switzerland) on two separate runs on an Illumina MiSeq (2x250 bp) (Illumina, San Diego, CA, USA) and the Miseq Kit v3 (Illumina) following the manufacturer’s instructions. Two negative extraction controls and one negative PCR control (12 replicates of ultrapure water) were amplified and sequenced to monitor for possible contaminants (Polanco Fernández <em>et al.</em>, 2020).</p> <p> </p> <p> </p> <p> </p>
Figure 4 from: Binh HT, Ngoc NV, Tagane S, Toyama H, Mase K, Mitsuyuki C, Strijk JS, Suyama Y, Yahara T (2018) A taxonomic study of Quercus langbianensis complex based on morphology, and DNA barcodes of classic and next generation sequences. PhytoKeys 95: 37-70. https://doi.org/10.3897/phytokeys.95.21126
Figure 4 Comparison of Q. langbianensis complex between NJ tree (left, Clade M3 of Fig. 3) and Bayesian tree (right: Clade 2 of Fig. 2).
Figure 1 from: Binh HT, Ngoc NV, Tagane S, Toyama H, Mase K, Mitsuyuki C, Strijk JS, Suyama Y, Yahara T (2018) A taxonomic study of Quercus langbianensis complex based on morphology, and DNA barcodes of classic and next generation sequences. PhytoKeys 95: 37-70. https://doi.org/10.3897/phytokeys.95.21126
Figure 1 Collection sites in Vietnam and Cambodia in this study, including eight national parks, four nature reserves and two conservation areas.
Figure 7 from: Binh HT, Ngoc NV, Tagane S, Toyama H, Mase K, Mitsuyuki C, Strijk JS, Suyama Y, Yahara T (2018) A taxonomic study of Quercus langbianensis complex based on morphology, and DNA barcodes of classic and next generation sequences. PhytoKeys 95: 37-70. https://doi.org/10.3897/phytokeys.95.21126
Figure 7 Quercus bidoupensis Binh & Ngoc. A Leafy twig B Abaxial side of mature leaf C, D Side view and base view of the cupule, respectively E Inside of cupule F Nut. Materials: A–F from Tagane et al. V4328.
Figure 11 from: Binh HT, Ngoc NV, Tagane S, Toyama H, Mase K, Mitsuyuki C, Strijk JS, Suyama Y, Yahara T (2018) A taxonomic study of Quercus langbianensis complex based on morphology, and DNA barcodes of classic and next generation sequences. PhytoKeys 95: 37-70. https://doi.org/10.3897/phytokeys.95.21126
Figure 11 Quercus donnaiensis A.Camus. A Leafy twig B Infructescence, young fruits and abaxial side of mature leaf C Dried specimen. Materials: A, B from Tagane S., Wai J. V4398 C from Ngoc et al. V3208.
Figure 13 from: Binh HT, Ngoc NV, Tagane S, Toyama H, Mase K, Mitsuyuki C, Strijk JS, Suyama Y, Yahara T (2018) A taxonomic study of Quercus langbianensis complex based on morphology, and DNA barcodes of classic and next generation sequences. PhytoKeys 95: 37-70. https://doi.org/10.3897/phytokeys.95.21126
Figure 13 Quercus langbianensis Hickel & A.Camus. A Leafy twig B Abaxial side of mature leaf C Infructescence and mature fruits D Apex of the nut E Basal scar of the nut F Inside of cupule. Materials: A, B from Tagane et al. V 4165 C–F from Tagane et al. V4166.
Figure 10 from: Binh HT, Ngoc NV, Tagane S, Toyama H, Mase K, Mitsuyuki C, Strijk JS, Suyama Y, Yahara T (2018) A taxonomic study of Quercus langbianensis complex based on morphology, and DNA barcodes of classic and next generation sequences. PhytoKeys 95: 37-70. https://doi.org/10.3897/phytokeys.95.21126
Figure 10 Quercus camusiae Trel. ex Hickel & A.Camus. A Branch with young fruit, B. Infructescence and young fruits C, D Abaxial side of young and mature leaf, E. Dried specimen. Materials: A–D from Tagane et al. V342 E from Toyama et al. V2173.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.