Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
484
datasets available to search
ShareScore release 0.9.0
Dataset results
484 results for “Next-Generation Sequencing”
Data from: Inferring the mode of origin of polyploid species from next-generation sequence data
Many eucaryote organisms are polyploid. However, despite their importance, evolutionary inference of polyploid origins and modes of inheritance has been limited by a need for analyses of allele segregation at multiple loci using crosses. The increasing availability of sequence data for non-model species now allows the application of established approaches for the analysis of genomic data in polyploids. Here, we ask whether approximate Bayesian computation (ABC), applied to realistic traditional and next-generation sequence data, allows correct inference of the evolutionary and demographic history of polyploids. Using simulations, we evaluate the robustness of evolutionary inference by ABC for tetraploid species as a function of the number of individuals and loci sampled, and the presence or absence of an outgroup. We find that ABC adequately retrieves the recent evolutionary history of polyploid species on the basis of both old and new sequencing technologies. Application of ABC to sequence data from diploid and polyploid species of the plant genus Capsella confirms its utility. Our analysis strongly supports an allopolyploid origin of C. bursa-pastoris about 80,000 years ago. This conclusion runs contrary to previous findings based on the same dataset but using an alternative approach and is in agreement with recent findings based on whole-genome sequencing. Our results indicate that ABC is a promising and powerful method for revealing the evolution of polyploid species, without the need to attribute alleles to a homeologous chromosome pair. The approach can readily be extended to more complex scenarios involving higher ploidy levels.
Data from: PCR-Free enrichment of mitochondrial DNA from human blood and cell lines for high quality next-generation DNA sequencing
Recent advances in sequencing technology allow for accurate detection of mitochondrial sequence variants, even those in low abundance at heteroplasmic sites. Considerable sequencing cost savings can be achieved by enriching samples for mitochondrial (relative to nuclear) DNA. Reduction in nuclear DNA (nDNA) content can also help to avoid false positive variants resulting from nuclear mitochondrial sequences (numts). We isolate intact mitochondrial organelles from both human cell lines and blood components using two separate methods: a magnetic bead binding protocol and differential centrifugation. DNA is extracted and further enriched for mitochondrial DNA (mtDNA) by an enzyme digest. Only 1 ng of the purified DNA is necessary for library preparation and next generation sequence (NGS) analysis. Enrichment methods are assessed and compared using mtDNA (versus nDNA) content as a metric, measured by using real-time quantitative PCR and NGS read analysis. Among the various strategies examined, the optimal is differential centrifugation isolation followed by exonuclease digest. This strategy yields >35% mtDNA reads in blood and cell lines, which corresponds to hundreds-fold enrichment over baseline. The strategy also avoids false variant calls that, as we show, can be induced by the long-range PCR approaches that are the current standard in enrichment procedures. This optimization procedure allows mtDNA enrichment for efficient and accurate massively parallel sequencing, enabling NGS from samples with small amounts of starting material. This will decrease costs by increasing the number of samples that may be multiplexed, ultimately facilitating efforts to better understand mitochondria-related diseases.
Data from: aTRAM - automated Target Restricted Assembly Method: a fast method for assembling loci across divergent taxa from next-generation sequencing data
Background: Assembling genes from next-generation sequencing data is not only time consuming but computationally difficult, particularly for taxa without a closely related reference genome. Assembling even a draft genome using de novo approaches can take days, even on a powerful computer, and these assemblies typically require data from a variety of genomic libraries. Here we describe software that will alleviate these issues by rapidly assembling genes from distantly related taxa using a single library of paired-end reads: aTRAM, automated Target Restricted Assembly Method. The aTRAM pipeline uses a reference sequence, BLAST, and an iterative approach to target and locally assemble the genes of interest. Results: Our results demonstrate that aTRAM rapidly assembles genes across distantly related taxa. In comparative tests with a closely related taxon, aTRAM assembled the same sequence as reference-based and de novo approaches taking on average < 1 min per gene. As a test case with divergent sequences, we assembled >1,000 genes from six taxa ranging from 25 – 110 million years divergent from the reference taxon. The gene recovery was between 97 – 99% from each taxon. Conclusions: aTRAM can quickly assemble genes across distantly-related taxa, obviating the need for draft genome assembly of all taxa of interest. Because aTRAM uses a targeted approach, loci can be assembled in minutes depending on the size of the target. Our results suggest that this software will be useful in rapidly assembling genes for phylogenomic projects covering a wide taxonomic range, as well as other applications. The software is freely available: http://www.github.com/juliema/aTRAM
Data from: Unforeseen consequences of excluding missing data from next-generation sequences: simulation study of RAD sequences
There is a lack of consensus on how next-generation sequence data should be considered for phylogenetic and phylogeographic estimates, with some studies excluding loci with missing data, while others include them, even when sequences are missing from a large number of individuals. Here we use simulations, focusing specifically on RAD sequences, to highlight some of the unforeseen consequence of excluding missing data from next-generation sequencing. Specifically, we show that in addition to the obvious effects associated with reducing the amount of data used to make historical inferences, the decisions we make about missing data (such as the minimum number of individuals with a sequence for a locus to be included in the study) also impact the types of loci sampled for a study. In particular, as the tolerance for missing data becomes more stringent, the mutational spectrum represented in the sampled loci becomes truncated such that loci with the highest mutation rates are disproportionately excluded. This effect is exacerbated further by factors involved in the preparation of the genomic library (i.e., the use of reduced representation libraries, as well as the coverage) and the taxonomic diversity represented in the library (i.e., the level of divergence among the individuals). We demonstrate that the intuitive appeals about being conservative by removing loci may be misguided.
Data from: Species level phylogeny and polyploid relationships in Hordeum (Poaceae) inferred by next-generation sequencing and in-silico cloning of multiple nuclear loci
Polyploidization is an important speciation mechanism in the barley genus Hordeum. To analyze evolutionary changes after allopolyploidization, knowledge of parental relationships is essential. One chloroplast and 12 nuclear single-copy loci were amplified by polymerase chain reaction (PCR) in all Hordeum plus six out-group species. Amplicons from each of 96 individuals were pooled, sheared, labeled with individual-specific barcodes and sequenced in a single run on a 454 platform. Reference sequences were obtained by cloning and Sanger sequencing of all loci for nine supplementary individuals. The 454 reads were assembled into contigs representing the 13 loci and, for polyploids, also homoeologues. Phylogenetic analyses were conducted for all loci separately and for a concatenated data matrix of all loci. For diploid taxa, a Bayesian concordance analysis and a coalescent-based dated species tree was inferred from all gene trees. Chloroplast matK was used to determine the maternal parent in allopolyploid taxa. The relative performance of different multilocus analyses in the presence of incomplete lineage sorting and hybridization was also assessed. The resulting multilocus phylogeny reveals for the first time species phylogeny and progenitor-derivative relationships of all di- and polyploid Hordeum taxa within a single analysis. Our study proves that it is possible to obtain a multilocus species-level phylogeny for di- and polyploid taxa by combining PCR with next-generation sequencing, without cloning and without creating a heavy load of sequence data.
Data from: A long PCR based approach for DNA enrichment prior to next-generation sequencing for systematic studies
Premise of the study: We present an alternative approach for molecular systematic studies that combines long PCR and next-generation sequencing (NGS). Our approach can be used to generate templates from any DNA source for NGS. Here we test our approach by amplifying complete chloroplast genomes and we present a set of 58 potentially universal primers for angiosperms to do so. Additionally, this approach is likely to be particularly useful for nuclear regions. Methods and Results: Chloroplast genomes of 30 species across angiosperms were amplified to test our approach. Amplification success varied depending on whether PCR conditions were optimized for a given taxon. To further test our approach, some amplicons were sequenced on an Illumina HiSeq 2000. Conclusions: Although here we tested this approach by sequencing plastomes, long PCR amplicons could be generated using DNA from any genome, expanding the possibilities of this approach for molecular systematic studies.
Mullus surmuletus environmental DNA intraspecific metabarcoding Next-Generation Sequencing data
<p>Four 250-liter aquariums were bleached clean one day prior to be used (filled with seawater; fish transfer) in Montpellier (France). Seawater collected by the French Research Institute for Exploitation of the Sea at Palavas-les-Flots (France) was first stored in a 1,000 L tank for two weeks, under UV treatment to avoid any contamination. The aquariums were then filled with 120 L of this water. Each aquarium had a closed-circuit water circulation and was equipped with an air bubbles exhauster in a tube that brought up the water on a neutral synthetic foam filter. The aquariums were thus oxygenated and the coarsest suspended matter was filtered out. The remaining seawater in the tank was used as a negative control (Aquarium 1). Nine to eleven fish were added to each of the four aquariums (Fig. 1). The aquarium water was sampled six hours after introducing the fish into the aquariums using an Athena peristaltic pump (SPYGEN, Le Bourget-du-Lac, France) with a nominal flow of 1.0 L/min to filter 30 L, and VigiDNA 0.22 μm crossflow filtration capsules (SPYGEN) with disposable sterile tubing. After filtration, 80 mL of CL1 conservation buffer (SPYGEN) was added before storing the samples at ambient temperature.</p> <p> </p> <p>We reanalyzed here two eDNA samples of 30 L replicate each, collected in the Mediterranean Sea, at Banyuls (France, coordinates: 42.41568, 3.17110) and Calvi (France, coordinates: 42.62964, 8.89161) published in a previous metabarcoding analysis and known to contain <em>M. surmuletus</em> sequences (detected with the metabarcode teleo 12S) (Boulanger <em>et al.</em> 2021). These two Mediterranean eDNA samples were amplified and sequenced using the primers developed for this study and then analyzed using the best-performing pipeline as determined by our evaluation. These two samples were used as proof of concept of the possibility to estimate within site variability in real conditions.</p> <p> </p> <p> </p> <p>DNA extraction and amplification from eDNA samples were performed by the company SPYGEN (Le Bourget du Lac, France) in separate, dedicated rooms following the protocol described by Polanco Fernández <em>et al.</em> (2020). The amplification was performed in a final volume of 25 μL including 1 U of AmpliTaq Gold DNA Polymerase (Applied Biosystems, Foster City, CA, USA), 10 mM of Tris-HCl, 50 mM of KCl, 2.5 mM of MgCl2, 0.2 mM of each dNTP, 0.2 μM of each primer, 0.2 μg/μL of bovine serum albumin (Roche Diagnostics, Basel, Switzerland) and 3 μL of DNA template. The PCR mixture was denatured at 95°C for 10 min, followed by 50 cycles of 30 s at 95°C, 30 s at 47°C and 1 min at 72°C and a final elongation step at 72°C for 7 min. The primers were 5’-labelled with an eight-nucleotide tag unique to each DNA sample, allowing each sequence to be assigned to the corresponding sample during the sequence analysis. Twelve replicate PCRs were run per sample. Two libraries were prepared using the MetaFast protocol (Fasteris 2020, <a href="https://www.fasteris.com/dna/">https://www.fasteris.com/dna/</a>) and the sequencing was performed by Fasteris (Geneva, Switzerland) on two separate runs on an Illumina MiSeq (2x250 bp) (Illumina, San Diego, CA, USA) and the Miseq Kit v3 (Illumina) following the manufacturer’s instructions. Two negative extraction controls and one negative PCR control (12 replicates of ultrapure water) were amplified and sequenced to monitor for possible contaminants (Polanco Fernández <em>et al.</em>, 2020).</p> <p> </p> <p> </p> <p> </p>
Data from: Next-generation museum genomics: phylogenetic relationships among palpimanoid spiders using sequence capture techniques (Araneae: Palpimanoidea)
Historical museum specimens are invaluable for morphological and taxonomic research, but typically the DNA is degraded making traditional sequencing techniques difficult to impossible for many specimens. Recent advances in Next-Generation Sequencing, specifically target capture, makes use of short fragment sizes typical of degraded DNA, opening up the possibilities for gathering genomic data from museum specimens. This study uses museum specimens and recent target capture sequencing techniques to sequence both Ultra-Conserved Elements (UCE) and exonic regions for lineages that span the modern spiders, Araneomorphae, with a focus on Palpimanoidea. While many previous studies have used target capture techniques on dried museum specimens (for example, skins, pinned insects), this study includes specimens that were collected over the last two decades and stored in 70% ethanol at room temperature. Our findings support the utility of target capture methods for examining deep relationships within Araneomorphae: sequences from both UCE and exonic loci were important for resolving relationships; a monophyletic Palpimanoidea was recovered in many analyses and there was strong support for family and generic-level palpimanoid relationships. Ancestral character state reconstructions reveal that the highly modified carapace observed in mecysmaucheniids and archaeids has evolved independently.
Supplementary material 1 from: Dong Y, Zhu L, Bai Y, Ou Y, Wang C (2016) Complete mitochondrial genomes of two flat-backed millipedes by next-generation sequencing (Diplopoda, Polydesmida). ZooKeys 637: 1-20. https://doi.org/10.3897/zookeys.637.9909
Supplementary tables :
Figure 4 from: Dong Y, Zhu L, Bai Y, Ou Y, Wang C (2016) Complete mitochondrial genomes of two flat-backed millipedes by next-generation sequencing (Diplopoda, Polydesmida). ZooKeys 637: 1-20. https://doi.org/10.3897/zookeys.637.9909
Figure 4 - Putative secondary structures of the 22 tRNA genes of Xystodesmus sp. Watson-Crick base-pairing is indicated by solid lines, and G–T pairs are indicated with plus signs.
Figure 3 from: Dong Y, Zhu L, Bai Y, Ou Y, Wang C (2016) Complete mitochondrial genomes of two flat-backed millipedes by next-generation sequencing (Diplopoda, Polydesmida). ZooKeys 637: 1-20. https://doi.org/10.3897/zookeys.637.9909
Figure 3 - Putative secondary structures of the 22 tRNA genes of Asiomorpha coarctata. Watson-Crick base-pairing is indicated by solid lines, and G–T pairs are indicated with plus signs.
Figure 2 from: Dong Y, Zhu L, Bai Y, Ou Y, Wang C (2016) Complete mitochondrial genomes of two flat-backed millipedes by next-generation sequencing (Diplopoda, Polydesmida). ZooKeys 637: 1-20. https://doi.org/10.3897/zookeys.637.9909
Figure 2 - Sequences of the non-coding region in Asiomorpha coarctata, primary structures of tandemly repeated regions (11.4 × 38 bp).
Figure 1 from: Dong Y, Zhu L, Bai Y, Ou Y, Wang C (2016) Complete mitochondrial genomes of two flat-backed millipedes by next-generation sequencing (Diplopoda, Polydesmida). ZooKeys 637: 1-20. https://doi.org/10.3897/zookeys.637.9909
Figure 1 - Mitochondrial genomes of the two millipedes sequenced in this study. A Asiomorpha coarctata B Xystodesmus sp. Circular maps were drawn with Geneious v9.1.2. Arrows indicate the orientation of gene transcription. Abbreviations of gene names are: atp6 and atp8 for ATP synthase subunits 6 and 8; cox1–3 for cytochrome oxidase subunits 1–3; cob for cytochrome b, nad1–6 and nad4L for NADH dehydrogenase subunits 1–6 and 4L; and lrRNA and srRNA for large and small rRNA subunits. tRNA genes are indicated with their one-letter corresponding amino acids. CR for control region. The GC content was plotted using a green sliding window and the AT content was blue.
Figure 6 from: Dong Y, Zhu L, Bai Y, Ou Y, Wang C (2016) Complete mitochondrial genomes of two flat-backed millipedes by next-generation sequencing (Diplopoda, Polydesmida). ZooKeys 637: 1-20. https://doi.org/10.3897/zookeys.637.9909
Figure 6 - Phylogenetic tree of the Arthropoda, including Myriapoda, Hexapoda, Crustacea and Chelicerata and outgroups reconstructed based on protein-coding genes from mtDNA genomes. Each group of four numbers indicates node confidence values (from top left): Bayesian posterior probabilities in percent (BPP) in amino acid and nucleotide datasets; maximum likelihood bootstrapping values (MLBP) in amino acid and nucleotide datasets.
Figure 5 from: Dong Y, Zhu L, Bai Y, Ou Y, Wang C (2016) Complete mitochondrial genomes of two flat-backed millipedes by next-generation sequencing (Diplopoda, Polydesmida). ZooKeys 637: 1-20. https://doi.org/10.3897/zookeys.637.9909
Figure 5 - Comparison of gene arrangements in mtDNA of the arthropod ground pattern. Gene segments are not drawn to scale. Genes shaded gray have different relative positions compared to the ground pattern. Underlining indicates the gene is encoded on the opposite strand, and arrows indicate translocation of trnT. CR: putative control region. Gene arrangements of two diplopods, Narceus annularus and Thyropygus sp. are similar and represented as one.
Figure 1 from: Yuhui X, Lijun Z, Yue H, Xiaoqi W, Chen Z, Huilun Z, Ruoran W, Da P, Hongying S (2017) Complete mitochondrial genomes from two species of Chinese freshwater crabs of the genus Sinopotamon recovered using next-generation sequencing reveal a novel gene order (Brachyura, Potamidae). ZooKeys 705: 41-60. https://doi.org/10.3897/zookeys.705.11852
Figure 1 - Mitochondrial genome sequenced in the present study. Gene order and sizes are shown relative to one another, including non-coding regions. Protein-coding genes encoded on the light strand are underlined. Transfer RNA (tRNA) genes encoded on the light strand are underlined. Each tRNA gene is designated by a single-letter amino acid code, except L1 (trnLeu (CUN)), L2 (trnLeu (UUR)), S1 (trnSer (AGN)) and S2 (trnSer (UCN)). Numbers inside circles represent the size of the non-coding region separating two adjacent genes or the amount of shared nucleotides between two overlapping genes. The translocations of gene or gene block are shaded gray.
Figure 2 from: Yuhui X, Lijun Z, Yue H, Xiaoqi W, Chen Z, Huilun Z, Ruoran W, Da P, Hongying S (2017) Complete mitochondrial genomes from two species of Chinese freshwater crabs of the genus Sinopotamon recovered using next-generation sequencing reveal a novel gene order (Brachyura, Potamidae). ZooKeys 705: 41-60. https://doi.org/10.3897/zookeys.705.11852
Figure 2 - Phylogenetic analyses derived for brachyurans using the maximum likelihood (ML) analyses and Bayesian inferences (BI) using dataset A (13 PCGs) and dataset B (13 PCGs + two rRNAs). Branch lengths and topologies came from ML analysis. Values at the branches represent BP (Bootstrap value)/BPP (Bayesian posterior probability). 100/1.00 is denoted by an asterisk. The horizontal line stands for BP under 50 or BPP under 0.9 ML analyses. The gene rearrangement is denoted by the block on (A): (I) the translocation of trnH shared by the Brachyura taxa sampled; (II) the transposition of trnQ shared by potamid species; (III) the five-gene block, (trnM-nad2-trnW-trnC-trnY), translocation shared by three Sinopotamon crabs sampled.
Figure 3 from: Shaoli M, Hao Y, Chao L, Yafu Z, Fuming S, Yuchao W (2018) The complete mitochondrial genome of Xizicus (Haploxizicus) maculatus revealed by next-generation sequencing and phylogenetic implication (Orthoptera, Meconematinae). ZooKeys 773: 57-67. https://doi.org/10.3897/zookeys.773.24156
Figure 3 Phylogenetic reconstruction of Tettigoniidea using mitochondrial PCGs and rRNA concatenated dataset. A Bayesian result, applicable posterior probability values are shown B Maximum likelihood result with applicable bootstrap values shown.
Figure 2 from: Shaoli M, Hao Y, Chao L, Yafu Z, Fuming S, Yuchao W (2018) The complete mitochondrial genome of Xizicus (Haploxizicus) maculatus revealed by next-generation sequencing and phylogenetic implication (Orthoptera, Meconematinae). ZooKeys 773: 57-67. https://doi.org/10.3897/zookeys.773.24156
Figure 2 Relative synonymous codon usage of X. (X.) fascipes, X. (E.) howardi, X. (H.) maculatus mitochondrial protein-coding genes. Condon families are provided on the x-axis.
Data from: Next-generation polyploid phylogenetics: rapid resolution of hybrid polyploid complexes using PacBio single-molecule sequencing
Difficulties in generating nuclear data for polyploids have impeded phylogenetic study of these groups. We describe a high-throughput protocol and an associated bioinformatics pipeline (PURC: "Pipeline for Untangling Reticulate Complexes") that is able to generate these data quickly and conveniently, and demonstrate its efficacy on accessions from the fern family Cystopteridaceae. We conclude with a demonstration of the downstream utility of these data by inferring a multilabeled species tree for a subset of our accessions. We amplified four ~1kb-long nuclear loci and sequenced them in a parallel-tagged amplicon sequencing approach using the PacBio platform. PURC infers the final sequences from the raw reads via an iterative approach that corrects PCR and sequencing errors and removes PCR-mediated recombinant sequences (chimeras). We generated data for all gene copies (homeologs, paralogs, and segregating alleles) present in each of three sets of 50 mostly-polyploid accessions, for four loci, in three PacBio runs (one run per set). From the raw sequencing reads PURC was able to accurately infer the underlying sequences. This approach makes it easy and economical to study the phylogenetics of polyploids, and in conjunction with recent analytical advances, facilitates investigation of broad patterns of polyploid evolution.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.