Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
190
datasets available to search
ShareScore release 0.9.0
Dataset results
190 results for “Targeted enrichment”
Data from: A phylogenomic perspective on the biogeography of skinks in the Plestiodon brevirostris group inferred from target enrichment of ultraconserved elements
Open the record for dataset details and reuse information.
Target enrichment of long open reading frames and ultraconserved elements to link microevolution and macroevolution in non-model organisms
Open the record for dataset details and reuse information.
Validating a target-enrichment design for capturing uniparental haplotypes in ancient domesticated animals
Open the record for dataset details and reuse information.
Data from: Targeted enrichment of large gene families for phylogenetic inference: phylogeny and molecular evolution of photosynthesis genes in the Portullugo clade (Caryophyllales)
Open the record for dataset details and reuse information.
A target enrichment probe set for resolving phylogenetic relationships in the coffee family, Rubiaceae
Open the record for dataset details and reuse information.
Data from: Evidence of intraspecific adaptive variation in the American pika (Ochotona princeps) on a continental scale using a target enrichment and mitochondrial genome skimming approach
Open the record for dataset details and reuse information.
Data from: Extensive allopolyploidy in the neotropical genus Lachemilla (Rosaceae) revealed by PCR ‐based target enrichment of the nuclear ribosomal DNA cistron and plastid phylogenomics
Open the record for dataset details and reuse information.
Raw target enrichment of conserved element sequence data for 24 black coral species
Open the record for dataset details and reuse information.
Data from: Phylogenomic analysis of target enrichment and transcriptome data uncovers rapid radiation and extensive hybridization in slipper orchid genus Cypripedium L.
Open the record for dataset details and reuse information.
A new pipeline for removing paralogs in target enrichment data
Open the record for dataset details and reuse information.
Transcriptome-based target-enrichment baits for stony corals (Cnidaria: Anthozoa: Scleractinia)
<p>Bait sets, sequences, trees and scripts</p>
MitoFinder: efficient automated large-scale extraction of mitogenomic data in target enrichment phylogenomics
<p><strong>MitoFinder: efficient automated large-scale extraction of mitogenomic data in target enrichment phylogenomics</strong></p> <p>Rémi Allio<sup>1</sup>, Alex Schomaker-Bastos<sup>2,†</sup>, Jonathan Romiguier<sup>1</sup>, Francisco Prosdocimi<sup>2</sup>, Benoit Nabholz<sup>1</sup>, and Frédéric Delsuc<sup>1</sup></p> <p><sup>1</sup><em>Institut des Sciences de l’Evolution de Montpellier (ISEM), CNRS, EPHE, IRD, Université de Montpellier, Montpellier, France.</em></p> <p><sup>2</sup><em>Laboratório Multidisciplinar para Análise de Dados (LAMPADA), Instituto de Bioquímica Médica Leopoldo de Meis, Universidade Federal do Rio de Janeiro, Rio de Janeiro, Brasil.</em></p> <p><sup>†</sup><em> In Memoriam (08/01/2015) </em></p> <p> </p> <p><em><strong>Correspondence</strong></em></p> <p>Rémi Allio</p> <p>Email: <a href="mailto:remi.allio@umontpelier.fr">remi.allio@umontpellier.fr</a></p> <p>Frédéric Delsuc</p> <p>Email: <a href="mailto:frederic.delsuc@umontpellier.fr">frederic.delsuc@umontpellier.fr</a></p> <p> </p> <p><strong><em>Running head</em></strong></p> <p>Mitochondrial signal from UCE capture data</p> <p> </p> <p><strong>Abstract</strong><strong> </strong></p> <p>Thanks to the development of high-throughput sequencing technologies, target enrichment sequencing of nuclear ultraconserved DNA elements (UCEs) now allows routinely inferring phylogenetic relationships from thousands of genomic markers. Recently, it has been shown that mitochondrial DNA (mtDNA) is frequently sequenced alongside the targeted loci in such capture experiments. Despite its broad evolutionary interest, mtDNA is rarely assembled and used in conjunction with nuclear markers in capture-based studies. Here, we developed MitoFinder, a user-friendly bioinformatic pipeline, to efficiently assemble and annotate mitogenomic data from hundreds of UCE libraries. As a case study, we used ants (Formicidae) for which 501 UCE libraries have been sequenced whereas only 29 mitogenomes are available. We compared the efficiency of four different assemblers (IDBA-UD, MEGAHIT, MetaSPAdes, and Trinity) for assembling both UCE and mtDNA loci. Using MitoFinder, we show that metagenomic assemblers, in particular MetaSPAdes, are well suited to assemble both UCEs and mtDNA. Mitogenomic signal was successfully extracted from all 501 UCE libraries allowing confirming species identification using COI barcoding. Moreover, our automated procedure retrieved 296 cases in which the mitochondrial genome was assembled in a single contig, thus increasing the number of available ant mitogenomes by an order of magnitude. By leveraging the power of metagenomic assemblers, MitoFinder provides an efficient tool to extract complementary mitogenomic data from UCE libraries, allowing testing for potential mito-nuclear discordance. Our approach is potentially applicable to other sequence capture methods, transcriptomic data, and whole genome shotgun sequencing in diverse taxa.</p> <p> </p> <p><strong><em>Figures & Tables</em></strong></p> <p><strong>Figure 1.</strong> Conceptualization of the pipeline used to assemble and extract UCE and mitochondrial signal from ultraconserved element sequencing data.</p> <p><strong>Figure 2</strong>. Comparison of the efficiency of the assemblers in terms of: A) computational time, B) number of potentially mitochondrial contigs identified, and C) number of mitochondrial genes annotated. Violin plots reflect the data distribution with a horizontal line indicating the median. Note that for the three metagenomic assemblers, 5 CPUs were used compared to 35 CPUs for Trinity. Plots were obtained using PlotsOfData (Postma & Goedhart 2019).</p> <p><strong>Figure 3.</strong> Phylogenomic relationships of ants (Formicidae). AA) Mito-nuclear phylogenetic differences among subfamily relationships based on the UCE and mtDNA supermatrices obtained with the assembler MetaSPAdes assembler. Clades corresponding to subfamilies were collapsed. Inter-subfamily relationships with UFBS < 95% were collapsed. Non-maximal node support values are reported. B) The topology obtained reflects the results of phylogenetic analyses based on the amino acid mitochondrial supermatrix (using MetaSPAdes as assembler). Histograms reflect the percent of UCEs (light grey) and mitochondrial genes (dark grey) recovered for each species. Illustrative pictures (*): <em>Diacamma sp</em>. (Ponerinae; top left), <em>Formica sp</em>. (Formicinae; top right), and <em>Messor barbarus </em>(Myrmicinae; bottom right).</p> <p><strong>Table 1. </strong>Summary statistics on assembly results according to the assembler used. The values are averages over the 501 assemblies, except for the assembly time, which is a median value. The two tables report specific statistics for A) ultraconserved elements data, and B) mitochondrial data. Note that 35 CPUs were used for Trinity whereas 5 CPUs were used for other assemblers.</p> <p><strong>Table 2.</strong> Statistical comparison between the performances of the different assemblers. Statistical significance was estimated with a paired non parametric test (paired wilcoxon test). *** = <em>p</em><0.001; ** = <em>p</em><0.01; * = <em>p</em><0.05; NS = <em>p</em>>0.05; and (+)/(-) is the result of the comparison between the row and the column.</p> <p> </p> <p><strong><em>Appendices</em></strong></p> <p><strong>Appendix S1.</strong> List of the 501 UCE libraries (SRA accessions) and associated metadata.</p> <p><strong>Appendix S2.</strong> Summary statistics on mitochondrial signal recovered per species and depending on the assembler used. The table provides the number of contigs and genes recovered with MitoFinder and the size of each annotated gene.</p> <p><strong>Appendix S3.</strong> Summary statistics of barcoding analyses. Detailed results for both BOLDsystem and Megablast analyses are provided for each CO1 recovered with MitoFinder using MetaSPAdes.</p> <p><strong>Appendix S4.</strong> Detailed results of tree distance analyses realized with Dquad (Ranwez, Criscuolo, & Douzery 2010). Trees obtained with each assembler with mitochondrial amino acid supermatrix, mitochondrial nucleotide supermatrix, and UCE nucleotide supermatrix were compared with each others.</p> <p><strong>Appendix S5</strong>. List of Genbank accession numbers for newly generated mitchondrial contigs.</p> <p> </p> <p><strong><em>Zenodo supplementary files</em></strong></p> <p><strong>Assembly_results.tar.gz</strong> Contains all contigs obtained for each species with the different assemblers implemented in MitoFinder.</p> <p><strong>MitoFinder_annotations.tar.gz</strong> Contains MitoFinder annotations for each species. (based on the contigs obtained with MetaSPAdes)</p> <p><strong>UCE_results.tar.gz</strong> Contains all annotated UCE obtained for each species after UCE identification with PHYLUCE. (MetaSPAdes)</p> <p><strong>Final_mtDNA_alignments.tar.gz</strong> Contains the final mitochondrial gene alignments. (MetaSPAdes)</p> <p><strong>Final_UCE_alignments.tar.gz</strong> Contains the final UCE alignments. (MetaSPAdes)</p> <p><strong>Final_mtDNA_matrices.tar.gz</strong> Contains the final mi tochondrial supermatrices (AA and NT) used for the phylogenetic analyses. (MetaSPAdes)</p> <p><strong>Metaspades_final_UCE_matrix.phy</strong> The final UCE supermatrix used for the phylogenetic analyses. (MetaSPAdes)</p>
Data from: An enhanced target-enrichment bait set for Hexacorallia provides phylogenomic resolution of the staghorn corals (Acroporidae) and close relatives
<p>Targeted enrichment of genomic DNA can profoundly increase the phylogenetic resolution of clades and inform taxonomy. Here, we redesign a custom bait set previously developed for the cnidarian class Anthozoa to more efficiently target and capture ultraconserved elements (UCEs) and exonic loci within the subclass Hexacorallia. We test this enhanced bait set (targeting 2,476 loci) on 99 specimens of scleractinian corals spanning both the "complex" (Acroporidae, Agariciidae) and "robust" (Fungiidae) clades.Focused sampling in the staghorn corals (genus <i>Acropora</i>)highlights the ability of sequence capture to inform the taxonomy of a clade previously deficient in molecular resolution. A mean of 1850 (± 298) loci were captured per taxon (955 UCEs, 894 exons), and a 75% complete concatenated alignment of 96 samples included 1792 loci (991 UCE, 801 exons) and ~1.87 million base pairs. Maximum likelihood and Bayesian analyses recovered robust molecular relationships and revealed that species-level relationships within the <i>Acropora</i>are incongruent with traditional morphological groupings. Both UCE and exon datasets delineated six well-supported clades within <i>Acropora.</i>The enhanced bait set will facilitate investigations of the evolutionary history of many important groups of reef corals, particularly where previous molecular marker development has been unsuccessful.</p>
Urban wastewater virome by viral metagenomics and target enrichment sequencing
<p>Metagenomic analysis of virus in raw sewage.</p>
Data from: Target enrichment of ultraconserved elements from arthropods provides a genomic perspective on relationships among Hymenoptera
Gaining a genomic perspective on phylogeny requires the collection of data from many putatively independent loci across the genome. Among insects, an increasingly common approach to collecting this class of data involves transcriptome sequencing, because few insects have high-quality genome sequences available; assembling new genomes remains a limiting factor; the transcribed portion of the genome is a reasonable, reduced subset of the genome to target; and the data collected from transcribed portions of the genome are similar in composition to the types of data with which biologists have traditionally worked (e.g. exons). However, molecular techniques requiring RNA as a template, including transcriptome sequencing, are limited to using very high-quality source materials, which are often unavailable from a large proportion of biologically important insect samples. Recent research suggests that DNA-based target enrichment of conserved genomic elements offers another path to collecting phylogenomic data across insect taxa, provided that conserved elements are present in and can be collected from insect genomes. Here, we identify a large set (n = 1510) of ultraconserved elements (UCEs) shared among the insect order Hymenoptera. We used in silico analyses to show that these loci accurately reconstruct relationships among genome-enabled hymenoptera, and we designed a set of RNA baits (n = 2749) for enriching these loci that researchers can use with DNA templates extracted from a variety of sources. We used our UCE bait set to enrich an average of 721 UCE loci from 30 hymenopteran taxa, and we used these UCE loci to reconstruct phylogenetic relationships spanning very old (≥220 Ma) to very young (≤1 Ma) divergences among hymenopteran lineages. In contrast to a recent study addressing hymenopteran phylogeny using transcriptome data, we found ants to be sister to all remaining aculeate lineages with complete support, although this result could be explained by factors such as taxon sampling. We discuss this approach and our results in the context of elucidating the evolutionary history of one of the most diverse and speciose animal orders.
Data from: Targeted gene enrichment and high-throughput sequencing for environmental biomonitoring: a case study using freshwater macroinvertebrates
Recent studies have advocated biomonitoring using DNA techniques. In this study, two high-throughput sequencing (HTS)-based methods were evaluated: amplicon metabarcoding of the cytochrome C oxidase subunit I (COI) mitochondrial gene and gene enrichment using MYbaits (targeting nine different genes including COI). The gene-enrichment method does not require PCR amplification and thus avoids biases associated with universal primers. Macroinvertebrate samples were collected from 12 New Zealand rivers. Macroinvertebrates were morphologically identified and enumerated, and their biomass determined. DNA was extracted from all macroinvertebrate samples and HTS undertaken using the illumina miseq platform. Macroinvertebrate communities were characterized from sequence data using either six genes (three of the original nine were not used) or just the COI gene in isolation. The gene-enrichment method (all genes) detected the highest number of taxa and obtained the strongest Spearman rank correlations between the number of sequence reads, abundance and biomass in 67% of the samples. Median detection rates across rare (<1% of the total abundance or biomass), moderately abundant (1–5%) and highly abundant (>5%) taxa were highest using the gene-enrichment method (all genes). Our data indicated primer biases occurred during amplicon metabarcoding with greater than 80% of sequence reads originating from one taxon in several samples. The accuracy and sensitivity of both HTS methods would be improved with more comprehensive reference sequence databases. The data from this study illustrate the challenges of using PCR amplification-based methods for biomonitoring and highlight the potential benefits of using approaches, such as gene enrichment, which circumvent the need for an initial PCR step.
Figure 10 in Target-enriched DNA sequencing from historical type material enables a partial revision of the Madagascar giant stream frogs (genus Mantidactylus)
Figure 10. Mantidactylus (Mantidactylus) radaka sp. nov. being prepared for human consumption. (a) Frogs and crabs are collected from broad streams. Then (b) the frogs are gutted and skinned, and the head, hands and feet removed. The frog is then rinsed in the stream, leaving (c) cleaned animals for cooking in a stew. Note the ovaries full with hundreds of eggs.
Figure 9 in Target-enriched DNA sequencing from historical type material enables a partial revision of the Madagascar giant stream frogs (genus Mantidactylus)
Figure 9. Preserved type specimens of the four nomina in the Mantidactylus subgenus Mantidactylus and one of the paralectotypes of Rana guttulata.
Figure 7 in Target-enriched DNA sequencing from historical type material enables a partial revision of the Madagascar giant stream frogs (genus Mantidactylus)
Figure 7. Photographs of living specimens of Mantidactylus (Mantidactylus) guttulatus, M. (M.) grandidieri, and of three candidate species. (a, b) M. (M.) guttulatus, female ZSM 1013/2003 (FGMV 2002.438) from Ranomafana. (c) Unidentified specimen from Ranomafana, assigned tentatively to M. (M.) guttulatus (no genetic evidence). (d, e) M. (M.) guttulatus, specimen KU 340853 (CRH729) from Ranomafana. (f) M. (M.) grandidieri, specimen ZSM 5077/2005 (ZCMV 2159) from Nosy Mangabe. (g) M. (M.) grandidieri, specimen ZSM 276/2005 (FGZC 2682) from Vohidrazana. (h) M. (M.) grandidieri, unidentified specimen (probably subadult) from Andranofotsy. (i, j) M. (M.) grandidieri, specimen KU
Figure 5. Per-base coverage plots for the 16S in Target-enriched DNA sequencing from historical type material enables a partial revision of the Madagascar giant stream frogs (genus Mantidactylus)
Figure 5. Per-base coverage plots for the 16S fragment in four Mantidactylus type specimens from the MNHN and BMNH collections. (a) BMNH 1947.2.25.48 (paralectotype of Rana guttulata); (b) BMNH 1947.2.25.51 (paralectotype of Rana guttulata); (c) MNHN 1895.255 (syntype of M. grandidieri); (d) MNHN 1883.520 (syntype of M. grandidieri).
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.