Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

190

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

190 results for “Targeted enrichment”

Learn how ShareScore rates datasets ↗
dryad36/100

Data from: A phylogenomic perspective on the biogeography of skinks in the Plestiodon brevirostris group inferred from target enrichment of ultraconserved elements

Open the record for dataset details and reuse information.

publicFeb 2018View details →
dryad36/100

Target enrichment of long open reading frames and ultraconserved elements to link microevolution and macroevolution in non-model organisms

Open the record for dataset details and reuse information.

publicNov 2022View details →
dryad36/100

Validating a target-enrichment design for capturing uniparental haplotypes in ancient domesticated animals

Open the record for dataset details and reuse information.

publicApr 2025View details →
dryad36/100

Data from: Targeted enrichment of large gene families for phylogenetic inference: phylogeny and molecular evolution of photosynthesis genes in the Portullugo clade (Caryophyllales)

Open the record for dataset details and reuse information.

publicSep 2017View details →
dryad36/100

A target enrichment probe set for resolving phylogenetic relationships in the coffee family, Rubiaceae

Open the record for dataset details and reuse information.

publicJan 2024View details →
dryad36/100

Data from: Evidence of intraspecific adaptive variation in the American pika (Ochotona princeps) on a continental scale using a target enrichment and mitochondrial genome skimming approach

Open the record for dataset details and reuse information.

publicOct 2024View details →
dryad36/100

Data from: Extensive allopolyploidy in the neotropical genus Lachemilla (Rosaceae) revealed by PCR ‐based target enrichment of the nuclear ribosomal DNA cistron and plastid phylogenomics

Open the record for dataset details and reuse information.

publicMar 2019View details →
dryad36/100

Raw target enrichment of conserved element sequence data for 24 black coral species

Open the record for dataset details and reuse information.

publicAug 2023View details →
dryad36/100

Data from: Phylogenomic analysis of target enrichment and transcriptome data uncovers rapid radiation and extensive hybridization in slipper orchid genus Cypripedium L.

Open the record for dataset details and reuse information.

publicApr 2025View details →
dryad36/100

A new pipeline for removing paralogs in target enrichment data

Open the record for dataset details and reuse information.

publicJun 2021View details →
zenodo32/100

Transcriptome-based target-enrichment baits for stony corals (Cnidaria: Anthozoa: Scleractinia)

<p>Bait sets, sequences, trees&nbsp;and scripts</p>

opencc-by-4.0Dec 2019View details →
zenodo32/100

MitoFinder: efficient automated large-scale extraction of mitogenomic data in target enrichment phylogenomics

<p><strong>MitoFinder: efficient automated large-scale extraction of mitogenomic data in target enrichment phylogenomics</strong></p> <p>R&eacute;mi Allio<sup>1</sup>, Alex Schomaker-Bastos<sup>2,&dagger;</sup>, Jonathan Romiguier<sup>1</sup>, Francisco Prosdocimi<sup>2</sup>, Benoit Nabholz<sup>1</sup>, and Fr&eacute;d&eacute;ric Delsuc<sup>1</sup></p> <p><sup>1</sup><em>Institut des Sciences de l&rsquo;Evolution de Montpellier (ISEM), CNRS, EPHE, IRD, Universit&eacute; de Montpellier, Montpellier, France.</em></p> <p><sup>2</sup><em>Laborat&oacute;rio Multidisciplinar para An&aacute;lise de Dados (LAMPADA), Instituto de Bioqu&iacute;mica M&eacute;dica Leopoldo de Meis, Universidade Federal do Rio de Janeiro, Rio de Janeiro, Brasil.</em></p> <p><sup>&dagger;</sup><em> In Memoriam (08/01/2015) </em></p> <p>&nbsp;</p> <p><em><strong>Correspondence</strong></em></p> <p>R&eacute;mi Allio</p> <p>Email: <a href="mailto:remi.allio@umontpelier.fr">remi.allio@umontpellier.fr</a></p> <p>Fr&eacute;d&eacute;ric Delsuc</p> <p>Email: <a href="mailto:frederic.delsuc@umontpellier.fr">frederic.delsuc@umontpellier.fr</a></p> <p>&nbsp;</p> <p><strong><em>Running head</em></strong></p> <p>Mitochondrial signal from UCE capture data</p> <p>&nbsp;</p> <p><strong>Abstract</strong><strong> </strong></p> <p>Thanks to the development of high-throughput sequencing technologies, target enrichment sequencing of nuclear ultraconserved DNA elements (UCEs) now allows routinely inferring phylogenetic relationships from thousands of genomic markers. Recently, it has been shown that mitochondrial DNA (mtDNA) is frequently sequenced alongside the targeted loci in such capture experiments. Despite its broad evolutionary interest, mtDNA is rarely assembled and used in conjunction with nuclear markers in capture-based studies. Here, we developed MitoFinder, a user-friendly bioinformatic pipeline, to efficiently assemble and annotate mitogenomic data from hundreds of UCE libraries. As a case study, we used ants (Formicidae) for which 501 UCE libraries have been sequenced whereas only 29 mitogenomes are available. We compared the efficiency of four different assemblers (IDBA-UD, MEGAHIT, MetaSPAdes, and Trinity) for assembling both UCE and mtDNA loci. Using MitoFinder, we show that metagenomic assemblers, in particular MetaSPAdes, are well suited to assemble both UCEs and mtDNA. Mitogenomic signal was successfully extracted from all 501 UCE libraries allowing confirming species identification using COI barcoding. Moreover, our automated procedure retrieved 296 cases in which the mitochondrial genome was assembled in a single contig, thus increasing the number of available ant mitogenomes by an order of magnitude. By leveraging the power of metagenomic assemblers, MitoFinder provides an efficient tool to extract complementary mitogenomic data from UCE libraries, allowing testing for potential mito-nuclear discordance. Our approach is potentially applicable to other sequence capture methods, transcriptomic data, and whole genome shotgun sequencing in diverse taxa.</p> <p>&nbsp;</p> <p><strong><em>Figures &amp; Tables</em></strong></p> <p><strong>Figure 1.</strong> Conceptualization of the pipeline used to assemble and extract UCE and mitochondrial signal from ultraconserved element sequencing data.</p> <p><strong>Figure 2</strong>. Comparison of the efficiency of the assemblers in terms of: A) computational time, B) number of potentially mitochondrial contigs identified, and C) number of mitochondrial genes annotated. Violin plots reflect the data distribution with a horizontal line indicating the median. Note that for the three metagenomic assemblers, 5 CPUs were used compared to 35 CPUs for Trinity. Plots were obtained using PlotsOfData (Postma &amp; Goedhart 2019).</p> <p><strong>Figure 3.</strong> Phylogenomic relationships of ants (Formicidae). AA) Mito-nuclear phylogenetic differences among subfamily relationships based on the UCE and mtDNA supermatrices obtained with the assembler MetaSPAdes assembler. Clades corresponding to subfamilies were collapsed. Inter-subfamily relationships with UFBS &lt; 95% were collapsed. Non-maximal node support values are reported. B) The topology obtained reflects the results of phylogenetic analyses based on the amino acid mitochondrial supermatrix (using MetaSPAdes as assembler). Histograms reflect the percent of UCEs (light grey) and mitochondrial genes (dark grey) recovered for each species. Illustrative pictures (*): <em>Diacamma sp</em>. (Ponerinae; top left), <em>Formica sp</em>. (Formicinae; top right), and <em>Messor barbarus </em>(Myrmicinae; bottom right).</p> <p><strong>Table 1. </strong>Summary statistics on assembly results according to the assembler used. The values are averages over the 501 assemblies, except for the assembly time, which is a median value. The two tables report specific statistics for A) ultraconserved elements data, and B) mitochondrial data. Note that 35 CPUs were used for Trinity whereas 5 CPUs were used for other assemblers.</p> <p><strong>Table 2.</strong> Statistical comparison between the performances of the different assemblers. Statistical significance was estimated with a paired non parametric test (paired wilcoxon test). *** = <em>p</em>&lt;0.001; ** = <em>p</em>&lt;0.01; * = <em>p</em>&lt;0.05; NS = <em>p</em>&gt;0.05; and (+)/(-) is the result of the comparison between the row and the column.</p> <p>&nbsp;</p> <p><strong><em>Appendices</em></strong></p> <p><strong>Appendix S1.</strong> List of the 501 UCE libraries (SRA accessions) and associated metadata.</p> <p><strong>Appendix S2.</strong> Summary statistics on mitochondrial signal recovered per species and depending on the assembler used. The table provides the number of contigs and genes recovered with MitoFinder and the size of each annotated gene.</p> <p><strong>Appendix S3.</strong> Summary statistics of barcoding analyses. Detailed results for both BOLDsystem and Megablast analyses are provided for each CO1 recovered with MitoFinder using MetaSPAdes.</p> <p><strong>Appendix S4.</strong> Detailed results of tree distance analyses realized with Dquad (Ranwez, Criscuolo, &amp; Douzery 2010). Trees obtained with each assembler with mitochondrial amino acid supermatrix, mitochondrial nucleotide supermatrix, and UCE nucleotide supermatrix were compared with each others.</p> <p><strong>Appendix S5</strong>. List of Genbank accession numbers for newly generated mitchondrial contigs.</p> <p>&nbsp;</p> <p><strong><em>Zenodo supplementary files</em></strong></p> <p><strong>Assembly_results.tar.gz</strong> Contains all contigs obtained for each species with the different assemblers implemented in MitoFinder.</p> <p><strong>MitoFinder_annotations.tar.gz</strong> Contains MitoFinder annotations for each species. (based on the contigs obtained with MetaSPAdes)</p> <p><strong>UCE_results.tar.gz</strong> Contains all annotated UCE obtained for each species after UCE identification with PHYLUCE. (MetaSPAdes)</p> <p><strong>Final_mtDNA_alignments.tar.gz</strong> Contains the final mitochondrial gene&nbsp;alignments. (MetaSPAdes)</p> <p><strong>Final_UCE_alignments.tar.gz</strong> Contains the final UCE alignments. (MetaSPAdes)</p> <p><strong>Final_mtDNA_matrices.tar.gz</strong> Contains the final mi&nbsp; tochondrial supermatrices (AA and NT) used for the phylogenetic analyses. (MetaSPAdes)</p> <p><strong>Metaspades_final_UCE_matrix.phy</strong> The final UCE supermatrix used for the phylogenetic analyses. (MetaSPAdes)</p>

opencc-by-4.0Sep 2019View details →
dryad32/100

Data from: An enhanced target-enrichment bait set for Hexacorallia provides phylogenomic resolution of the staghorn corals (Acroporidae) and close relatives

<p>Targeted enrichment of genomic DNA can profoundly increase the phylogenetic resolution of clades and inform taxonomy. Here, we redesign a custom bait set previously developed for the cnidarian class Anthozoa to more efficiently target and capture ultraconserved elements (UCEs) and exonic loci within the subclass Hexacorallia. We test this enhanced bait set (targeting 2,476 loci) on 99 specimens of scleractinian corals spanning both the "complex" (Acroporidae, Agariciidae) and "robust" (Fungiidae) clades.Focused sampling in the staghorn corals (genus <i>Acropora</i>)highlights the ability of sequence capture to inform the taxonomy of a clade previously deficient in molecular resolution. A mean of 1850 (± 298) loci were captured per taxon (955 UCEs, 894 exons), and a 75% complete concatenated alignment of 96 samples included 1792 loci (991 UCE, 801 exons) and ~1.87 million base pairs. Maximum likelihood and Bayesian analyses recovered robust molecular relationships and revealed that species-level relationships within the <i>Acropora</i>are incongruent with traditional morphological groupings. Both UCE and exon datasets delineated six well-supported clades within <i>Acropora.</i>The enhanced bait set will facilitate investigations of the evolutionary history of many important groups of reef corals, particularly where previous molecular marker development has been unsuccessful.</p>

opencc-zeroAug 2020View details →
zenodo32/100

Urban wastewater virome by viral metagenomics and target enrichment sequencing

<p>Metagenomic analysis of virus in raw sewage.</p>

opencc-by-4.0Jan 2021View details →
dryad32/100

Data from: Target enrichment of ultraconserved elements from arthropods provides a genomic perspective on relationships among Hymenoptera

Gaining a genomic perspective on phylogeny requires the collection of data from many putatively independent loci across the genome. Among insects, an increasingly common approach to collecting this class of data involves transcriptome sequencing, because few insects have high-quality genome sequences available; assembling new genomes remains a limiting factor; the transcribed portion of the genome is a reasonable, reduced subset of the genome to target; and the data collected from transcribed portions of the genome are similar in composition to the types of data with which biologists have traditionally worked (e.g. exons). However, molecular techniques requiring RNA as a template, including transcriptome sequencing, are limited to using very high-quality source materials, which are often unavailable from a large proportion of biologically important insect samples. Recent research suggests that DNA-based target enrichment of conserved genomic elements offers another path to collecting phylogenomic data across insect taxa, provided that conserved elements are present in and can be collected from insect genomes. Here, we identify a large set (n = 1510) of ultraconserved elements (UCEs) shared among the insect order Hymenoptera. We used in silico analyses to show that these loci accurately reconstruct relationships among genome-enabled hymenoptera, and we designed a set of RNA baits (n = 2749) for enriching these loci that researchers can use with DNA templates extracted from a variety of sources. We used our UCE bait set to enrich an average of 721 UCE loci from 30 hymenopteran taxa, and we used these UCE loci to reconstruct phylogenetic relationships spanning very old (≥220 Ma) to very young (≤1 Ma) divergences among hymenopteran lineages. In contrast to a recent study addressing hymenopteran phylogeny using transcriptome data, we found ants to be sister to all remaining aculeate lineages with complete support, although this result could be explained by factors such as taxon sampling. We discuss this approach and our results in the context of elucidating the evolutionary history of one of the most diverse and speciose animal orders.

opencc-zeroDec 2013View details →
dryad32/100

Data from: Targeted gene enrichment and high-throughput sequencing for environmental biomonitoring: a case study using freshwater macroinvertebrates

Recent studies have advocated biomonitoring using DNA techniques. In this study, two high-throughput sequencing (HTS)-based methods were evaluated: amplicon metabarcoding of the cytochrome C oxidase subunit I (COI) mitochondrial gene and gene enrichment using MYbaits (targeting nine different genes including COI). The gene-enrichment method does not require PCR amplification and thus avoids biases associated with universal primers. Macroinvertebrate samples were collected from 12 New Zealand rivers. Macroinvertebrates were morphologically identified and enumerated, and their biomass determined. DNA was extracted from all macroinvertebrate samples and HTS undertaken using the illumina miseq platform. Macroinvertebrate communities were characterized from sequence data using either six genes (three of the original nine were not used) or just the COI gene in isolation. The gene-enrichment method (all genes) detected the highest number of taxa and obtained the strongest Spearman rank correlations between the number of sequence reads, abundance and biomass in 67% of the samples. Median detection rates across rare (&lt;1% of the total abundance or biomass), moderately abundant (1–5%) and highly abundant (&gt;5%) taxa were highest using the gene-enrichment method (all genes). Our data indicated primer biases occurred during amplicon metabarcoding with greater than 80% of sequence reads originating from one taxon in several samples. The accuracy and sensitivity of both HTS methods would be improved with more comprehensive reference sequence databases. The data from this study illustrate the challenges of using PCR amplification-based methods for biomonitoring and highlight the potential benefits of using approaches, such as gene enrichment, which circumvent the need for an initial PCR step.

opencc-zeroDec 2014View details →
zenodo32/100

Figure 10 in Target-enriched DNA sequencing from historical type material enables a partial revision of the Madagascar giant stream frogs (genus Mantidactylus)

Figure 10. Mantidactylus (Mantidactylus) radaka sp. nov. being prepared for human consumption. (a) Frogs and crabs are collected from broad streams. Then (b) the frogs are gutted and skinned, and the head, hands and feet removed. The frog is then rinsed in the stream, leaving (c) cleaned animals for cooking in a stew. Note the ovaries full with hundreds of eggs.

opennotspecifiedMay 2020View details →
zenodo32/100

Figure 9 in Target-enriched DNA sequencing from historical type material enables a partial revision of the Madagascar giant stream frogs (genus Mantidactylus)

Figure 9. Preserved type specimens of the four nomina in the Mantidactylus subgenus Mantidactylus and one of the paralectotypes of Rana guttulata.

opennotspecifiedMay 2020View details →
zenodo32/100

Figure 7 in Target-enriched DNA sequencing from historical type material enables a partial revision of the Madagascar giant stream frogs (genus Mantidactylus)

Figure 7. Photographs of living specimens of Mantidactylus (Mantidactylus) guttulatus, M. (M.) grandidieri, and of three candidate species. (a, b) M. (M.) guttulatus, female ZSM 1013/2003 (FGMV 2002.438) from Ranomafana. (c) Unidentified specimen from Ranomafana, assigned tentatively to M. (M.) guttulatus (no genetic evidence). (d, e) M. (M.) guttulatus, specimen KU 340853 (CRH729) from Ranomafana. (f) M. (M.) grandidieri, specimen ZSM 5077/2005 (ZCMV 2159) from Nosy Mangabe. (g) M. (M.) grandidieri, specimen ZSM 276/2005 (FGZC 2682) from Vohidrazana. (h) M. (M.) grandidieri, unidentified specimen (probably subadult) from Andranofotsy. (i, j) M. (M.) grandidieri, specimen KU

opennotspecifiedMay 2020View details →
zenodo32/100

Figure 5. Per-base coverage plots for the 16S in Target-enriched DNA sequencing from historical type material enables a partial revision of the Madagascar giant stream frogs (genus Mantidactylus)

Figure 5. Per-base coverage plots for the 16S fragment in four Mantidactylus type specimens from the MNHN and BMNH collections. (a) BMNH 1947.2.25.48 (paralectotype of Rana guttulata); (b) BMNH 1947.2.25.51 (paralectotype of Rana guttulata); (c) MNHN 1895.255 (syntype of M. grandidieri); (d) MNHN 1883.520 (syntype of M. grandidieri).

opennotspecifiedMay 2020View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record