Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

98

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

98 results for “Targeted capture”

Learn how ShareScore rates datasets ↗
dryad32/100

A new approach using targeted sequence capture for phylogenomic studies across Cactaceae

<p>Relationships within the major clades of Cactaceae are relatively well known based on DNA sequence data mostly from the chloroplast genome. Nevertheless, some nodes along the backbone of the phylogeny, and especially generic and species-level relationships, remain poorly resolved and are in need of more informative genetic markers. In this study, we propose a new approach to solve the relationships within Cactaceae, applying a targeted sequence capture pipeline. We designed a custom probe set for Cactaceae using MarkerMiner and complemented it with the Angiosperms353 probe set. We then tested both probe sets against 36 different transcriptomes using Hybpiper preferentially retaining phylogenetically informative loci and reconstructed the relationships using RAxML-NG and Astral. Finally, we tested each probe set through sequencing 96 accessions, representing 88 species across Cactaceae. Our preliminary analyses recovered a well-supported phylogeny across Cactaceae with a near identical topology among major clade relationships as that recovered with plastome data. As expected, however, we found incongruences in relationships when comparing our nuclear probe set results to plastome datasets, especially at the generic level. Our results reveal great potential for the combination of Cactaceae-specific and Angiosperm353 probe set application to improve phylogenetic resolution for Cactaceae and for other studies.</p>

opencc-zeroMar 2022View details →
dryad32/100

Data from: Phylogenomic resolution of the cetacean tree of life using target sequence capture

The evolution of the cetaceans, from their early transition to an aquatic lifestyle to their subsequent diversification, has been the subject of numerous studies. However, while the higher-level relationships among cetacean families have been largely settled, several aspects of the systematics within these groups remain unresolved. Problematic clades include the oceanic dolphins (37 spp.), which have experienced a recent rapid radiation, and the beaked whales (22 spp.), which have not been investigated in detail using nuclear loci. The combined application of high-throughput sequencing with techniques that target specific genomic sequences provide a powerful means of rapidly generating large volumes of orthologous sequence data for use in phylogenomic studies. To elucidate the phylogenetic relationships within the Cetacea, we combined sequence capture with Illumina sequencing to generate data for ~3200 protein-coding genes for 68 cetacean species and their close relatives including the pygmy hippopotamus. By combining data from &gt;38,000 exons with existing sequences from 11 cetaceans and seven outgroup taxa, we produced the first comprehensive comparative genomic dataset for cetaceans, spanning 6,527,596 aligned base pairs and 89 taxa. Phylogenetic trees reconstructed with maximum likelihood and Bayesian inference of concatenated loci, as well as with coalescence analyses of individual gene trees, produced mostly concordant and well-supported trees. Our results completely resolve the relationships among beaked whales as well as the contentious relationships among ocean dolphins, especially the problematic subfamily Delphininae, which includes the common and bottlenose dolphins. We performed Bayesian estimation of species divergence times using MCMCtree, integrating recently described fossils as calibration points (e.g., Mystacodon selenensis) that have not been used before. Integration of new fossil dates in the context of autocorrelated rates indicate that the diversification of Crown Cetacea began before the Late Eocene and the divergence of Crown Delphinidae as early as the Middle Miocene.

opencc-zeroOct 2019View details →
dryad32/100

Assessment of targeted enrichment locus capture across time and museums using odonate specimens

<p><span>The use of gDNA isolated from museum specimens for high throughput sequencing, especially targeted sequencing in the context of phylogenetics, is a common practice. Yet, little understanding has been focused on comparing the quality of DNA and the results of sequencing museum DNAs. Dragonflies and damselflies are ubiquitous in freshwater ecosystems and are commonly collected and preserved insects in museum collections hence their use in this study. However, the history of odonate preservation across time and museums have resulted in wide variability in the success of viable DNA extraction, necessitating an assessment of their usefulness in genetic studies. Using Anchored Hybrid Enrichment probes, we sequenced DNA from samples at two museums, 48 from the American Museum of Natural History (AMNH) in NYC, USA, and 46 from the Naturalis Biodiversity Center (RMNH) in Leiden, Netherlands ranging from global collection localities and across a 120-year time span. We recovered at least 4 loci out of an &gt;1000 locus probe set for all samples, with the average capture being ~385 loci. Neither specimen age nor size was a good predictor of locus capture, but recapture rates differed significantly between museums. Samples from the AMNH had lower overall locus capture than the RMNH, perhaps due to differences in specimen storage over time.</span></p>

opencc-zeroMay 2023View details →
dryad32/100

Data from: Evaluating the performance of targeted sequence capture, RNA-Seq, and degenerate-primer PCR cloning for sequencing the largest mammalian multigene family

Open the record for dataset details and reuse information.

publicSep 2019View details →
dryad32/100

Data from: Brooding phylogenomics: Target-capture probe sets for the analysis of ultraconserved elements in the Peracarida

Open the record for dataset details and reuse information.

publicNov 2025View details →
dryad32/100

Data from: A Phylogenomic analysis of Genipa (Rubiaceae) using target sequence capture data

Open the record for dataset details and reuse information.

publicDec 2024View details →
dryad32/100

Data from: Phylogenomics of horned lizards (genus: Phrynosoma) using targeted sequence capture data

Open the record for dataset details and reuse information.

publicJun 2016View details →
dryad32/100

Alignments from: Gene count from target sequence capture places three whole genome duplication events in Hibiscus L. (Malvaceae)

Open the record for dataset details and reuse information.

publicJan 2021View details →
dryad32/100

Data from: Phylogenomic analyses of Sabal (Arecaceae) species relationships using targeted sequence capture

Open the record for dataset details and reuse information.

publicMay 2015View details →
dryad32/100

Data from: Phylogenomic resolution of the cetacean tree of life using target sequence capture

Open the record for dataset details and reuse information.

publicOct 2019View details →
dryad32/100

A new approach using targeted sequence capture for phylogenomic studies across Cactaceae

Open the record for dataset details and reuse information.

publicMar 2022View details →
dryad32/100

Assessment of targeted enrichment locus capture across time and museums using odonate specimens

Open the record for dataset details and reuse information.

publicMay 2023View details →
dryad32/100

Data from: Phylogenetics of moth-like butterflies (Papilionoidea: Hedylidae) based on a new 13-locus target capture probe set

Open the record for dataset details and reuse information.

publicSep 2018View details →
dryad32/100

Data from: Comparison of taxon-specific versus general locus sets for targeted sequence capture for plant phylogenomics

Open the record for dataset details and reuse information.

publicJan 2019View details →
dryad28/100

Target-capture phylogenomics provide insights on gene and species tree discordances in Old World Treefrogs (Anura: Rhacophoridae)

<p>Genome-scale data have greatly facilitated the resolution of recalcitrant nodes that Sanger-based datasets have been unable to resolve. However, phylogenomic studies continue to utilize traditional methods such as bootstrapping to estimate branch support; and high bootstrap values are still interpreted as providing strong support for the correct topology. Furthermore, relatively little attention is given to assessing discordances between gene and species trees, and the underlying processes that produce phylogenetic conflict. We generated novel genomic datasets to characterize and determine the causes of discordance in Old World Treefrogs (Family: Rhacophoridae)—a group that is fraught with conflicting and poorly supported topologies among major clades. We showed that incomplete lineage sorting was present at all nodes that exhibited high levels of discordance, which was caused by extremely short internal branches. We also clearly demonstrate that bootstrap values do not reflect uncertainty or confidence for the correct topology, and hence, should not be used as a measure of branch support in phylogenomic datasets. Overall, we showed that species tree inference can be improved using a total-evidence and multi-faceted approach that utilizes the most amount of data and considers results from different analytical methods and datasets.</p>

opencc-zeroAug 2020View details →
dryad28/100

Data from: A dedicated target capture approach reveals variable genetic markers across micro- and macro-evolutionary time scales in palms

Understanding the genetics of biological diversification across micro- and macro-evolutionary time scales is a vibrant field of research for molecular ecologists as rapid advances in sequencing technologies promise to overcome former limitations. In palms, an emblematic, economically and ecologically important plant family with high diversity in the tropics, studies of diversification at the population and species levels are still hampered by a lack of genomic markers suitable for the genotyping of large numbers of recently diverged taxa. To fill this gap, we used a whole genome sequencing approach to develop target sequencing for molecular markers in 4,184 genome regions, including 4,051 genes and 133 non-genic putatively neutral regions. These markers were chosen to cover a wide range of evolutionary rates allowing future studies at the family, genus, species and population levels. Special emphasis was given to the avoidance of copy number variation during marker selection. In addition, a set of 149 well-known sequence regions previously used as phylogenetic markers by the palm biological research community were included in the target regions, to open the possibility to combine and jointly analyse already available data sets with genomic data to be produced with this new toolkit. The bait set was effective for species belonging to all three palm subfamilies tested (Arecoideae, Ceroxyloideae and Coryphoideae), with high mapping rates, specificity and efficiency. The number of high quality Single Nucleotide Polymorphisms (SNPs) detected at both the subfamily and population levels facilitates efficient analyses of genomic diversity across micro- and macro-evolutionary time scales.

opencc-zeroDec 2017View details →
dryad28/100

Data from: Comparison of target-capture and restriction-site associated DNA sequencing for phylogenomics: a test in cardinalid tanagers (Aves, genus: Piranga)

Restriction-site associated DNA sequencing (RAD-seq) and target capture of specific genomic regions, such as ultraconserved elements (UCEs), are emerging as two of the most popular methods for phylogenomics using reduced-representation genomic datasets. These two methods were designed to target different evolutionary timescales: RAD-seq was designed for population-genomic level questions and UCEs for deeper phylogenetics. The utility of both datasets to infer phylogenies across a variety of taxonomic levels has not been adequately compared within the same taxonomic system. Additionally, the effects of uninformative gene trees on species tree analyses (for target capture data) have not been explored. Here, we utilize RAD-seq and UCE data to infer a phylogeny of the bird genus Piranga. The group has a range of divergence dates (0.5 my – 6 my), contains eleven recognized species, and lacks a resolved phylogeny. We compared two species tree methods for the RAD-seq data and six species tree methods for the UCE data. Additionally, in the UCE data, we analyzed a complete matrix as well as datasets with only highly informative loci. A complete matrix of 189 UCE loci with ten or more parsimony informative (PI) sites, and an ~80% complete matrix of 1128 PI SNPs (from RAD-seq) yield the same fully resolved phylogeny of Piranga. We inferred non-monophyletic relationships of P. lutea individuals, with all other a priori species identified as monophyletic. Finally, we found that species tree analyses that included predominantly uninformative gene trees provided strong support for different topologies, with consistent phylogenetic results when limiting species tree analyses to highly informative loci or only using less informative loci with concatenation or methods meant for SNPs alone.

opencc-zeroDec 2015View details →
dryad28/100

Data from: Targeted capture of complete coding regions across divergent species

Despite continued advances in sequencing technologies, there is a need for methods that can efficiently sequence large numbers of genes from diverse species. One approach to accomplish this is targeted capture (hybrid enrichment). While these methods are well established for genome resequencing projects, cross-species capture strategies are still being developed and generally focus on the capture of conserved regions, rather than complete coding regions from specific genes of interest. The resulting data is thus useful for phylogenetic studies, but the wealth of comparative data that could be used for evolutionary and functional studies is lost. Here we design and implement a targeted capture method that enables recovery of complete coding regions across broad taxonomic scales. Capture probes were designed from multiple reference species and extensively tiled in order to facilitate cross-species capture. Using novel bioinformatics pipelines we were able to recover nearly all of the targeted genes with high completeness from species that were up to 200 myr divergent. Increased probe diversity and tiling for a subset of genes had a large positive effect on both recovery and completeness. The resulting data produced an accurate species tree, but importantly this same data can also be applied to studies of molecular evolution and function that will allow researchers to ask larger questions in broader phylogenetic contexts. Our method demonstrates the utility of cross-species approaches for the capture of full length coding sequences, and will substantially improve the ability for researchers to conduct large-scale comparative studies of molecular evolution and function.

opencc-zeroDec 2016View details →
dryad28/100

Data from: Target capture and massively parallel sequencing of ultraconserved elements for comparative studies at shallow evolutionary time scales

Comparative genetic studies of non-model organisms are transforming rapidly due to major advances in sequencing technology. A limiting factor in these studies has been the identification and screening of orthologous loci across an evolutionarily distant set of taxa. Here, we evaluate the efficacy of genomic markers targeting ultraconserved DNA elements (UCEs) for analyses at shallow evolutionary timescales. Using sequence capture and massively parallel sequencing to generate UCE data for five co-distributed Neotropical rainforest bird species, we recovered 776–1516 UCE loci across the five species. Across species, 53–77% of the loci were polymorphic, containing between 2.0 and 3.2 variable sites per polymorphic locus, on average. We performed species tree construction, coalescent modeling, and species delimitation, and we found that the five co-distributed species exhibited discordant phylogeographic histories. We also found that species trees and divergence times estimated from UCEs were similar to the parameters obtained from mtDNA. The species that inhabit the understory had older divergence times across barriers, contained a higher number of cryptic species, and exhibited larger effective population sizes relative to the species inhabiting the canopy. Because orthologous UCEs can be obtained from a wide array of taxa, are polymorphic at shallow evolutionary timescales, and can be generated rapidly at low cost, they are an effective genetic marker for studies investigating evolutionary patterns and processes at shallow timescales.

opencc-zeroDec 2013View details →
dryad28/100

Data from: Targeted sequence capture and resequencing implies a predominant role of regulatory regions in the divergence of a sympatric lake whitefish species pair (Coregonus clupeaformis)

Latest technological developments in evolutionary biology bring new challenges in documenting the intricate genetic architecture of species in the process of divergence. Sympatric populations of lake whitefish represent one of the key systems to investigate this issue. Despite the value of random genotype-by-sequencing methods and decreasing cost of sequencing technologies, it remains challenging to investigate variation in coding regions, especially in the case of recently duplicated genomes as in salmonids, as this greatly complicates whole genome resequencing. We thus designed a sequence capture array targeting 2773 annotated genes to document the nature and the extent of genomic divergence between sympatric dwarf and normal whitefish. Among the 2728 genes successfully captured, a total of 2182 coding and 10 415 noncoding putative single-nucleotide polymorphisms (SNPs) were identified after applying a first set of basic filters. A genome scan with a quality-refined selection of 2203 SNPs identified 267 outlier SNPs in 210 candidate genes located in genomic regions potentially involved in whitefish divergence and reproductive isolation. We found highly heterogeneous FST estimates among SNP loci. There was an overall low level of coding polymorphism, with a predominance of noncoding mutations among outliers. The heterogeneous patterns of divergence among loci confirm the porous nature of genomes during speciation with gene flow. Considering that few protein-coding mutations were identified as highly divergent, our results, along with previous transcriptomic studies, imply that changes in regulatory regions most likely had a greater role in the process of whitefish population divergence than protein-coding mutations. This study is the first to demonstrate the efficiency of large-scale targeted resequencing for a nonmodel species with such a large and unsequenced genome.

opencc-zeroDec 2012View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record