Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

42

datasets available to search

ShareScore release 0.7.1

Reset

Dataset results

42 results for “targeted sequence capture”

Learn how ShareScore rates datasets ↗
zenodo44/100

Whole-genome capture and sequencing of Mycobacterium tuberculosis directly from clinical samples - Design of RNA oligonucleotide baits for Agilent Technologies' SureSelect target enrichment

<p>This dataset comprises the sequence of <strong>44&nbsp;278&nbsp;RNA oligonucleotide &quot;baits&quot; (120 bp each) </strong>designed to perform&nbsp;<strong>whole-genome capture and sequencing of <em>Mycobacterium tuberculosis</em>&nbsp;directly from clinical samples</strong>&nbsp;(DNA)&nbsp;using Agilent Technologies&rsquo; SureSelect target enrichment system following the Illumina paired-end multiplexed sequencing library protocol.&nbsp;</p> <p>RNA oligonucleotide &ldquo;baits&rdquo; were designed to span the &sim;4.5 Mb of the <em>M. tuberculosis</em> genome. In brief, the reference genome sequence of the MTBC H37Rv strain (Genbank #AL123456) was <em>in silico</em> fragmented into 120 bp sequences twice, to ensure an overlap of 60 bp between sequences. Due to their rich GC content, which could interfere with DNA capture, all MTBC genes of the PE, PPE and PE-PGRS family were also independently fragmented into 120 bp sequences, in order to increase capture sensitivity. All resulting sequences were BLASTn searched against the Human Genomic + Transcript database to excluded homologous sequences to the human genome. Overall, a total of 42,278 RNA probes were generated and this custom bait library was then uploaded to the SureDesign software (https://earray.chem.agilent.com/suredesign) and synthesized by Agilent Technologies. During synthesis, the 2198 sequences complementary to the PE, PPE and PE-PGRS family were unbalanced 8:1 to potentiate capture.</p> <p>More details can be found in the following publication:</p> <p>- Macedo, R., Isidro, J., Ferreira, R., Pinto, M., Borges, V., Duarte, S., Vieira, L., &amp; Gomes, J. P. (2023). Molecular Capture of&nbsp;<em>Mycobacterium tuberculosis</em>&nbsp;Genomes Directly from Clinical Samples: A Potential Backup Approach for Epidemiological and Drug Susceptibility Inferences.&nbsp;<em>International journal of molecular sciences</em>,&nbsp;<em>24</em>(3), 2912. https://doi.org/10.3390/ijms24032912</p>

opencc-by-4.0Jan 2023View details →
dryad40/100

Morphological and DNA sequence data generated by Sanger sequencing and target capture methods for moss plants in the genus Fissidens from herbarium specimens

<p><span>Morphological evolution in mosses has long been hypothesized to accompany shifts in microhabitats and can be tested using comparative phylogenetics. These lines of inquiry have developed substantially, in part, by target capture sequencing allowing for phylogenomic scale data generated from herbarium specimens. In the present study, we test the relationship between taxonomically important morphological characters in the moss genus <em>Fissidens</em>, using both a 400-locus dataset generated using a target-capture approach as well as a three-locus phylogeny generated using sanger sequencing. Phylogenetic trees were generated using ASTRAL and Bayesian Inference and used to test the monophyly of subgenera/sections and provided the basis for ancestral character reconstruction and phylogenetic correlation analyses among five morphological characters as well as habitat moisture scored from literature. The characters <em>axillary hyaline nodules</em>, <em>limbidium</em>, <em>costa</em>, and <em>peristome morphology</em> as well as <em>sexual system</em>, <em>minimum habitat moisture</em>, <em>average habitat moisture</em>, <em>maximum habitat moisture</em>, and <em>habitat moisture niche breadth</em> each exhibit statistically significant phylogenetic signal. Significant correlations were found between the limbidium (phyllid/leaf border) and habitat moisture niche breadth, which could be interpreted as a more extensive <em>limbidium</em> enabling species to survive across a wider variety of habitats. Correlations were also found between <em>costa anatomy</em> and the <em>limbidum</em> of the gametophyte and sporophyte <em>peristome</em> <em>morphology</em>, as well as <em>average habitat moisture</em> and <em>sexual system</em>. Continued exploration of the relationships between morphological evolution, life history, and habitat will enable us to expand our understanding of functional morphology in mosses.</span></p>

opencc-zeroJun 2022View details →
dryad40/100

Morphological and DNA sequence data generated by Sanger sequencing and target capture methods for moss plants in the genus Fissidens from herbarium specimens

Open the record for dataset details and reuse information.

publicJun 2022View details →
zenodo36/100

BRAVO target sequence capture V3

<p>This data set contains the bait sequences for a sequence capture library targeting specific genes in Brassica ssp.. Source sequences have been manually selected and processed with BaitLibraryBuilder (<a href="https://github.com/steuernb/BaitLibraryBuilder">https://github.com/steuernb/BaitLibraryBuilder</a>).</p>

opencc-by-4.0Jan 2021View details →
zenodo36/100

Data in support of Using target sequence capture to improve the phylogenetic resolution of a rapid radiation in New Zealand Veronica

<p>Includes alignments and trees for the analysis found in Thomas et al. 2021, Using target sequence capture to improve the phylogenetic resolution of a rapid radiation in New Zealand Veronica;&nbsp;American Journal of Botany, Special Issue: Exploring Angiosperms353: a Universal Toolkit for Flowering Plant Phylogenomics. Alignments comprise subsets of Angiosperms353 genes given each filtering scheme (full, intersection, sortadate_BP, sortadate_TL) and gene type/subset (exons, introns, supercontigs), and for markers downloaded from GenBank, as explained in the Methods section of Thomas et al. 2021. Trees were included for each of these alignments from IQtree and Astral; SVDquartets tree was only estimated for the full set of supercontigs. Gene trees were generated with IQtree. Tree files are named differently than the final manuscript; refer to the number of genes specified in Fig 1 of Thomas et al, 2021 and specified in each filename to identify filtering scheme.&nbsp;Raw sequence reads are available on the Sequence Read Archive at <a href="http://www.ncbi.nlm.nih.gov/bioproject/715342">http://www.ncbi.nlm.nih.gov/bioproject/715342</a>.</p>

opencc-by-4.0Jun 2022View details →
dryad32/100

Alignments from: Gene count from target sequence capture places three whole genome duplication events in Hibiscus L. (Malvaceae)

<p class="BodyA"><span><b>Background:</b> The great diversity in plant genome size and chromosome number is partly due to polyploidization (i.e., genome doubling events). The differences in genome size and chromosome number among diploid plant species can be a window into the intriguing phenomenon of past genome doubling that may be obscured through time by the process of diploidization. The genus <i>Hibiscus </i>L. (Malvaceae) has a wide diversity of chromosome numbers and a complex genomic history. <i>Hibiscus </i>is ideal for exploring past genomic events because although two ancient genome duplication events have been identified, more are likely to be found due to its diversity of chromosome numbers. To reappraise the history of whole genome duplication events, we tested  three alternative scenarios describing different polyploidization events.</span></p> <p class="BodyA"><span><b>Results:</b> Using target sequence capture, we designed a new probe set for <i>Hibiscus </i>and generated 87 orthologous genes from four diploid species. We detected paralogues in &gt;54% putative single-copy genes. 34 of these genes were selected for testing three different genome duplication scenarios using gene counting. All species of <i>Hibiscus</i> sampled shared one genome duplication with <i>H. syriacus</i> and one whole genome duplication occurred along the branch leading to <i>H. syriacus</i>.</span></p> <p class="BodyA"><span><b>Conclusions:</b> Here, we corroborated the independent genome doubling previously found in the lineage leading to <i>H. syriacus </i>and a shared genome doubling of this lineage and the remainder of <i>Hibiscus</i>. Additionally, we found a previously undiscovered genome duplication shared by the /Pavonia and /Malvaviscus clades (both nested within <i>Hibiscus</i>) with the occurrences of two copies in what were otherwise single-copy genes. Our results highlight the complexity of genomic diversity in some plant groups, which makes orthology assessment and accurate phylogenomic inference difficult.</span></p>

opencc-zeroJan 2021View details →
dryad32/100

Data from: Comparison of taxon-specific versus general locus sets for targeted sequence capture for plant phylogenomics

Premise of the study: Targeted sequence capture can be used to efficiently gather sequence data for large numbers of loci, such as single-copy nuclear loci. Most published studies in plants have used taxon-specific locus sets developed individually for a clade using multiple genomic and transcriptomic resources. General locus sets can also be developed from loci that have been identified as single-copy and having orthologs in large clades of plants. Methods: We identify and compare a taxon-specific locus set and three general locus sets (COSII, APVO SSC, PPR) for targeted sequence capture in Buddleja (Scrophulariaceae) and outgroups. We evaluate their performance in terms of assembly success, sequence variability, and resolution and support of inferred phylogenetic trees. Results: The taxon-specific locus set had the most target loci. Assembly success was high for all locus sets in Buddleja samples. For outgroups, general locus sets had greater assembly success. Taxon-specific and PPR loci had the highest average variability. The taxon-specific dataset produced the best supported tree, but all datasets showed improved resolution over previous non-sequence capture datasets. Discussion: General loci can be a useful source of sequence capture targets, especially if multiple genomic resources are not available for a taxon.

opencc-zeroDec 2017View details →
dryad32/100

Data from: Phylogenomics of horned lizards (genus: Phrynosoma) using targeted sequence capture data

New genome sequencing techniques are enabling phylogenetic studies to scale-up from using a handful of loci to hundreds or thousands of loci from throughout the genome. In this study, we use targeted sequence capture (TSC) data from 540 ultraconserved elements and 44 protein-coding genes to estimate the phylogenetic relationships among all 17 species of horned lizards in the genus Phrynosoma. Previous molecular phylogenetic analyses of Phrynosoma based on a few nuclear genes, restriction site associated DNA (RAD) sequencing, or mitochondrial DNA (mtDNA) have produced conflicting relationships. Some of these conflicts are likely the result of rapid speciation at the start of Phrynosoma diversification, whereas other examples of gene tree discordance appear to be caused by active and residual traces of hybridization. Concatenation and coalescent-based species tree phylogenetic analyses of these new TSC data support the same topology, and a divergence dating analysis suggests that the Phrynosoma crown group is up to 30 million years old. The new phylogenomic tree supports the recognition of four main clades within Phrynosoma, including Anota (P. mcallii, P. solare, and the P. coronatum complex), Doliosaurus (P. modestum, P. goodei, and P. platyrhinos), Tapaja (P. ditmarsi, P. douglasii, P. hernandesi, and P. orbiculare), and Brevicauda (P. braconnieri, P. sherbrookei, and P. taurus). The phylogeny provides strong support for the relationships among all species of Phrynosoma and provides a robust new framework for conducting comparative analyses.

opencc-zeroDec 2014View details →
dryad32/100

Data from: Phylogenomic analyses of Sabal (Arecaceae) species relationships using targeted sequence capture

With the increasing availability of high-throughput sequencing, phylogenetic analyses are no longer constrained by the limited availability of a few loci. Here, we describe a sequence capture methodology, which we used to collect data for analyses of diversification within Sabal (Arecaceae), a palm genus native to the south-eastern USA, Caribbean, Bermuda and Central America. RNA probes were developed and used to enrich DNA samples for putatively low copy nuclear genes and the plastomes for all Sabal species and two outgroup species. Sequence data were generated on an Illumina MiSeq sequencer and target sequences were assembled using custom workflows. Both coalescence and supermatrix analyses of 133 nuclear genes were used to estimate species trees relationships. Plastid genomes were also analysed, yielding generally poor resolution with regard to species relationships. Species relationships described in both nuclear gene and plastome sequences largely reflect the biogeography of the group and, to a lesser extent, previous morphology-based hypotheses. Beyond the biological implications, this research validates a high-throughput methodology for generating a large number of genes for coalescence-based phylogenetic analyses in plant lineages.

opencc-zeroDec 2014View details →
dryad32/100

A new approach using targeted sequence capture for phylogenomic studies across Cactaceae

<p>Relationships within the major clades of Cactaceae are relatively well known based on DNA sequence data mostly from the chloroplast genome. Nevertheless, some nodes along the backbone of the phylogeny, and especially generic and species-level relationships, remain poorly resolved and are in need of more informative genetic markers. In this study, we propose a new approach to solve the relationships within Cactaceae, applying a targeted sequence capture pipeline. We designed a custom probe set for Cactaceae using MarkerMiner and complemented it with the Angiosperms353 probe set. We then tested both probe sets against 36 different transcriptomes using Hybpiper preferentially retaining phylogenetically informative loci and reconstructed the relationships using RAxML-NG and Astral. Finally, we tested each probe set through sequencing 96 accessions, representing 88 species across Cactaceae. Our preliminary analyses recovered a well-supported phylogeny across Cactaceae with a near identical topology among major clade relationships as that recovered with plastome data. As expected, however, we found incongruences in relationships when comparing our nuclear probe set results to plastome datasets, especially at the generic level. Our results reveal great potential for the combination of Cactaceae-specific and Angiosperm353 probe set application to improve phylogenetic resolution for Cactaceae and for other studies.</p>

opencc-zeroMar 2022View details →
dryad32/100

Data from: Phylogenomic resolution of the cetacean tree of life using target sequence capture

The evolution of the cetaceans, from their early transition to an aquatic lifestyle to their subsequent diversification, has been the subject of numerous studies. However, while the higher-level relationships among cetacean families have been largely settled, several aspects of the systematics within these groups remain unresolved. Problematic clades include the oceanic dolphins (37 spp.), which have experienced a recent rapid radiation, and the beaked whales (22 spp.), which have not been investigated in detail using nuclear loci. The combined application of high-throughput sequencing with techniques that target specific genomic sequences provide a powerful means of rapidly generating large volumes of orthologous sequence data for use in phylogenomic studies. To elucidate the phylogenetic relationships within the Cetacea, we combined sequence capture with Illumina sequencing to generate data for ~3200 protein-coding genes for 68 cetacean species and their close relatives including the pygmy hippopotamus. By combining data from &gt;38,000 exons with existing sequences from 11 cetaceans and seven outgroup taxa, we produced the first comprehensive comparative genomic dataset for cetaceans, spanning 6,527,596 aligned base pairs and 89 taxa. Phylogenetic trees reconstructed with maximum likelihood and Bayesian inference of concatenated loci, as well as with coalescence analyses of individual gene trees, produced mostly concordant and well-supported trees. Our results completely resolve the relationships among beaked whales as well as the contentious relationships among ocean dolphins, especially the problematic subfamily Delphininae, which includes the common and bottlenose dolphins. We performed Bayesian estimation of species divergence times using MCMCtree, integrating recently described fossils as calibration points (e.g., Mystacodon selenensis) that have not been used before. Integration of new fossil dates in the context of autocorrelated rates indicate that the diversification of Crown Cetacea began before the Late Eocene and the divergence of Crown Delphinidae as early as the Middle Miocene.

opencc-zeroOct 2019View details →
dryad32/100

Data from: Evaluating the performance of targeted sequence capture, RNA-Seq, and degenerate-primer PCR cloning for sequencing the largest mammalian multigene family

Open the record for dataset details and reuse information.

publicSep 2019View details →
dryad32/100

Data from: A Phylogenomic analysis of Genipa (Rubiaceae) using target sequence capture data

Open the record for dataset details and reuse information.

publicDec 2024View details →
dryad32/100

Data from: Phylogenomics of horned lizards (genus: Phrynosoma) using targeted sequence capture data

Open the record for dataset details and reuse information.

publicJun 2016View details →
dryad32/100

Alignments from: Gene count from target sequence capture places three whole genome duplication events in Hibiscus L. (Malvaceae)

Open the record for dataset details and reuse information.

publicJan 2021View details →
dryad32/100

Data from: Phylogenomic analyses of Sabal (Arecaceae) species relationships using targeted sequence capture

Open the record for dataset details and reuse information.

publicMay 2015View details →
dryad32/100

Data from: Phylogenomic resolution of the cetacean tree of life using target sequence capture

Open the record for dataset details and reuse information.

publicOct 2019View details →
dryad32/100

A new approach using targeted sequence capture for phylogenomic studies across Cactaceae

Open the record for dataset details and reuse information.

publicMar 2022View details →
dryad32/100

Data from: Comparison of taxon-specific versus general locus sets for targeted sequence capture for plant phylogenomics

Open the record for dataset details and reuse information.

publicJan 2019View details →
dryad28/100

Data from: Comparison of target-capture and restriction-site associated DNA sequencing for phylogenomics: a test in cardinalid tanagers (Aves, genus: Piranga)

Restriction-site associated DNA sequencing (RAD-seq) and target capture of specific genomic regions, such as ultraconserved elements (UCEs), are emerging as two of the most popular methods for phylogenomics using reduced-representation genomic datasets. These two methods were designed to target different evolutionary timescales: RAD-seq was designed for population-genomic level questions and UCEs for deeper phylogenetics. The utility of both datasets to infer phylogenies across a variety of taxonomic levels has not been adequately compared within the same taxonomic system. Additionally, the effects of uninformative gene trees on species tree analyses (for target capture data) have not been explored. Here, we utilize RAD-seq and UCE data to infer a phylogeny of the bird genus Piranga. The group has a range of divergence dates (0.5 my – 6 my), contains eleven recognized species, and lacks a resolved phylogeny. We compared two species tree methods for the RAD-seq data and six species tree methods for the UCE data. Additionally, in the UCE data, we analyzed a complete matrix as well as datasets with only highly informative loci. A complete matrix of 189 UCE loci with ten or more parsimony informative (PI) sites, and an ~80% complete matrix of 1128 PI SNPs (from RAD-seq) yield the same fully resolved phylogeny of Piranga. We inferred non-monophyletic relationships of P. lutea individuals, with all other a priori species identified as monophyletic. Finally, we found that species tree analyses that included predominantly uninformative gene trees provided strong support for different topologies, with consistent phylogenetic results when limiting species tree analyses to highly informative loci or only using less informative loci with concatenation or methods meant for SNPs alone.

opencc-zeroDec 2015View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record