Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
98
datasets available to search
ShareScore release 0.9.0
Dataset results
98 results for “Targeted capture”
Patterns of Speciation in a Parapatric Pair of Saturnia Moths as Revealed by Target Capture
<p>This is the dataset for the manuscript entitled Patterns of Speciation in a Parapatric Pair of Saturnia Moths as Revealed by Target Capture. This study helps in the delimitation of a parapatric pair of two species of moths in a complex distribution considering their evolutionary history with the help of the Target Capture method.</p>
Phylogenomics of Gesneriaceae using targeted capture of nuclear genes
<p>Gesneriaceae (ca. 3400 species) is a pantropical plant family with a wide range of growth form and floral morphology that are associated with repeated adaptations to different environments and pollinators. Although Gesneriaceae systematics has been largely improved by the use of Sanger sequencing data, our understanding of the evolutionary history of the group is still far from complete due to the limited number of informative characters provided by this type of data. To overcome this limitation, we developed here a Gesneriaceae-specific gene capture kit targeting 830 single-copy loci (776,754 bp in total), including 279 genes from the Universal Angiosperm-353 kit. With an average of 557,600 reads and 87.8% gene recovery, our target capture was successful across the family Gesneriaceae and also in other families of Lamiales. From our bait set, we selected the most informative 418 loci to resolve phylogenetic relationships across the entire Gesneriaceae family using maximum likelihood and coalescent-based methods. Upon testing the phylogenetic performance of our baits on 78 taxa representing 20 out of 24 subtribes within the family, we showed that our data provided high support for the phylogenetic relationships among the major lineages, and were able to provide high resolution within more recent radiations. Overall, the molecular resources we developed here open new perspectives for the study of Gesneriaceae phylogeny at different taxonomical levels and the identification of the factors underlying the diversification of this plant group. </p>
Whole-genome capture and sequencing of Mycobacterium tuberculosis directly from clinical samples - Design of RNA oligonucleotide baits for Agilent Technologies' SureSelect target enrichment
<p>This dataset comprises the sequence of <strong>44 278 RNA oligonucleotide "baits" (120 bp each) </strong>designed to perform <strong>whole-genome capture and sequencing of <em>Mycobacterium tuberculosis</em> directly from clinical samples</strong> (DNA) using Agilent Technologies’ SureSelect target enrichment system following the Illumina paired-end multiplexed sequencing library protocol. </p> <p>RNA oligonucleotide “baits” were designed to span the ∼4.5 Mb of the <em>M. tuberculosis</em> genome. In brief, the reference genome sequence of the MTBC H37Rv strain (Genbank #AL123456) was <em>in silico</em> fragmented into 120 bp sequences twice, to ensure an overlap of 60 bp between sequences. Due to their rich GC content, which could interfere with DNA capture, all MTBC genes of the PE, PPE and PE-PGRS family were also independently fragmented into 120 bp sequences, in order to increase capture sensitivity. All resulting sequences were BLASTn searched against the Human Genomic + Transcript database to excluded homologous sequences to the human genome. Overall, a total of 42,278 RNA probes were generated and this custom bait library was then uploaded to the SureDesign software (https://earray.chem.agilent.com/suredesign) and synthesized by Agilent Technologies. During synthesis, the 2198 sequences complementary to the PE, PPE and PE-PGRS family were unbalanced 8:1 to potentiate capture.</p> <p>More details can be found in the following publication:</p> <p>- Macedo, R., Isidro, J., Ferreira, R., Pinto, M., Borges, V., Duarte, S., Vieira, L., & Gomes, J. P. (2023). Molecular Capture of <em>Mycobacterium tuberculosis</em> Genomes Directly from Clinical Samples: A Potential Backup Approach for Epidemiological and Drug Susceptibility Inferences. <em>International journal of molecular sciences</em>, <em>24</em>(3), 2912. https://doi.org/10.3390/ijms24032912</p>
A two-tier bioinformatic pipeline to develop probes for target capture of nuclear loci with applications in Melastomataceae
<p><b><i>Premise of the study</i></b><b>: </b>Putatively single-copy nuclear (SCN) loci, identified using genomic resources of closely related species, are ideal for phylogenomic inference. However, suitable genomic resources are not available for many clades, including Melastomataceae. We introduce a versatile approach to identify SCN loci for clades with few genomic resources and use it to develop probes for target enrichment<i> </i>in the distantly related <i>Memecylon</i> and <i>Tibouchina</i> (Melastomataceae).</p> <p><b><i>Methods</i></b>: We present a two-tiered pipeline. First, we identified putatively SCN loci using MarkerMiner and transcriptomes from distantly related species in Melastomataceae. Published loci and genes of functional significance were added (384 total loci). Second, using HybPiper, we retrieved 689 homologous template sequences for these loci using genome-skimming data from within the focal clades.</p> <p><b><i>Results</i></b>: We sequenced 193 loci from both <i>Memecylon</i> and <i>Tibouchina</i>, with probes designed from 56 template sequences successfully targeting sequences in both clades. Probes designed from genome-skimming data within a focal clade were more successful than probes designed from other sources.</p> <p><b><i>Discussion: </i></b>Our pipeline successfully identified and targeted SCN loci in <i>Memecylon </i>and <i>Tibouchina</i>, enabling phylogenomic studies in both clades and potentially across Melastomataceae. This pipeline could be easily applied to other clades with few genomic resources. </p>
Taxon-specific or universal? Using target capture to study the evolutionary history of a rapid radiation
<p>Target capture emerged as an important tool for phylogenetics and population genetics in non-model taxa. Whereas developing taxon-specific capture probes requires sustained efforts, available universal kits may have a lower power to reconstruct relationships at shallow phylogenetic scales and within rapidly radiating clades. We present here a newly-developed target capture set for Bromeliaceae, a large and ecologically-diverse plant family with highly variable diversification rates. The set targets 1,776 coding regions, including genes putatively involved in key innovations, with the aim to empower testing of a wide range of evolutionary hypotheses. We compare the relative power of this taxon-specific set, Bromeliad1776, to the universal Angiosperms353 kit. The taxon-specific set results in higher enrichment success across the entire family, however, the overall performance of both kits to reconstruct phylogenetic trees is relatively comparable, highlighting the vast potential of universal kits for resolving evolutionary relationships. For more detailed phylogenetic or population genetic analyses, e.g. the exploration of gene tree concordance, nucleotide diversity or population structure, the taxon-specific capture set presents clear benefits. We discuss the potential lessons that this comparative study provides for future phylogenetic and population genetic investigations, in particular for the study of evolutionary radiations.</p>
Morphological and DNA sequence data generated by Sanger sequencing and target capture methods for moss plants in the genus Fissidens from herbarium specimens
<p><span>Morphological evolution in mosses has long been hypothesized to accompany shifts in microhabitats and can be tested using comparative phylogenetics. These lines of inquiry have developed substantially, in part, by target capture sequencing allowing for phylogenomic scale data generated from herbarium specimens. In the present study, we test the relationship between taxonomically important morphological characters in the moss genus <em>Fissidens</em>, using both a 400-locus dataset generated using a target-capture approach as well as a three-locus phylogeny generated using sanger sequencing. Phylogenetic trees were generated using ASTRAL and Bayesian Inference and used to test the monophyly of subgenera/sections and provided the basis for ancestral character reconstruction and phylogenetic correlation analyses among five morphological characters as well as habitat moisture scored from literature. The characters <em>axillary hyaline nodules</em>, <em>limbidium</em>, <em>costa</em>, and <em>peristome morphology</em> as well as <em>sexual system</em>, <em>minimum habitat moisture</em>, <em>average habitat moisture</em>, <em>maximum habitat moisture</em>, and <em>habitat moisture niche breadth</em> each exhibit statistically significant phylogenetic signal. Significant correlations were found between the limbidium (phyllid/leaf border) and habitat moisture niche breadth, which could be interpreted as a more extensive <em>limbidium</em> enabling species to survive across a wider variety of habitats. Correlations were also found between <em>costa anatomy</em> and the <em>limbidum</em> of the gametophyte and sporophyte <em>peristome</em> <em>morphology</em>, as well as <em>average habitat moisture</em> and <em>sexual system</em>. Continued exploration of the relationships between morphological evolution, life history, and habitat will enable us to expand our understanding of functional morphology in mosses.</span></p>
A target capture approach for phylogenomic analyses at multiple evolutionary timescales in rosewoods (Dalbergia spp.) and the legume family (Fabaceae)
<p>Understanding the genetic changes associated with the evolution of biological diversity is of fundamental interest to molecular ecologists. The assessment of genetic variation at hundreds or thousands of unlinked genetic loci forms a sound basis to address questions ranging from micro- to macro-evolutionary timescales, and is now possible thanks to advances in sequencing technology. Major difficulties are associated with i) the lack of genomic resources for many taxa, especially from tropical biodiversity hotspots, ii) scaling the numbers of individuals analyzed and loci sequenced, and iii) building tools for reproducible bioinformatic analyses of such datasets. To address these challenges, we developed a set of target capture probes for phylogenomic studies of the highly diverse, pantropically distributed and economically significant rosewoods (<em>Dalbergia</em> spp.), explored the performance of an overlapping probe set for target capture across the legume family (Fabaceae), and built a general-purpose bioinformatics pipeline. Phylogenomic analyses of <em>Dalbergia</em> species from Madagascar yielded highly resolved and well supported hypotheses of evolutionary relationships. Population genomic analyses identified differences between closely related species and revealed the existence of a potentially new species, suggesting that the diversity of Malagasy <em>Dalbergia</em> species has been underestimated. Analyses at the family level corroborated previous findings by the recovery of monophyletic subfamilies and many well-known clades, as well as high levels of gene tree discordance, especially near the root of the family. The new genomic and bioinformatics resources will hopefully advance systematics and ecological genetics research in legumes, and promote conservation of the highly diverse and endangered <em>Dalbergia</em> rosewoods.</p>
Supplementary materials to: Nano-Strainer: a workflow for identification of single-copy nuclear loci for plant systematic studies, using target capture kits and Oxford Nanopore long reads
<p>In the paper associated with this dataset, a workflow is presented which enables the identification of single-/low-copy nuclear molecular markers for a plant group of interest, by mining data from a small representative target capture experiment done using a commercial probe kit and Oxford Nanopore long-read sequencing. The proposed pipeline first assesses sequence variability contained in the data from targeted loci and assigns reads to their respective genes, via a combined BLAST/clustering procedure. Cluster consensus sequences are then examined based on four pre-defined criteria presumably indicative for absence of paralogy. This is done by calculating four specialized indices; loci are ranked according to their performance in these indices, and top-scoring loci are considered putatively single- or low-copy. The approach can be applied to any probe set. As it relies on long reads, the contribution also provides template workflows for processing Nanopore-based target capture data. Identified loci can be used for NGS amplicon sequencing. For detection of possibly remaining paralogy in these data, which might occur in groups with rampant paralogy, the long-read assembly tool CANU is employed. The presented workflow can be useful for researchers dealing with reticulate or polyploidization phylogenetic histories in plants.</p> <p>The present dataset contains several documents supplementing the original paper. Its most important elements are a detailed description (alongside two graphical workflow figures) of all methods employed in the study, suitable for reproducing the steps of the workflow and also the wet-lab work. The workflow employs a collection of BASH, Python and R scripts which is available here, together with a detailed account on command line use in Linux. Also, reference sequences for the identified markers can be found as well as sequence alignments derived from the amplicon sequencing.</p>
Supplementary materials to: Nano-Strainer: a workflow for identification of single-copy nuclear loci for plant systematic studies, using target capture kits and Oxford Nanopore long reads
Open the record for dataset details and reuse information.
A target capture approach for phylogenomic analyses at multiple evolutionary timescales in rosewoods (Dalbergia spp.) and the legume family (Fabaceae)
Open the record for dataset details and reuse information.
A two-tier bioinformatic pipeline to develop probes for target capture of nuclear loci with applications in Melastomataceae
Open the record for dataset details and reuse information.
Morphological and DNA sequence data generated by Sanger sequencing and target capture methods for moss plants in the genus Fissidens from herbarium specimens
Open the record for dataset details and reuse information.
Taxon-specific or universal? Using target capture to study the evolutionary history of a rapid radiation
Open the record for dataset details and reuse information.
Canis lupus capture targets
Open the record for dataset details and reuse information.
Data from: Development and validation of a RAD-Seq target-capture based genotyping assay for routine application in advanced black tiger shrimp (Penaeus monodon) breeding programs
<p><i><span>Background</span></i></p> <p><span>The development of genome-wide genotyping resources has provided terrestrial livestock and crop industries with the unique ability to accurately assess genomic relationships between individuals, uncover the genetic architecture of commercial traits, as well as identify superior individuals for selection based on their specific genetic profile. Utilising recent advancements in <i>de-novo</i> genome-wide genotyping technologies, it is now possible to provide aquaculture industries with these same important genotyping resources, even in the absence of existing genome assemblies. Here, we present the development of a genome-wide SNP assay for the Black Tiger shrimp (<i>Penaeus monodon</i>) through utilisation of a reduced-representation whole-genome genotyping approach (DArTseq).</span></p> <p><i><span>Results</span></i></p> <p><span>Based on a single reduced-representation library, 31,262 polymorphic SNPs were identified across 650 individuals obtained from Australian wild stocks and commercial aquaculture populations. After filtering to remove SNPs with low read depth, low MAF, low call rate, deviation from HWE, and non-Mendelian inheritance, 7,542 high-quality SNPs were retained. From these, 4,236 high-quality genome-wide loci were selected for bates-probe development and 4,194 SNPs were included within a finalized target-capture genotype-by-sequence assay (DArTcap). This assay was designed for routine and cost effective commercial application in large scale breeding programs, and demonstrates higher confidence in genotype calls through increased call rate (from 80.2 </span>± 14.7 to 93.0% ± 3.5%<span>), </span>increased read depth (from 20.4 ± 15.6 to 80.0 ± 88.7<span>), as well as a 3-fold reduction in cost over traditional genotype-by-sequencing approaches.</span></p> <p><i><span>Conclusion</span></i></p> <p><span>Importantly, this assay equips the <em>P. monodon</em> industry with the ability to simultaneously assign parentage of communally reared animals, undertake genomic relationship analysis, manage mate pairings between cryptic family lines, as well as undertake advance studies of genome and trait architecture. Critically this assay can be cost effectively applied as <em>P. monodon</em> breeding programs transition to undertaking genomic selection.</span></p>
Genome-scale target capture of mitochondrial and nuclear environmental DNA from water samples
<p>Environmental DNA (eDNA) provides a promising supplement to traditional sampling methods for population genetic inferences, but current studies have almost entirely focused on short mitochondrial markers. Here, we develop one mitochondrial and one nuclear set of target capture probes for the whale shark (<i>Rhincodon typus</i>) and test them on seawater samples collected in Qatar to investigate the potential of target capture for eDNA-based population studies. The mitochondrial target capture successfully retrieved ~235x (90x-352x per base position) coverage of the whale shark mitogenome. Using a minor allele frequency of 5%, we find 29 variable sites throughout the mitogenome, indicative of at least five contributing individuals. We also retrieved numerous mitochondrial reads from an abundant non-target species mackerel tuna<i> </i>(<i>Euthynnus affinis</i>), showing a clear relation between sequence similarity to the capture probes and the number of captured reads. The nuclear target capture probes retrieved only few reads and polymorphic variants from the whale shark, but we successfully obtained millions of reads and thousands of polymorphic variants with different allele frequencies from <i>E</i>. <i>affinis</i>. We demonstrate that target capture of complete mitochondrial genomes and thousands of nuclear loci is possible from aquatic eDNA samples. Our results highlight that careful probe design, taking into account the range of divergence between target and non-target sequences as well as presence of non-target species at the sampling site, is crucial to consider. Environmental DNA sampling coupled with target capture approaches provide an efficient means with which to retrieve population genomic data from aggregating and spawning aquatic species.</p>
BRAVO target sequence capture V3
<p>This data set contains the bait sequences for a sequence capture library targeting specific genes in Brassica ssp.. Source sequences have been manually selected and processed with BaitLibraryBuilder (<a href="https://github.com/steuernb/BaitLibraryBuilder">https://github.com/steuernb/BaitLibraryBuilder</a>).</p>
Development of a genus-specific antigen capture ELISA for orthopoxviruses. Target selection and optimized screening
<p><strong>Raw data for quantification of anti surface protein antibody binding to vaccinia virus.</strong></p> <p>Method description</p> <p>Immuno-negative staining and electron microscopy were performed as described elsewhere (Laue, 2010). Briefly, purified VACV<sub>NYCBOH</sub> particles were inactivated by incubation in freshly prepared 2% PFA in 0.05 M HEPES (pH 7.2), sonicated and immobilized on sample supports for transmission electron microscopy. Biotinylated pAbs were titrated on BSA coated grids, until detection with 5 nm gold nanoparticle coupled streptavidin (British Biocell, Cardiff, United Kingdom) resulted in the same mean background labelling density of ~10 particles per view field at a, 87,000-fold magnification (anti-A27: 0.7 µg/mL; anti-D8: 2 µg/mL; anti-H3: 0.9 µg/mL; anti-L1: 1.9 µg/mL). Negative staining was performed with either 0.1 or 0.5% uranyl acetate solution. For quantification, only IMV particles of the mulberry form, which were found isolated from other particles, were analyzed. Randomized sampling was done in 22 evenly distributed mesh areas with five viral particles analyzed per area. Imaging was done with a Tecnai 12 BioTwin (FEI Corp.) at 120 kV and a 1k digital CCD camera (Megaview III, Olympus Soft Imaging Solutions).</p>
Data in support of Using target sequence capture to improve the phylogenetic resolution of a rapid radiation in New Zealand Veronica
<p>Includes alignments and trees for the analysis found in Thomas et al. 2021, Using target sequence capture to improve the phylogenetic resolution of a rapid radiation in New Zealand Veronica; American Journal of Botany, Special Issue: Exploring Angiosperms353: a Universal Toolkit for Flowering Plant Phylogenomics. Alignments comprise subsets of Angiosperms353 genes given each filtering scheme (full, intersection, sortadate_BP, sortadate_TL) and gene type/subset (exons, introns, supercontigs), and for markers downloaded from GenBank, as explained in the Methods section of Thomas et al. 2021. Trees were included for each of these alignments from IQtree and Astral; SVDquartets tree was only estimated for the full set of supercontigs. Gene trees were generated with IQtree. Tree files are named differently than the final manuscript; refer to the number of genes specified in Fig 1 of Thomas et al, 2021 and specified in each filename to identify filtering scheme. Raw sequence reads are available on the Sequence Read Archive at <a href="http://www.ncbi.nlm.nih.gov/bioproject/715342">http://www.ncbi.nlm.nih.gov/bioproject/715342</a>.</p>
Target capture data resolve recalcitrant relationships in the coffee family (Rubioideae, Rubiaceae)
<p class="MsoNormal"><span>Subfamily Rubioideae is the largest of the main lineages in the coffee family (Rubiaceae), with over 8,000 species and 29 tribes. Phylogenetic relationships among tribes and other major clades within this group of plants are still only partly resolved despite considerable efforts. While previous studies have mainly utilized data from the organellar genomes and nuclear ribosomal DNA, we here use a large number of low-copy nuclear genes obtained via a target capture approach to infer phylogenetic relationships within Rubioideae. We included 101 Rubioideae species representing all but two (the monogeneric tribes Foonchewieae and Aitchinsonieae) of the currently recognized tribes, and all but one non-monogeneric tribe were represented by more than one genus. Using data from the 353 genes targeted with the universal Angiosperms353 probe set we investigated the impact of data type, analytical approach, and potential paralogs on phylogenetic reconstruction. We inferred a robust phylogenetic hypothesis of Rubioideae with the vast majority (or all) nodes being highly supported across all analyses and datasets and few incongruences between the inferred topologies. The results were similar to those of previous studies but novel relationships were also identified. We found that supercontigs (coding sequence [CDS] + noncoding sequence) clearly outperformed CDS data in levels of support and gene tree congruence. The full datasets (353 genes) outperformed the datasets with potential paralogous genes removed (186 genes) in levels of support but increased gene tree incongruence slightly. The pattern of gene tree conflict at short internal branches was often consistent with high levels of incomplete lineage sorting (ILS) due to rapid speciation in the group. While concatenation- and coalescence-based trees mainly agreed, the observed phylogenetic discordance between the two approaches may be best explained by their differences in accounting for ILS. The use of target capture data greatly improved our confidence and understanding of the Rubioideae phylogeny, highlighted by the increased support for previously uncertain relationships and the increased possibility to explore sources of underlying phylogenetic discordance.</span></p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.