Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
198
datasets available to search
ShareScore release 0.9.0
Dataset results
198 results for “sequence capture”
Whole-genome capture and sequencing of Mycobacterium tuberculosis directly from clinical samples - Design of RNA oligonucleotide baits for Agilent Technologies' SureSelect target enrichment
<p>This dataset comprises the sequence of <strong>44 278 RNA oligonucleotide "baits" (120 bp each) </strong>designed to perform <strong>whole-genome capture and sequencing of <em>Mycobacterium tuberculosis</em> directly from clinical samples</strong> (DNA) using Agilent Technologies’ SureSelect target enrichment system following the Illumina paired-end multiplexed sequencing library protocol. </p> <p>RNA oligonucleotide “baits” were designed to span the ∼4.5 Mb of the <em>M. tuberculosis</em> genome. In brief, the reference genome sequence of the MTBC H37Rv strain (Genbank #AL123456) was <em>in silico</em> fragmented into 120 bp sequences twice, to ensure an overlap of 60 bp between sequences. Due to their rich GC content, which could interfere with DNA capture, all MTBC genes of the PE, PPE and PE-PGRS family were also independently fragmented into 120 bp sequences, in order to increase capture sensitivity. All resulting sequences were BLASTn searched against the Human Genomic + Transcript database to excluded homologous sequences to the human genome. Overall, a total of 42,278 RNA probes were generated and this custom bait library was then uploaded to the SureDesign software (https://earray.chem.agilent.com/suredesign) and synthesized by Agilent Technologies. During synthesis, the 2198 sequences complementary to the PE, PPE and PE-PGRS family were unbalanced 8:1 to potentiate capture.</p> <p>More details can be found in the following publication:</p> <p>- Macedo, R., Isidro, J., Ferreira, R., Pinto, M., Borges, V., Duarte, S., Vieira, L., & Gomes, J. P. (2023). Molecular Capture of <em>Mycobacterium tuberculosis</em> Genomes Directly from Clinical Samples: A Potential Backup Approach for Epidemiological and Drug Susceptibility Inferences. <em>International journal of molecular sciences</em>, <em>24</em>(3), 2912. https://doi.org/10.3390/ijms24032912</p>
Data for: Regularized sequence-context mutational trees capture variation in mutation rates across the human genome
<p>Additional data on output models from Bayer as reported in:</p> <p>Regularized sequence-context mutational trees capture variation in mutation rates across the human genome</p> <p>Adams CJ, Conery M, Auerbach BJ, Jensen ST, Mathieson I, Voight BF. BioRxiv https://doi.org/10.1101/2022.10.14.512160</p> <p>Accepted, PLoS Genetics. </p> <p>Code Available at: https://github.com/bvoightlab/Baymer</p>
Morphological and DNA sequence data generated by Sanger sequencing and target capture methods for moss plants in the genus Fissidens from herbarium specimens
<p><span>Morphological evolution in mosses has long been hypothesized to accompany shifts in microhabitats and can be tested using comparative phylogenetics. These lines of inquiry have developed substantially, in part, by target capture sequencing allowing for phylogenomic scale data generated from herbarium specimens. In the present study, we test the relationship between taxonomically important morphological characters in the moss genus <em>Fissidens</em>, using both a 400-locus dataset generated using a target-capture approach as well as a three-locus phylogeny generated using sanger sequencing. Phylogenetic trees were generated using ASTRAL and Bayesian Inference and used to test the monophyly of subgenera/sections and provided the basis for ancestral character reconstruction and phylogenetic correlation analyses among five morphological characters as well as habitat moisture scored from literature. The characters <em>axillary hyaline nodules</em>, <em>limbidium</em>, <em>costa</em>, and <em>peristome morphology</em> as well as <em>sexual system</em>, <em>minimum habitat moisture</em>, <em>average habitat moisture</em>, <em>maximum habitat moisture</em>, and <em>habitat moisture niche breadth</em> each exhibit statistically significant phylogenetic signal. Significant correlations were found between the limbidium (phyllid/leaf border) and habitat moisture niche breadth, which could be interpreted as a more extensive <em>limbidium</em> enabling species to survive across a wider variety of habitats. Correlations were also found between <em>costa anatomy</em> and the <em>limbidum</em> of the gametophyte and sporophyte <em>peristome</em> <em>morphology</em>, as well as <em>average habitat moisture</em> and <em>sexual system</em>. Continued exploration of the relationships between morphological evolution, life history, and habitat will enable us to expand our understanding of functional morphology in mosses.</span></p>
AlphaFold2-Based Characterization of Apo and Holo Protein Structures and Conformational Ensembles Using Randomized Alanine Sequence Scanning Adaptation: Capturing Shared Signature Dynamics and Ligand-Induced Conformational Changes
<p>Proteins often exist in multiple conformational states, influenced by the binding of ligands or substrates. The study of these states, particularly the apo (unbound) and holo (ligand-bound) forms, is crucial for understanding protein function, dynamics, and interactions. In the current study, we use AlphaFold2 that combines<span> randomized</span> <span><span> </span>alanine<span> </span>sequence masking<span> </span>with shallow multiple sequence alignment<span> </span>subsampling to expand the conformational diversity of the predicted structural<span> </span>ensembles and<span> </span>capture conformational changes between apo and holo protein forms. Using several well-established datasets of<span> </span>structurally diverse apo-holo protein pairs, the proposed approach </span><span>enables<span> </span>robust predictions of apo and holo structures and conformational ensembles, while also displaying notably similar dynamics distributions. These observations are consistent with<span> </span>the view </span><span> </span>that the intrinsic dynamics of allosteric proteins is defined by the structural topology of the fold and favors conserved conformational motions driven by soft modes among orthologs. We also found<span> </span>a significant <span>correlation </span>between conformational flexibility and <span> </span>AlphaFold2 metric of statistical significance pLDDT for the apo-holo pairs in which ligand binding induced local moderate conformational changes. For apo-holo pairs exhibiting larger structural changes, this relationship<span> </span>becomes nonlinear, reflecting inability of AlphaFold2 confidence metrics to identify high energy functional conformations. Our findings support the notion that AlphaFold2 approaches can yield reasonable accuracy in predicting minor conformational adjustments between apo and holo states, especially for proteins with <span> </span>moderate localized changes upon ligand binding. However, for large, hinge-like domain movements, AF2 tends to predict the most stable domain orientation which is typically the apo form rather than the full range of functional conformations characteristic of the holo ensemble. These results indicate that modeling of multiple functional states of proteins may require more accurate detection of flexible region conformations and cannot solely rely on the pLDDT metric as the major determinant of the prediction accuracy in reproducing functional conformational ensembles.<span> </span></p>
Raw Counts: A protocol for low-input RNA-sequencing of patients with febrile neutropenia captures relevant immunological information
<p>Raw counts for scientific article: </p> <p><em>"A protocol for low-input RNA-sequencing of patients with febrile neutropenia captures relevant immunological information"</em></p> <p>Victoria Probst*<sup>1</sup>, Lotte Møller Smedegaard*<sup>2</sup>, Arman Simonyan<sup>1</sup>, Yuliu Guo<sup>1</sup>, Olga Østrup<sup>1</sup>, Kia Hee Schultz Dungu<sup>2</sup><sub>, </sub>Nadja Hawwa Vissing<sup>2</sup><sub>, </sub>Ulrikka Nygaard<sup>2</sup><sub> </sub>and<sub> </sub>Frederik Otzen Bagger<sup>1</sup></p> <p><sup>*Shared first authorship</sup></p> <p>Data description: </p> <p>The raw counts are from 88 samples of 22 patients with leukaemia and suspected infection sequenced by a low-input protocol (Takara SMART-seq HT) (96% succeeded) and 15 of these were also processed by the standard protocol (Truseq).</p> <p> </p> <p><sup>CLI.CSV: Raw counts of control samples processed using a low input RNA sequencing protocol. 15 samples processed by the low-input protocol. </sup></p> <p><sup>CRNA.CSV: Raw counts of control samples processed using a standard RNA sequencing protocol. 15 samples processed by the standard protocol. </sup></p> <p><sup>FEB.CSV: Raw gene counts from patients. 88 samples processed by the low input protocol. 4 samples failed sequencing.</sup></p> <p> </p> <p> </p> <p> </p>
Data from: Using transcriptome sequencing and pooled exome capture to study local adaptation in the giga-genome of Pinus cembra
Open the record for dataset details and reuse information.
Morphological and DNA sequence data generated by Sanger sequencing and target capture methods for moss plants in the genus Fissidens from herbarium specimens
Open the record for dataset details and reuse information.
BRAVO target sequence capture V3
<p>This data set contains the bait sequences for a sequence capture library targeting specific genes in Brassica ssp.. Source sequences have been manually selected and processed with BaitLibraryBuilder (<a href="https://github.com/steuernb/BaitLibraryBuilder">https://github.com/steuernb/BaitLibraryBuilder</a>).</p>
Data from: A stable phylogenomic classification of Travunioidea (Arachnida, Opiliones, Laniatores) based on sequence capture of ultraconserved elements
Molecular phylogenetics has transitioned into the phylogenomic era, with data derived from next-generation sequencing technologies allowing unprecedented phylogenetic resolution in all animal groups, including understudied invertebrate taxa. Within the most diverse harvestmen suborder, Laniatores, most relationships at all taxonomic levels have yet to be explored from a phylogenomics perspective. Travunioidea is an early-diverging lineage of laniatorean harvestmen with a Laurasian distribution, with species distributed in eastern Asia, eastern and western North America, and south-central Europe. This clade has had a challenging taxonomic history, but the current classification consists of ~77 species in three families, the Travuniidae, Paranonychidae, and Nippononychidae. Travunioidea classification has traditionally been based on structure of the tarsal claws of the hind legs. However, it is now clear that tarsal claw structure is a poor taxonomic character due to homoplasy at all taxonomic levels. Here, we utilize DNA sequences derived from capture of ultraconserved elements (UCEs) to reconstruct travunioid relationships. Data matrices consisting of 317–677 loci were used in maximum likelihood, Bayesian, and species tree analyses. Resulting phylogenies recover four consistent and highly supported clades; the phylogenetic position and taxonomic status of the enigmatic genus Yuria is less certain. Based on the resulting phylogenies, a revision of Travunioidea is proposed, now consisting of the Travuniidae, Cladonychiidae, Paranonychidae (Nippononychidae is synonymized), and the new family Cryptomastridae Derkarabetian & Hedin, fam. n., diagnosed here. The phylogenetic utility and diagnostic features of the intestinal complex and male genitalia are discussed in light of phylogenomic results, and the inappropriateness of the tarsal claw in diagnosing higher-level taxa is further corroborated.
Data in support of Using target sequence capture to improve the phylogenetic resolution of a rapid radiation in New Zealand Veronica
<p>Includes alignments and trees for the analysis found in Thomas et al. 2021, Using target sequence capture to improve the phylogenetic resolution of a rapid radiation in New Zealand Veronica; American Journal of Botany, Special Issue: Exploring Angiosperms353: a Universal Toolkit for Flowering Plant Phylogenomics. Alignments comprise subsets of Angiosperms353 genes given each filtering scheme (full, intersection, sortadate_BP, sortadate_TL) and gene type/subset (exons, introns, supercontigs), and for markers downloaded from GenBank, as explained in the Methods section of Thomas et al. 2021. Trees were included for each of these alignments from IQtree and Astral; SVDquartets tree was only estimated for the full set of supercontigs. Gene trees were generated with IQtree. Tree files are named differently than the final manuscript; refer to the number of genes specified in Fig 1 of Thomas et al, 2021 and specified in each filename to identify filtering scheme. Raw sequence reads are available on the Sequence Read Archive at <a href="http://www.ncbi.nlm.nih.gov/bioproject/715342">http://www.ncbi.nlm.nih.gov/bioproject/715342</a>.</p>
Solidago hybrid-sequence capture probe set
<p><em>Premise of the study</em>: The phylogenetic relationships among the ca. 138 species of goldenrods (<em>Solidago</em>; Asteraceae) have been difficult to infer due to species richness, and shallow interspecific genetic divergences. This study aims to overcome these obstacles by combining extensive sampling of goldenrod herbarium specimens with the use of a custom <em>Solidago</em> hybrid-sequence capture probe set.</p> <p><em>Methods</em>: A set of tissues from herbarium samples comprising ca. 90% of <em>Solidago</em> species was assembled, and DNA was extracted. A custom hybrid-sequence capture probe set was designed, and data from 854 nuclear regions were obtained and analyzed from 209 specimens. Maximum likelihood and coalescent approaches were used to estimate the genus phylogeny for 157 diploid samples.</p> <p><em>Key results</em>: Although DNAs from older specimens were both more fragmented and produced fewer sequencing reads, there was no relationship between specimen age and our ability to obtain sufficient data at the target loci. The <em>Solidago</em> phylogeny was generally well supported, with 88/155 (57%) nodes receiving ≥95% bootstrap support. <em>Solidago</em> was supported as monophyletic, with <em>Chrysoma</em> <em>pauciflosculosa</em> identified as sister. A clade comprising <em>Solidago</em> <em>ericameriodes</em>, <em>Solidago</em> <em>odora</em>, and <em>Solidago</em> <em>chapmanii</em> was identified as the earliest diverging <em>Solidago</em> lineage. The previously segregated genera <em>Brintonia</em> and <em>Oligoneuron</em> were identified as placed well within <em>Solidago</em>. These and other phylogenetic results were used to establish four subgenera and fifteen sections within the genus.</p> <p><em>Conclusions</em>: The combination of expansive herbarium sampling and hybrid-sequence capture data allowed us to quickly and rigorously establish the evolutionary relationships within this difficult, species-rich group.</p>
Whole-genome capture and sequencing of Francisella tularensis directly from clinical samples
<p>This dataset comprises:</p> <p><strong>1. The design of RNA oligonucleotide baits for Agilent Technologies’ SureSelect target enrichment - <strong>54756 RNA oligonucleotide "baits" (120 bp each) </strong></strong>designed to perform <strong>whole-genome capture and sequencing of </strong><strong>Francisella tularensis<strong> directly from clinical samples</strong></strong> (DNA) using Agilent Technologies’ SureSelect target enrichment system following the Illumina paired-end multiplexed sequencing library protocol.</p> <p>RNA oligonucleotide “baits” were designed to span the <em>Francisella tularensis </em>chromosome and plasmid, accounting for the genetic variability among publicly available genome sequences. All resulting sequences were BLASTn searched against the Human Genomic + Transcript database to excluded homologous sequences to the human genome. Overall, a total of 54756 RNA probes were generated and this custom bait library was then uploaded to the SureDesign software (https://earray.chem.agilent.com/suredesign) and synthesized by Agilent Technologies.</p> <p><strong>2. Genome assemblies of 17 Francisella tularensis samples generated in the context of validation and application of SureSelect target enrichment for Whole-genome capture and sequencing of Francisella tularensis directly from clinical samples </strong></p> <p>File “<strong>Ft_assembly_metadata.xlsx</strong>” contains the genome assembly statistics for each isolate, including European Nucleotide Archive accession numbers.</p> <p>The archive “<strong>Ft_assemblies.zip</strong>” contains all the genome assemblies (.fasta format) of each isolate presented in the metadata file.</p> <p>More details can be found in the following publication: (available soon)</p>
Solidago hybrid-sequence capture probe set
Open the record for dataset details and reuse information.
Data from: Genomic sequence capture of haemosporidian parasites: methods and prospects for enhanced study of host-parasite evolution
Open the record for dataset details and reuse information.
Data from: A stable phylogenomic classification of Travunioidea (Arachnida, Opiliones, Laniatores) based on sequence capture of ultraconserved elements
Open the record for dataset details and reuse information.
A multi-tiered sequence capture strategy spanning broad evolutionary scales: application for phylogenetic and phylogeographic studies of orchids
<p><span><span><span><span><span><span><span><span><span><span><span>With over 25,000 species, the drivers of diversity in the Orchidaceae remain to be fully understood. Here we outline a multi-tiered sequence capture strategy aimed at capturing 100's of loci to enable phylogenetic resolution from subtribe to subspecific levels in orchids of the tribe Diurideae. For the probe design, we mined subsets of 18 transcriptomes, to give five target sequence sets aimed at the tribe (Sets 1 & 2), subtribe (Set 3), and within subtribe levels (Sets 4 & 5). Analysis included alternative <i>de novo </i>and reference-guided assembly, before target sequence extraction, annotation and alignment, and application of a homology-aware <i>k-mer</i> block phylogenomic approach, prior to phylogenetic inference using maximum-likelihood. Our evaluation considered 87 taxa in two test datasets: 67 samples spanning the tribe, and 72samples involving 24 closely related <i>Caladenia</i> species. The tiered design achieved high target loci recovery (>89%), with the median number of recovered loci in Sets 1–5 as follows: 212, 219, 816, 1024, and 1009, respectively. Interestingly, as a first test of the homologous <i>k</i>-mer approach for targeted sequence capture data, our study revealed its potential for enabling robust phylogenetic species tree inferences. Specifically, we found matching, and in one case improved phylogenetic resolution within species complexes, compared to conventional phylogenetic analysis involving target gene extraction. Our findings indicate that a customised multi-tiered sequence capture strategy, in combination with promising yet under-utilized phylogenomic approaches, will be effective for groups where interspecific divergence is recent, but information on deeper phylogenetic relationships is also required.</span></span></span></span></span></span></span></span></span></span></span></p>
Alignments from: Gene count from target sequence capture places three whole genome duplication events in Hibiscus L. (Malvaceae)
<p class="BodyA"><span><b>Background:</b> The great diversity in plant genome size and chromosome number is partly due to polyploidization (i.e., genome doubling events). The differences in genome size and chromosome number among diploid plant species can be a window into the intriguing phenomenon of past genome doubling that may be obscured through time by the process of diploidization. The genus <i>Hibiscus </i>L. (Malvaceae) has a wide diversity of chromosome numbers and a complex genomic history. <i>Hibiscus </i>is ideal for exploring past genomic events because although two ancient genome duplication events have been identified, more are likely to be found due to its diversity of chromosome numbers. To reappraise the history of whole genome duplication events, we tested three alternative scenarios describing different polyploidization events.</span></p> <p class="BodyA"><span><b>Results:</b> Using target sequence capture, we designed a new probe set for <i>Hibiscus </i>and generated 87 orthologous genes from four diploid species. We detected paralogues in >54% putative single-copy genes. 34 of these genes were selected for testing three different genome duplication scenarios using gene counting. All species of <i>Hibiscus</i> sampled shared one genome duplication with <i>H. syriacus</i> and one whole genome duplication occurred along the branch leading to <i>H. syriacus</i>.</span></p> <p class="BodyA"><span><b>Conclusions:</b> Here, we corroborated the independent genome doubling previously found in the lineage leading to <i>H. syriacus </i>and a shared genome doubling of this lineage and the remainder of <i>Hibiscus</i>. Additionally, we found a previously undiscovered genome duplication shared by the /Pavonia and /Malvaviscus clades (both nested within <i>Hibiscus</i>) with the occurrences of two copies in what were otherwise single-copy genes. Our results highlight the complexity of genomic diversity in some plant groups, which makes orthology assessment and accurate phylogenomic inference difficult.</span></p>
Supplementary information for integrating sequence capture and restriction-site associated DNA sequencing to resolve recent radiations of Pelagic seabirds
<p><b>The diversification of modern birds has been shaped by a number of radiations. Rapid diversification events make reconstructing the evolutionary relationships among taxa challenging due to the convoluted effects of incomplete lineage sorting (ILS) and introgression. Phylogenomic datasets have the potential to detect patterns of phylogenetic incongruence, and to address their causes. However, the footprints of ILS and introgression on sequence data can vary between different phylogenomic markers at different phylogenetic scales depending on factors such as their evolutionary rates or their selection pressures. We show that combining phylogenomic markers that evolve at different rates, such as paired-end double-digest restriction site-associated DNA (PE-ddRAD) and ultraconserved elements (UCEs), allows a comprehensive exploration of the causes of phylogenetic discordance associated with short internodes at different timescales. We used thousands of UCE and PE-ddRAD markers to produce the first well-resolved phylogeny of shearwaters, a group of medium-sized pelagic seabirds amongst the most phylogenetically controversial and endangered bird groups. We found that phylogenomic conflict was mainly derived from high levels of ILS due to rapid speciation events. We also documented a case of introgression, despite the high philopatry of shearwaters to their breeding sites, which typically limits gene flow. We integrated state-of-the-art concatenated and coalescent-based approaches to expand on previous comparisons of UCE and RAD-Seq datasets for phylogenetics, divergence time estimation and inference of introgression, and we propose a strategy to optimise RAD-Seq data for phylogenetic analyses. Our results highlight the usefulness of combining phylogenomic markers evolving at different rates to understand the causes of phylogenetic discordance at different timescales.</b></p>
Data from: Comparison of taxon-specific versus general locus sets for targeted sequence capture for plant phylogenomics
Premise of the study: Targeted sequence capture can be used to efficiently gather sequence data for large numbers of loci, such as single-copy nuclear loci. Most published studies in plants have used taxon-specific locus sets developed individually for a clade using multiple genomic and transcriptomic resources. General locus sets can also be developed from loci that have been identified as single-copy and having orthologs in large clades of plants. Methods: We identify and compare a taxon-specific locus set and three general locus sets (COSII, APVO SSC, PPR) for targeted sequence capture in Buddleja (Scrophulariaceae) and outgroups. We evaluate their performance in terms of assembly success, sequence variability, and resolution and support of inferred phylogenetic trees. Results: The taxon-specific locus set had the most target loci. Assembly success was high for all locus sets in Buddleja samples. For outgroups, general locus sets had greater assembly success. Taxon-specific and PPR loci had the highest average variability. The taxon-specific dataset produced the best supported tree, but all datasets showed improved resolution over previous non-sequence capture datasets. Discussion: General loci can be a useful source of sequence capture targets, especially if multiple genomic resources are not available for a taxon.
Data from: A high-density exome capture genotype-by-sequencing panel for forestry breeding in Pinus radiata
Development of genome-wide resources for application in genomic selection or genome-wide association studies, in the absences of full reference genomes, present a challenge to the forestry industry, where longer breeding cycles could benefit from the accelerated selection possible through marker-based breeding value predictions. In particular, large conifer megagenomes require a strategy to reduce complexity, whilst ensuring genome-wide coverage is achieved. Using a transcriptome-based reference template, we have successfully developed a high density exome capture genotype-by-sequencing panel for radiata pine (Pinus radiata D.Don), capable of capturing in excess of 80,000 single nucleotide polymorphism (SNP) markers with a minor allele frequency above 0.03 in the population tested. This represents approximately 29,000 gene models from a core set of 48,914 probes. A set of 704 SMP markers capable of pedigree reconstruction and differentiating individual genotypes were tested within two full-sib mapping populations. While as few as 70 markers could reconstruct parentage in almost all cases, the impact of missing genotypes was noticeable in several offspring. Therefore, sets of 60 sets of 110 randomly selected SNP markers were compared for both parentage reconstruction and clone differentiation. The performance in parentage reconstruction showed little variation over 60 iterations. However, there was notable variation in discriminatory power between closely related individuals, indicating a higher density SNP marker panel may be required to elucidate hidden relationships in complex pedigrees.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.