Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
49
datasets available to search
ShareScore release 0.9.0
Dataset results
49 results for “targeted genome sequencing”
GWAS Summary Statistics for Publication: Identifying novel genetic and phenotypic associations to genomic features by leveraging off-target reads in exome sequencing data
<p>This dataset contains summary statistics for genome-wide association studies (GWAS) conducted on genomic features derived from off-target reads in whole-exome sequencing (WES) data. The study utilized tools like Seeing Beyond the Target (SBT) and ImReP to construct novel phenotypic features from unmapped reads in ~50,000 participants in the UK Biobank. Features include mitochondrial DNA (mtDNA) copy number, ribosomal DNA (rDNA) copy number (5S, 18S, 28S), immune repertoire metrics (e.g., T-cell receptor alpha diversity), and microvial genome load (viral and fungal).</p> <p>Summary statistics can be used for replication studies, meta-analyses, or further exploration of these phenotypes.</p>
Whole-genome capture and sequencing of Mycobacterium tuberculosis directly from clinical samples - Design of RNA oligonucleotide baits for Agilent Technologies' SureSelect target enrichment
<p>This dataset comprises the sequence of <strong>44 278 RNA oligonucleotide "baits" (120 bp each) </strong>designed to perform <strong>whole-genome capture and sequencing of <em>Mycobacterium tuberculosis</em> directly from clinical samples</strong> (DNA) using Agilent Technologies’ SureSelect target enrichment system following the Illumina paired-end multiplexed sequencing library protocol. </p> <p>RNA oligonucleotide “baits” were designed to span the ∼4.5 Mb of the <em>M. tuberculosis</em> genome. In brief, the reference genome sequence of the MTBC H37Rv strain (Genbank #AL123456) was <em>in silico</em> fragmented into 120 bp sequences twice, to ensure an overlap of 60 bp between sequences. Due to their rich GC content, which could interfere with DNA capture, all MTBC genes of the PE, PPE and PE-PGRS family were also independently fragmented into 120 bp sequences, in order to increase capture sensitivity. All resulting sequences were BLASTn searched against the Human Genomic + Transcript database to excluded homologous sequences to the human genome. Overall, a total of 42,278 RNA probes were generated and this custom bait library was then uploaded to the SureDesign software (https://earray.chem.agilent.com/suredesign) and synthesized by Agilent Technologies. During synthesis, the 2198 sequences complementary to the PE, PPE and PE-PGRS family were unbalanced 8:1 to potentiate capture.</p> <p>More details can be found in the following publication:</p> <p>- Macedo, R., Isidro, J., Ferreira, R., Pinto, M., Borges, V., Duarte, S., Vieira, L., & Gomes, J. P. (2023). Molecular Capture of <em>Mycobacterium tuberculosis</em> Genomes Directly from Clinical Samples: A Potential Backup Approach for Epidemiological and Drug Susceptibility Inferences. <em>International journal of molecular sciences</em>, <em>24</em>(3), 2912. https://doi.org/10.3390/ijms24032912</p>
MarkerScan: Separation and assembly of cobionts sequenced alongside target species in biodiversity genomics projects
<p>This dataset represents all results files described in the paper 'MarkerScan: Separation and assembly of cobionts sequenced alongside target species in biodiversity genomics projects'.</p>
Alignments from: Gene count from target sequence capture places three whole genome duplication events in Hibiscus L. (Malvaceae)
<p class="BodyA"><span><b>Background:</b> The great diversity in plant genome size and chromosome number is partly due to polyploidization (i.e., genome doubling events). The differences in genome size and chromosome number among diploid plant species can be a window into the intriguing phenomenon of past genome doubling that may be obscured through time by the process of diploidization. The genus <i>Hibiscus </i>L. (Malvaceae) has a wide diversity of chromosome numbers and a complex genomic history. <i>Hibiscus </i>is ideal for exploring past genomic events because although two ancient genome duplication events have been identified, more are likely to be found due to its diversity of chromosome numbers. To reappraise the history of whole genome duplication events, we tested three alternative scenarios describing different polyploidization events.</span></p> <p class="BodyA"><span><b>Results:</b> Using target sequence capture, we designed a new probe set for <i>Hibiscus </i>and generated 87 orthologous genes from four diploid species. We detected paralogues in >54% putative single-copy genes. 34 of these genes were selected for testing three different genome duplication scenarios using gene counting. All species of <i>Hibiscus</i> sampled shared one genome duplication with <i>H. syriacus</i> and one whole genome duplication occurred along the branch leading to <i>H. syriacus</i>.</span></p> <p class="BodyA"><span><b>Conclusions:</b> Here, we corroborated the independent genome doubling previously found in the lineage leading to <i>H. syriacus </i>and a shared genome doubling of this lineage and the remainder of <i>Hibiscus</i>. Additionally, we found a previously undiscovered genome duplication shared by the /Pavonia and /Malvaviscus clades (both nested within <i>Hibiscus</i>) with the occurrences of two copies in what were otherwise single-copy genes. Our results highlight the complexity of genomic diversity in some plant groups, which makes orthology assessment and accurate phylogenomic inference difficult.</span></p>
Sanger sequencing of target and off-target genomic regions for gene-edited iPSC clones with SETBP1 genetic variants
<p>This data set includes chromatograms generated using sanger sequencing of targeted regions of genomic DNA from clonal iPSC lines. The iPSC lines include clones generated using CRISPR/Cas9 homology directed repair to introduce genetic variants into <em>SETBP1,</em> and their wild-type controls. Additional files have been included in the data set to link chromatogram (ab1) files to specific iPSC clones for genomic regions across the variant in <em>SETBP1 (</em>SETBP1 clones genetic variant sanger sequencing.xslx)<em> </em>and top<em> </em>off-target sites (SETBP1 clones off-target sanger sequencing.xlsx). </p>
Alignments from: Gene count from target sequence capture places three whole genome duplication events in Hibiscus L. (Malvaceae)
Open the record for dataset details and reuse information.
Targeted sequencing and iterative assembly of near-complete genomes
Open the record for dataset details and reuse information.
Data from: Targeted sequencing of venom genes from cone snail genomes improves understanding of conotoxin molecular evolution
To expand our capacity to discover venom sequences from the genomes of venomous organisms, we applied targeted sequencing techniques to selectively recover venom gene superfamilies and non-toxin loci from the genomes of 32 cone snail species (family, Conidae), a diverse group of marine gastropods that capture their prey using a cocktail of neurotoxic peptides (conotoxins). We were able to successfully recover conotoxin gene superfamilies across all species with high confidence (> 100X coverage) and used these data to provide new insights into conotoxin evolution. First, we found that conotoxin gene superfamilies are composed of 1-6 exons and are typically short in length (mean = ~85bp). Second, we expanded our understanding of the following genetic features of conotoxin evolution: (a) positive selection, where exons coding the mature toxin region were often three times more divergent than their adjacent noncoding regions, (b) expression regulation, with comparisons to transcriptome data showing that cone snails only express a fraction of the genes available in their genome (24%-63%), and (c) extensive gene turnover, where Conidae species varied from 120-859 conotoxin gene copies. Finally, using comparative phylogenetic methods, we found that while diet specificity did not predict patterns of conotoxin evolution, dietary breadth was positively correlated with total conotoxin gene diversity. Overall, the targeted sequencing technique demonstrated here has the potential to radically increase the pace at which venom gene families are sequenced and studied, reshaping our ability to understand the impact of genetic changes on ecologically relevant phenotypes and subsequent diversification.
Feasibility Clinical Study of Targeted and Genome-Wide Sequencing
ClinicalTrials.gov study NCT01345513. IPD Sharing: Not stated. Countries: 1. Publications: 0.
Data from: Targeted multiplex next-generation sequencing: Advances in techniques of mitochondrial and nuclear DNA sequencing for population genomics
Open the record for dataset details and reuse information.
Data from: Targeted sequencing of venom genes from cone snail genomes improves understanding of conotoxin molecular evolution
Open the record for dataset details and reuse information.
Next Gen(etics): targeted genome enrichment and next-generation sequencing enhances phenotype-driven forward genetics and gene-driven reverse genetics
GEO Series GSE22024. Rattus norvegicus. 30 samples. Type: Genome variation profiling by high throughput sequencing.
Sensitive and unbiased genome-wide profiling of base-editor-induced off-target activity with CHANGE-seq-BE [Hybrid Capture Sequencing for CBE and ABE]
GEO Series GSE308237. Homo sapiens. 36 samples. Type: Other.
Integration of whole genome sequencing analysis with unique patient-derived models reveals clinically relevant drug targets in TFCP2 fusion-defined intraosseous rhabdomyosarcoma
GEO Series GSE287694. Homo sapiens. 2 samples. Type: Expression profiling by high throughput sequencing.
Distinct structural and functional heterochromatin partitioning of lamin B1 and lamin B2 revealed using genome-wide Nicking Enzyme Epitope targeted DNA sequencing.
GEO Series GSE261834. Homo sapiens. 51 samples. Type: Genome binding/occupancy profiling by high throughput sequencing; Other.
Genome-wide survey of YY1 targets in C2C12 cells by High-throughput sequencing
GEO Series GSE45875. Mus musculus. 8 samples. Type: Genome binding/occupancy profiling by high throughput sequencing.
Genome-wide identification of vegetative phase transition-associated microRNAs and target predictions using degradome sequencing in Malus hupehensis
GEO Series GSE63373. Malus hupehensis. 3 samples. Type: Other; Non-coding RNA profiling by high throughput sequencing.
Alteration of genome folding via engineered transposon insertion [Target-enriched sequencing]
GEO Series GSE137371. Homo sapiens. 2 samples. Type: Other.
Targeted, long-read RNA sequencing of non-coding genomic regions associated with neuropsychiatric functions.
GEO Series GSE118158. Homo sapiens. 20 samples. Type: Expression profiling by high throughput sequencing.
Targeted genomic sequencing of the TCRγ locus in ILC2s and γδT cells
GEO Series GSE152726. Mus musculus. 11 samples. Type: Other.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.