Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

590

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

590 results for “targeted sequencing”

Learn how ShareScore rates datasets ↗
zenodo48/100

Graphic Illustration of Neal Platt's Talk: Targeted sequencing of pathogen DNA from museum specimens

<p><a href="https://lib.ku.edu/people/courtney-foat" target="_blank" rel="noopener">Courtney Foat</a>, Advisor for Strategic Initiatives &amp; Organizational Engagement at the University of Kansas, graphically recorded this invited talk by Neal Platt at an NSF-supported Workshop: &nbsp;Digital Collections Data and Tracking Disease.</p>

opencc-by-4.0May 2024View details →
zenodo44/100

GWAS Summary Statistics for Publication: Identifying novel genetic and phenotypic associations to genomic features by leveraging off-target reads in exome sequencing data

<p>This dataset contains summary statistics for genome-wide association studies (GWAS) conducted on genomic features derived from off-target reads in whole-exome sequencing (WES) data. The study utilized tools like Seeing Beyond the Target (SBT) and ImReP to construct novel phenotypic features from unmapped reads in ~50,000 participants in the UK Biobank. Features include mitochondrial DNA (mtDNA) copy number, ribosomal DNA (rDNA) copy number (5S, 18S, 28S), immune repertoire metrics (e.g., T-cell receptor alpha diversity), and microvial genome load (viral and fungal).</p> <p>Summary statistics can be used for replication studies, meta-analyses, or further exploration of these phenotypes.</p>

opencc-by-4.0Nov 2024View details →
zenodo44/100

Targeted Re-sequencing Identifies Candidate Fusiform Rust Resistance Genes in Loblolly Pine

<p>A fasta file containing the subset of the v2.01 Pita genome in addition to the novel NLR genes that were targeted by hybridization probes.&nbsp;</p> <p>A bed file describing the intervals targeted by the hybridization&nbsp;probes.</p> <p>Trinity assemblies of the 30 RNAseq libraries along with predictions by transdecoder of CDS and peptide sequences from those trinity assemblies.&nbsp;&nbsp;</p>

opencc-by-4.0Oct 2020View details →
zenodo44/100

Human Family With Sequence Similarity 83 Member B (FAM83B); A Target Enabling Package

<p>FAM83A-H are newly identified oncogenes characterised by a conserved DUF1669 domain. FAM83B can substitute for RAS to promote malignant transformation. Ablation of FAM83B or mutation of Lys230 inhibits malignant phenotypes, implicating FAM83B as potential therapeutic target. As part of this TEP, we solved the first crystal structures from the FAM83 family, including FAM83A and FAM83B. The structures of the DUF1669 domain reveal a phospholipase D-like fold lacking conservation of key catalytic residues. We deorphanise the FAM83 DUF1669 domain as a critical docking scaffold for binding of casein kinase 1 isoforms. Finally, using XChem fragment screening we report chemical fragments that bind to Lys230 in the central pocket of the DUF1669 and form starting points for potential drug development.</p>

opencc-by-4.0Aug 2018View details →
zenodo44/100

Whole-genome capture and sequencing of Mycobacterium tuberculosis directly from clinical samples - Design of RNA oligonucleotide baits for Agilent Technologies' SureSelect target enrichment

<p>This dataset comprises the sequence of <strong>44&nbsp;278&nbsp;RNA oligonucleotide &quot;baits&quot; (120 bp each) </strong>designed to perform&nbsp;<strong>whole-genome capture and sequencing of <em>Mycobacterium tuberculosis</em>&nbsp;directly from clinical samples</strong>&nbsp;(DNA)&nbsp;using Agilent Technologies&rsquo; SureSelect target enrichment system following the Illumina paired-end multiplexed sequencing library protocol.&nbsp;</p> <p>RNA oligonucleotide &ldquo;baits&rdquo; were designed to span the &sim;4.5 Mb of the <em>M. tuberculosis</em> genome. In brief, the reference genome sequence of the MTBC H37Rv strain (Genbank #AL123456) was <em>in silico</em> fragmented into 120 bp sequences twice, to ensure an overlap of 60 bp between sequences. Due to their rich GC content, which could interfere with DNA capture, all MTBC genes of the PE, PPE and PE-PGRS family were also independently fragmented into 120 bp sequences, in order to increase capture sensitivity. All resulting sequences were BLASTn searched against the Human Genomic + Transcript database to excluded homologous sequences to the human genome. Overall, a total of 42,278 RNA probes were generated and this custom bait library was then uploaded to the SureDesign software (https://earray.chem.agilent.com/suredesign) and synthesized by Agilent Technologies. During synthesis, the 2198 sequences complementary to the PE, PPE and PE-PGRS family were unbalanced 8:1 to potentiate capture.</p> <p>More details can be found in the following publication:</p> <p>- Macedo, R., Isidro, J., Ferreira, R., Pinto, M., Borges, V., Duarte, S., Vieira, L., &amp; Gomes, J. P. (2023). Molecular Capture of&nbsp;<em>Mycobacterium tuberculosis</em>&nbsp;Genomes Directly from Clinical Samples: A Potential Backup Approach for Epidemiological and Drug Susceptibility Inferences.&nbsp;<em>International journal of molecular sciences</em>,&nbsp;<em>24</em>(3), 2912. https://doi.org/10.3390/ijms24032912</p>

opencc-by-4.0Jan 2023View details →
zenodo40/100

Sequence data for the article "Targeted reduction of highly abundant transcripts with pseudo-random primers"

<p>Sequence data (Illumina MiSeq runs) for the article "Targeted reduction of highly abundant transcripts with pseudo-random primers". File names indicate unique run identifiers. In the manuscript, shorter names are used:</p> <ul> <li>NC12: 140602_M00528_0019_000000000-A88YD</li> <li>NC17: 140918_M00528_0047_000000000-A8GHU</li> <li>NC22b: 141105_M00528_0062_000000000-AAPC2</li> <li>NCki: 140207_M00528_0069_000000000-A5TY9</li> </ul> <p> </p>

opencc-zeroMar 2016View details →
dryad40/100

Data from: A cost-effective blood DNA methylation-based age estimation method in domestic cats, Tsushima leopard cats (Prionailurus bengalensis euptilurus), and Panthera species, using targeted bisulfite sequencing and machine learning models

<p><span>Knowledge of individual age can help both in-situ and ex-situ conservation programs to design more efficient and suitable management plans for targeted wildlife species. DNA methylation is one of the epigenetic aging markers that has emerged as a promising tool that can estimate age with high accuracy using only a tiny amount of biological material, which can be collected in a minimally invasive way. Here, we sequenced five targeted genetic regions and used </span><span>8–23</span><span> selected CpG sites to build age estimation models with machine learning methods </span><span>with about only $3–7 per sample</span><span>, using blood samples of seven Felidae species—ranging from small to big, and domestic to endangered species: domestic cats (<em>Felis catus</em>, 139 samples), Tsushima leopard cats (<em>Prionailurus bengalensis euptilurus</em>, 84 samples), and five<em> Panthera </em>species (96 samples). </span><span>The models built achieved satisfactory accuracy—the mean absolute error of the best models was 1.966, 1.348, and 1.552 years in domestic cats, Tsushima leopard cats, and <em>Panthera</em> spp., respectively.</span><span> Our models in domestic cats and Tsushima leopard cats were applicable to individuals regardless of health conditions, indicating the high applicability of our models to samples collected from diverse situations, e.g., rescued individuals in the context of conservation. We also showed the possibility of developing universal age estimation models for the five<em> Panthera</em> spp. using two of the five genetic regions, suggesting an even lower cost to use our models for future applications.</span></p>

opencc-zeroJan 2024View details →
zenodo40/100

Code for generating figures and analyzing amplicon sequencing of human mRNA and reporter mRNA targeted with type III-A CRISPR complex from Streptococcus thermophiles

<p>This dataset contains code for analyzing amplicon sequencing data and generating figures in the manuscript by Anna Nemudraia, Artem Nemudryi, and Blake Wiedenheft (2024), "Repair of CRISPR-guided RNA breaks enables site-specific RNA excision in human cells."&nbsp;</p> <p>Amplicon sequencing data has been deposited to NCBI Sequence Read Archive (SRA) under BioProject PRJNA1099688. The description of read files deposited to SRA can be found in the spreadsheet ./code_for_sequencing_data_analysis/SRA_read_files_description.xlsx</p> <p>The code for analyzing amplicon sequencing data can be found in the archive "code_for_sequencing_data_analysis.tar.gz." Output files from this analysis were used to generate figures. Figures were generated using the ggplot2 package in R and finalized in CorelDRAW.</p> <p>Code for generating figures can be found in the archive "code_for_generating_figures.tar.gz".&nbsp;</p> <p>Any questions or requests regarding the data or the code should be addressed to Dr. Artem Nemudryi at artem.nemudryi@gmail.com.</p>

opencc-by-4.0Apr 2024View details →
zenodo40/100

Comprehensive discovery of CRISPR-targeted terminally redundant sequences in the human gut metagenome: viruses, plasmids, and more

<p>S1 Data</p> <p>Dataset including the discovered CRISPR spacers, direct repeats, protospacers, co-occurrence-based spacer clustering results, predicted protein sequences, built HMMs, database comparison results, phylogenetic analysis results, predicted targeting hosts, and CRISPR-targeted TR sequences.</p>

opencc-by-4.0Sep 2021View details →
dryad40/100

Morphological and DNA sequence data generated by Sanger sequencing and target capture methods for moss plants in the genus Fissidens from herbarium specimens

<p><span>Morphological evolution in mosses has long been hypothesized to accompany shifts in microhabitats and can be tested using comparative phylogenetics. These lines of inquiry have developed substantially, in part, by target capture sequencing allowing for phylogenomic scale data generated from herbarium specimens. In the present study, we test the relationship between taxonomically important morphological characters in the moss genus <em>Fissidens</em>, using both a 400-locus dataset generated using a target-capture approach as well as a three-locus phylogeny generated using sanger sequencing. Phylogenetic trees were generated using ASTRAL and Bayesian Inference and used to test the monophyly of subgenera/sections and provided the basis for ancestral character reconstruction and phylogenetic correlation analyses among five morphological characters as well as habitat moisture scored from literature. The characters <em>axillary hyaline nodules</em>, <em>limbidium</em>, <em>costa</em>, and <em>peristome morphology</em> as well as <em>sexual system</em>, <em>minimum habitat moisture</em>, <em>average habitat moisture</em>, <em>maximum habitat moisture</em>, and <em>habitat moisture niche breadth</em> each exhibit statistically significant phylogenetic signal. Significant correlations were found between the limbidium (phyllid/leaf border) and habitat moisture niche breadth, which could be interpreted as a more extensive <em>limbidium</em> enabling species to survive across a wider variety of habitats. Correlations were also found between <em>costa anatomy</em> and the <em>limbidum</em> of the gametophyte and sporophyte <em>peristome</em> <em>morphology</em>, as well as <em>average habitat moisture</em> and <em>sexual system</em>. Continued exploration of the relationships between morphological evolution, life history, and habitat will enable us to expand our understanding of functional morphology in mosses.</span></p>

opencc-zeroJun 2022View details →
zenodo40/100

Direct RNA targeted in situ sequencing for transcriptomic profiling in tissue

<p>You can find here the Direct RNA In Situ Sequencing (HybISS-based)&nbsp;maps generated using the Hight Sensitivity kit from CARTANA AB. They include half a mouse brain coronal section, targeting 50 genes. Genes were targeted in a sequential manner. Both reads, DAPI staining and segmented cells are included. The analysis of the same cells, but using 10X magnification are also provided in an anndata object.</p>

opencc-by-4.0Jul 2022View details →
zenodo40/100

Data from: Bratzel et al. (2022) Target-enrichment sequencing reveals for the first time a well-resolved phylogeny of the core Bromelioideae (Bromeliaceae). Taxon

<p>DNA sequence alignments used for phylogenetic analyses in Bratzel et al. (2022) Target-enrichment sequencing reveals for the first time a well-resolved phylogeny of the core Bromelioideae (Bromeliaceae). Taxon.</p>

opencc-by-4.0Oct 2022View details →
zenodo40/100

Targeted Gene Panel Sequencing Data of RELN paper - Meyer Children's Hospital IRCCS

<h3>Dataset description</h3> <p>&nbsp;</p> <p>The dataset has been prepared according the Minimal Information about a high throughput SEQuencing Experiment (MINSEQE) as reported in: <a href="https://doi.org/10.5281/zenodo.5706412">https://doi.org/10.5281/zenodo.5706412</a></p> <p>This dataset includes:</p> <ul> <li>The Targeted Gene Panel Sequencing Raw Data (FASTQ files) from two individuals harbouring RELN variants</li> <li>The &lsquo;final&rsquo; processed data,&nbsp;submitted both as VCF and TXT files, and obtained from the ANNOVAR annotations of the two patients</li> </ul> <p>The gene panel list used in the targeted capture and the essential experimental and data processing protocols has been reported in the RELN paper.</p> <h3>Identifiers</h3> <p>The 444D indentifier correspond to&nbsp;<strong>DN1 patient</strong> in the RELN paper.</p> <p>Tissue: peripheral blood sample</p> <p>Sex: female</p> <p>Age at sequencing: 21 years</p> <p>The 528T indentifier correspond to <strong>DN2 patient </strong>in the RELN paper.</p> <p>Tissue: peripheral blood sample</p> <p>Sex: female</p> <p>Age at sequencing: 1.5 years</p>

opencc-by-4.0May 2024View details →
zenodo40/100

Linked collectors and determiners for: Combining target enrichment and Sanger sequencing data to clarify the systematics of the diverse Neotropical butterfly subtribe Euptychiina (Nymphalidae, Satyrinae).

Natural history specimen data linked to collectors and determiners held within, "Combining target enrichment and Sanger sequencing data to clarify the systematics of the diverse Neotropical butterfly subtribe Euptychiina (Nymphalidae, Satyrinae)". Claims or attributions were made on Bionomia by volunteer Scribes, <a href="https://bionomia.net/dataset/bfb878f3-8a74-46d3-a104-36485c32aaba">https://bionomia.net/dataset/bfb878f3-8a74-46d3-a104-36485c32aaba</a> using specimen data from the dataset aggregated by the Global Biodiversity Information Facility, <a href="https://gbif.org/dataset/bfb878f3-8a74-46d3-a104-36485c32aaba">https://gbif.org/dataset/bfb878f3-8a74-46d3-a104-36485c32aaba</a>. Formatted as a Frictionless Data package.

opencc-zeroJan 2024View details →
zenodo40/100

TaqMan array sequence set for 52-target respiratory pathogen array

<p>Sequence list for probes and primers used in the respiratory Taqman array card&nbsp;investigated in project, protocol&nbsp;10.5281/zenodo.5081880</p>

opencc-by-4.0Sep 2021View details →
dryad40/100

Data from: A cost-effective blood DNA methylation-based age estimation method in domestic cats, Tsushima leopard cats (Prionailurus bengalensis euptilurus), and Panthera species, using targeted bisulfite sequencing and machine learning models

Open the record for dataset details and reuse information.

publicJan 2024View details →
dryad40/100

Comparative assessment of line-probe assays and targeted next-generation sequencing in drug-resistant tuberculosis diagnosis

Open the record for dataset details and reuse information.

publicAug 2025View details →
dryad40/100

Morphological and DNA sequence data generated by Sanger sequencing and target capture methods for moss plants in the genus Fissidens from herbarium specimens

Open the record for dataset details and reuse information.

publicJun 2022View details →
zenodo36/100

BRAVO target sequence capture V3

<p>This data set contains the bait sequences for a sequence capture library targeting specific genes in Brassica ssp.. Source sequences have been manually selected and processed with BaitLibraryBuilder (<a href="https://github.com/steuernb/BaitLibraryBuilder">https://github.com/steuernb/BaitLibraryBuilder</a>).</p>

opencc-by-4.0Jan 2021View details →
zenodo36/100

MarkerScan: Separation and assembly of cobionts sequenced alongside target species in biodiversity genomics projects

<p>This dataset represents all results files described in the paper 'MarkerScan: Separation and assembly of cobionts sequenced alongside target species in biodiversity genomics projects'.</p>

opencc-by-4.0Dec 2023View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record