Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

590

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

590 results for “targeted sequencing”

Learn how ShareScore rates datasets ↗
dryad36/100

Targeted sequencing of T-DNA borders in OCP1xOGC transgenic lines of Camelina

Open the record for dataset details and reuse information.

publicApr 2022View details →
dryad36/100

Data from: Evaluating culture-free targeted next-generation sequencing for diagnosing drug-resistant tuberculosis: A multicentre clinical study of two end-to-end commercial workflows

Open the record for dataset details and reuse information.

publicAug 2025View details →
zenodo32/100

Urban wastewater virome by viral metagenomics and target enrichment sequencing

<p>Metagenomic analysis of virus in raw sewage.</p>

opencc-by-4.0Jan 2021View details →
dryad32/100

Alignments from: Gene count from target sequence capture places three whole genome duplication events in Hibiscus L. (Malvaceae)

<p class="BodyA"><span><b>Background:</b> The great diversity in plant genome size and chromosome number is partly due to polyploidization (i.e., genome doubling events). The differences in genome size and chromosome number among diploid plant species can be a window into the intriguing phenomenon of past genome doubling that may be obscured through time by the process of diploidization. The genus <i>Hibiscus </i>L. (Malvaceae) has a wide diversity of chromosome numbers and a complex genomic history. <i>Hibiscus </i>is ideal for exploring past genomic events because although two ancient genome duplication events have been identified, more are likely to be found due to its diversity of chromosome numbers. To reappraise the history of whole genome duplication events, we tested  three alternative scenarios describing different polyploidization events.</span></p> <p class="BodyA"><span><b>Results:</b> Using target sequence capture, we designed a new probe set for <i>Hibiscus </i>and generated 87 orthologous genes from four diploid species. We detected paralogues in &gt;54% putative single-copy genes. 34 of these genes were selected for testing three different genome duplication scenarios using gene counting. All species of <i>Hibiscus</i> sampled shared one genome duplication with <i>H. syriacus</i> and one whole genome duplication occurred along the branch leading to <i>H. syriacus</i>.</span></p> <p class="BodyA"><span><b>Conclusions:</b> Here, we corroborated the independent genome doubling previously found in the lineage leading to <i>H. syriacus </i>and a shared genome doubling of this lineage and the remainder of <i>Hibiscus</i>. Additionally, we found a previously undiscovered genome duplication shared by the /Pavonia and /Malvaviscus clades (both nested within <i>Hibiscus</i>) with the occurrences of two copies in what were otherwise single-copy genes. Our results highlight the complexity of genomic diversity in some plant groups, which makes orthology assessment and accurate phylogenomic inference difficult.</span></p>

opencc-zeroJan 2021View details →
dryad32/100

Data from: Comparison of taxon-specific versus general locus sets for targeted sequence capture for plant phylogenomics

Premise of the study: Targeted sequence capture can be used to efficiently gather sequence data for large numbers of loci, such as single-copy nuclear loci. Most published studies in plants have used taxon-specific locus sets developed individually for a clade using multiple genomic and transcriptomic resources. General locus sets can also be developed from loci that have been identified as single-copy and having orthologs in large clades of plants. Methods: We identify and compare a taxon-specific locus set and three general locus sets (COSII, APVO SSC, PPR) for targeted sequence capture in Buddleja (Scrophulariaceae) and outgroups. We evaluate their performance in terms of assembly success, sequence variability, and resolution and support of inferred phylogenetic trees. Results: The taxon-specific locus set had the most target loci. Assembly success was high for all locus sets in Buddleja samples. For outgroups, general locus sets had greater assembly success. Taxon-specific and PPR loci had the highest average variability. The taxon-specific dataset produced the best supported tree, but all datasets showed improved resolution over previous non-sequence capture datasets. Discussion: General loci can be a useful source of sequence capture targets, especially if multiple genomic resources are not available for a taxon.

opencc-zeroDec 2017View details →
dryad32/100

Data from: Targeted next-generation sequencing panels in the diagnosis of Charcot Marie Tooth disease

Objective: To investigate the effectiveness of targeted NGS panels in achieving a molecular diagnosis in CMT and related disorders in a clinical setting Methods: We prospectively enrolled 220 patients from two tertiary referral centres, one in London, UK (n=120) and one in Iowa, US (n=100) in whom a targeted CMT NGS panel had been requested as a diagnostic test. PMP22 duplication/deletion was previously excluded in demyelinating cases. We reviewed the genetic and clinical data upon completion of the diagnostic process. Results: After targeted NGS sequencing a definite molecular diagnosis, defined as a pathogenic or likely pathogenic variant, was reached in 30% of cases (n=67). The diagnostic rate was similar in London (32%) and Iowa (29%). Variants of unknown significance were found in an additional 33% of cases. Mutations in GJB1, MFN2, MPZ accounted for 39% of cases who received genetic confirmation, while the remainder of positive cases had mutations in diverse genes, including SH3TC2, GDAP1, IGHMBP2, LRSAM1, FDG4, GARS and another 12 less common genes. Copy number changes in PMP22, MPZ, MFN2, SH3TC2 and FDG4 were also accurately detected. A definite genetic diagnosis was more likely in cases with an early onset, a positive family history of neuropathy or consanguinity and a demyelinating neuropathy. Conclusions: NGS panels are effective tools in the diagnosis of CMT leading to the genetic confirmation in one third cases negative for PMP22 duplication/deletion, thus highlighting how rarer and previously undiagnosed subtypes represent today a relevant part of the genetic landscape of CMT.

opencc-zeroDec 2019View details →
dryad32/100

Data from: Phylogenomics of horned lizards (genus: Phrynosoma) using targeted sequence capture data

New genome sequencing techniques are enabling phylogenetic studies to scale-up from using a handful of loci to hundreds or thousands of loci from throughout the genome. In this study, we use targeted sequence capture (TSC) data from 540 ultraconserved elements and 44 protein-coding genes to estimate the phylogenetic relationships among all 17 species of horned lizards in the genus Phrynosoma. Previous molecular phylogenetic analyses of Phrynosoma based on a few nuclear genes, restriction site associated DNA (RAD) sequencing, or mitochondrial DNA (mtDNA) have produced conflicting relationships. Some of these conflicts are likely the result of rapid speciation at the start of Phrynosoma diversification, whereas other examples of gene tree discordance appear to be caused by active and residual traces of hybridization. Concatenation and coalescent-based species tree phylogenetic analyses of these new TSC data support the same topology, and a divergence dating analysis suggests that the Phrynosoma crown group is up to 30 million years old. The new phylogenomic tree supports the recognition of four main clades within Phrynosoma, including Anota (P. mcallii, P. solare, and the P. coronatum complex), Doliosaurus (P. modestum, P. goodei, and P. platyrhinos), Tapaja (P. ditmarsi, P. douglasii, P. hernandesi, and P. orbiculare), and Brevicauda (P. braconnieri, P. sherbrookei, and P. taurus). The phylogeny provides strong support for the relationships among all species of Phrynosoma and provides a robust new framework for conducting comparative analyses.

opencc-zeroDec 2014View details →
dryad32/100

Data from: Allele phasing has minimal impact on phylogenetic reconstruction from targeted nuclear gene sequences in a case study of Artocarpus

Premise of the study: Untapped information about allelic diversity within populations and individuals (i.e. heterozygosity) could improve phylogenetic resolution and accuracy. Many phylogenetic reconstructions ignore heterozygosity because it is difficult to assemble allele sequences and combine allelic data across unlinked loci and it is unclear how reconstruction methods accommodate variable sequences. We review the common methods of including heterozygosity in phylogenetic studies and present a novel method for assembling allele sequences from target enriched Illumina sequencing libraries. Methods: We perform supermatrix phylogeny reconstruction and species tree estimation of Artocarpus based on three methods of accounting for heterozygous sequences: a consensus method based on de novo sequence assembly, the use of ambiguity characters, and a novel method for phasing alleles. We characterize the extent to which highly heterozygous sequences impeded phylogeny reconstruction and determine whether the use of allele sequences improves resolution or decreases topological uncertainty. Key Results: We show that it is possible to infer phased alleles from target enriched Illumina libraries. We find that highly heterozygous sequences do not contribute disproportionately to poor phylogenetic resolution and that the use of allele sequences for phylogeny reconstruction does not have a clear effect on phylogenetic resolution or topological consistency. Conclusions: We provide a framework for inferring phased alleles from target enrichment data and for assessing the contribution of allelic diversity to phylogenetic reconstruction. In our dataset, the impact of allele phasing on phylogeny is minimal compared to the impact of using phylogenetic reconstruction methods that account for gene tree incongruence.

opencc-zeroDec 2017View details →
dryad32/100

Data from: Detection of somatic epigenetic variation in Norway spruce via targeted bisulfite sequencing

Epigenetic mechanisms represent a possible mechanism for achieving a rapid response of long‐lived trees to changing environmental conditions. However, our knowledge on plant epigenetics is largely limited to a few model species. With increasing availability of genomic resources for many tree species, it is now possible to adopt approaches from model species that permit to obtain single‐base pair resolution data on methylation at a reasonable cost. Here, we used targeted bisulfite sequencing (TBS) to study methylation patterns in the conifer species Norway spruce (Picea abies). To circumvent the challenge of disentangling epigenetic and genetic differences, we focused on four clone pairs, where clone members were growing in different climatic conditions for 24 years. We targeted &gt;26.000 genes using TBS and determined the performance and reproducibility of this approach. We characterized gene body methylation and compared methylation patterns between environments. We found highly comparable capture efficiency and coverage across libraries. Methylation levels were relatively constant across gene bodies, with 21.3 ± 0.3%, 11.0 ± 0.4% and 1.3 ± 0.2% in the CG, CHG, and CHH context, respectively. The variance in methylation profiles did not reveal consistent changes between environments, yet we could identify 334 differentially methylated positions (DMPs) between environments. This supports that changes in methylation patterns are a possible pathway for a plant to respond to environmental change. After this successful application of TBS in Norway spruce, we are confident that this approach can contribute to broaden our knowledge of methylation patterns in natural tree populations.

opencc-zeroDec 2017View details →
dryad32/100

Data from: Targeted gene enrichment and high-throughput sequencing for environmental biomonitoring: a case study using freshwater macroinvertebrates

Recent studies have advocated biomonitoring using DNA techniques. In this study, two high-throughput sequencing (HTS)-based methods were evaluated: amplicon metabarcoding of the cytochrome C oxidase subunit I (COI) mitochondrial gene and gene enrichment using MYbaits (targeting nine different genes including COI). The gene-enrichment method does not require PCR amplification and thus avoids biases associated with universal primers. Macroinvertebrate samples were collected from 12 New Zealand rivers. Macroinvertebrates were morphologically identified and enumerated, and their biomass determined. DNA was extracted from all macroinvertebrate samples and HTS undertaken using the illumina miseq platform. Macroinvertebrate communities were characterized from sequence data using either six genes (three of the original nine were not used) or just the COI gene in isolation. The gene-enrichment method (all genes) detected the highest number of taxa and obtained the strongest Spearman rank correlations between the number of sequence reads, abundance and biomass in 67% of the samples. Median detection rates across rare (&lt;1% of the total abundance or biomass), moderately abundant (1–5%) and highly abundant (&gt;5%) taxa were highest using the gene-enrichment method (all genes). Our data indicated primer biases occurred during amplicon metabarcoding with greater than 80% of sequence reads originating from one taxon in several samples. The accuracy and sensitivity of both HTS methods would be improved with more comprehensive reference sequence databases. The data from this study illustrate the challenges of using PCR amplification-based methods for biomonitoring and highlight the potential benefits of using approaches, such as gene enrichment, which circumvent the need for an initial PCR step.

opencc-zeroDec 2014View details →
dryad32/100

Data from: Phylogenomic analyses of Sabal (Arecaceae) species relationships using targeted sequence capture

With the increasing availability of high-throughput sequencing, phylogenetic analyses are no longer constrained by the limited availability of a few loci. Here, we describe a sequence capture methodology, which we used to collect data for analyses of diversification within Sabal (Arecaceae), a palm genus native to the south-eastern USA, Caribbean, Bermuda and Central America. RNA probes were developed and used to enrich DNA samples for putatively low copy nuclear genes and the plastomes for all Sabal species and two outgroup species. Sequence data were generated on an Illumina MiSeq sequencer and target sequences were assembled using custom workflows. Both coalescence and supermatrix analyses of 133 nuclear genes were used to estimate species trees relationships. Plastid genomes were also analysed, yielding generally poor resolution with regard to species relationships. Species relationships described in both nuclear gene and plastome sequences largely reflect the biogeography of the group and, to a lesser extent, previous morphology-based hypotheses. Beyond the biological implications, this research validates a high-throughput methodology for generating a large number of genes for coalescence-based phylogenetic analyses in plant lineages.

opencc-zeroDec 2014View details →
dryad32/100

Data from: A universal probe set for targeted sequencing of 353 nuclear genes from any flowering plant designed using k-medoids clustering

Sequencing of target-enriched libraries is an efficient and cost-effective method for obtaining DNA sequence data from hundreds of nuclear loci for phylogeny reconstruction. Much of the cost of developing targeted sequencing approaches is associated with the generation of preliminary data needed for the identification of orthologous loci for probe design. In plants, identifying orthologous loci has proven difficult due to a large number of whole-genome duplication events, especially in the angiosperms (flowering plants). We used multiple sequence alignments from over 600 angiosperms for 353 putatively single-copy protein-coding genes identified by the One Thousand Plant Transcriptomes Initiative to design a set of targeted sequencing probes for phylogenetic studies of any angiosperm group. To maximize the phylogenetic potential of the probes while minimizing the cost of production, we introduce a k-medoids clustering approach to identify the minimum number of sequences necessary to represent each coding sequence in the final probe set. Using this method, five to 15 representative sequences were selected per orthologous locus, representing the sequence diversity of angiosperms more efficiently than if probes were designed using available sequenced genomes alone. To test our approximately 80,000 probes, we hybridized libraries from 42 species spanning all higher-order groups of angiosperms, with a focus on taxa not present in the sequence alignments used to design the probes. Out of a possible 353 coding sequences, we recovered an average of 283 per species and at least 100 in all species. Differences among taxa in sequence recovery could not be explained by relatedness to the representative taxa selected for probe design, suggesting that there is no phylogenetic bias in the probe set. Our probe set, which targeted 260 kbp of coding sequence, achieved a median recovery of 137 kbp per taxon in coding regions, a maximum recovery of 250 kbp, and an additional median of 212 kbp per taxon in flanking non-coding regions across all species. These results suggest that the Angiosperms353 probe set described here is effective for any group of flowering plants and would be useful for phylogenetic studies from the species level to higher-order groups, including the entire angiosperm clade itself.

opencc-zeroDec 2017View details →
zenodo32/100

Figure 10 in Target-enriched DNA sequencing from historical type material enables a partial revision of the Madagascar giant stream frogs (genus Mantidactylus)

Figure 10. Mantidactylus (Mantidactylus) radaka sp. nov. being prepared for human consumption. (a) Frogs and crabs are collected from broad streams. Then (b) the frogs are gutted and skinned, and the head, hands and feet removed. The frog is then rinsed in the stream, leaving (c) cleaned animals for cooking in a stew. Note the ovaries full with hundreds of eggs.

opennotspecifiedMay 2020View details →
zenodo32/100

Figure 9 in Target-enriched DNA sequencing from historical type material enables a partial revision of the Madagascar giant stream frogs (genus Mantidactylus)

Figure 9. Preserved type specimens of the four nomina in the Mantidactylus subgenus Mantidactylus and one of the paralectotypes of Rana guttulata.

opennotspecifiedMay 2020View details →
zenodo32/100

Figure 7 in Target-enriched DNA sequencing from historical type material enables a partial revision of the Madagascar giant stream frogs (genus Mantidactylus)

Figure 7. Photographs of living specimens of Mantidactylus (Mantidactylus) guttulatus, M. (M.) grandidieri, and of three candidate species. (a, b) M. (M.) guttulatus, female ZSM 1013/2003 (FGMV 2002.438) from Ranomafana. (c) Unidentified specimen from Ranomafana, assigned tentatively to M. (M.) guttulatus (no genetic evidence). (d, e) M. (M.) guttulatus, specimen KU 340853 (CRH729) from Ranomafana. (f) M. (M.) grandidieri, specimen ZSM 5077/2005 (ZCMV 2159) from Nosy Mangabe. (g) M. (M.) grandidieri, specimen ZSM 276/2005 (FGZC 2682) from Vohidrazana. (h) M. (M.) grandidieri, unidentified specimen (probably subadult) from Andranofotsy. (i, j) M. (M.) grandidieri, specimen KU

opennotspecifiedMay 2020View details →
zenodo32/100

Figure 5. Per-base coverage plots for the 16S in Target-enriched DNA sequencing from historical type material enables a partial revision of the Madagascar giant stream frogs (genus Mantidactylus)

Figure 5. Per-base coverage plots for the 16S fragment in four Mantidactylus type specimens from the MNHN and BMNH collections. (a) BMNH 1947.2.25.48 (paralectotype of Rana guttulata); (b) BMNH 1947.2.25.51 (paralectotype of Rana guttulata); (c) MNHN 1895.255 (syntype of M. grandidieri); (d) MNHN 1883.520 (syntype of M. grandidieri).

opennotspecifiedMay 2020View details →
zenodo32/100

Figure 3 in Target-enriched DNA sequencing from historical type material enables a partial revision of the Madagascar giant stream frogs (genus Mantidactylus)

Figure 3. Haplotype network of the subgenus Mantidactylus based on 1227 bp of the nuclear RAG-1 gene from 39 samples. Small black dots represent additional mutational steps.

opennotspecifiedMay 2020View details →
zenodo32/100

Figure 2 in Target-enriched DNA sequencing from historical type material enables a partial revision of the Madagascar giant stream frogs (genus Mantidactylus)

Figure 2. Diagonal matrix visualising the mean uncorrected genetic distances (p-distances) in the mitochondrial 16S rRNA gene between the different lineages in the subgenus Mantidactylus, calculated from 514 bp of the 16S mitochondrial gene.

opennotspecifiedMay 2020View details →
zenodo32/100

Figure 1. Maximum likelihood phylogenetic tree obtained from 514 in Target-enriched DNA sequencing from historical type material enables a partial revision of the Madagascar giant stream frogs (genus Mantidactylus)

Figure 1. Maximum likelihood phylogenetic tree obtained from 514 bp of the mitochondrial 16S rRNA gene. The values at the nodes are the bootstrap supports (not given for intra-lineage nodes for improved clarity). The type specimens of M. guttulatus and M. grandidieri from the London and Paris museum collections are highlighted in red and brown, respectively.

opennotspecifiedMay 2020View details →
zenodo32/100

Figure 4 in Target-enriched DNA sequencing from historical type material enables a partial revision of the Madagascar giant stream frogs (genus Mantidactylus)

Figure 4. Stacked barplots showing the number of reads uniquely matching different reference sequences for the three targeted mitochondrial genes with a similarity threshold of 98%. The Rana pigra type was not included because the number of reads was too low.

opennotspecifiedMay 2020View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record