Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
190
datasets available to search
ShareScore release 0.9.0
Dataset results
190 results for “targeted enrichment”
Data from: Microfluidic PCR-based target enrichment: a case study in two rapid radiations of Commiphora (Burseraceae) from Madagascar
Open the record for dataset details and reuse information.
Data from: Targeted gene enrichment and high-throughput sequencing for environmental biomonitoring: a case study using freshwater macroinvertebrates
Open the record for dataset details and reuse information.
Building a robust backbone for Astragalus (Fabaceae) using a clade-specific target enrichment bait set
Open the record for dataset details and reuse information.
A target enrichment probe set for resolving the flagellate land plant tree of life
Open the record for dataset details and reuse information.
Data from: Tracking temporal shifts in area, biomes, and pollinators in the radiation of Salvia (sages) across continents: leveraging anchored hybrid enrichment and targeted sequence data
Open the record for dataset details and reuse information.
Data from: An enhanced target-enrichment bait set for Hexacorallia provides phylogenomic resolution of the staghorn corals (Acroporidae) and close relatives
Open the record for dataset details and reuse information.
Analysis of paralogs in target enrichment data pinpoints multiple ancient polyploidy events in Alchemilla s.l. (Rosaceae)
Open the record for dataset details and reuse information.
Target enrichment sequence data of Odonata, supporting Bybee et al. 2021
Open the record for dataset details and reuse information.
Data from: Phylogenetic marker development for target enrichment from transcriptome and genome skim data: the pipeline and its application in southern African Oxalis (Oxalidaceae)
Open the record for dataset details and reuse information.
Assessment of targeted enrichment locus capture across time and museums using odonate specimens
Open the record for dataset details and reuse information.
Data from: Resolving phylogenetic relationships of the recently radiated carnivorous plant genus Sarracenia using target enrichment
Open the record for dataset details and reuse information.
Data from: Using targeted enrichment of nuclear genes to increase phylogenetic resolution in the neotropical rain forest genus Inga (Leguminosae: Mimosoideae)
Open the record for dataset details and reuse information.
Data from: A phylogeny of birds based on over 1,500 loci collected by target enrichment and high-throughput sequencing
Open the record for dataset details and reuse information.
A phylogenomic perspective on gene tree conflict and character evolution in Caprifoliaceae using target enrichment data, with Zabelioideae recognized as a new subfamily
Open the record for dataset details and reuse information.
covtobed test files - 2 BAM files produced by target enrichment experiments
<p>2 BAM files produced by target enrichment experiments (sequences have been masked) used to test and benchmark "covtobed"</p> <p>https://github.com/telatin/covtobed</p>
Data from: Schneider et al. (2020). Phylogenomics of the tropical plant family Ochnaceae using targeted enrichment of nuclear genes and 250+ taxa. Taxon.
<p>DNA sequence alignments with all loci concatenated (e.g., "Alignment_concatenated_xxx_dataset") or for each locus separated (see folders "Gene_alignments_xxx_dataset"). The individual gene alignments are identified by their locus number. Numbers in the sequence headers of the fasta files correspond to the Lab IDs of specimens (see related publication for detailed voucher information). These alignments were used for the phylogenetic analyses in the related publication. The bait set contains the probe sequences used for the targeted enrichment of nuclear loci of Ochnaceae.</p>
Data from: A phylogenomic approach based on PCR target enrichment and high throughput sequencing: resolving the diversity within the South American species of Bartsia l. (Orobanchaceae)
Advances in high-throughput sequencing (HTS) have allowed researchers to obtain large amounts of biological sequence information at speeds and costs unimaginable only a decade ago. Phylogenetics, and the study of evolution in general, is quickly migrating towards using HTS to generate larger and more complex molecular datasets. In this paper, we present a method that utilizes microfluidic PCR and HTS to generate large amounts of sequence data suitable for phylogenetic analyses. The approach uses the Fluidigm Access Array System (Fluidigm, San Francisco, CA, USA) and two sets of PCR primers to simultaneously amplify 48 target regions across 48 samples, incorporating sample-specific barcodes and HTS adapters (2,304 unique amplicons per Access Array). The final product is a pooled set of amplicons ready to be sequenced, and thus, there is no need to construct separate, costly genomic libraries for each sample. Further, we present a bioinformatics pipeline to process the raw HTS reads to either generate consensus sequences (with or without ambiguities) for every locus in every sample or—more importantly—recover the separate alleles from heterozygous target regions in each sample. This is important because it adds allelic information that is well suited for coalescent-based phylogenetic analyses that are becoming very common in conservation and evolutionary biology. To test our approach and bioinformatics pipeline, we sequenced 576 samples across 96 target regions belonging to the South American clade of the genus Bartsia L. in the plant family Orobanchaceae. After sequencing cleanup and alignment, the experiment resulted in ~25,300bp across 486 samples for a set of 48 primer pairs targeting the plastome, and ~13,500bp for 363 samples for a set of primers targeting regions in the nuclear genome. Finally, we constructed a combined concatenated matrix from all 96 primer combinations, resulting in a combined aligned length of ~40,500bp for 349 samples.
Data from: A target enrichment method for gathering phylogenetic information from hundreds of loci: an example from the Compositae
Premise of the study: The Compositae (Asteraceae) are a large and diverse family of plants, and the most comprehensive phylogeny to date is a meta-tree based on 10 chloroplast loci that has several major unresolved nodes. We describe the development of an approach that enables the rapid sequencing of large numbers of orthologous nuclear loci to facilitate efficient phylogenomic analyses. Methods and Results: We designed a set of sequence capture probes that target conserved orthologous sequences in the Compositae. We also developed a bioinformatic and phylogenetic workflow for processing and analyzing the resulting data. Application of our approach to 15 species from across the Compositae resulted in the production of phylogenetically informative sequence data from 763 loci and the successful reconstruction of known phylogenetic relationships across the family. Conclusions: These methods should be of great use to members of the broader Compositae community, and the general approach should also be of use to researchers studying other families.
Data from: Ultraconserved elements anchor thousands of genetic markers for target enrichment spanning multiple evolutionary timescales
Although massively parallel sequencing has facilitated large-scale DNA sequencing, comparisons among distantly related species rely upon small portions of the genome that are easily aligned. Methods are needed to efficiently obtain comparable DNA fragments prior to massively parallel sequencing, particularly for biologists working with non-model organisms. We introduce a new class of molecular marker, anchored by ultraconserved genomic elements (UCEs), that universally enable target enrichment and sequencing of thousands of orthologous loci across species separated by hundreds of millions of years of evolution. Our analyses here focus on use of UCE markers in Amniota, because UCEs and phylogenetic relationships are well known in some amniotes. We perform an in silico experiment to demonstrate that sequence flanking 2,030 UCEs contains information sufficient to enable unambiguous recovery of the established primate phylogeny. We extend this experiment by performing an in vitro enrichment of 2,386 UCE-anchored loci from nine, non-model avian species. We then use alignments of 854 of these loci to unambiguously recover the established evolutionary relationships within and among three ancient bird lineages. Because many organismal lineages have UCEs, this type of genetic marker and the analytical framework we outline can be applied across the tree of life, potentially reshaping our understanding of phylogeny at many taxonomic levels.
Fig. 3 in Phylogenomics of the tropical plant family Ochnaceae using targeted enrichment of nuclear genes and 250+ taxa
Fig. 3. Continued. RAxML tree based on the concatenated 83 nuclear loci of the LEO dataset. Numbers on the branches are bootstrap values (BS)>50%; additionally, LPP and quartet support values (QSV) from MSC analysis (see suppl. Fig. S4) are given in the order BS/LLP/QSV for nodes along the backbone of Ochneae. The indicated classification of subfamilies and tribes follows Schneider & al. (2014). Numbers in parentheses after species names correspond to the specimen IDs (only for species with multiple accessions).
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.