Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
41
datasets available to search
ShareScore release 0.9.0
Dataset results
41 results for “Genome Skimming”
A genome-skimmed phylogeny of a widespread bryozoan family, Adeonidae
<p>Understanding the phylogenetic relationships among species is one of the main goals of systematic biology. Simultaneously, credible phylogenetic hypotheses are often the first requirement for unveiling the evolutionary history of traits and for modelling macroevolutionary processes. However, many non-model taxa have not yet been sequenced to an extent such that statistically well-supported molecular phylogenies can be constructed for these purposes. Here, we use a genome-skimming approach to extract sequence information for 15 mitochondrial and 2 ribosomal operon genes from the cheilostome bryozoan family, the Adeonidae, Busk, 1884, whose current systematics is based purely on morphological traits. The members of the Adeonidae are, like all cheilostome bryozoans, benthic, colonial, marine organisms. Adeonids are also geographically widely-distributed, often locally common, and are sometimes important habitat-builders. Results We successfully genome-skimmed 35 adeonid colonies representing 6 genera (Adeona, Adeonellopsis, Bracebridgia, Adeonella, Laminopora and Cucullipora). We also contributed 16 new, circularised mitochondrial genomes to the eight previously published for cheilostome bryozoans. Using the aforementioned mitochondrial and ribosomal genes, we inferred the relationships among these 35 samples. Contrary to some previous suggestions, the Adeonidae is a robustly supported monophyletic clade. However, the genera Adeonella and Laminopora are in need of revision: Adeonella is polyphyletic and Laminopora paraphyletically forms a clade with some Adeonella species. Additionally, we assign a sequence clustering identity using cox1 barcoding region of 99% at the species and 83% at the genus level. Conclusions We provide sequence data, obtained via genome-skimming, that greatly increases the resolution of the phylogenetic relationships within the adeonids. We present a highly-supported topology based on 17 genes and substantially increase availability of circularised cheilostome mitochondrial genomes, and highlight how we can extend our pipeline to other bryozoans.</p>
Capturing single-copy nuclear genes, organellar genomes, and nuclear ribosomal DNA from deep genome skimming data for plant phylogenetics: A case study in Vitaceae
<p>With the decreasing cost and availability of many newly developed bioinformatics pipelines, next-generation sequencing (NGS) has revolutionized plant systematics in recent years. Genome skimming has been widely used to obtain high-copy fractions of the genomes, including plastomes, mitochondrial DNA (mtDNA), and nuclear ribosomal DNA (nrDNA). In this study, through simulations, we evaluated the optimal (minimum) sequencing depth and performance for recovering single-copy nuclear genes (SCNs) from genome skimming data, by subsampling genome resequencing data and generating 10 datasets with different sequencing coverage <i>in silico</i>. We tested the performance of four datasets (plastome, nrDNA, mtDNA, and SCNs) obtained from genome skimming based on phylogenetic analyses of the <i>Vitis</i> clade at the genus level and Vitaceae at the family level, respectively. Our results showed that optimal minimum sequencing depth for high-quality SCNs assembly via genome skimming was about 10× coverage. Without the steps of synthesizing baits and enrichment experiments, coupled with incredibly low sequencing costs, we showcase that deep genome skimming (DGS) is as effective for capturing large datasets of SCNs as the widely used Hyb-Seq approach, in addition to capturing plastomes, mtDNA, and entire nrDNA repeats. DGS may serve as an efficient and economical alternative and may be superior to the popular target enrichment/Hyb-Seq approach.</p>
A genome-skimmed phylogeny of a widespread bryozoan family, Adeonidae
Open the record for dataset details and reuse information.
Capturing single-copy nuclear genes, organellar genomes, and nuclear ribosomal DNA from deep genome skimming data for plant phylogenetics: A case study in Vitaceae
Open the record for dataset details and reuse information.
The postglacial history of Euphrasia micrantha in Scotland: evidence from genome skimming
Open the record for dataset details and reuse information.
Data from: Advancing Pyrus phylogeny: Deep genome skimming-based inference coupled with paralogy analysis yields a robust phylogenetic backbone and an updated infrageneric classification of the pear genus (Maleae, Rosaceae)
Open the record for dataset details and reuse information.
Data from: Evidence of intraspecific adaptive variation in the American pika (Ochotona princeps) on a continental scale using a target enrichment and mitochondrial genome skimming approach
Open the record for dataset details and reuse information.
A snakemake toolkit for the batch assembly, annotation, and phylogenetic analysis of mitochondrial genomes and ribosomal genes from genome skims of museum collections
Open the record for dataset details and reuse information.
The impact of contaminants on the accuracy of genome skimming and the effectiveness of exclusion read filters
Open the record for dataset details and reuse information.
Species delimitation, classical taxonomy, and genome skimming: a review of the ground beetle genus Lionepha (Coleoptera: Carabidae)
<p>The western North American genus <i>Lionepha</i> is shown to contain at least 11 species through a combination of eight-gene species delimitation analyses and morphological study. In order to confirm the names of several species, we sequence DNA of primary types of several names, including a LeConte lectotype collected in the 1850s, using next-generation sequencing. We examine chromosomes of eight species, and show that all have 12 pairs of autosomes and an X0/XX sex-chromosome system. The following species are described as new: <i>Lionepha australerasa</i>, <i>L. kavanaughi</i>, <i>L. lindrothi</i>, and <i>L. tuulukwa</i>. The name <i>Lionepha erasa</i> is shown to belong to a relatively rare, western species ranging from Oregon through Alaska; the common, widespread species previously known as <i>Lionepha erasa</i> now takes the name <i>L. probata</i>. <i>Bembidion lindrothellus</i>, <i>B. chintimini</i>, and <i>B. lummi </i>are synonymized with<i> L. erasa.</i> We provide tools to identify specimens to species, including illustrations and diagnoses.</p>
Data from: Lessons from genome skimming of arthropod-preserving ethanol
Field-collected specimens of invertebrates are regularly killed and preserved in ethanol, prior to DNA extraction from the specimens, while the ethanol fraction is usually discarded. However, DNA may be released from the specimens into the ethanol, which can potentially be exploited to study species diversity in the sample without the need for DNA extraction from tissue. We used shallow shotgun sequencing of the total DNA to characterize the preservative ethanol from two pools of insects (from a freshwater habitat and terrestrial habitat) to evaluate the efficiency of DNA transfer from the specimens to the ethanol. In parallel, the specimens themselves were subjected to bulk DNA extraction and shotgun sequencing, followed by assembly of mitochondrial genomes for 39 of 40 species in the two pools. Shotgun sequencing from the ethanol fraction and read-matching to the mitogenomes detected ~40% of the arthropod species in the ethanol, confirming the transfer of DNA whose quantity was correlated to the biomass of specimens. The comparison of diversity profiles of microbiota in specimen and ethanol samples showed that 'closed association' (internal tissue) bacterial species tend to be more abundant in DNA extracted from the specimens, while 'open association' symbionts were enriched in the preservative fluid. The vomiting reflex of many insects also ensures that gut content is released into the ethanol, which provides easy access to DNA from prey items. Shotgun sequencing of DNA from preservative ethanol provides novel opportunities for characterizing the functional or ecological components of an ecosystem and their trophic interactions.
Data from: The unexpected depths of genome-skimming data: a case study examining Goodeniaceae floral symmetry genes
Premise of the study: The use of genome skimming allows systematists to quickly generate large data sets, particularly of sequences in high abundance (e.g., plastomes); however, researchers may be overlooking data in low abundance that could be used for phylogenetic or evo-devo studies. Here, we present a bioinformatics approach that explores the low-abundance portion of genome-skimming next-generation sequencing libraries in the fan-flowered Goodeniaceae. Methods: Twenty-four previously constructed Goodeniaceae genome-skimming Illumina libraries were examined for their utility in mining low-copy nuclear genes involved in floral symmetry, specifically the CYCLOIDEA (CYC)-like genes. De novo assemblies were generated using multiple assemblers, and BLAST searches were performed for CYC1, CYC2, and CYC3 genes. Results: Overall Trinity, SOAPdenovo-Trans, and SOAPdenovo implementing lower k-mer values uncovered the most data, although no assembler consistently outperformed the others. Using SOAPdenovo-Trans across all 24 data sets, we recovered four CYC-like gene groups (CYC1, CYC2, CYC3A, and CYC3B) from a majority of the species. Alignments of the fragments included the entire coding sequence as well as upstream and downstream regions. Discussion: Genome-skimming data sets can provide a significant source of low-copy nuclear gene sequence data that may be used for multiple downstream applications.
Sequences of Staudtia kamerunensis obtained through low coverage whole genome skimming
<p>The impact of Pleistocene climatic oscillations on the biodiversity of African tropical rain forests remains poorly understood, and the Congo Basin is particularly understudied. We aim to elucidate how Pleistocene climatic oscillations shaped lowland tropical rain forests by investigating the intraspecific diversity and evolutionary history of a widespread tree species, <em>Staudtia kamerunensis</em> Warb.</p> <p>We sequenced 88 individuals of <em>Staudtia kamerunensis</em> and 1 of <em>Staudtia pterocarpa</em> using a genome skimming approach. We used maximum likelihood and Bayesian inference to infer the plastid phylogeny. We estimated the time of speciation and differentiation, genetic diversity, and we employed a continuous phylogeographic approach to infer the dispersal history of its plastid lineages.</p> <p>We identified five plastid lineages that diverged during the Early or Middle Pleistocene and are parapatric, suggesting past population fragmentation. Four lineages are endemic to Lower Guinea, and one spans the Congo Basin. We found contrasting patterns of expansion in the two regions, with a rapid and recent range expansion of the Congolian lineage in the last 200,000 years, while the spread of the Lower Guinean lineages was substantially slower.</p> <p>The contrasting demographic histories between eastern and western lineages, associated with contrasted levels of plant species richness and rates of endemism, suggest that forest cover was more stable in Lower Guinea during the Late Pleistocene than in Congolia, where the biodiversity might have been eroded before the forest re-expanded in the Congo basin. This study illustrates how a continuous phylogeographic inference approach, mostly applied so far for inferring the spread of fast-evolving pathogens over months or years, can provide new insights to reconstruct the dispersal history of tropical tree species over thousands or millions of years.</p>
Genome Skimming of Thysanoptera (Arthropoda, Insecta) and Its Taxonomic and Systematic Applications
<p>High-throughput sequencing has transformed molecular systematics. This study presents a semi-automated pipeline for genome skimming in Thysanoptera, an insect order known for challenging species identification and cryptic relationships. By efficiently obtaining mitochondrial genomes and nuclear genes from multiple thrips specimens, the study evaluates the limitations of traditional barcoding and the data required for accurate species delimitation. The results highlight the importance of the sequencing data volume and this pipeline in reconstructing Thysanoptera phylogeny. This research also showcases the potential of advanced sequencing techniques for species delimitation and phylogenetics.</p>
An evolutionary framework of Acanthaceae based on transcriptomes and genome skims
<p>Acanthaceae is a family of tropical flowering plants with approximately 4000 species. Despite remarkable variation in morphological traits, research on patterns of character evolution has been limited by uncertain relationships among some of the major lineages. We sampled from these major lineages to estimate a phylogenomic framework using a combination of newly sequenced shotgun genome skims plus new and publicly available transcriptomes. We used OrthoFinder2 to infer a species tree with strong branch support. Except for the placement of Crabbea, our results corroborate the most recent chloroplast and nrITS sequence-based topology. Of 587 single copy loci, 10 were recovered for all 16 species; a RAxML tree estimated from these 10 loci resulted in the same topology as other datasets assembled in this study, with the exception of relationships among three sampled species of Barleria; however, branch support was lower compared to the tree reconstructed using more data. ABBA-BABA tests were conducted to investigate patterns of introgression involving Crabbea; few nucleotides supported alternative topologies. SplitsTree networks of the 587 loci and 6,136 trees revealed conflict among the branches leading to Andrographideae, Whitfieldieae, and Neuracanthus. A principal components analysis in treespace found no distinct clusters of trees. Our results strongly corroborate the previously published chloroplast and nr-ITS-based phylogeny of Acanthaceae with increased resolution among Barlerieae, Andrographideae, Whitfieldieae, and Neuracanthus. We propose that the tree presented here is the best estimate to date of Acanthaceae phylogeny. This advance in our knowledge of relationships will allow us to investigate character evolution and other phenomena within this diverse group of plants.</p>
Data from: Testing genome skimming for species discrimination in the large and taxonomically difficult genus Rhododendron
<p>Standard plant DNA barcodes based on 2-3 plastid regions, and nrDNA ITS show variable levels of resolution, and fail to discriminate among species in many plant groups. Genome skimming to recover complete plastid genome sequences and nrDNA arrays has been proposed as a solution to address these resolution limitations. However, few studies have empirically tested what gains are achieved in practice. Of particular interest is whether adding substantially more plastid and nrDNA characters will lead to an increase in discriminatory power, or whether the resolution limitations of standard plants barcodes are fundamentally due to plastid genomes and nrDNA not tracking species boundaries. To address this, we used genome skimming to recover near-complete plastid genomes and nuclear ribosomal DNA from <i>Rhododendron </i>species and compared discrimination success with standard plant barcodes<i>. </i>We sampled 218 individuals representing 145 species of this species-rich and taxonomically difficult genus, focusing on the global biodiversity hotspots of the Himalaya-Hengduan Mountains. Only 33% of species were distinguished using ITS+<i>matK</i>+<i>rbcL</i>+<i>trnH-psbA. </i>In contrast, 55% of species were distinguished using plastid genome and nrDNA sequences. The vast majority of this increase is due to the additional plastid characters. Thus, despite previous studies showing an asymptote in discrimination success beyond 3-4 plastid regions, these results show that a demonstrable increase in discriminatory power is possible with extensive plastid genome data. However, despite these gains, many species remain unresolved, and these results also reinforce the need to access multiple unlinked nuclear loci to obtain transformative gains in species discrimination in plants.</p>
Data from: The unexpected depths of genome-skimming data: a case study examining Goodeniaceae floral symmetry genes
Open the record for dataset details and reuse information.
Data from: Genome skimming reveals the origin of the Jerusalem Artichoke tuber crop species: neither from Jerusalem nor an Artichoke
Open the record for dataset details and reuse information.
Species delimitation, classical taxonomy, and genome skimming: a review of the ground beetle genus Lionepha (Coleoptera: Carabidae)
Open the record for dataset details and reuse information.
An evolutionary framework of Acanthaceae based on transcriptomes and genome skims
Open the record for dataset details and reuse information.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.