Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
46
datasets available to search
ShareScore release 0.9.0
Dataset results
46 results for “Plastid genome”
Variation of and associations with the depth and evenness of sequencing coverage in a sample of archived plastid genomes
<p>Depth and evenness of sequencing coverage are considered potential indicators of genome assembly quality. In plastid genomics, where new data generation has outpaced the development of suitable assembly quality indicators, these coverage metrics could offer insights into the quality of plastomes of different sizes, structures, or taxonomic origins. However, the typical variation of sequencing depth and evenness among archived plastid genomes, their variability between plastome partitions, and any association with methodological factors have yet to be evaluated. This study explores the variation of sequencing depth and evenness across a sample of publicly accessible plastid genomes and their potential associations with plastome structure, assembly accuracy, and the methodological provenance of the genome data using statistical tests. Our results indicate significant differences in sequencing depth across the four structural partitions as well as between the coding and non-coding sections of the genomes, a significant correlation between sequencing evenness and the number of ambiguous nucleotides, and a significant difference in sequencing evenness between several DNA sequencing platforms. These findings highlight that many publicly accessible plastid genomes are based on sequence data with highly variable sequencing depth and evenness and that this variation is influenced, at least partially, by genome structure and methodological factors.</p>
Fig. 3 in Plastid genome of Aster altaicus var. uchiyamae Kitam., an endanger species of Korean asterids
Fig. 3. Comparison of chloroplast genomes of Aster altaicus var. uchiyamae and A. spathulifolius using mVISTA program. Grey arrows and thick black lines above the alignment indicate genes with their orientation and the position of the IRs, respectively. The Y-scale represents the percent identity between 50-100%. Genome regions are color-coded: Coding regions in blue; noncoding sequences (CNS) in red.
Text-fig. 2. Representative most parsimonious trees obtained after addition of Liliacidites to (A) the D&E tree (Text-fig. 1) and (B) the J/M tree, with relationships among major clades based on the plastid genome analyses of Jansen et al. (2007) and Moore et al. (2007). Thicker lines indicate all most parsimonious (MP), one step less parsimonious (MP+1), and two step less parsimonious (MP+2) positions for Liliacidites. Abbreviations as in Text-fig. 1. in Early Cretaceous Monocots: A Phylogenetic Evaluation
Text-fig. 2. Representative most parsimonious trees obtained after addition of Liliacidites to (A) the D&E tree (Text-fig. 1) and (B) the J/M tree, with relationships among major clades based on the plastid genome analyses of Jansen et al. (2007) and Moore et al. (2007). Thicker lines indicate all most parsimonious (MP), one step less parsimonious (MP+1), and two step less parsimonious (MP+2) positions for Liliacidites. Abbreviations as in Text-fig. 1.
Data from: NOVOWrap: an automated solution for plastid genome assembly and structure standardization
<p>Plastid genomes play an important role in genomics and evolutionary biology. Next-generation sequencing has revolutionized plastid genomic data acquisition to the point that genome assembly has become a bottlenecks for widespread utilization of plastid genome data. To solve this problem, we developed an open-source, cross-platform tool known as, NOVOWrap, which includes both command-line and graphical interfaces for automatically assembling plastid genomes on personal computers. With minimal inputs, settings, and user intervention, NOVOWrap can automatically assemble plastid genomes, validate results and standardize the structure using affordable computer resources. The performance of this software has been successfully benchmarked against the plastid genomes of 11 species belonging to lycopods, gymnosperms, and angiosperms. This program is expected to liberate researchers from laborious and cumbersome computer manipulations and create reliable and standardized genomic data.</p>
Data from: Evolution of the plastid genomes in diatoms
Diatoms are a monophyletic group of eukaryotic, single-celled heterokont algae. Despite years of phylogenetic research, relationships among major groups of diatoms remain uncertain. Here we assess diatom phylogenetic relationships using the plastid genome (plastome). The 22 previously published diatom plastomes showed variable genome size, gene content and extensive rearrangement. We report another 18 diatom plastome sequences ranging in size from 119,120 to 201,816 bp. Plagiogramma staurophorum had the largest plastome sequenced so far due to large inverted repeats and a 2971 bp group II intron insertion in petD. The previously reported loss of psaE, psaI and psaM genes in Rhizosolenia imbricata also occurred in the closely related species Rhizosolenia fallax. In the largest genome-scale phylogeny yet published for diatoms based on 103 shared plastid-coding genes from 40 diatoms and Triparma laevis as the outgroup, Leptocylindrus was recovered as sister to the remaining diatoms and the clade of Attheya plus Biddulphia was recovered as sister to pennate diatoms, strongly rejecting monophyly of two of the three proposed classes of diatoms. Our study also revealed extensive gene loss and a strong positive correlation between sequence divergence and gene order change in diatom plastomes.
the supplementary of Novel plastid genome characteristics in Fugacium kawagutii and accelerated evolution of plastid proteins in dinoflagellates
Open the record for dataset details and reuse information.
Data from: Remarkably conserved plastid genomes of Quercus Group Cerris in China: comparative and phylogenetic analyses
Quercus is one of the most important genera for considering its economic and ecological values, with approximately 500 species worldwide. Quercus group Cerris is endemic to Eurasia (including 11 species), and three species (Quercus acutissima, Quercus chenii and Quercus variabilis) are widely distributed in China. Here, we sequenced the complete plastid genomes of Q. acutissima and Q. chenii by Illumina pair-end sequencing, and obtained an additional plastome of Q. variabilis from GenBank. Although geographically distant sampling, the three plastomes in group Cerris were remarkably conserved with regard to genome size, gene organization, GC content, and IR/SC boundary regions. The phylogenetic analysis showed that group Cerris nested in group Ilex, forming a Cerris-Ilex clade. The current study provided plastid genomic-scale data for the less intensively studied group Cerris, which would be useful for studying speciation processes, geographical structure and phylogeny within the group Cerris in the future.
Polytomous radiation revealed in phylogenomic analysis of Allium (Amaryllidaceae) plastid genomes
<p>Alignment of 115 <em>Allium </em>chloroplast genomes plus three outgroups with all sites with missing data masked.</p>
<em>Terniopsis chanthaburiensis</em> (Podostemaceae), a new record for China and its complete plastid genome
Open the record for dataset details and reuse information.
Data from: NOVOWrap: an automated solution for plastid genome assembly and structure standardization
Open the record for dataset details and reuse information.
Data from: Remarkably conserved plastid genomes of Quercus Group Cerris in China: comparative and phylogenetic analyses
Open the record for dataset details and reuse information.
Data from: Evolution of the plastid genomes in diatoms
Open the record for dataset details and reuse information.
Data from: Rapid diversification rates in Amazonian Chrysobalanaceae inferred from plastid genome phylogenetics
<p>We studied the evolutionary history of Chrysobalanaceae with phylogenetic analyses of complete plastid genomes from 156 species to assess the tempo of diversification in the Neotropics and help to unravel the causes of Amazonian plant diversification. These plastid genomes had a mean length of 162,204 base pairs, and the nearly complete DNA sequence matrix, with reliable fossils, was used to estimate a phylogenetic tree. Chrysobalanaceae diversified from 38.9 Mya (95% highest posterior density, 95%HPD: 34.2-43.9 Mya). A single clade containing almost all Neotropical species arose after a single dispersal event from the Palaeotropics into the Amazonian biome <i>c.</i> 29.1 Mya (95%HPD: 25.5-32.6 Mya), with subsequent dispersals into other Neotropical biomes. All Neotropical genera diversified from 10 to 14 Mya, lending clear support to the role of Andean orogeny as a major cause of diversification in Chrysobalanaceae. In particular, the understory genus <i>Hirtella</i> diversified extremely rapidly, producing > 100 species in the last 6 My (95% HPD: 4.9-7.4 My). Our study suggests that a large fraction of the Amazonian tree flora has been assembled <i>in situ</i> within the last 15 My. </p>
Dissection for floral micromorphology and plastid genome of valuable medicinal borages Arnebia and Lithospermum (Boraginaceae)
<p>The genera <em>Arnebia </em>and <em>Lithospermum </em>(Lithospermeae-Boraginaceae) comprise 25–30 and 50–60 species, respectively. Some of them are economically valuable, as their roots frequently contain a purple-red dye used in the cosmetic industry. Furthermore, dried roots of <em>Arnebia euchroma</em>, <em>A. guttata</em>, and <em>Lithospermum erythrorhizon</em>, which have been designated Lithospermi Radix, are used as traditional Korean herbal medicine. This study is the first report on the floral micromorphology and complete chloroplast (cp) genome sequences of <em>A. guttata </em>(including <em>A. tibetana</em>), <em>A. euchroma</em>, and <em>L. erythrorhizon</em>. We reveal great diversity in floral epidermal cell patterns, gynoecium, and structure of trichomes. The cp genomes were 149,361–150,465 bp in length, with conserved quadripartite structures. In total, 112 genes were identified, including 78 protein-coding regions, 30 tRNA genes, and four rRNA genes. Gene order, content, and orientation were highly conserved and were consistent with the general structure of angiosperm cp genomes. Comparison of the four cp genomes revealed locally divergent regions, mainly within intergenic spacer regions (<em>atpH-atpI, petN-psbM, rbcL-psaI, ycf4-cemA, ndhF-rpl32, </em>and <em>ndhC-trnV-UAC</em>). To facilitate species identification, we developed molecular markers <em>psaA- ycf3 </em>(PSY), <em>trnI-CAU- ycf2 </em>(TCY), and <em>ndhC-trnV-UAC</em> (NCTV) based on divergence hotspots. High-resolution phylogenetic analysis revealed clear clustering and a close relationship of <em>Arnebia </em>to its <em>Lithospermum </em>sister group, which was supported by strong bootstrap values and posterior probabilities. Overall, gynoecium characteristics and genetic distance of cp genomes suggest that <em>A. tibetana</em>, might be recognized as an independent species rather than a synonym of <em>A. guttata</em>. The present morphological and cp genomic results provide useful information for future studies, such as taxonomic, phylogenetic, and evolutionary analysis of Boraginaceae.</p>
Data from: From algae to angiosperms–inferring the phylogeny of green plants (Viridiplantae) from 360 plastid genomes
Background: Next-generation sequencing has provided a wealth of plastid genome sequence data from an increasingly diverse set of green plants (Viridiplantae). Although these data have been useful for reconstructing the phylogeny of numerous clades of photosynthetic organisms (e.g., green algae, angiosperms, and gymnosperms), their utility for inferring relationships across all green plants is uncertain. Viridiplantae originated 700-1500 million years ago and may comprise as many as 500,000 species. This clade represents a major source of photosynthetic carbon and contains an immense diversity of life forms, including some of the smallest and largest eukaryotes. Here we explore the limits and challenges of inferring a comprehensive green plant phylogeny from available complete or nearly complete plastid genome data. Results: We assembled protein-coding sequence data for 78 genes from 360 diverse green plant taxa with complete or nearly complete plastid genome sequences available from GenBank. Phylogenetic analyses of the plastid data recovered well-supported backbone relationships and strong support for relationships that were not observed in previous analyses of major subclades within Viridiplantae. However, there also is evidence of systematic error in some analyses. In several instances we obtained strongly supported but conflicting topologies from analyses of nucleotides versus amino acid characters, and the considerable variation in GC content among lineages and within single genomes affected the phylogenetic placement of several taxa. Conclusions: Analyses of the plastid data recovered a strongly supported framework of relationships for green plants. This includes the placement of Zygnematophyceace as sister to land plants (Embryophyta) and a clade of extant gymnosperms (Acrogymnospermae) with cycads + Ginkgo sister to remaining members and with gnetophytes (Gnetophyta) sister to non-Pinaceae conifers (Gnecup trees); within the monilophyte clade (Monilophyta), relationships are strongly supported with Equisetales + Psilotales sister to Marattiales + leptosporangiate ferns. We also highlight the challenges of using plastid genome sequences in deep-level phylogenomic analyses and provide suggestions for future analyses that will likely incorporate plastid genome data for thousands of species. We particularly emphasize the importance of exploring the effects of different partitioning and character coding protocols for the entire data set as well as subsets of the data.
Data from: Recombination-dependent replication and gene conversion homogenize repeat sequences and diversify plastid genome structure
PREMISE OF THE STUDY: There is a misinterpretation in the literature regarding the variable orientation of the small single copy region of plastid genomes (plastomes). The common phenomenon of small and large single copy inversion, hypothesized to occur through intramolecular recombination between inverted repeats (IR) in a circular, single unit-genome, in fact more likely occurs through recombination-dependent replication (RDR) of linear plastome templates. If RDR can be primed through both intra- and intermolecular recombination, then this mechanism could not only create inversion isomers of so-called single copy regions, but also an array of alternative sequence arrangements. METHODS: We used Illumina paired-end and PacBio single-molecule real-time (SMRT) sequences to characterize repeat structure in the plastome of Monsonia emarginata L'Hér. (Geraniaceae). We used OrgConv and inspected nucleotide alignments to infer ancestral nucleotides and identify gene conversion among repeats and mapped long (>1 kb) SMRT reads against the unit-genome assembly to identify alternative sequence arrangements. RESULTS: Although M. emarginata lacks the canonical IR, we found that large repeats (>1 kilobase; kb) represent ~22% of the plastome nucleotide content. Among the largest repeats (>2 kb) we identified GC-biased gene conversion and mapping filtered, long SMRT reads to the M. emarginata unit-genome assembly revealed alternative, substoichiometric sequence arrangements. CONCLUSION: We offer a model based on RDR and gene conversion between long repeated sequences in the M. emarginata plastome, and provide support that both intra-and intermolecular recombination between large repeats, particularly in repeat-rich plastomes, varies unit-genome structure while homogenizing the nucleotide sequence of repeats.
Hoarding and horizontal transfer led to an expanded gene and intron repertoire in the plastid genome of the diatom, Toxarium undulatum (Bacillariophyta)
<p>Multiple sequence alignments used to produce Figure 2</p>
Plastid genome evolution in subtribe Gentianinae (Gentianaceae)
<p>We investigated plastome evolution in Subtribe Gentianinae of the Gentianaceae, which encompasses ca. 450 species distributed around the world, particularly in alpine and subalpine environments. We sequenced, assembled and annotated the plastomes of 41 species, representing all six genera in subtribe Gentianinae as well as all 14 sections of the species-rich genus <i>Gentiana</i>. Here, we upload (1) original alignments, Gblock alignments and tree topology file in phylogenetic analysis, and (2) results about a 5 kb insertion in <em>Gentiana</em> <i>cuneibarba</i>, including Sanger sequencing results which verified its two boundaries as well as the middle, and annotation results.</p>
Plastid genome random walks
<p>A common genome composition pattern in eubacteria is an asymmetry between the leading and lagging strands resulting in opposite skew patterns in the two replichores that lie between the origin and terminus of replication. Although this pattern has been reported for a couple of isolated plastid genomes, it is not clear how widespread it is overall in this chromosome. Using a random walk approach, we examine plastid genomes outside of the land plants, which are excluded since they are known not to initiate replication at a single site, for such a pattern of asymmetry. Although it is not a common feature, we find that it is detectable in the plastid genome of species from several diverse lineages. The euglenozoa in particular show a strong skew pattern as do several rhodophytes. There is a weaker pattern in some chlorophytes but it is not apparent in other lineages. The ramifications of this for analyses of plastid evolution are discussed.</p> <p>This dataset contains the random walk plots for the genomes analyzed.</p>
Osmanthus plastid genome sequence for: Plastid genomes reveal evolutionary shifts in elevational range and flowering time of Osmanthus (Oleaceae)
<p><span>Species of <em>Osmanthus</em> are economically important ornamental trees, yet information regarding their plastid genomes (plastomes) has rarely been reported, thus hindering taxonomic and evolutionary studies of this small but enigmatic genus. Here, we performed comparative genomics and evolutionary analyses on plastomes of 16 of the 28 currently accepted species, with 11 plastomes newly sequenced. Phylogenetic studies identified four main lineages within the genus that are here designated: 'Caucasian <em>Osmanthus</em>' (corresponding to <em>O</em>. <em>decorus</em>), '<em>Siphosmanthus</em>' (corresponding to <em>O</em>. sect. <em>Siphosmanthus</em>), '<em>O</em>. <em>serrulatus</em> + <em>O</em>. <em>yunnanensis</em>', and 'Core <em>Osmanthus</em>' (corresponding to <em>O</em>. sect. <em>Osmanthus</em> + <em>O</em>. sect. <em>Linocieroides</em>). Molecular clock analysis suggested that <em>Osmanthus</em> split from its sister clade c. 15.83 Ma. The estimated crown ages of the lineages were the following: genus <em>Osmanthus</em> at 12.66 Ma; '<em>Siphosmanthus</em>' clade at 5.85 Ma; '<em>O. serrulatus </em>+<em> O. yunnanensis</em>' at 4.89 Ma; 'Core <em>Osmanthus</em>' clade at 6.2 Ma. Ancestral state reconstructions and trait mapping showed that ancestors of <em>Osmanthus</em> were spring-flowering and originated at lower elevations. Phylogenetic principal component analysis clearly distinguished spring-flowering species from autumn-flowering species, suggesting that flowering time differentiation is related to the difference in ecological niches. Nucleotide substitution rates of 80 common genes showed a slow evolutionary pace and low nucleotide variations, all genes being subjected to purifying selection.</span></p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.