Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
41
datasets available to search
ShareScore release 0.7.1
Dataset results
41 results for “codon usage”
Supplementary data for Willemsen et al., 2024 "Novel high-quality amoeba genomes reveal widespread codon usage mismatch between giant viruses and their hosts".
<p>Supplementary data for Willemsen et al., 2024 "Novel high-quality amoeba genomes reveal widespread codon usage mismatch between giant viruses and their hosts". The data set consists of five folders: “Codon_usage_amoebae_and_viruses”, "Genome_annotations_amoebae", "Phylogenetic_trees_18S_amoebae", “Phylogenomic_trees_amoebae”, and "Viral_integration_detection_amoebae". The “Codon_usage_amoebae_and_viruses” folder contains for each amoeba host the calculated codon usage tables in the subfolder "codon_usage_table_host", the calculated codon usage preferences using different scores in the subfolder "codon_usage_scores_host", and the calculated codon usage preferences of giant viruses versus each host in the subfolder "codon_usage_scores_viruses_vs_host". The giant viruses in the subfolder "codon_usage_scores_viruses_vs_host" are organised by viral family and genus in separate sub-subfolders. The "Genome_annotations_amoebae" folder contains the generated genome annotations in different formats and the manually curated mitochondrial genome annotations for each amoeba host. The "Phylogenetic_trees_18S_amoebae" contains for the eukaryotic phyla <em>Discosea</em>, <em>Heterolobosea</em>, and <em>Tubulinea, </em>the 18S rRNA nucleotide alignments, distance matrices, and computed phylogenetic trees. The folder "Phylogenomic_trees_amoebae" contains for the eukaryotic clades <em>Amoebozoa</em> and <em>Discoba, </em>the protein alignment matrices and computed phylogenomic trees. The folder "Viral_integration_detection_amoebae" contains the MCP databases used (fasta file, alignment file, HMM profile and DIAMOND BLASTX database) and the MCP sequences detected in this study and the blast results of these. </p>
Modeling and measuring how codon usage modulates the relationship between burden and yield during protein overexpression in bacteria
<p>Additional data from experiments associated with paper revisions.</p> <p>Also added codon optimizer script.</p>
Transcriptome-wide meta-analysis of codon usage in Escherichia coli
<p>Data generated by the CUBseq pipeline on Escherichia coli RNA-seq data.</p>
Determinants of associations between codon and amino acid usage patterns of microbial communities and the environment inferred based on a cross-biome metagenomic analysis
<p>Raw data set for npj Bioflims and Microbiome article: “Determinants of associations between codon and amino acid usage patterns of microbial communities and the environment inferred based on a cross-biome metagenomic analysis”</p>
Simulation of the evolution of codon usage in cpDNA
<p>The codon usage of the Angiosperm <i>psbA</i> gene is atypical for flowering plant chloroplast genes but similar to the codon usage observed in highly expressed plastid genes from some other Plantae, particularly Chlorobionta, lineages. The pattern of codon bias in these genes is suggestive of selection for a set of translationally optimal codons but the degree of bias towards these optimal codons is much weaker in the flowering plant <i>psbA</i> gene than in high expression plastid genes from lineages such as certain green algal groups. Two scenarios have been proposed to explain these observations. One is that the flowering plant <i>psbA</i> gene is currently under weak selective constraints for translation efficiency, the other is that there are no current selective constraints and we are observing the remnants of an ancestral codon adaptation that is decaying under mutational pressure. We test these two models using simulations studies that incorporate the context-dependent mutational properties of plant chloroplast DNA. We first reconstruct ancestral sequences and then simulate their evolution in the absence of selection on codon usage by using mutation dynamics estimated from intergenic regions. The results show that <i>psbA</i> has a significantly higher level of codon adaptation than expected while other chloroplast genes are within the range predicted by the simulations. These results suggest that there have been selective constraints on the codon usage of the flowering plant <i>psbA</i> gene during Angiosperm evolution.</p>
Simulation of the evolution of codon usage in cpDNA
Open the record for dataset details and reuse information.
Data from: Genomic analysis of codon usage shows influence of mutation pressure, natural selection, and host features on Marburg virus evolution
Background. The Marburg virus (MARV) has a negative-sense single-stranded RNA genome, belongs to the family Filoviridae, and is responsible for several outbreaks of highly fatal hemorrhagic fever. Codon usage patterns of viruses reflect a series of evolutionary changes that enable viruses to shape their survival rates and fitness toward the external environment and, most importantly, their hosts. To understand the evolution of MARV at the codon level, we report a comprehensive analysis of synonymous codon usage patterns in MARV genomes. Multiple codon analysis approaches and statistical methods were performed to determine overall codon usage patterns, biases in codon usage, and influence of various factors, including mutation pressure, natural selection, and its two hosts, Homo sapiens and Rousettus aegyptiacus. Results. Nucleotide composition and relative synonymous codon usage (RSCU) analysis revealed that MARV shows mutation bias and prefers U- and A-ended codons to code amino acids. Effective number of codons analysis indicated that overall codon usage among MARV genomes is slightly biased. The Parity Rule 2 plot analysis showed that GC and AU nucleotides were not used proportionally which accounts for the presence of natural selection. Codon usage patterns of MARV were also found to be influenced by its hosts. This indicates that MARV have evolved codon usage patterns that are specific to both of its hosts. Moreover, selection pressure from R. aegyptiacus on the MARV RSCU patterns was found to be dominant compared with that from H. sapiens. Overall, mutation pressure was found to be the most important and dominant force that shapes codon usage patterns in MARV. Conclusions. To our knowledge, this is the first detailed codon usage analysis of MARV and extends our understanding of the mechanisms that contribute to codon usage and evolution of MARV.
Data from: Gene expression levels are correlated with synonymous codon usage, amino acid composition and gene architecture in the red flour beetle, Tribolium castaneum
Gene expression levels correlate with multiple aspects of gene sequence and gene structure in phylogenetically diverse taxa suggesting an important role of gene expression levels in the evolution of protein-coding genes. Here we present results of a genome-wide study of the influence of gene expression on synonymous codon usage, amino acid composition and gene structure in the red flour beetle, Tribolium castaneum. Consistent with the action of translational selection, we find that synonymous codon usage bias increases with gene expression. However, the correspondence between tRNA gene copy number and optimal codons is weak. At the amino acid level, translational selection is suggested by the positive correlation between tRNA gene numbers and amino acid usage which is stronger for highly expressed genes. In addition, there is a clear trend for increased use of metabolically cheaper, less complex, amino acids as gene expression increases. tRNA gene numbers also correlate negatively with amino acid size/complexity score indicating the coupling between translational selection and selection to minimize the use of large/complex amino acids. Interestingly, the correlation between tRNA gene numbers and amino acid size/complexity score appears to be widespread given our analyses of 10 additional genomes and might be explained by selection against negative consequences of protein misfolding. At the level of gene structure, three major trends are detected 1) CDS length increases across low and intermediate expression levels but decreases in highly expressed genes; 2) the average intron size shows the opposite trend, first decreasing with expression, followed by a slight increase in highly expressed genes and 3) intron density remains nearly constant across all expression levels. These changes in gene architecture are only in partial agreement with selection favoring reduced cost of biosynthesis.
Data from: Antagonistic relationships between intron content and codon usage bias of genes in three mosquito species: functional and evolutionary implications
Genome biology of mosquitoes holds potential in developing knowledge-based control strategies against vector-borne diseases such as malaria, dengue, West Nile Virus and others. Although the genomes of three major vector mosquitoes have been sequenced, attempts to elucidate the relationship between intron and codon usage bias across species in phylogenetic contexts are limited. In this study, we investigated the relationship between intron content and codon bias of orthologous genes among three vector mosquito species. We found an antagonistic relationship between codon usage bias and the intron number of genes in each mosquito species. The pattern is further evident among the intronless and the intron-containing orthologous genes associated with either low or high codon bias among the three species. Furthermore, the co-variance between codon bias and intron number has a directional component associated with the species phylogeny when compared with other non-mosquito insects. By applying a maximum likelihood based continuous regression method, we show that codon bias and intron content of genes vary among the insects in a phylogeny dependent manner but with no evidence of adaptive radiation or species-specific adaptation. We discuss the functional and evolutionary significance of antagonistic relationships between intron content and codon bias.
Data from: Translational selection frequently overcomes genetic drift in shaping synonymous codon usage patterns in vertebrates
Synonymous codon usage (SCU) patterns are shaped by a balance between mutation, drift, and natural selection. To date, detection of translational selection in vertebrates has proven to be a challenging task, obscured by small long-term effective population sizes in larger animals and the existence of isochores in some species. The consensus is that, in such species, natural selection is either completely ineffective at overcoming mutational pressures and genetic drift or perhaps is effective but so weak that it is not detectable. The aim of this research is to understand the interplay between mutation, selection, and genetic drift in vertebrates. We observe that although variation in mutational bias is undoubtedly the dominant force influencing codon usage, translational selection acts as a weak additional factor influencing synonymous codon usage. These observations indicate that translational selection is a widespread phenomenon in vertebrates and is not limited to a few species.
Data from: Mitochondrial phylogenomics of early land plants: mitigating the effects of saturation, compositional heterogeneity, and codon-usage bias
Phylogenetic analyses using concatenation of genomic-scale data have been seen as the panacea to resolving the incongruences among inferences from few or single genes. However, phylogenomics may also suffer from systematic errors, due to the, perhaps cumulative, effects of saturation, among-taxa compositional (GC content) heterogeneity, or codon-usage bias plaguing the individual nucleotide loci that are concatenated. Here we provide an example of how these factors affect the inferences of the phylogeny of early land plants based on mitochondrial genomic data. Mitochondrial sequences evolve slowly in plants and hence are thought to be suitable for resolving deep relationships. We newly assembled mitochondrial genomes from 20 bryophytes, complemented these with 40 other streptophytes (land plants plus algal outgroups), compiling a data matrix of 60 taxa and 41 mitochondrial genes. Homogeneous analyses of the concatenated nucleotide data resolve mosses as sister-group to the remaining land plants. However, the corresponding translated amino acid data support the liverwort lineage in this position. Both results receive weak to moderate support in maximum likelihood analyses, but strong support in Bayesian inferences. Tests of alternative hypotheses using either nucleotide or amino-acid data provide implicit support for the respective optimal topologies. By analyzing the nucleotide data, we found that the 3rd codon positions are more saturated than the 1st and 2nd codon positions, and excluding these from the analyses leads to a topology congruent with that obtained using amino-acid data. Further, we determined that land plant lineages differ in their nucleotide composition, and in their usage of synonymous codon variants. Composition heterogeneous Bayesian analyses employing a non-stationary model that accounts for variation in among-lineage composition, and inferences from degenerated nucleotide data that avoids the effects of synonymous mutations that underlie codon-usage bias, again recovered liverworts being sister to the remaining land plants. These analyses indicate that the discrepancy between the nucleotide-based and the amino acid-based trees is caused by the lineage specific, parallel compositional bias, or synonymous mutations driving codon-usage bias, as well as saturation in the 3rd codon positions. While genomic data may generate highly supported phylogenetic trees, these inferences may be artifacts. We suggest that phylogenomic analyses should assess the possible impact of potential biases through comparisons of protein coding gene data and their amino-acids translations, by analyzing data modeling compositional bias, and by excluding nucleotide noisy signals due to saturation or codon-usage bias. We caution against relying on any one presentation of the data (nucleotide or amino acid) or any one type of analysis even when analyzing large-scale data sets, no matter how well-supported, without fully exploring the effects of substitution models.
Variability in codon usage in Coronaviruses is mainly driven by mutational bias and selective constraints on CpG dinucleotide
<p>Supplementary Figures and Tables for the article called: " Variability in codon usage in Coronaviruses<em> </em>is mainly driven by mutational bias and selective constraints on CpG dinucleotide<sup>"</sup></p>
Data from: Genomic analysis of codon usage shows influence of mutation pressure, natural selection, and host features on Marburg virus evolution
Open the record for dataset details and reuse information.
Data from: Translational selection frequently overcomes genetic drift in shaping synonymous codon usage patterns in vertebrates
Open the record for dataset details and reuse information.
Data from: Mitochondrial phylogenomics of early land plants: mitigating the effects of saturation, compositional heterogeneity, and codon-usage bias
Open the record for dataset details and reuse information.
Data from: Serine codon-usage bias in deep phylogenomics: pancrustacean relationships as a case study
Open the record for dataset details and reuse information.
Data from: Antagonistic relationships between intron content and codon usage bias of genes in three mosquito species: functional and evolutionary implications
Open the record for dataset details and reuse information.
Data from: Gene expression levels are correlated with synonymous codon usage, amino acid composition and gene architecture in the red flour beetle, Tribolium castaneum
Open the record for dataset details and reuse information.
Codon usage influences the local rate of translation elongation to regulate co-translational protein folding
GEO Series GSE71032. Neurospora crassa. 8 samples. Type: Expression profiling by high throughput sequencing; Other.
Synonymous codon usage regulates translation initiation
GEO Series GSE202900. Homo sapiens. 12 samples. Type: Expression profiling by high throughput sequencing; Other.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.