Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

231

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

231 results for “codon”

Learn how ShareScore rates datasets ↗
zenodo32/100

Distribution of alternative start codons for each microbial Refseq genome

<p>Distributions of alternative start codons for each microbial Refseq genome.&nbsp;</p>

opencc-zeroDec 2014View details →
zenodo32/100

Data for Standard Codon Substitution Models Overestimate Purifying Selection for Non-Stationary Data

<p>Codon-aligned, filtered alignments for Kaehler et al. (2016) (https://peerj.com/preprints/2218/). Please refer to preprint for preparation details.</p> <p>Data obtained from Ensembl (http://www.ensembl.org/) and antbase (http://antbase.org).</p> <p> </p> <p> </p>

opencc-by-nc-4.0Dec 2016View details →
dryad32/100

Data from: The twenty amino acids are identified by unique numbers assigned to the uracil, cytosine, adenine, and guanine found in the three base positions of the sixty-four messenger RNA genetic codons

<p>A codon's three bases consist of any combination of uracil, cytosine, adenine, or guanine and these encode the twenty amino acids. When the codon's first two bases are given specific values, and those values are multiplied, then the third base of the codon is used during translation only when the product is greater than three. Here we show that those values plus more variables within the ribosomal decoding site results in specific flow values for each of the twenty amino acid groups. These results are demonstrated in a flow chart showing the unidirectional flow which is expected during the translation process. All twenty amino acids can be represented by numbers that describe their relationship to each other and to the decoding site. We anticipate our findings will increase discussion about using a number system to better understand the translation process.</p>

opencc-zeroJan 2024View details →
dryad32/100

Simulation of the evolution of codon usage in cpDNA

<p>The codon usage of the Angiosperm <i>psbA</i> gene is atypical for flowering plant chloroplast genes but similar to the codon usage observed in highly expressed plastid genes from some other Plantae, particularly Chlorobionta, lineages. The pattern of codon bias in these genes is suggestive of selection for a set of translationally optimal codons but the degree of bias towards these optimal codons is much weaker in the flowering plant <i>psbA</i> gene than in high expression plastid genes from lineages such as certain green algal groups. Two scenarios have been proposed to explain these observations. One is that the flowering plant <i>psbA</i> gene is currently under weak selective constraints for translation efficiency, the other is that there are no current selective constraints and we are observing the remnants of an ancestral codon adaptation that is decaying under mutational pressure. We test these two models using simulations studies that incorporate the context-dependent mutational properties of plant chloroplast DNA. We first reconstruct ancestral sequences and then simulate their evolution in the absence of selection on codon usage by using mutation dynamics estimated from intergenic regions. The results show that <i>psbA</i> has a significantly higher level of codon adaptation than expected while other chloroplast genes are within the range predicted by the simulations. These results suggest that there have been selective constraints on the codon usage of the flowering plant <i>psbA</i> gene during Angiosperm evolution.</p>

opencc-zeroOct 2021View details →
zenodo32/100

Translational recoding by chemical modification of non-AUG start codon bases: Simulation dataset

<p>Dataset of molecular dynamics simulations for RNA binding in ribosomal pre-initiation complex.<br> NAMD version 2.13 (multi-core) was used for the simulations.</p>

opencc-by-4.0Jun 2024View details →
zenodo32/100

Ultradeep characterisation of translational sequence determinants refutes rare-codon hypothesis and unveils quadruplet base pairing of initiator tRNA and transcript

<p>## Overview</p> <p>This repository contains the data and R-scripts to reproduce the figures in the main text of the manuscript &quot;Ultradeep characterisation of translational sequence determinants refutes rare-codon hypothesis and unveils quadruplet base pairing of initiator tRNA and transcript&quot;, which can be found here: (https://doi.org/10.1093/nar/gkad040).</p> <p>&nbsp;</p> <p>## Additional information</p> <p>The data can also be found on github via https://github.com/JeschekLab/uASPIre_UTR_CDS. The github repository also includes updates, scripts for NGS data analysis and additional code for data processing.</p> <p>&nbsp;</p>

opencc-by-nc-4.0Feb 2023View details →
ClinicalTrials.gov32/100

Melpida: Recombinant Adeno-associated Virus (serotype 9) Encoding a Codon Optimized Human AP4M1 Transgene (hAP4M1opt)

ClinicalTrials.gov study NCT05518188. IPD Sharing: YES. Countries: 1. Publications: 1.

controlledIPD-YESFeb 2026View details →
ClinicalTrials.gov32/100

Six Month Study of Gentamicin in Duchenne Muscular Dystrophy With Stop Codons

ClinicalTrials.gov study NCT00451074. IPD Sharing: Not stated. Countries: 1. Publications: 2.

restrictedIPD-UNDECIDEDFeb 2026View details →
dryad32/100

Data from: Mitochondrial gene diversity associated with the atp9 stop codon in natural populations of wild carrot (Daucus carota ssp. carota)

Open the record for dataset details and reuse information.

publicNov 2011View details →
dryad32/100

Simulation of the evolution of codon usage in cpDNA

Open the record for dataset details and reuse information.

publicNov 2021View details →
dryad32/100

Data from: Ribosome profiling reveals pervasive and regulated stop codon readthrough in Drosophila melanogaster

Open the record for dataset details and reuse information.

publicOct 2014View details →
dryad32/100

Data from: A phenotype-genotype codon model for detecting adaptive evolution

Open the record for dataset details and reuse information.

publicNov 2019View details →
dryad32/100

Data from: The twenty amino acids are identified by unique numbers assigned to the uracil, cytosine, adenine, and guanine found in the three base positions of the sixty-four messenger RNA genetic codons

Open the record for dataset details and reuse information.

publicJan 2024View details →
zenodo28/100

Mutations in the initiation codon of homo sapiens and capra hirucs genome

<p>Dataset obtained by using the script <em>ensemblMining.pl</em> available in the dev branch of the following repository:&nbsp;&nbsp;<a href="https://github.com/fanavarro/hemodonacion">https://github.com/fanavarro/hemodonacion</a>. It containes several features that refer to mutations in the initiation codon of both human and goat genome. This dataset has been used for the following work:&nbsp;<a href="https://github.com/JavierCastellD/PredictorMutacionCodonInicio">https://github.com/JavierCastellD/PredictorMutacionCodonInicio</a>.</p>

opencc-by-4.0Jun 2020View details →
dryad28/100

Data from: Alternative translation initiation codons for the plastid maturase MatK: unraveling the pseudogene misconception in the Orchidaceae

Background: The plastid maturase MatK has been implicated as a possible model for the evolutionary "missing link" between prokaryotic and eukaryotic splicing machinery. This evolutionary implication has sparked investigations concerning the function of this unusual maturase. Intron targets of MatK activity suggest that this is an essential enzyme for plastid function. The matK gene, however, is described as a pseudogene in many photosynthetic orchid species due to presence of premature stop codons in translations, and its high rate of nucleotide and amino acid substitution. Results: Sequence analysis of the matK gene from orchids identified an out-of-frame alternative AUG initiation codon upstream from the consensus initiation codon used for translation in other angiosperms. We demonstrate translation from the alternative initiation codon generates a conserved MatK reading frame. We confirm that MatK protein is expressed and functions in sample orchids currently described as having a matK pseudogene using immunodetection and reverse-transcription methods. We demonstrate using phylogenetic analysis that this alternative initiation codon emerged de novo within the Orchidaceae, with several reversal events at the basal lineage and deep in orchid history. Conclusion: These findings suggest a novel evolutionary shift for expression of matK in the Orchidaceae and support the function of MatK as a group II intron maturase in the plastid genome of land plants including the orchids.

opencc-zeroDec 2014View details →
dryad28/100

Data from: Genomic analysis of codon usage shows influence of mutation pressure, natural selection, and host features on Marburg virus evolution

Background. The Marburg virus (MARV) has a negative-sense single-stranded RNA genome, belongs to the family Filoviridae, and is responsible for several outbreaks of highly fatal hemorrhagic fever. Codon usage patterns of viruses reflect a series of evolutionary changes that enable viruses to shape their survival rates and fitness toward the external environment and, most importantly, their hosts. To understand the evolution of MARV at the codon level, we report a comprehensive analysis of synonymous codon usage patterns in MARV genomes. Multiple codon analysis approaches and statistical methods were performed to determine overall codon usage patterns, biases in codon usage, and influence of various factors, including mutation pressure, natural selection, and its two hosts, Homo sapiens and Rousettus aegyptiacus. Results. Nucleotide composition and relative synonymous codon usage (RSCU) analysis revealed that MARV shows mutation bias and prefers U- and A-ended codons to code amino acids. Effective number of codons analysis indicated that overall codon usage among MARV genomes is slightly biased. The Parity Rule 2 plot analysis showed that GC and AU nucleotides were not used proportionally which accounts for the presence of natural selection. Codon usage patterns of MARV were also found to be influenced by its hosts. This indicates that MARV have evolved codon usage patterns that are specific to both of its hosts. Moreover, selection pressure from R. aegyptiacus on the MARV RSCU patterns was found to be dominant compared with that from H. sapiens. Overall, mutation pressure was found to be the most important and dominant force that shapes codon usage patterns in MARV. Conclusions. To our knowledge, this is the first detailed codon usage analysis of MARV and extends our understanding of the mechanisms that contribute to codon usage and evolution of MARV.

opencc-zeroDec 2014View details →
dryad28/100

Data from: Gene expression levels are correlated with synonymous codon usage, amino acid composition and gene architecture in the red flour beetle, Tribolium castaneum

Gene expression levels correlate with multiple aspects of gene sequence and gene structure in phylogenetically diverse taxa suggesting an important role of gene expression levels in the evolution of protein-coding genes. Here we present results of a genome-wide study of the influence of gene expression on synonymous codon usage, amino acid composition and gene structure in the red flour beetle, Tribolium castaneum. Consistent with the action of translational selection, we find that synonymous codon usage bias increases with gene expression. However, the correspondence between tRNA gene copy number and optimal codons is weak. At the amino acid level, translational selection is suggested by the positive correlation between tRNA gene numbers and amino acid usage which is stronger for highly expressed genes. In addition, there is a clear trend for increased use of metabolically cheaper, less complex, amino acids as gene expression increases. tRNA gene numbers also correlate negatively with amino acid size/complexity score indicating the coupling between translational selection and selection to minimize the use of large/complex amino acids. Interestingly, the correlation between tRNA gene numbers and amino acid size/complexity score appears to be widespread given our analyses of 10 additional genomes and might be explained by selection against negative consequences of protein misfolding. At the level of gene structure, three major trends are detected 1) CDS length increases across low and intermediate expression levels but decreases in highly expressed genes; 2) the average intron size shows the opposite trend, first decreasing with expression, followed by a slight increase in highly expressed genes and 3) intron density remains nearly constant across all expression levels. These changes in gene architecture are only in partial agreement with selection favoring reduced cost of biosynthesis.

opencc-zeroDec 2012View details →
dryad28/100

Data from: Antagonistic relationships between intron content and codon usage bias of genes in three mosquito species: functional and evolutionary implications

Genome biology of mosquitoes holds potential in developing knowledge-based control strategies against vector-borne diseases such as malaria, dengue, West Nile Virus and others. Although the genomes of three major vector mosquitoes have been sequenced, attempts to elucidate the relationship between intron and codon usage bias across species in phylogenetic contexts are limited. In this study, we investigated the relationship between intron content and codon bias of orthologous genes among three vector mosquito species. We found an antagonistic relationship between codon usage bias and the intron number of genes in each mosquito species. The pattern is further evident among the intronless and the intron-containing orthologous genes associated with either low or high codon bias among the three species. Furthermore, the co-variance between codon bias and intron number has a directional component associated with the species phylogeny when compared with other non-mosquito insects. By applying a maximum likelihood based continuous regression method, we show that codon bias and intron content of genes vary among the insects in a phylogeny dependent manner but with no evidence of adaptive radiation or species-specific adaptation. We discuss the functional and evolutionary significance of antagonistic relationships between intron content and codon bias.

opencc-zeroDec 2012View details →
dryad28/100

Data from: Translational selection frequently overcomes genetic drift in shaping synonymous codon usage patterns in vertebrates

Synonymous codon usage (SCU) patterns are shaped by a balance between mutation, drift, and natural selection. To date, detection of translational selection in vertebrates has proven to be a challenging task, obscured by small long-term effective population sizes in larger animals and the existence of isochores in some species. The consensus is that, in such species, natural selection is either completely ineffective at overcoming mutational pressures and genetic drift or perhaps is effective but so weak that it is not detectable. The aim of this research is to understand the interplay between mutation, selection, and genetic drift in vertebrates. We observe that although variation in mutational bias is undoubtedly the dominant force influencing codon usage, translational selection acts as a weak additional factor influencing synonymous codon usage. These observations indicate that translational selection is a widespread phenomenon in vertebrates and is not limited to a few species.

opencc-zeroDec 2012View details →
dryad28/100

Data from: The fitness landscape of the codon space across environments

Fitness landscapes map the relationship between genotypes and fitness. However, most fitness landscape studies ignore the genetic architecture imposed by the codon table and thereby neglect the potential role of synonymous mutations. To quantify the fitness effects of synonymous mutations and their potential impact on adaptation on a fitness landscape, we use a new software based on Bayesian Monte Carlo Markov Chain methods and re-estimate selection coefficients of all possible codon mutations across 9 amino-acid positions in Saccharomyces cerevisiae Hsp90 across 6 environments. We quantify the distribution of fitness effects of synonymous mutations and show that it is dominated by many mutations of small or no effect and few mutations of larger effect. We then compare the shape of the codon fitness landscape across amino-acid positions and environments, and quantify how the consideration of synonymous fitness effects changes the evolutionary dynamics on these fitness landscapes. Together these results highlight a possible role of synonymous mutations in adaptation and indicate the potential mis-inference when they are neglected in fitness landscape studies.

opencc-zeroDec 2017View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record