Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
231
datasets available to search
ShareScore release 0.9.0
Dataset results
231 results for “codon”
TAILVAR (Terminal codon Analysis and Improved prediction of Lengthened VARiants)
<p>This dataset includes relevant files for developing the TAILVAR score designed to assess the functional impact of <strong>stop-loss variants</strong> occurring at stop codons (TAA, TGA, TAG). <strong>TAILVAR</strong> is built using a Random Forest model that predicts the pathogenicity of <strong>stop-loss variants</strong>. By integrating a combination of in-silico prediction scores, transcript features, and protein context information, <strong>TAILVAR</strong> provides a score ranging from 0 to 1, indicating the probability of a variant being pathogenic.</p> <p>For more information, please visit <a href="https://github.com/dr-yoon/TAILVAR">https://github.com/dr-yoon/TAILVAR</a></p>
Supplementary data for Willemsen et al., 2024 "Novel high-quality amoeba genomes reveal widespread codon usage mismatch between giant viruses and their hosts".
<p>Supplementary data for Willemsen et al., 2024 "Novel high-quality amoeba genomes reveal widespread codon usage mismatch between giant viruses and their hosts". The data set consists of five folders: “Codon_usage_amoebae_and_viruses”, "Genome_annotations_amoebae", "Phylogenetic_trees_18S_amoebae", “Phylogenomic_trees_amoebae”, and "Viral_integration_detection_amoebae". The “Codon_usage_amoebae_and_viruses” folder contains for each amoeba host the calculated codon usage tables in the subfolder "codon_usage_table_host", the calculated codon usage preferences using different scores in the subfolder "codon_usage_scores_host", and the calculated codon usage preferences of giant viruses versus each host in the subfolder "codon_usage_scores_viruses_vs_host". The giant viruses in the subfolder "codon_usage_scores_viruses_vs_host" are organised by viral family and genus in separate sub-subfolders. The "Genome_annotations_amoebae" folder contains the generated genome annotations in different formats and the manually curated mitochondrial genome annotations for each amoeba host. The "Phylogenetic_trees_18S_amoebae" contains for the eukaryotic phyla <em>Discosea</em>, <em>Heterolobosea</em>, and <em>Tubulinea, </em>the 18S rRNA nucleotide alignments, distance matrices, and computed phylogenetic trees. The folder "Phylogenomic_trees_amoebae" contains for the eukaryotic clades <em>Amoebozoa</em> and <em>Discoba, </em>the protein alignment matrices and computed phylogenomic trees. The folder "Viral_integration_detection_amoebae" contains the MCP databases used (fasta file, alignment file, HMM profile and DIAMOND BLASTX database) and the MCP sequences detected in this study and the blast results of these. </p>
Codon similarity data in ATTED-II ver 8.0 (Bra, Mtr)
<p>Codon similarity data in ATTED-II ver 8.0</p> <p>The gene-to-gene codon similarity data is organized in the form of tables, each named according to the Entrez Gene ID of a particular query gene. Each table encompasses three columns, specifying: the Entrez Gene ID of a corresponding gene, an MR (Mutual Rank) value (where a smaller number signifies a stronger relationship), and a Pearson correlation coefficient (where a larger number suggests a stronger association).</p> <p>Protein-coding sequences utilized in this study were retrieved from NCBI's RefSeq database. For each gene, a 61-dimensional vector was derived from the count of codons in the protein-coding sequence. In instances where multiple RefSeq sequences were associated with a single gene, the longest sequence was selected for the codon usage calculation. Pearson correlation coefficients (PCCs) were determined between the vectors of any two given genes. These PCCs were subsequently converted into MRs, employed as an index to evaluate the similarity in codon usage between the genes.</p>
Codon similarity data in ATTED-II ver 8.0 (Ath, Gma, Osa, Sly, Vvi)
<p>Codon similarity data in ATTED-II ver 8.0</p> <p>The gene-to-gene codon similarity data is organized in the form of tables, each named according to the Entrez Gene ID of a particular query gene. Each table encompasses three columns, specifying: the Entrez Gene ID of a corresponding gene, an MR (Mutual Rank) value (where a smaller number signifies a stronger relationship), and a Pearson correlation coefficient (where a larger number suggests a stronger association).</p> <p>Protein-coding sequences utilized in this study were retrieved from NCBI's RefSeq database. For each gene, a 61-dimensional vector was derived from the count of codons in the protein-coding sequence. In instances where multiple RefSeq sequences were associated with a single gene, the longest sequence was selected for the codon usage calculation. Pearson correlation coefficients (PCCs) were determined between the vectors of any two given genes. These PCCs were subsequently converted into MRs, employed as an index to evaluate the similarity in codon usage between the genes.</p> <p> </p>
Modeling and measuring how codon usage modulates the relationship between burden and yield during protein overexpression in bacteria
<p>Additional data from experiments associated with paper revisions.</p> <p>Also added codon optimizer script.</p>
Annex 2: Supplementary Tables to 'Understanding Chinese hamster translation at sub-codon resolution'
<p>Annex 2 to PhD thesis titled 'Understanding Chinese hamster ovary cell translation at sub-codon resolution'. </p>
Vibrio cholerae codon biased genes list
<p>Files called top25: For each codon, 25% of genes with highest differential codon usage are listed. </p> <p>Files called top: For each codon, top genes with highest differential codon usage are listed, as indicated by violin plots in pdf files. </p> <p>Excel sheet with assigned codon usage for each codon and each gene.</p> <p>Whole genome of Vibrio cholerae N16961</p>
datset related to article "THE NOVEL I213S MUTATION IN PSEN1 GENE IS LOCATED IN A HOTSPOT CODON ASSOCIATED WITH FAMILIAL EARLY-ONSET ALZHEIMER'S DISEASE"
<p><strong>Electropherogram of the proband psen1 exon 7</strong></p> <p><strong>ngs analysis of causal and risk genes associated to dementia</strong></p>
Codon similarity data in ATTED-II ver 8.0 (Ptr, Zma)
<p>Codon similarity data in ATTED-II ver 8.0</p> <p>The gene-to-gene codon similarity data is organized in the form of tables, each named according to the Entrez Gene ID of a particular query gene. Each table encompasses three columns, specifying: the Entrez Gene ID of a corresponding gene, an MR (Mutual Rank) value (where a smaller number signifies a stronger relationship), and a Pearson correlation coefficient (where a larger number suggests a stronger association).</p> <p>Protein-coding sequences utilized in this study were retrieved from NCBI's RefSeq database. For each gene, a 61-dimensional vector was derived from the count of codons in the protein-coding sequence. In instances where multiple RefSeq sequences were associated with a single gene, the longest sequence was selected for the codon usage calculation. Pearson correlation coefficients (PCCs) were determined between the vectors of any two given genes. These PCCs were subsequently converted into MRs, employed as an index to evaluate the similarity in codon usage between the genes.</p>
Transcriptome-wide meta-analysis of codon usage in Escherichia coli
<p>Data generated by the CUBseq pipeline on Escherichia coli RNA-seq data.</p>
A codon model for associating phenotypic traits with altered selective patterns of sequence evolution
<p>Detecting the signature of selection in coding sequences and associating it with shifts in phenotypic states can unveil genes underlying complex traits. Of the various signatures of selection exhibited at the molecular level, changes in the pattern of selection at protein coding genes have been of main interest. To this end, phylogenetic branch-site codon models are routinely applied to detect changes in selective patterns along specific branches of the phylogeny. Many of these methods rely on a pre-specified partition of the phylogeny to branch categories, thus treating the course of trait evolution as fully resolved and assuming that phenotypic transitions have occurred only at speciation events. Here we present TraitRELAX, a new phylogenetic model that alleviates these strong assumptions by explicitly accounting for the uncertainty in the evolution of both trait and coding sequences. This joint statistical framework enables the detection of changes in selection intensity upon repeated trait transitions. We evaluated the performance of TraitRELAX using simulations and then applied it to two case studies. Using TraitRELAX, we found an intensification of selection in the primate SEMG2 gene in polygynandrous species compared to species of other mating forms, as well as changes in the intensity of purifying selection operating on sixteen bacterial genes upon transitioning from a free-living to an endosymbiotic lifestyle.</p>
Stop-codon recoding in bacteriophages may regulate translation of lytic genes
<p><strong>This has some basic datasets for bacteriophages that use alternative genetic codes, and their close standard code relatives. </strong></p> <p>I have included the following:</p> <p>- Genomes for all alternatively coded phages and their relatives</p> <p>- Predicted proteins all alternatively coded phages and their relatives</p> <p>- A sheet with some basic information about these phages</p> <p>- Terminase treefile (Figure 2A) </p> <p>- Genomes for crAss-like phages used in the alternative code bias analysis (Figure 4)</p> <p>- Genomes for Agate phages used in the alternative code bias analysis (Figure 4) as well as the ANI analysis (Figure 3A)</p> <p>- Untrimmed lysogenic contigs for prophages (Like those shown in Figure 5)</p> <p> </p>
Data for Bernabeu-Herrero et al, Mutations causing premature termination codons discriminate and generate cellular and clinical variability in HHT
<p>This dataset is for the 2024 manuscript<strong>: </strong></p> <p><strong>Bernabéu-Herrero ME, Patel D, Bielowka A, Zhu J, Jain K, Mackay IS, Chaves Guerrero P, Emanuelli G, Jovine L, Noseda M, Marciniak SJ, Aldred MA, Shovlin CL. </strong></p> <p><strong>Mutations causing premature termination codons discriminate and generate cellular and clinical variability in HHT. </strong></p> <p><strong>Blood. 2024 May 30;143(22):2314-2331. </strong></p> <p><strong>doi: 10.1182/blood.2023021777. PMID: 38457357; PMCID: PMC11181359.</strong></p> <p>It was originally uploaded in 2021 ahead of an earlier manuscript submission<br>- see https://www.biorxiv.org/content/10.1101/2021.12.05.471269v1</p>
Determinants of associations between codon and amino acid usage patterns of microbial communities and the environment inferred based on a cross-biome metagenomic analysis
<p>Raw data set for npj Bioflims and Microbiome article: “Determinants of associations between codon and amino acid usage patterns of microbial communities and the environment inferred based on a cross-biome metagenomic analysis”</p>
Data from: Antibody production relies on the tRNA inosine wobble modification to meet biased codon demand
Open the record for dataset details and reuse information.
A codon model for associating phenotypic traits with altered selective patterns of sequence evolution
Open the record for dataset details and reuse information.
Mitogenomes reveal alternative initiation codons and lineage-specific gene order conservation in echinoderms
<p>Sample information, gene alignments and concatenated matrix</p>
Data from: Mitochondrial gene diversity associated with the atp9 stop codon in natural populations of wild carrot (Daucus carota ssp. carota)
Mitochondrial genomes extracted from wild populations of Daucus carota have been used as a genetic resource by breeders of cultivated carrot, yet little is known concerning the extent of their diversity in nature. Of special interest is a SNP in the putative stop codon of the mitochondrial gene atp9 that has been associated previously with male-sterile and male-fertile phenotypic variants. In this study either sequence or PCR/RFLP genotypes were obtained from the mitochondrial genes atp1, atp9 and cox1 found in D. carota individuals collected from 24 populations in the eastern U.S. More than half of the 128 individuals surveyed had a CAA or AAA, rather than TAA, genotype at the position usually thought to function as an atp9 stop codon in this species. We also found no evidence for mitochondrial RNA editing (Cytosine to Uridine) of the CAA stop codon in either floral or leaf tissue. Evidence for intra-genic recombination, as opposed the more common inter-genic recombination in plant mitochondrial genomes, in our data set is presented. Indel and SNP variants elsewhere in atp9, and in the other two genes surveyed, were non-randomly associated with the three atp9 stop codon variants, though further analysis suggested that multi-locus genotypic diversity had been enhanced by recombination. Overall the mitochondrial genetic diversity was only modestly structured among populations with an Fst of 0.34.
Data from: A phenotype-genotype codon model for detecting adaptive evolution
A central objective in biology is to link adaptive evolution in a gene to structural and/or functional phenotypic novelties. Yet most analytic methods make inferences mainly from either phenotypic data or genetic data alone. A small number of models have been developed to infer correlations between the rate of molecular evolution and changes in a discrete or continuous life history trait. But such correlations are not necessarily evidence of adaptation. Here we present a novel approach called the phenotype-genotype branch-site model (PG-BSM) designed to detect evidence of adaptive codon evolution associated with discrete-state phenotype evolution. An episode of adaptation is inferred under standard codon substitution models when there is evidence of positive selection in the form of an elevation in the nonsynonymous-to-synonymous rate ratio ω to a value ω > 1. As it is becoming increasingly clear that ω > 1 can occur without adaptation, the PG-BSM was formulated to infer an instance of adaptive evolution without appealing to evidence of positive selection. The null model makes use of a covarion-like component to account for general heterotachy (i.e., random changes in the evolutionary rate at a site over time). The alternative model employs samples of the phenotypic evolutionary history to test for phenomenological patterns of heterotachy consistent with specific mechanisms of molecular adaptation. These include (i) a persistent increase/decrease in ω at a site following a change in phenotype (the pattern) consistent with an increase/decrease in the functional importance of the site (the mechanism); and (ii) a transient increase in ω at a site along a branch over which the phenotype changed (the pattern) consistent with a change in the site's optimal amino acid (the mechanism). Rejection of the null is followed by post hoc analyses to identify sites with strongest evidence for adaptation in association with changes in the phenotype as well as the most likely evolutionary history of the phenotype. Simulation studies based on a novel method for generating mechanistically realistic signatures of molecular adaptation show that the PG-BSM has good statistical properties. Analyses of three real alignments show that site patterns identified post hoc are consistent with the specific mechanisms of adaptation included in the alternate model. Further simulation studies show that the covarion-like component of the PG-BSM plays a crucial role in mitigating recently discovered statistical pathologies associated with confounding by accounting for heterotachy-by-any-means.
Distributions of alternative start codons for all microbial Refseq genomes
<p>Distributions of alternative start codons for all microbial Refseq genomes</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.