Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
167
datasets available to search
ShareScore release 0.9.0
Dataset results
167 results for “coding sequences”
Data and code for, "Predicting self-assembly of sequence-controlled copoly- mers with stochastic sequence variation"
Open the record for dataset details and reuse information.
FIGURE 3 in Sansevieria (Asparagaceae, Nolinoideae) is a herbaceous clade within Dracaena: inference from non-coding plastid and nuclear DNA sequence data
FIGURE 3. Bayesian maximum clade reliability trees based on combined nuclear At103 and chloroplast rps16, trnL-F datasets for Dracaena, Sansevieria, and selected outgroups. The values above the branch represent the maximum parsimony bootstrap percentage (BS), and the ones below are the Bayesian posterior probability (PP). Bold branches indicate strong support, interpreted as ≥ 70 BS and ≥ 95 PP. Long branches were shortened by half their length (indicated by \\).
FIGURE 2 in Sansevieria (Asparagaceae, Nolinoideae) is a herbaceous clade within Dracaena: inference from non-coding plastid and nuclear DNA sequence data
FIGURE 2. Bayesian maximum clade credibility trees based on nuclear At103 (A) and chloroplast rps16, trnL-F (B) datasets for Dracaena and Sansevieria. Outgroups were trimmed from the Figure. The values above the branch represent the maximum parsimony bootstrap percentage (BS), and the ones below are the Bayesian posterior probability (PP). Bold branches indicate strong support, interpreted as ≥ 70 BS and ≥ 95 PP.
FIGURE 1 in Sansevieria (Asparagaceae, Nolinoideae) is a herbaceous clade within Dracaena: inference from non-coding plastid and nuclear DNA sequence data
FIGURE 1. Representative morphological diversity in the dracaenoid genera, Dracaena and Sansevieria. A, Dracaena draco subsp. draco, Spain, Canary Islands, Tenerife, Icod de los Vinos; B, D. konaensis, origin: USA, Hawai'i, Big Island, Kona coast, in cultivation at Kew (Acc. No. 2008-239); C, D. arborea, Gabon, Woleu-Ntem Rd, Mitzic to Njole; D, D. laxissima, São Tomé and Príncipe, São Nicolau; E, D. goldieana, origin: Gabon, in cultivation at Kew (Acc. No. 1990-2300); F, D. aubryana, Gabon, Woleu-Ntem Rd Mitzic to Njole; G, Sansevieria frequens, Kenya, Laikipia District, Ngare Ndare Farm (type locality); H, S. aethiopica, Namibia, 74 km from Windhoek, on road to Walvis Bay; I, S. fischeri, Kenya, Munda, 18.9 km NE of Mwatate on Taveta road; J, S. pinguicula, Kenya, by Kowi airstrip, north bank of Tiva Lugga; K, S. ascendens, Kenya, Coast Province, Kwale District, around base of Taru Hill (type locality); L, S. kirkii var. pulchra, in cultivation (private collection, Miami, FL). Photographs by A, L. Mucina; B, I. Willey; C, E–F, T.H.J. Damen; D, J.J.F.E. de Wilde; G-K, L. E. Newton; L, S. Zona.
16s rRNA sequences, R code used for amplicon analysis and example code for NMGS analysis
<p>This submission contains the following data presented in: "Selection processes of Arctic seasonal glacier snowpack bacterial communities" by Keuschnig et al.</p> <p>the R code used to analyze the 16S rRNA amplicon data</p> <p>the script used for NMGS analysis</p> <p>the sequences obtained from snow samples</p>
Bayesian phylodynamic and phylogeographic analyses of invasive, hypervirulent Streptococcus agalactiae sequence type 283 dataset and R code
<p>Supplementary dataset S1, R code, and subsampled trees</p>
Data from: Targeted re-sequencing of coding DNA sequences for SNP discovery in non-model species
Open the record for dataset details and reuse information.
Data from: A hypervariable mitochondrial protein coding sequence associated with geographical origin in a cosmopolitan bloom-forming alga, Heterosigma akashiwo
Open the record for dataset details and reuse information.
Data from: The importance of being genomic: non-coding and coding sequences suggest different models of toxin multi-gene family evolution
Open the record for dataset details and reuse information.
Data from: Deconstruction of archaeal genome depict strategic consensus in core pathways coding sequence assembly
A comprehensive in silico analysis of 71 species representing the different taxonomic classes and physiological genre of the domain Archaea was performed. These organisms differed in their physiological attributes, particularly oxygen tolerance and energy metabolism. We explored the diversity and similarity in the codon usage pattern in the genes and genomes of these organisms, emphasizing on their core cellular pathways. Our thrust was to figure out whether there is any underlying similarity in the design of core pathways within these organisms. Analyses of codon utilization pattern, construction of hierarchical linear models of codon usage, expression pattern and codon pair preference pointed to the fact that, in the archaea there is a trend towards biased use of synonymous codons in the core cellular pathways and the Nc-plots appeared to display the physiological variations present within the different species. Our analyses revealed that aerobic species of archaea possessed a larger degree of freedom in regulating expression levels than could be accounted for by codon usage bias alone. This feature might be a consequence of their enhanced metabolic activities as a result of their adaptation to the relatively O2-rich environment. Species of archaea, which are related from the taxonomical viewpoint, were found to have striking similarities in their ORF structuring pattern. In the anaerobic species of archaea, codon bias was found to be a major determinant of gene expression. We have also detected a significant difference in the codon pair usage pattern between the whole genome and the genes related to vital cellular pathways, and it was not only species-specific but pathway specific too. This hints towards the structuring of ORFs with better decoding accuracy during translation. Finally, a codon-pathway interaction in shaping the codon design of pathways was observed where the transcription pathway exhibited a significantly different coding frequency signature.
Data from: Successful recovery of nuclear protein-coding genes from small insects in museums using illumina sequencing
In this paper we explore high-throughput Illumina sequencing of nuclear protein-coding, ribosomal, and mitochondrial genes in small, dried insects stored in natural history collections. We sequenced one tenebrionid beetle and 12 carabid beetles ranging in size from 3.7 to 9.7 mm in length that have been stored in various museums for 4 to 84 years. Although we chose a number of old, small specimens for which we expected low sequence recovery, we successfully recovered at least some low-copy nuclear protein-coding genes from all specimens. For example, in one 56-year-old beetle, 4.4 mm in length, our de novo assembly recovered about 63% of approximately 41,900 nucleotides in a target suite of 67 nuclear protein-coding gene fragments, and 70% using a reference-based assembly. Even in the least successfully sequenced carabid specimen, reference-based assembly yielded fragments that were at least 50% of the target length for 34 of 67 nuclear protein-coding gene fragments. Exploration of alternative references for reference-based assembly revealed few signs of bias created by the reference. For all specimens we recovered almost complete copies of ribosomal and mitochondrial genes. We verified the general accuracy of the sequences through comparisons with sequences obtained from PCR and Sanger sequencing, including of conspecific, fresh specimens, and through phylogenetic analysis that tested the placement of sequences in predicted regions. A few possible inaccuracies in the sequences were detected, but these rarely affected the phylogenetic placement of the samples. Although our sample sizes are low, an exploratory regression study suggests that the dominant factor in predicting success at recovering nuclear protein-coding genes is a high number of Illumina reads, with success at PCR of COI and killing by immersion in ethanol being secondary factors; in analyses of only high-read samples, the primary significant explanatory variable was body length, with small beetles being more successfully sequenced.
Learning the sequence code for protein abundance in human immune cells
<p><span>Protein abundance is defined by </span><span>transcriptional, post-transcriptional and post-translational regulatory mechanisms. Understanding the code for gene expression could inform novel therapies. Here, we developed a machine learning pipeline, termed SONAR, to decipher the endogenous sequence code that defines the abundance of protein in human cells. SONAR predicts up to 63% of protein abundance independently of promoter or enhancer information. Our analysis reveals a strong - yet dynamic - </span><span>cell-type specific sequence code<span>. The deep knowledge of SONAR provides a map of biologically active sequence features (SFs), which we leveraged to </span></span><span>manipulate protein expression and tailored to a specific cell-type. </span><span>Beyond providing fundamental insights in gene expression regulation, our study offers novel means to improve therapeutic and biotechnology applications.</span></p> <p><span>Datasets of the protein models (with gamma=1) are included in this repository. Please refer to GitHub for more details on how the models were generated. </span></p>
Complementary evolution of coding and noncoding sequence underlies mammalian hairlessness
<p><span>Body hair is a defining mammalian characteristic, but several mammals, such as whales, naked mole-rats, and humans, have notably less hair than others. To find the genetic basis of reduced hair quantity, we used our evolutionary-rates-based method, RERconverge, to identify coding and noncoding sequences that evolve at significantly different rates in so-called hairless mammals compared to hairy mammals. Using RERconverge, we performed an unbiased, genome-wide scan over 62 mammal species using 19,149 genes and 343,598 conserved noncoding regions to find genetic elements that evolve at significantly different rates in hairless mammals compared to hairy mammals. We show that these rate shifts resulted from relaxation of evolutionary constraint on hair-related sequences in hairless species. In addition to detecting known and potential novel hair-related genes, we also discovered hundreds of putative hair-related regulatory elements. Computational investigation revealed that genes and their associated noncoding regions show different evolutionary patterns and influence different aspects of hair growth and development. Many genes under accelerated evolution are associated with the structure of the hair shaft itself, while evolutionary rate shifts in noncoding regions also included the dermal papilla and matrix regions of the hair follicle that contribute to hair growth and cycling. Genes that were top-ranked for coding sequence acceleration included known hair and skin genes KRT2, KRT35, PKP1, and PTPRM that surprisingly showed no signals of evolutionary rate shifts in nearby noncoding regions. Conversely, accelerated noncoding regions are most strongly enriched near regulatory hair-related genes and microRNAs, such as mir205, ELF3, and FOXC1, that themselves do not show rate shifts in their protein-coding sequences. Such dichotomy highlights the interplay between the evolution of protein sequence and regulatory sequence to contribute to the emergence of a convergent phenotype.</span></p>
Dataset and code for Fast sequence alignment for centromere with RaMA
<p>Dataset and code for Fast sequence alignment for centromere with RaMA</p>
Mass cytometry and integration sequencing data and code from "Quantification of intrinsic regulatory factors refines human hematopoietic progenitor definitions and reveals early erythroid lineage priming"
Open the record for dataset details and reuse information.
Dastset and code for Fast sequence alignment for centromere with RaMA
<p>Dastset and code for Fast sequence alignment for centromere with RaMA</p>
Data and code for, "Large language models design sequence-defined macromolecules via evolutionary optimization"
<div> <pre># Codes and data for "Large language models design sequence-defined macromolecules via evolutionary optimization"<br><br>Note this repository contains codes and data files for the manuscript. This is a snapshot of the repository, frozen at the time of submission.<br><br># Codes<br><br>## LLM codes<br>- `run_claude.py` - the routine for performing LLM-based rollouts; intended for command line execution using argparse<br>- `message_utils.py` - utilities for constructing and parsing messages for LLM I/O<br>- `model_utils.py` - lightweight utilities for retrieving formatted predictions from the RNN ensemble<br>- `target_defs.py` - defines the sequence, locations, and natural language descriptions of the target structures<br>- `ask_about_oracle.ipynb` - asks the LLM to speculate about the nature of the optimization task<br><br>## other algorithms<br>- `active_learning.ipynb` - use EI acquisition with RF surrogate to label new sequences; includes an unused tokenization scheme<br>- `evolutionary_algorithm.ipynb` - use DEAP library to perform evolutionary optimization<br>- `random_sampling.ipynb` - sample sequences randomly from all possible sequences<br><br>## postprocessing<br>- `process_aggregated_logs.py` - reads data from the raw log files and prepares them for visualization<br>- `process_sample_rollouts.py` - reads data from the raw log files and prepares individual rollouts<br><br>## visualization<br>- `figure1b.ipynb` - renders panel b of Fig. 1<br>- `figure1efg.ipynb` - renders the last row of Fig. 1 (panels e-g)<br>- `figure2.ipynb` - renders all of Fig. 2<br>- `figure_si.ipynb` - renders Figs. S1 and S2<br>- `figure_md_validation.ipynb` - renders Fig. S3<br><br># Data files<br><br>- `prompts/`<br> - `prompt-scientific-v4.4.yml` - the full text of the scientific prompt, to be read by `run_claude.py`<br> - `prompt-oracle-v4.4.yml` - the full text of the oracle prompt, to be read by `run_claude.py`<br>- `models/` - the TorchScript RNN models used to make predictions<br>- `data/`<br> - `embeddings` - calculated embeddings for a collection of sequences from our prior work<br> - `llm-logs` - the raw logs obtained from the Claude 3.5 Sonnet LLM (other algorithms made to look like the LLM logs after the fact)<br> - `llm-logs-opus` - the raw logs obtained from the Claude 3.0 Opus LLM (used in the first draft of the article, replaced by Claude 3.5 Sonnet) <br> - `all-rollouts-kltd.csv` - postprocessed logs for all the rollouts using the "top $k < d^*$" metric<br> - `all-rollouts-topkd.csv` - postprocessed logs for all the rollouts using the "mean $d$ for top $k$" metric<br> - `sample-rollout-membranes-x-3.csv` - postprocessed logs for a single rollout replica, `x` = each algorithm type<br> - `snapshots` - png snapshots of MD simulation results at different locations in the manifold</pre> </div>
Data from: Timing transantarctic disjunctions in the Atherospermataceae (Laurales): evidence from coding and noncoding chloroplast sequences
Open the record for dataset details and reuse information.
Data from: Widespread position-specific conservation of synonymous rare codons within coding sequences
Open the record for dataset details and reuse information.
Complementary evolution of coding and noncoding sequence underlies mammalian hairlessness
Open the record for dataset details and reuse information.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.