Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
1,344
datasets available to search
ShareScore release 0.9.0
Dataset results
1,344 results for “ribosome”
Figure 3 in Phylogenetic structure of the Sphaeriinae, a global clade of freshwater bivalve molluscs, inferred from nuclear (ITS-1) and mitochondrial (16S) ribosomal gene sequences
Figure 3. The single most-parsimonious tree (L = 951; CI = 0.568; RI = 0.793) obtained from the maximum parsimony analysis of combined (16S + ITS1) sequence dataset. Maximum likelihood analysis produced a largely congruent topology (HKY model; Ln likelihood = - 7034.61154) with the only difference being Pisidium dubium sister to Sphaerium/Musculium clade. Taxonomic names are arranged according to suggested sphaeriinid taxonomy in the present study and five major monophyletic lineages are indicated. Two Eupera species, E. cubensis and E. platensis, were designated as outgroups. MP bootstrap values are shown to the left of the slash and decay index values to the right above the branches. Numbers below the branches indicate ML bootstrap values.
Figure 1 in Phylogenetic structure of the Sphaeriinae, a global clade of freshwater bivalve molluscs, inferred from nuclear (ITS-1) and mitochondrial (16S) ribosomal gene sequences
Figure 1. Strict consensus of the four equally most parsimonious trees (L = 526; CI = 0.447; RI = 0.743) obtained from the phylogenetic analysis of sphaeriid mitochondrial 16S rDNA sequences. Two Eupera species, E. cubensis and E. platensis, were designated as outgroups and inferred sequence gaps were considered as missing data. Numbers above the branches represent bootstrap values and numbers below indicate decay index values.
mTAGs: taxonomic profiling using degenerate consensus reference sequences of ribosomal RNA gene
<p>mTAGs is a tool for the taxonomic profiling of metagenomes. It detects sequencing reads belonging to the small subunit of the ribosomal RNA (SSU-rRNA) gene and annotates them through the alignment to full-length degenerate consensus SSU-rRNA reference sequences. The tool is capable of processing single-end and pair-end metagenomic reads, takes advantage of the information contained in any region of the SSU-rRNA gene and provides relative abundance profiles at multiple taxonomic ranks (Domain, Phylum, Class, Order, Family, Genus and OTUs defined at a 97% sequence identity cutoff).</p>
Data from: Feasibility of nuclear ribosomal region ITS1 over ITS2 in barcoding taxonomically challenging genera of subtribe Cassiinae (Fabaceae)
Premise of the Study The internal transcribed spacer (ITS) region is situated between 18S and 26S in a polycistronic rRNA precursor transcript. It had been proved to be the most commonly sequenced region across plant species to resolve phylogenetic relationships ranging from shallow to deep taxonomic levels. Despite several taxonomical revisions in Cassiinae, a stable phylogeny remains elusive at the molecular level, particularly concerning the delineation of species in the genera Cassia, Senna and Chamaecrista. This study addresses the comparative potential of ITS datasets (ITS1, ITS2 and concatenated) in resolving the underlying morphological disparity in the highly complex genera, to assess their discriminatory power as potential barcode candidates in Cassiinae. Methodology A combination of experimental data and an in-silico approach based on threshold genetic distances, sequence similarity based and hierarchical tree-based methods was performed to decipher the discriminating power of ITS datasets on 18 different species of Cassiinae complex. Lab-generated sequences were compared against those available in the GenBank using BLAST and were aligned through MUSCLE 3.8.31 and analysed in PAUP 4.0 and BEAST1.8 using parsimony ratchet, maximum likelihood and Bayesian inference (BI) methods of gene and species tree reconciliation with bootstrapping. DNA barcoding gap was realized based on the Kimura two-parameter distance model (K2P) in TaxonDNA and MEGA. Principal Findings Based on the K2P distance, significant divergences between the inter- and intra-specific genetic distances were observed, while the presence of a DNA barcoding gap was obvious. The ITS1 region efficiently identified 81.63% and 90% of species using TaxonDNA and BI methods, respectively. The PWG-distance method based on simple pairwise matching indicated the significance of ITS1 whereby highest number of variable (210) and informative sites (206) were obtained. The BI tree-based methods outperformed the similarity-based methods producing well-resolved phylogenetic trees with many nodes well supported by bootstrap analyses. Conclusion The reticulated phylogenetic hypothesis using the ITS1 region mainly supported the relationship between the species of Cassiinae established by traditional morphological methods. The ITS1 region showed a higher discrimination power and desirable characteristics as compared to ITS2 and ITS1 + 2, thereby concluding to be the locus of choice. Considering the complexity of the group and the underlying biological ambiguities, the results presented here are encouraging for developing DNA barcoding as a useful tool for resolving taxonomical challenges in corroboration with morphological framework.
The nucleotides absent in genes of SARS-CoV-2 non-canonical subgenomic RNAs generate new Programmed -1 Ribosomal Frameshifting
<p>The data correspond to the article entitled: "dNTPs and adjuvant reagent solutions in 3’ RACE improve the characterization of noncanonical RNA SARS-CoV-2 genomes"</p> <p>R1. RACE 3’ Primer Blast Alignment. Contains BLAST alignments against the GenBank database using the consensus nucleotide sequence from the 3’ end of the SARS-CoV-2 genome and the polylinker. In addition, an illustration of the restriction enzyme pattern of the 3' RACE primer RV30AkCOVID19 and its synthesis by MALDI-TOF is included. The red box indicates the nucleotide sequence of the polylinker and the yellow box represents the 3' RACE primer along with the result of primer synthesis and purification.</p> <p>Graphic representation of the procedure for SARS-CoV-2 genome cDNA synthesis and design of the 3’ RACE RV30AkCOVID19 primer. The rectangle with vertical lines and the dots represents the 3’ RACE RV30AkCOVID19 primer and the polylinker, respectively, in the region complementary to the 3’ UTR end. The arrow represents the reverse transcriptase during complementary strand synthesis. The scissors represent RNases used in purification. The black spheres and magnets indicate the purification process using magnetism.</p> <p>R2. Reads and assembles SARS-CoV-2 genomes.</p> <p>The folder "1) Reads - Ion torrent" contains the reads obtained from sequencing via Ion Torrent technology and the reagents used in this study.</p> <p>The folder named "2) FastQC" contains the results of Ion Torrent sequencing. In the file name, the number indicates the sample, and the letters "RNA" indicate the sequencing according to the IonTorrent protocol. The cDNA synthesis procedures for this study correspond to the following nomenclature: dNTPs-R = dNTPs SARS-CoV-2 solution, DES-R = denaturation reagent, and COM PRO = commercial procedure.</p> <p>The folders named "3) IRMA" and "4) Bowtie2" contain the assemblies of the genomes.</p> <p>Regions and/or codons with loss of genomes 07dN120320 and 27sT122620.</p> <p>Mutations and amino acid substitutions of the SARS-CoV-2 genomes.</p> <p>In addition, an Excel document with the nucleotide ratios of each characterized genome is included from SARS-CoV-2.</p> <p>R3. BLAST alignment of assembled SARS-CoV-2 genomes. Contains two folders named "BLAST - IRMA" and "BLAST - Bowtie2," which contain plain text documents with the results of the BLAST alignment for the genomes obtained with each of the assemblies.</p> <p>R4. Pangolin v1.16 and Nextclade v2.9.1 lineages for SARS-CoV-2 genomes. Contains the folders "Pangolin and Nextclade (Bowtie2)" and "Pangolin and Nextclade (IRMA)." Each folder shows the data obtained with the Pangolin v1.16 and Nextclade v2.9.1 software for the classification of the genomes reported in this study, which were assembled with the IRMA and Bowtie2 software.</p> <p>R5. Reference genome alignment and assembled genomes. Contains the folders "1) IRMA genomes," "2) Bowtie2 genomes," and "3) Genomes 07dN120320 and 27St122620." The files show the sequences and alignments of the examined genomes (the file name indicates the analyzed genome) relative to the SARS-CoV-2 reference genome both in FASTA and Clustal W formats.</p> <p>R6. Programmed −1 Ribosomal Frameshifting Structure. The folder "1) Gibbs free energy 2D" contains a plain text document indicating the secondary structures of the open reading frame stimulation element in dot-bracket format. The folder "2) modeling Data Modeling 3D" contains the information for generating the structure of folder 1 in 3D.</p> <p>R7. SARS-CoV-2 Database.</p> <p>1) GISAID_sequences.zip contains a Zip file that contains a folder named GISAID, which in turn contains plain text documents with the genomes of each variant indicated in the filename of each document.</p> <p>2) The depuration of sequences_GISAID contains two subfolders. The first subfolder, named "1) SARS-CoV-2 complete genome" contains plain text documents with the genomes downloaded from GISAID without undetermined nucleotides. The file name of each document corresponds to the analyzed variant. The subfolder "2) SARS-CoV-2 eliminate genome" contains the sequences eliminated from subfolder 1 because they differed from the majority of the analyzed sequences.</p> <p>3) SARS-CoV-2 consensus variants. Contains plain text documents with consensus sequences for each variant, with frequency thresholds of 20 and 100 indicated in the file name of each document.</p> <p>4) SARS-CoV-2 alignment consensus variants. Contains two subfolders, with the number indicating the alignment frequency threshold. The "Alignment 20_" subfolder contains four documents named "with Ns," which correspond to fasta and Clustal formats with undetermined nucleotides, whereas the files named "without" do not have undetermined nucleotides. The "100_" folder has the same file pattern as the previous folder.</p> <p>5) SARS-CoV-2 codons alignment consensus variants and nc-sgRNA. Contains a document with the alignment of the genomes characterized in this study with the reference genome of SARS-CoV-2. A subfolder named “SARS-CoV-2 codons nc-sgRNA” shows each of the nc-sgRNA obtained in this study with the reference genome, and the file name corresponds to the nc-sgRNAs. The subfolder “SARS-CoV-2 Geneious Prime” contains 4 documents. Each document includes the graphical representation of the alignment of the nc-sgRNA obtained with each treatment for the synthesis of SARS-CoV-2 cDNA with respect to the reference genome. The following three documents indicated with the numbers 25, 50, and 100 correspond to the percentage of identity with respect to the number of annotations relative to the reference genome, which is indicated in the title of each document.</p> <p>6) Variant Alignment – Ns. Contains eight documents corresponding to the fasta and clustal formats with SARS-CoV-2 genomes obtained in this study from the reference genome and from genomes containing undetermined nucleotides of the Gamma, Lambda, Mu and Omicron variants.</p> <p>R8. Phylogeny SARS-CoV-2. Contains two subfolders with the results of the phylogenetic analyses conducted via the maximum likelihood method of the genomes characterized in this study compared to the variants. The subfolder named "Phylogeny with Ns" indicates the analysis of genomes containing undetermined nucleotides, whereas "Phylogeny without Ns" corresponds to the analysis of complete genomes.</p>
The alignments of chloroplast genome sequences and nuclear ribosomal DNA fragments of six oak species sampled in the hot-dry valley of the Jinsha River, southwestern China
<p>Both chloroplast (cp) genome sequences and nuclear ribosomal (nr) DNA were assembled using GetOrganelle v.1.7.6.1 for 18 oak trees sampled in the Panzhihua Cycad National Nature Reserve, Sichuan Province, China. These trees belong to six oak species, including Quercus cocciferoides, Q. dolicholepis, Q. franchetii, Q. griffithii, Q. longispica, and Q. variabilis. We used PhyloSuite v.1.1.152 to extract coding sequences (CDSs), tRNA genes, rRNA genes, introns, and intergenic spacers (IGSs) of the 18 oak cp genomes. These sequences were aligned separately using MAFFT v.7.3.13 and manually adjusted with BioEdit v.7.2.5. Length variations in mononucleotide repeats were excluded and inversions were replaced with their reverse complements because of their tendency for homoplasy. Other indels were coded as binary characters according to the simple gap coding method using GapCoder. Separate assignments were concatenated according to their respective positions in the cp genome to obtain the alignments of LSC, SSC, IRb, and the whole cp genome.</p>
Scripts and tomoDRGN models used in rapid processing of E.coli ribosomes by sub-tomogram averaging
<p>This repository contains scripts and tomoDRGN model analyses described in our preprint: <br> Rapid structural analysis of bacterial ribosomes in situ<br> Barrett M. Powell, Tyler S. Brant, Joseph H. Davis, Shyamal Mosalaganti<br> bioRxiv 2024.03.22.586148; doi: https://doi.org/10.1101/2024.03.22.586148 </p>
E. coli ribosome phenotypes for antimicrobial susceptibility testing
<p>This repository contains the anonymised images, masks, and metadata for "<strong>Ribosome Phenotypes Enable Rapid Antibiotic Susceptibility Testing in Escherichia coli". </strong>Code for training neural networks using these images can be found at <a href="https://github.com/KapanidisLab/ribosome_phenotype_classification">https://github.com/KapanidisLab/ribosome_phenotype_classification</a>.</p> <p>The data contains fluorescence microscopy images of <em>E. coli</em> MG1655 and clinical isolates, details of which can be found in the manuscript. </p>
Substrate recognition and cryo-EM structure of the ribosome-bound TAC toxin of Mycobacterium tuberculosis
<p>Datasets for the Figures 2 and S2 of the manuscript "Substrate recognition and cryo-EM structure of the ribosome-bound TAC toxin of Mycobacterium tuberculosis".</p> <p> </p> <p>The HTML files describe the analysis and the raw counts after nEMOTE-conv treatment.</p> <p>There are 5 files for each MMEMOTExx dataset:</p> <p>EmoteBarcodesReport.csv = summary<br> UnambNegTable.csv = counts of unique cuts on the reverse strand<br> UnambPosTable.csv = counts of unique cuts on the forward strand<br> AmbPosTable.csv = counts of all cuts on the forward strand<br> AmbNegTable.csv = counts of all cuts on the reverse strand</p> <p> </p>
An in vivo stable isotope labeling method to investigate individual matrix protein synthesis, ribosomal biogenesis, and chondrocyte proliferation in murine articular cartilage
<p>These experiments have used a stable-isotope method using <em>in vivo</em> deuterium oxide labeling and mass spectrometry to measure protein concentration, protein half-life, cell proliferation, and ribosomal biogenesis in a single sample of murine articular cartilage. We hypothesized that a 60-day labeling period would capture age-related declines in cartilage matrix protein content, protein synthesis rates, and chondrocyte proliferation. Knee cartilage was isolated from 25- and 90-week-old female C57BL/6J mice treated with deuterium oxide for 15, 30, 45 and 60 days. We measured protein abundance and half-lives using high resolution accurate mass spectrometry (HRAM) and d2ome data processing software. </p>
Supplementary material 1 from: Scacchetti P, Pansonato-Alves J, Utsunomia R, Oliveira C, Foresti F (2011) Karyotypic diversity in four species of the genus Gymnotus Linnaeus, 1758 (Teleostei, Gymnotiformes, Gymnotidae): physical mapping of ribosomal genes and telomeric sequences. Comparative Cytogenetics 5(3): 223-235. https://doi.org/10.3897/compcytogen.v5i3.1375
Nexus file of aligned COI and COII nucleotide sequences.
Nuc/Mito Ribosomal Ratio as an index of growth rate
<p>The dataset (NMRR.csv) consists of information regarding <em>Daphnia magna</em> in Nuc/Mito ribosomal ratio experiment. Information included treatments, ribosomal ratio, somatic growth rate, size, culture duration, rRNA counts and total RNA concentration. Total 166 samples recorded in this data set. </p> <p>R script file (NMRR.R) consists of R code for all the statistical analysis conducted in this study.</p>
Capturing single-copy nuclear genes, organellar genomes, and nuclear ribosomal DNA from deep genome skimming data for plant phylogenetics: A case study in Vitaceae
<p>With the decreasing cost and availability of many newly developed bioinformatics pipelines, next-generation sequencing (NGS) has revolutionized plant systematics in recent years. Genome skimming has been widely used to obtain high-copy fractions of the genomes, including plastomes, mitochondrial DNA (mtDNA), and nuclear ribosomal DNA (nrDNA). In this study, through simulations, we evaluated the optimal (minimum) sequencing depth and performance for recovering single-copy nuclear genes (SCNs) from genome skimming data, by subsampling genome resequencing data and generating 10 datasets with different sequencing coverage <i>in silico</i>. We tested the performance of four datasets (plastome, nrDNA, mtDNA, and SCNs) obtained from genome skimming based on phylogenetic analyses of the <i>Vitis</i> clade at the genus level and Vitaceae at the family level, respectively. Our results showed that optimal minimum sequencing depth for high-quality SCNs assembly via genome skimming was about 10× coverage. Without the steps of synthesizing baits and enrichment experiments, coupled with incredibly low sequencing costs, we showcase that deep genome skimming (DGS) is as effective for capturing large datasets of SCNs as the widely used Hyb-Seq approach, in addition to capturing plastomes, mtDNA, and entire nrDNA repeats. DGS may serve as an efficient and economical alternative and may be superior to the popular target enrichment/Hyb-Seq approach.</p>
Fig. 4 in Some Unusual Small-Subunit Ribosomal RNA Sequences of Metazoans
Fig. 4. Variable region (V7) of the 18S rRNA locus of 17 species of centipedes.
Supplemental information and Data for: Colloidal physics modeling reveals how per-ribosome productivity increases with growth rate in E. coli
<p>Faster growing cells must synthesize proteins more quickly. Increased ribosome abundance only partly accounts for increases in total protein synthesis rates. The productivity of individual ribosomes must increase too, almost doubling by an unknown mechanism. Prior models point to diffusive transport as a limiting factor but surface a paradox: faster growing cells are more crowded, yet crowding slows diffusion. We suspected physical crowding, transport, and stoichiometry, considered together, might reveal a more nuanced explanation. To investigate, we built a first-principles physics-based model of <em>E. coli</em> cytoplasm in which Brownian motion and diffusion arise directly from physical interactions between individual molecules of finite size, density, and physiological abundance. Using our microscopically-detailed model, we predict that physical transport of individual ternary complexes accounts for ~80% of translation elongation latency. We also find that volumetric crowding increases at faster growth even as cytoplasmic mass density remains relatively constant. Despite slowed diffusion, we predict that improved proximity between ternary complexes and ribosomes wins out, illustrating a simple physics-based mechanism for how individual elongating ribosomes become more productive. We speculate how crowding imposes a physical limit on growth rate and undergirds cellular behavior more broadly. Unfitted colloidal-scale modeling offers systems biology a complementary "physics engine" for exploring how cellular-scale behaviors arise from physical transport and reactions among individual molecules.</p>
Comparison between ribosomal assembly and machine learning tools for microbial identification of organisms with different characteristics
<p><strong>DNABERT+DeLUCS_notebooks.zip </strong></p> <ul> <li>Code notebooks for running DNABERT and DeLUCS</li> </ul> <p> </p> <p><strong>images-20230519T015235Z-001.zip </strong></p> <ul> <li>Heatmaps</li> <li>Factor plots</li> </ul> <p> </p> <p><strong>Data-20230518T191203Z-003.zip </strong></p> <ul> <li>MBARC and Hot Springs datasets <ul> <li>Reference genomes</li> <li>16S sequences (barrnap)</li> </ul> </li> </ul> <p> </p> <p><strong>hot-springs-reads.gz </strong></p> <ul> <li>Reads data for Hot Springs dataset</li> </ul> <p> </p> <p><strong>mbarc-reads-download.txt</strong></p> <ul> <li>Reads data for MBARC dataset <ul> <li>Link to download from NCBI</li> </ul> </li> </ul> <p> </p> <p><strong>assemblies.zip</strong></p> <ul> <li>Megahit and MetaSPAdes assemblies for both MBARC and Hot Springs</li> </ul>
KsgA facilitates ribosomal small subunit maturation by proofreading a key structural lesion
<p>Weights, z files, poses, indices, CTF info, and particles for cryoDRGN training on KsgA-treated and untreated datasets in "KsgA facilitates ribosomal small subunit maturation by proofreading a key structural lesion"</p>
New insights into infrageneric relationships of Lonicera (Caprifoliaceae) as revealed by nuclear ribosomal DNA cistron data and plastid phylogenomics
<p>The discontinuous geographic distribution pattern of plants in the north temperate zone has been a focus of biogeographic research, especially concerning the mechanisms behind the formation of such a pattern and the spatial and temporal evolution of this intermittent distribution pattern. Hypotheses of boreotropical origin, land bridge migration, and out-of-Tibet have been proposed to explain the formation of the discontinuous distribution pattern. The distribution of <em>Lonicera</em> shows a typical Europe-Asia-North America discontinuous distribution, which makes for a good case study to investigate the above three hypotheses. In this study, we inferred the phylogeny based on plastid genomes and a nuclear data set with broad taxon sampling, covering 83 species representing two subgenera and four sections. Both nuclear and plastid phylogenetic analyses found section <em>Isika</em> polyphyletic, while sections <em>Nintooa</em>, <em>Isoxylosteum</em>, and <em>Coelxylosteum</em> were monophyletic in subgenus <em>Chamaecerasus</em>. Based on the nuclear and chloroplast phylogeny, we suggest transferring L. <em>maximowiczii</em> and L. <em>tangutica</em> into section <em>Nintooa</em>. Reconstruction of ancestral areas suggests that <em>Lonicera</em> originated in the Qinghai-Tibetan Plateau (QTP) and/or Asia, and subsequently dispersed to other regions. The aridification of the Asian interior may have facilitated the rapid radiation of <em>Lonicera</em> in the region. At the same time, the uplifts of the Tibetan Plateau appear to have triggered the spread and recent rapid diversification of the genus on the QTP and adjacent areas. Overall, our results deepen the understanding of the evolutionary diversification history of <em>Lonicera</em>.</p>
Data from: Spatial and temporal distribution of ribosomes in single cells reveals aging differences between old and new daughters of Escherichia coli
Open the record for dataset details and reuse information.
Data from: Feasibility of nuclear ribosomal region ITS1 over ITS2 in barcoding taxonomically challenging genera of subtribe Cassiinae (Fabaceae)
Open the record for dataset details and reuse information.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.