Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
904
datasets available to search
ShareScore release 0.9.0
Dataset results
904 results for “Biosynthesis”
Harnessing photoenzymatic reactions for unnatural biosynthesis in microorganisms
Open the record for dataset details and reuse information.
The weaken-fill-repair model for cell budding: Linking cell wall biosynthesis with mechanics
Open the record for dataset details and reuse information.
Ancient gene clusters govern the initiation of monoterpenoid indole alkaloid biosynthesis and C3 stereochemistry inversion
Open the record for dataset details and reuse information.
Inferred ancestry of scytonemin biosynthesis proteins in cyanobacteria indicates a response to Paleoproterozoic oxygenation
Open the record for dataset details and reuse information.
Ether lipid biosynthesis promotes lifespan extension and enables diverse prolongevity paradigms in Caenorhabditis elegans
Open the record for dataset details and reuse information.
Data from: Biosynthesis of the redox coenzyme F420 in Thermomicrobia involves reduction by standalone nitroreductase superfamily enzymes.
Open the record for dataset details and reuse information.
Compartmentalized sesquiterpenoid biosynthesis and functionalization in the Chlamydomonas reinhardtii plastid
Open the record for dataset details and reuse information.
Data from: Genome assembly of Chiococca alba uncovers key enzymes involved in the biosynthesis of unusual terpenoids
<p>Chiococca alba (L.) Hitchc. (snowberry), a member of the Rubiaceae, has been used as a folk remedy for a range of health issues including inflammation and rheumatism and produces a wealth of specialized metabolites including terpenes, alkaloids, and flavonoids. We generated a 558 Mb draft genome assembly for snowberry which encodes 28,707 high confidence genes. Comparative analyses with other angiosperm genomes revealed enrichment in snowberry of lineage-specific genes involved in specialized metabolism. Synteny between snowberry and Coffea canepehora Pierre ex A. Froehner (coffee) was evident, including the chromosomal region encoding caffeine biosynthesis in coffee, albeit syntelogs of N-methyltransferase were absent in snowberry. A total of 27 putative terpene synthase genes were identified, including 10 that encode diterpene synthases. Functional validation of a subset of putative terpene synthases revealed that combinations of diterpene synthases yielded access to products of both general and specialized metabolism. Specifically, we identified plausible intermediates in the biosynthesis of merilactone and ribenone, structurally unique antimicrobial diterpene natural products. Access to the C. alba genome will enable additional characterization of biosynthetic pathways responsible for health-promoting compounds in this medicinal species.</p>
Data from: Long read reference genome-free reconstruction of a full-length transcriptome from Astragalus membranaceus reveals transcript variants involved in bioactive compound biosynthesis
Astragalus membranaceus, also known as Huangqi in China, is one of the most widely used medicinal herbs in Traditional Chinese Medicine. Traditional Chinese Medicine formulations from Astragalus membranaceus have been used to treat a wide range of illnesses, such as cardiovascular disease, type 2 diabetes, nephritis and cancers. Pharmacological studies have shown that immunomodulating, anti-hyperglycemic, anti-inflammatory, antioxidant and antiviral activities exist in the extract of Astragalus membranaceus. Therefore, characterising the biosynthesis of bioactive compounds in Astragalus membranaceus, such as Astragalosides, Calycosin and Calycosin-7-O-β-D-glucoside, is of particular importance for further genetic studies of Astragalus membranaceus. In this study, we reconstructed the Astragalus membranaceus full-length transcriptomes from leaf and root tissues using PacBio Iso-Seq long reads. We identified 27 975 and 22 343 full-length unique transcript models in each tissue respectively. Compared with previous studies that used short read sequencing, our reconstructed transcripts are longer, and are more likely to be full-length and include numerous transcript variants. Moreover, we also re-characterised and identified potential transcript variants of genes involved in Astragalosides, Calycosin and Calycosin-7-O-β-D-glucoside biosynthesis. In conclusion, our study provides a practical pipeline to characterise the full-length transcriptome for species without a reference genome and a useful genomic resource for exploring the biosynthesis of active compounds in Astragalus membranaceus.
Data from: Nutritional control of body size through FoxO-Ultraspiracle mediated ecdysone biosynthesis
Despite their fundamental importance for body size regulation, the mechanisms that stop growth are poorly understood. In Drosophila melanogaster, growth ceases in response to a peak of the molting hormone ecdysone that coincides with a nutrition-dependent checkpoint, critical weight. Previous studies indicate that insulin/insulin-like growth factor signaling (IIS)/Target of Rapamycin (TOR) signaling in the prothoracic glands (PGs) regulates ecdysone biosynthesis and critical weight. Here we elucidate a mechanism through which this occurs. We show that Forkhead Box class O (FoxO), a negative regulator of IIS/TOR, directly interacts with Ultraspiracle (Usp), part of the ecdysone receptor. While overexpressing FoxO in the PGs delays ecdysone biosynthesis and critical weight, disrupting FoxO–Usp binding reduces these delays. Further, feeding ecdysone to larvae eliminates the effects of critical weight. Thus, nutrition controls ecdysone biosynthesis partially via FoxO–Usp prior to critical weight, ensuring that growth only stops once larvae have achieved a target nutritional status.
Fig. 5 in Limonoid biosynthesis 3: Functional characterization of crucial genes involved in neem limonoid biosynthesis
Fig. 5. Transient expression of AiTTS1 in neem twig leaves. A) Real-time qPCR analysis showing relative expression levels of AiTTS1 in pRI101-ANand pRI101- AN_AiTTS1 infiltrated leaves. Expression levels of these genes were normalized to actin and are represented incomparison with pRI101-AN control. Data represent means ±SE of two independent biological replicates. B) Limonoid profiling of neem transient leaves between control pRI101-AN and pRI101-AN-AiTTS1. Most of the neem limonoids production is increased as compared to vector control which clearly shows that AiTTS1 involved in limonoid biosynthesis.
Fig. 3 in Limonoid biosynthesis 3: Functional characterization of crucial genes involved in neem limonoid biosynthesis
Fig. 3. Phylogenetic analysis of neem triterpene synthases. Black colour represent the protosteryl cation, red colour representation for dammarenyl cation, green colour represents form luponyl cation, blue colour for multi-product forming and violet colour represents for germannyl cation stabilizing triterpene synthases. (For interpretation of the references to colour in this figure legend, the reader is referred to the Web version of this article.)
Fig. 2 in Limonoid biosynthesis 3: Functional characterization of crucial genes involved in neem limonoid biosynthesis
Fig. 2. Heat map for the RPKM values of transcripts which involved in isopreniod biosynthesis. Triterpenoid biosynthesis related genes were highly expressed in kernel and pericarp, which is in line with total triterpenoid profiling.
Fig. 4. A in Limonoid biosynthesis 3: Functional characterization of crucial genes involved in neem limonoid biosynthesis
Fig. 4. A) 1H and 13C NMR Chemical Shifts of tirucalla-7,24-dien-3β-ol. 13Cδ are shown in blue colour and 1Hδ are shown in black colour. B) RT-PCR analysis of triterpene synthases and squalene epoxidase in different tissues of neem. I) RT-PCR for AiTTS1 shows higher expression in kernel, II) RT-PCR for AiTTS2 shows higher expression in pericarp and kernel, III) RT-PCR for AiCAS shows higher expression in Kernel and flower and IV) RT-PCR for AiSQE1 shows higher expression in kernel and pericarp. (For interpretation of the references to colour in this figure legend, the reader is referred to the Web version of this article.)
Pigmentation biosynthesis influences the microbiome associated with sea urchins
<p>Organisms living on the seafloor are subject to encrustations by a wide variety of animal, plants, and microbes. Sea urchins, however, thwart this covering. Despite having a sophisticated immune system, there is no clear mechanism that allows sea urchins to remain clean. Here, by using CRISPR/Cas9, we test the hypothesis that pigmentation biosynthesis in sea urchin spines influences their interactions with microbes in vivo. We have three primary findings. First, the microbiome of sea urchin spines are species-specific and that much of this community is lost in captivity. Second, different color morphs associate with bacterial communities that are similar in composition and structure. Lastly, the gene activity of the pigmentation biosynthesis genes polyketide synthase and flavin-dependent mono-oxygenase induces a shift in which bacterial taxa colonize the tissues of sea urchin spines. We, therefore, find it plausible that host pigments are involved in host-microbe interactions and potentially in symbiotic homeostasis.</p>
Transcriptomic analysis reveals potential candidate pathways and genes involved in toxin biosynthesis in true toads
<p>Synthesized chemical defenses have broadly evolved across countless taxa and are important in 30 shaping evolutionary and ecological interactions within ecosystems. However, the underlying 31 genomic mechanisms by which these organisms synthesize and utilize their toxins are relatively 32 unknown. Herein, we use comparative transcriptomics to uncover potential toxin synthesizing 33 genes and pathways, as well as interspecific patterns of toxin synthesizing genes across ten 34 species of North American true toads (Bufonidae). Upon assembly and annotation of the ten 35 transcriptomes, we explored patterns of relative gene expression and possible protein-protein 36 interactions across the species to determine what genes and/or pathways may be responsible for 37 toxin synthesis. We also tested our transcriptome dataset for signatures of positive selection to 38 reveal how selection may be acting upon potential toxin producing genes. We assembled high 39 quality transcriptomes of the bufonid parotoid gland, a tissue not often investigated in other 40 bufonid related RNAseq studies. We found several genes involved in metabolic and biosynthetic 41 pathways (e.g. steroid biosynthesis, terpenoid backbone biosynthesis, isoquinoline biosynthesis, 42 glucosinolate biosynthesis) that were functionally enriched and/or relatively expressed across the 43 ten focal species that may be involved in the synthesis of alkaloid and steroid toxins, as well as 44 other small metabolic compounds that cause distastefulness in bufonids. We hope that our study 45 lays a foundation for future studies to explore the genomic underpinnings and specific pathways 46 of toxin synthesis in toads, as well as at the macroevolutionary scale across numerous taxa that 47 produce their own defensive toxins.</p>
Emergence and evolution of heterocyte glycolipid biosynthesis enabled specialized nitrogen fixation in cyanobacteria
<h1><strong>Abstract </strong></h1> <p>Paleontological and phylogenomic observations have shed light on the evolution of cyanobacteria. Nevertheless, the emergence of heterocytes, specialized cells for nitrogen fixation, remains unclear. Heterocytes are surrounded by heterocyte glycolipids (HGs), which contribute to protection of the nitrogenase enzyme from oxygen. Here, by comprehensive HG identification and screening of HG biosynthesis genes throughout cyanobacteria, we identify HG analogs produced by specific and distantly related non-heterocytous cyanobacteria. These structurally less complex molecules probably acted as precursors of HGs, suggesting that HGs arose after a genomic reorganization and expansion of ancestral biosynthetic machinery, enabling the rise of cyanobacterial heterocytes in an increasingly oxygenated atmosphere. Subsequently, HG chemical structure evolved convergently in response to environmental pressures. Our results open a new chapter in the potential use of diagenetic products of HGs and HG analogs as fossils for reconstructing the evolution of multicellularity and division of labor in cyanobacteria.</p> <h2><strong>Here we supply:</strong></h2> <div> <ul> <li><strong><span>Supplementary Data 1. Selected cyanobacterial genomes from the PATRIC genome database (now part of the BV-BRC database).</span></strong><span> Files called ‘selected_Cyanogenomes.genome_*.20220430.txt’ are sourced from the PATRIC File Transfer Protocol server (ftp.patricbrc.org). ‘gtdbtk.bac120.summary.tsv’ is the GTDB-Tk output file, and ‘qa.summary_extended.txt’ the CheckM output file.</span></li> <li><strong><span>Supplementary Data 2. HG biosynthetic gene clusters in selected PATRIC genomes and 14 newly sequenced genomes.</span></strong><span> The file ‘islands_on_contigs.3_ORFs_in_between.expanded_island_with_nucleotide_positions.txt’ contains the location of all hits to <em>Anabaena</em> sp. PCC 7120 HG biosynthesis genes. ORFs were predicted with Prodigal. The structure of a contig is as follows: “genome | contig”. The structure of a hit is as follows: “ORF number on contig | query (<em>e</em>-value; bit-score; start of alignment in query; end of alignment in query; query coverage per subject; start of alignment in subject; end of alignment in subject; subject coverage) [nucleotide position on contig start; nucleotide position on contig end; direction]”. Non-overlapping hits on the same ORF (see Online Methods) are connected with ‘&&&’ characters. An asterisk (‘*’) indicates that the hit is located at most three ORFs from a contig edge. Clusters of hits that are at most three open reading frames (ORFs) apart are connected with ‘~~~’ characters. The file ‘Supplementary_table.script_1.txt’ contains a summary of all identified <em>hgl</em> islands (i.e. clusters containing at least seven unique HG biosynthesis gene hits).</span></li> <li><strong><span>Supplementary Data 3. HG biosynthetic gene clusters in 255,388 prokaryotic genomes from the PATRIC genome database (now part of the BV-BRC database).</span></strong><span> The file called ‘PATRIC_20230120.selection_c50_c10.txt’ contains information on the selected PATRIC genomes based on data sourced from the PATRIC File Transfer Protocol server (<a>ftp.patricbrc.org</a>). The file ‘all_tree_of_life_genomes.islands_on_contigs.3_ORFs_in_between.expanded_island_with_nucleotide_positions.txt’ contains the location of all hits to <em>Anabaena</em> sp. PCC 7120 HG biosynthesis genes. ORFs were predicted with Prodigal. The structure of a contig is as follows: “genome | contig”. The structure of a hit is as follows: “ORF number on contig | query (<em>e</em>-value; bit-score; start of alignment in query; end of alignment in query; query coverage per subject; start of alignment in subject; end of alignment in subject; subject coverage) [nucleotide position on contig start; nucleotide position on contig end; direction]”. Non-overlapping hits on the same ORF (see Online Methods) are connected with ‘&&&’ characters. An asterisk (‘*’) indicates that the hit is located at most three ORFs from a contig edge. Clusters of hits that are at most three open reading frames (ORFs) apart are connected with ‘~~~’ characters.</span></li> <li><strong><span>Supplementary Data 4. Phylogeny of representative cyanobacterial genomes based on a core gene superalignment.</span></strong><span> The folder contains the files used to generate Fig. 2a. The directory ‘IQ-TREE’ contains the tree file and iTOL annotation files. The file ‘dRep.representative_to_cluster.txt’ contains the dRep clusters. Note that the manually defined subclades in the iTOL annotation file ‘iTOL_annotation.manually_defined_clades.DATASET_STYLE.txt’ have a different numbering from the paper: subclades 0 and 1 are the ‘heterocytous sister clades’, and subclades 2-10 in the annotation file are heterocytous subclades 1-9 in the paper, respectively.</span></li> <li><strong><span>Supplementary Data 5. Lipid data files.</span></strong><span> The folder contains all the UHPLC-HRMS<em><sup>n</sup></em> (Orbitrap) datafiles used in this study. The directory ‘CCY strains’ includes 24 heterocytous cyanobacterial cultures corresponding to 23 strains grown in nitrogen-deficient media, the resulting data are shown in Supplementary Table 10. Directory ‘HglT mutant’ contains the datafiles used to generate Supplementary Table 15. The directory ‘LEGE strains’ includes the UHPLC-HRMS<em><sup>n</sup></em> (Orbitrap) and GC-MS datafiles corresponding to eight cultures of two non-heterocytous strains grown in media with and without nitrogen for 38 to 77 days, the resulting data are shown in Supplementary Tables 10, 17 and 18.</span></li> <li><strong><span>Supplementary Data 6. Plasmid maps. </span></strong><span>GenBank and FASTA files of plasmids generated in this study. ‘HglT deletion’ directory contains the genomic region surrounding <em>hglT</em> in the <em>wild-type </em>strain and after deletion used to generate Supplementary Fig. 15. pAM5404 is shown in Supplementary Fig. 16 and p(A)RP0XX are shown in Supplementary Fig. 17.<span><span></span></span></span></li> <li><strong><span>Supplementary Data 7. Phylogenies of seven <em>hgl</em> island genes and of a concatenated alignment of these genes.</span></strong><span> The folder contains the files used to generate Supplementary Fig. 18 (in the directory ‘gene_trees_hgl_islands’), and Fig. 4 and related figures (in the directory ‘gene_trees_hgl_islands_4’. The directories contain the alignments and trimmed alignments, IQ-TREE output files, and iTOL annotation files. The file ‘gene_trees_hgl_islands/analysis_individual_gene_trees/explore_individual_gene_clusters.ipynb’ contains the code to identify the five <em>hgl</em> islands that contain genes with incongruent evolutionary histories.</span></li> <li><strong><span>Supplementary Data 8. Phylogeny of <em>hglE<sub>A</sub></em> homologs.</span></strong><span> The folder contains the files used to generate Supplementary Fig. 24 and related figures. The file ‘selected_hglE_hits.txt’ contains the selected <em>hglE<sub>A</sub></em> hits and the genomic cluster on which they are located. The folder contains the alignment and trimmed alignment, IQ-TREE output files, and iTOL annotation files.</span></li> <li><strong><span>All the code used in this publication </span></strong><span>including scripts used for: genome assemblies, download of genomes from public repositories, quality and contamination checks, genome analysis, construction of the phylogenetic trees, <em>hgl</em> island identification, etc. The shell script 'commands.sh' within each directory contains all the code used to generate the content in the directory.</span></li> <li><strong><span>All the figures used in this publication </span></strong><span>including the figures in the Supplementary Information file.</span></li> </ul> <div> <div> <div></div> </div> </div> </div>
Interactions between methylerythritol phosphate (MEP) pathway metabolites and Escherichia coli K-12 fatty acid biosynthesis enzymes
<p>Raw data files from native mass spectrometry analyses examining the interactions between methylerythritol phosphate (MEP) pathway metabolites and Escherichia coli K-12 fatty acid biosynthesis enzymes. Recorded in positive mode direct injection.</p>
Fig. 6 in Gene identification using RNA-seq in two sweetpotato genotypes and the use of mining to analyze carotenoid biosynthesis
Fig. 6. qRT-PCR analysis of the selected genes associated with carotenoid and terpenoid backbone biosynthesis pathways. Roots, stems and leaves were collected, and the total RNAs was extracted, respectively, pooled and used in qRT-PCR analysis of the genes FDPS, PSY, HMGCR, HMGCS1, CRTISO and HDR. Relative expression levels were calculated using the 2−ΔΔCt method, with actin as an internal control.
Fig. 4 in Gene identification using RNA-seq in two sweetpotato genotypes and the use of mining to analyze carotenoid biosynthesis
Fig. 4. Distribution and frequency of the EST-SSRs in the sweetpotato transcriptome. The y-axis shows the SSR motif, and the x-axis shows the number of the motif.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.