Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
23
datasets available to search
ShareScore release 0.9.0
Dataset results
23 results for “biosynthetic gene cluster”
The Minimum Information about a Biosynthetic Gene Cluster (MIBiG) data repository
<p>This dataset was originally published alongside the Minimum Information about a Biosynthetic Gene Cluster (MIBiG) data standard publication(s).</p> <p>It contains JSON files following the MIBiG data standard. Additional information on proteins/genes associated to biosynthetic gene clusters described by MIBiG can be found in the GenBank (gbk) and fasta files.</p> <p>This dataset was uploaded with permission from the corresponding author(s).</p> <p>For more information, see https://mibig.secondarymetabolites.org/.</p>
Uncovering a miltiradiene biosynthetic gene cluster in the Lamiaceae reveals a dynamic evolutionary trajectory
<p><span>The spatial organization of genes within plant genomes can drive evolution of specialized metabolic pathways. In this study we investigated the origin and subsequent evolution of a diterpenoid biosynthetic gene cluster (BGC) present throughout the Lamiaceae (mint) family. Terpenoids are important specialized metabolites in plants with </span><span>diverse</span><span> adaptive functions that enable environmental interactions, such as chemical defense. Based on core genes found in the BGCs of all species examined across the Lamiaceae, we predict a simplified version of this cluster evolved in an early Lamiaceae ancestor. The current composition of the extant BGCs highlights the dynamic nature of its evolution. We elucidate the terpene backbones made by the </span><span>Callicarpa americana</span><span> BGC enzymes, including miltiradiene and the novel terpene (+)-kaurene, and show oxidization activities of BGC cytochrome P450s. Our work reveals the fluid nature of BGC assembly and the importance of genome structure in contributing to the origin of novel metabolites.</span></p>
Streptomyces Biosynthetic Gene Clusters
<p>A collection of genbank files 11975 biosynthetic gene clusters from generated using antiSMASH v6.0.1</p>
Uncovering a miltiradiene biosynthetic gene cluster in the Lamiaceae reveals a dynamic evolutionary trajectory
Open the record for dataset details and reuse information.
Figures for: Biosynthetic gene clusters carried by plasmids may enhance adaptations to a changing ocean
<p>Figures created for the short communication: Biosynthetic gene clusters carried by plasmids may enhance adaptations to a changing ocean.</p>
Data for: Tomato root specialized metabolites evolved through gene duplication and regulatory divergence within a biosynthetic gene cluster
<p>Tremendous plant metabolic diversity arises from phylogenetically-restricted specialized metabolic pathways. Specialized metabolites are synthesized in dedicated cells or tissues, with pathway genes sometimes colocalizing in biosynthetic gene clusters (BGCs). However, the mechanisms by which spatial expression patterns arise and the role of BGCs in pathway evolution remain underappreciated. In this study, we investigated the mechanisms driving acylsugar evolution in the Solanaceae. Previously thought to be restricted to glandular trichomes, acyl sugars were recently discovered in cultivated tomato roots. We demonstrated that acyl sugars in cultivated tomato roots and trichomes have different sugar cores, identified root-enriched paralogs of trichome acyl sugar pathway genes, and characterized a key paralog required for root acyl sugar biosynthesis, <em>SlASAT1-LIKE</em> (<em>SlASAT1-L</em>), which is nested within a previously-reported trichome acyl sugar BGC. Finally, we provided evidence that <em>ASAT1-L</em> arose through duplication of its paralog, <em>ASAT1</em>, and was trichome-expressed before acquiring root-specific expression in the <em>Solanum</em> genus. Our results illuminate the genomic context and molecular mechanisms underpinning metabolic diversity in plants.</p>
Datasets for "Micromonosporaceae Biosynthetic Gene Cluster Diversity Highlights the Need for Broad Spectrum Investigation"
<p>In this data collection is:<br><strong>Data S1</strong>: A folder with all the fasta files, representing the 42 strains (41 <em>Micromonosporaceae</em>, 1 <em>Streptomycetaceae</em>).<br><strong>Data S2</strong>: A folder with all the .gbk files for the BGC regions predicted by antiSMASH v5.1.1. These files were used as inputs for BiG-SCAPE and BiG-SLiCE.<br><strong>Data S3</strong>: A folder with all the .gbk files for the BGC regions predicted by antiSMASH v6.1.0.<br><strong>Data S4</strong>: A folder containing all the Quast outputs for the 42 strains.<br><strong>Data S5</strong>: A folder containing all the BUSCO outputs for the 42 strains. Example scripts are provided for scraping relevant information from the individual BUSCO outputs.<br><strong>Data S6</strong>: A folder containing GTDB (Genome Taxonomy Database) classification results, and species-level grouping results using FastANI (95% cutoff).<br><strong>Data S7</strong>: A folder containing an Interactive Tree of Life (iTOL)-compatible bar chart annotation using antiSMASH v5.1.1 BGC region information.<br><strong>Data S8</strong>: A folder containing a word document that describes the parameters used with Ubuntu WSL (Windows Subsystem for Linux) on the command line for programs antiSMASH v6.1.2, BiG-SCAPE v1.1.2, and BiG-SLiCE v1.1.1. Also included are parameters for MDSC in python. An example script is also provided for batch queries of BGCs against BiG-SLiCE v1.1.1’s pre-processed dataset of ~1.2 million BGCs.<br><strong>Data S9</strong>: A folder containing the BiG-SCAPE visualization of the 38 <em>Micromonosporaceae</em> (post-QC filtering, excluding WMMA1363, WMMB482, WMMB486, and WMMC500) in Cytoscape.<br><strong>Data S10</strong>: A folder containing:<br>The pre-processed dataset of 1.2 million BGCs from BiG-SLiCE.<br>All report folders generated by BiG-SLiCE for the 779 <em>Micromonosporaceae </em>BGCs queried against the 1.2 million BGCs.<br>The results data.db and associated folders for the pre-processed dataset of 1.2 million BGCs.<br><strong>Data S11</strong>: A folder containing the scripts necessary to regenerate the figures and perform independent analyses, and the relevant data used for the analyses.<br><strong>Data S12: </strong>A folder containing the results of the nucleotide blast of WMMA1947.region12's siderophore contig against WMMD1120.region14's siderophore contig.</p> <p><strong>Supplementary Information: </strong>Supplementary Table S1 and Supplementary Figures S1-S181.</p>
IsoAnalyst: A System-Wide Stable Isotopic Labeling Aproach for Connecting Natural Products to Their Cognate Biosynthetic Gene Clusters
<p>Processed mass spectrometry data input used to develop and validate the IsoAnalyst program and the output from these analyses. The S_erythraea folder contains all MS input and output data for <em>Saccharopolyspora erythraea</em>. The Micromonospora_RL09050HVFA folder contains MS input and output data for <em>Micromonopora sp.</em> RL09050HVFA, as well as the full antiSMASH output for the <em>Micromonopora sp.</em> RL09050HVFA genome.</p> <p>Version 2 contains updated IsoAnalyst results for both organisms. </p>
Supporting data for the manuscript "Nerpa: a tool for discovering biosynthetic gene clusters of nonribosomal peptides"
<p>Preprocessed structures of nonribosomal peptides [NRPs] and genomic sequences (reference and representative genomes, biosynthetic gene clusters [BGCs]) used in the benchmark experiments in the Nerpa paper.</p> <p><strong>Files description</strong></p> <ul> <li><strong>mibig_nrp_bacteria_preprocessed.tar.gz</strong> contains the preprocessed dataset of 194 bacterial NRP BGCs from the MIBiG database.</li> <li><strong>mibig_nrp_bacteria_summary.tsv</strong> contains metadata for the MIBiG-NRP dataset.</li> <li><strong>bacterial_ref_and_repr_genomes_20210604_preprocessed.tar.gz</strong> contains the preprocessed dataset of 13,399 reference and representative bacterial genomes from the NCBI RefSeq database (retrieved on 2021/06/04).</li> <li><strong>bacterial_ref_and_repr_genomes_20210604_summary.txt</strong> contains metadata for the RefSeq dataset.</li> <li><strong>pnrpdb_preprocessed.info</strong> contains the Nerpa-preprocessed pNRPdb database, a database of 8,368 known and putative NRP structures.</li> <li><strong>pnrpdb_summary.tsv</strong> contains the pNRPdb database metadata.<br> </li> </ul>
Phylogenomics of the psychoactive mushroom genus Psilocybe and evolution of the psilocybin biosynthetic gene cluster
<p>Psychoactive mushrooms in the genus <em>Psilocybe</em> have immense cultural value and have been used for centuries in Mesoamerica. Despite a recent surge in interest in these mushrooms due to emerging evidence that psilocybin, the main psychoactive compound, is a promising therapeutic for a variety of mental illnesses, their phylogeny and taxonomy remain substantially incomplete. Moreover, the recent elucidation of the psilocybin biosynthetic gene cluster is known for only five species of <em>Psilocybe</em>, four of which belong to only one of two major clades. We set out to improve the phylogeny for <em>Psilocybe</em> using shotgun sequencing of 71 fungarium specimens, including 23 types, and conducting phylogenomic analysis using 2,983 single-copy gene families to generate a fully supported phylogeny. Molecular clock analysis suggests the stem lineage arose ~67 mya and diversified ~56 mya. We also show that psilocybin biosynthesis first arose in <em>Psilocybe</em>, with 4–5 possible horizontal transfers to other mushrooms between 40 and 9mya. Moreover, predicted orthologs of the psilocybin biosynthetic genes revealed two distinct gene orders within the cluster that corresponds to a deep split within the genus, possibly consistent with the independent acquisition of the cluster. By mining genomic data beyond markers for phylogenetic inference, we gained novel insights into the evolutionary origins of psilocybin biosynthesis that have implications for understanding the functional role of this powerful chemical and can inform translational applications for human well-being.</p>
Data from: A high-quality genome assembly of the tetraploid Teucrium chamaedrys unveils a recent whole genome duplication and a large biosynthetic gene cluster for diterpenoid metabolism
Open the record for dataset details and reuse information.
Data for: Tomato root specialized metabolites evolved through gene duplication and regulatory divergence within a biosynthetic gene cluster
Open the record for dataset details and reuse information.
Metagenomic bins and biosynthetic gene clusters in gut bacteria of turtle ants
Open the record for dataset details and reuse information.
Phylogenomics of the psychoactive mushroom genus Psilocybe and evolution of the psilocybin biosynthetic gene cluster
Open the record for dataset details and reuse information.
Genome sequences of Rhizopogon roseolus, Mariannaea elegans, Myrothecium verrucaria, and Sphaerostilbella broomeana and the identification of biosynthetic gene clusters for fungal peptide natural products
<p>Data to accompany the paper, detailed in manifest.txt</p>
Sub-functionalization and epigenetic regulation of a biosynthetic gene cluster in Solanaceae
<p>Datasets for the code analysis in <em>Priego-Cubero, Knoch et al., 2024 </em>(https://doi.org/10.1101/2024.10.02.615186)</p>
Diversity of biosynthetic gene clusters in gut bacteria of turtle ants
<p class="Standard"><span>In insect-microbe nutritional symbioses the gut symbionts supplement the host diet with nutrients by producing amino acids and vitamins or by degrading lignin or polysaccharides macromolecules. In multipartite mutualisms composed of multiple symbionts from different taxonomical orders, it has been suggested that beside the genes involved in the nutritional symbiosis the symbionts maintain genes responsible for the production of metabolites putatively playing a role in the maintenance and interaction of the bacterial communities living in close proximity. To test this hypothesis, we investigated the diversity of biosynthetic gene clusters (BGCs) producing different non-primary metabolites in the genomes and metagenomes of the conserved gut symbionts associated with the herbivorous turtle ants (genus: <i>Cephalotes</i>). We studied 17 <i>Cephalotes</i> species collected across several geographical areas to reveal that (i) mining metagenomes and genomes show complementary results demonstrating the robustness of this approach to retrieve BGCs, (ii) the conserved gut symbionts involved in the nutritional symbiosis have a high diversity of BGCs of different chemical families, (iii) the phylogenetic analysis of BGCs encoding the production of arylpolyenes, non-ribosomal peptides (NRP), polyketides (PK), and siderophores shows high similarity between BGCs of a single symbiont across different ant host species, and between BGCs originated from different bacterial orders within a single host species. These findings together suggest multiple mechanisms of bacterial genome conservation and evolution of BGCs. We document the occurrence, diversity, and similarity of BGCs in the genome of obligate gut bacteria involved in multipartite mutualisms across the phylogeny of turtle ants.</span></p>
Biosynthetic Gene Cluster Synteny - Orthologous Polyketide Synthases in Hypogymnia physodes, Hypogymnia tubulosa and Parmelia sulcata
<p>Supplementary Material of Publication</p>
Diversity of biosynthetic gene clusters in gut bacteria of turtle ants
Open the record for dataset details and reuse information.
Multilevel regulation of gene expression in brasilicardin biosynthetic gene cluster
GEO Series GSE271981. Nocardia terpenica. 8 samples. Type: Other.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.