Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

23

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

23 results for “biosynthetic gene cluster”

Learn how ShareScore rates datasets ↗
zenodo44/100

The Minimum Information about a Biosynthetic Gene Cluster (MIBiG) data repository

<p>This dataset was originally published alongside the Minimum Information about a Biosynthetic Gene Cluster (MIBiG) data standard publication(s).</p> <p>It contains JSON files following the MIBiG data standard. Additional information on proteins/genes associated to biosynthetic gene clusters described by MIBiG can be found in the GenBank (gbk) and fasta files.</p> <p>This dataset was uploaded with permission from the corresponding author(s).</p> <p>For more information, see https://mibig.secondarymetabolites.org/.</p>

opencc-by-4.0Jun 2015View details →
dryad40/100

Uncovering a miltiradiene biosynthetic gene cluster in the Lamiaceae reveals a dynamic evolutionary trajectory

<p><span>The spatial organization of genes within plant genomes can drive evolution of specialized metabolic pathways. In this study we investigated the origin and subsequent evolution of a diterpenoid biosynthetic gene cluster (BGC) present throughout the Lamiaceae (mint) family. Terpenoids are important specialized metabolites in plants with </span><span>diverse</span><span> adaptive functions that enable environmental interactions, such as chemical defense. Based on core genes found in the BGCs of all species examined across the Lamiaceae, we predict a simplified version of this cluster evolved in an early Lamiaceae ancestor. The current composition of the extant BGCs highlights the dynamic nature of its evolution. We elucidate the terpene backbones made by the </span><span>Callicarpa americana</span><span> BGC enzymes, including miltiradiene and the novel terpene (+)-kaurene, and show oxidization activities of BGC cytochrome P450s. Our work reveals the fluid nature of BGC assembly and the importance of genome structure in contributing to the origin of novel metabolites.</span></p>

opencc-zeroMay 2022View details →
zenodo40/100

Streptomyces Biosynthetic Gene Clusters

<p>A collection of genbank files 11975 biosynthetic gene clusters from generated using antiSMASH v6.0.1</p>

opencc-by-4.0Feb 2023View details →
dryad40/100

Uncovering a miltiradiene biosynthetic gene cluster in the Lamiaceae reveals a dynamic evolutionary trajectory

Open the record for dataset details and reuse information.

publicJan 2023View details →
zenodo36/100

Figures for: Biosynthetic gene clusters carried by plasmids may enhance adaptations to a changing ocean

<p>Figures created for the short communication: Biosynthetic gene clusters carried by plasmids may enhance adaptations to a changing ocean.</p>

opencc-by-4.0Nov 2023View details →
dryad36/100

Data for: Tomato root specialized metabolites evolved through gene duplication and regulatory divergence within a biosynthetic gene cluster

<p>Tremendous plant metabolic diversity arises from phylogenetically-restricted specialized metabolic pathways. Specialized metabolites are synthesized in dedicated cells or tissues, with pathway genes sometimes colocalizing in biosynthetic gene clusters (BGCs). However, the mechanisms by which spatial expression patterns arise and the role of BGCs in pathway evolution remain underappreciated. In this study, we investigated the mechanisms driving acylsugar evolution in the Solanaceae. Previously thought to be restricted to glandular trichomes, acyl sugars were recently discovered in cultivated tomato roots. We demonstrated that acyl sugars in cultivated tomato roots and trichomes have different sugar cores, identified root-enriched paralogs of trichome acyl sugar pathway genes, and characterized a key paralog required for root acyl sugar biosynthesis, <em>SlASAT1-LIKE</em> (<em>SlASAT1-L</em>), which is nested within a previously-reported trichome acyl sugar BGC. Finally, we provided evidence that <em>ASAT1-L</em> arose through duplication of its paralog, <em>ASAT1</em>, and was trichome-expressed before acquiring root-specific expression in the <em>Solanum</em> genus. Our results illuminate the genomic context and molecular mechanisms underpinning metabolic diversity in plants.</p>

opencc-zeroApr 2024View details →
zenodo36/100

Datasets for "Micromonosporaceae Biosynthetic Gene Cluster Diversity Highlights the Need for Broad Spectrum Investigation"

<p>In this data collection is:<br><strong>Data S1</strong>: A folder with all the fasta files, representing the 42 strains (41 <em>Micromonosporaceae</em>, 1 <em>Streptomycetaceae</em>).<br><strong>Data S2</strong>: A folder with all the .gbk files for the BGC regions predicted by antiSMASH v5.1.1. These files were used as inputs for BiG-SCAPE and BiG-SLiCE.<br><strong>Data S3</strong>: A folder with all the .gbk files for the BGC regions predicted by antiSMASH v6.1.0.<br><strong>Data S4</strong>: A folder containing all the Quast outputs for the 42 strains.<br><strong>Data S5</strong>: A folder containing all the BUSCO outputs for the 42 strains. Example scripts are provided for scraping relevant information from the individual BUSCO outputs.<br><strong>Data S6</strong>: A folder containing GTDB (Genome Taxonomy Database) classification results, and species-level grouping results using FastANI (95% cutoff).<br><strong>Data S7</strong>: A folder containing an Interactive Tree of Life (iTOL)-compatible bar chart annotation using antiSMASH v5.1.1 BGC region information.<br><strong>Data S8</strong>: A folder containing a word document that describes the parameters used with Ubuntu WSL (Windows Subsystem for Linux) on the command line for programs antiSMASH v6.1.2, BiG-SCAPE v1.1.2, and BiG-SLiCE v1.1.1. Also included are parameters for MDSC in python. An example script is also provided for batch queries of BGCs against BiG-SLiCE v1.1.1&rsquo;s pre-processed dataset of ~1.2 million BGCs.<br><strong>Data S9</strong>: A folder containing the BiG-SCAPE visualization of the 38 <em>Micromonosporaceae</em> (post-QC filtering, excluding WMMA1363, WMMB482, WMMB486, and WMMC500) in Cytoscape.<br><strong>Data S10</strong>: A folder containing:<br>The pre-processed dataset of 1.2 million BGCs from BiG-SLiCE.<br>All report folders generated by BiG-SLiCE for the 779 <em>Micromonosporaceae </em>BGCs queried against the 1.2 million BGCs.<br>The results data.db and associated folders for the pre-processed dataset of 1.2 million BGCs.<br><strong>Data S11</strong>: A folder containing the scripts necessary to regenerate the figures and perform independent analyses, and the relevant data used for the analyses.<br><strong>Data S12: </strong>A folder containing the results of the nucleotide blast of WMMA1947.region12's siderophore contig against WMMD1120.region14's siderophore contig.</p> <p><strong>Supplementary Information:&nbsp;</strong>Supplementary Table S1 and Supplementary Figures S1-S181.</p>

opencc-by-4.0Jul 2023View details →
zenodo36/100

IsoAnalyst: A System-Wide Stable Isotopic Labeling Aproach for Connecting Natural Products to Their Cognate Biosynthetic Gene Clusters

<p>Processed mass spectrometry data input used to develop and validate the IsoAnalyst program and the output from these analyses. The S_erythraea folder contains all MS input and output data for <em>Saccharopolyspora erythraea</em>. The Micromonospora_RL09050HVFA folder contains MS input and output data for <em>Micromonopora sp.</em> RL09050HVFA, as well as the full antiSMASH output for the <em>Micromonopora sp.</em> RL09050HVFA genome.</p> <p>Version 2 contains updated IsoAnalyst results for both organisms.&nbsp;</p>

opencc-by-4.0Apr 2021View details →
zenodo36/100

Supporting data for the manuscript "Nerpa: a tool for discovering biosynthetic gene clusters of nonribosomal peptides"

<p>Preprocessed structures of nonribosomal peptides [NRPs] and genomic sequences (reference and representative genomes, biosynthetic gene clusters [BGCs]) used in the benchmark experiments in the Nerpa paper.</p> <p><strong>Files description</strong></p> <ul> <li><strong>mibig_nrp_bacteria_preprocessed.tar.gz</strong>&nbsp;contains the preprocessed dataset of 194 bacterial NRP BGCs from the MIBiG database.</li> <li><strong>mibig_nrp_bacteria_summary.tsv</strong>&nbsp;contains metadata for the MIBiG-NRP dataset.</li> <li><strong>bacterial_ref_and_repr_genomes_20210604_preprocessed.tar.gz</strong>&nbsp;contains the preprocessed dataset of 13,399 reference and representative bacterial genomes from the NCBI RefSeq database (retrieved on 2021/06/04).</li> <li><strong>bacterial_ref_and_repr_genomes_20210604_summary.txt</strong>&nbsp;contains metadata for the RefSeq dataset.</li> <li><strong>pnrpdb_preprocessed.info</strong>&nbsp;contains the Nerpa-preprocessed pNRPdb database, a database of 8,368 known and putative NRP structures.</li> <li><strong>pnrpdb_summary.tsv</strong>&nbsp;contains the pNRPdb database metadata.<br> &nbsp;</li> </ul>

opencc-by-4.0Sep 2021View details →
dryad36/100

Phylogenomics of the psychoactive mushroom genus Psilocybe and evolution of the psilocybin biosynthetic gene cluster

<p>Psychoactive mushrooms in the genus <em>Psilocybe</em> have immense cultural value and have been used for centuries in Mesoamerica. Despite a recent surge in interest in these mushrooms due to emerging evidence that psilocybin, the main psychoactive compound, is a promising therapeutic for a variety of mental illnesses, their phylogeny and taxonomy remain substantially incomplete. Moreover, the recent elucidation of the psilocybin biosynthetic gene cluster is known for only five species of <em>Psilocybe</em>, four of which belong to only one of two major clades. We set out to improve the phylogeny for <em>Psilocybe</em> using shotgun sequencing of 71 fungarium specimens, including 23 types, and conducting phylogenomic analysis using 2,983 single-copy gene families to generate a fully supported phylogeny. Molecular clock analysis suggests the stem lineage arose ~67 mya and diversified ~56 mya. We also show that psilocybin biosynthesis first arose in <em>Psilocybe</em>, with 4–5 possible horizontal transfers to other mushrooms between 40 and 9mya. Moreover, predicted orthologs of the psilocybin biosynthetic genes revealed two distinct gene orders within the cluster that corresponds to a deep split within the genus, possibly consistent with the independent acquisition of the cluster. By mining genomic data beyond markers for phylogenetic inference, we gained novel insights into the evolutionary origins of psilocybin biosynthesis that have implications for understanding the functional role of this powerful chemical and can inform translational applications for human well-being.</p>

opencc-zeroMay 2023View details →
dryad36/100

Data from: A high-quality genome assembly of the tetraploid Teucrium chamaedrys unveils a recent whole genome duplication and a large biosynthetic gene cluster for diterpenoid metabolism

Open the record for dataset details and reuse information.

publicJul 2025View details →
dryad36/100

Data for: Tomato root specialized metabolites evolved through gene duplication and regulatory divergence within a biosynthetic gene cluster

Open the record for dataset details and reuse information.

publicApr 2024View details →
dryad36/100

Metagenomic bins and biosynthetic gene clusters in gut bacteria of turtle ants

Open the record for dataset details and reuse information.

publicMar 2021View details →
dryad36/100

Phylogenomics of the psychoactive mushroom genus Psilocybe and evolution of the psilocybin biosynthetic gene cluster

Open the record for dataset details and reuse information.

publicNov 2023View details →
zenodo32/100

Genome sequences of Rhizopogon roseolus, Mariannaea elegans, Myrothecium verrucaria, and Sphaerostilbella broomeana and the identification of biosynthetic gene clusters for fungal peptide natural products

<p>Data to accompany the paper, detailed in manifest.txt</p>

opencc-by-4.0Apr 2022View details →
zenodo32/100

Sub-functionalization and epigenetic regulation of a biosynthetic gene cluster in Solanaceae

<p>Datasets for the code analysis in <em>Priego-Cubero, Knoch et al., 2024 </em>(https://doi.org/10.1101/2024.10.02.615186)</p>

opencc-by-4.0Oct 2024View details →
dryad28/100

Diversity of biosynthetic gene clusters in gut bacteria of turtle ants

<p class="Standard"><span>In insect-microbe nutritional symbioses the gut symbionts supplement the host diet with nutrients by producing amino acids and vitamins or by degrading lignin or polysaccharides macromolecules. In multipartite mutualisms composed of multiple symbionts from different taxonomical orders, it has been suggested that beside the genes involved in the nutritional symbiosis the symbionts maintain genes responsible for the production of metabolites putatively playing a role in the maintenance and interaction of the bacterial communities living in close proximity. To test this hypothesis, we investigated the diversity of biosynthetic gene clusters (BGCs) producing different non-primary metabolites in the genomes and metagenomes of the conserved gut symbionts associated with the herbivorous turtle ants (genus: <i>Cephalotes</i>). We studied 17 <i>Cephalotes</i> species collected across several geographical areas to reveal that (i) mining metagenomes and genomes show complementary results demonstrating the robustness of this approach to retrieve BGCs, (ii) the conserved gut symbionts involved in the nutritional symbiosis have a high diversity of BGCs of different chemical families, (iii) the phylogenetic analysis of BGCs encoding the production of arylpolyenes, non-ribosomal peptides (NRP), polyketides (PK), and siderophores shows high similarity between BGCs of a single symbiont across different ant host species, and between BGCs originated from different bacterial orders within a single host species. These findings together suggest multiple mechanisms of bacterial genome conservation and evolution of BGCs. We document the occurrence, diversity, and similarity of BGCs in the genome of obligate gut bacteria involved in multipartite mutualisms across the phylogeny of turtle ants.</span></p>

opencc-zeroJan 2021View details →
zenodo28/100

Biosynthetic Gene Cluster Synteny - Orthologous Polyketide Synthases in Hypogymnia physodes, Hypogymnia tubulosa and Parmelia sulcata

<p>Supplementary Material of Publication</p>

opencc-by-4.0Aug 2023View details →
dryad28/100

Diversity of biosynthetic gene clusters in gut bacteria of turtle ants

Open the record for dataset details and reuse information.

publicJan 2021View details →
geo24/100

Multilevel regulation of gene expression in brasilicardin biosynthetic gene cluster

GEO Series GSE271981. Nocardia terpenica. 8 samples. Type: Other.

openGEO-OpenJul 2025View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record