Datasets for "Micromonosporaceae Biosynthetic Gene Cluster Diversity Highlights the Need for Broad Spectrum Investigation"
<p>In this data collection is:<br><strong>Data S1</strong>: A folder with all the fasta files, representing the 42 strains (41 <em>Micromonosporaceae</em>, 1 <em>Streptomycetaceae</em>).<br><strong>Data S2</strong>: A folder with all the .gbk files for the BGC regions predicted by antiSMASH v5.1.1. These files were used as inputs for BiG-SCAPE and BiG-SLiCE.<br><strong>Data S3</strong>: A folder with all the .gbk files for the BGC regions predicted by antiSMASH v6.1.0.<br><strong>Data S4</strong>: A folder containing all the Quast outputs for the 42 strains.<br><strong>Data S5</strong>: A folder containing all the BUSCO outputs for the 42 strains. Example scripts are provided for scraping relevant information from the individual BUSCO outputs.<br><strong>Data S6</strong>: A folder containing GTDB (Genome Taxonomy Database) classification results, and species-level grouping results using FastANI (95% cutoff).<br><strong>Data S7</strong>: A folder containing an Interactive Tree of Life (iTOL)-compatible bar chart annotation using antiSMASH v5.1.1 BGC region information.<br><strong>Data S8</strong>: A folder containing a word document that describes the parameters used with Ubuntu WSL (Windows Subsystem for Linux) on the command line for programs antiSMASH v6.1.2, BiG-SCAPE v1.1.2, and BiG-SLiCE v1.1.1. Also included are parameters for MDSC in python. An example script is also provided for batch queries of BGCs against BiG-SLiCE v1.1.1’s pre-processed dataset of ~1.2 million BGCs.<br><strong>Data S9</strong>: A folder containing the BiG-SCAPE visualization of the 38 <em>Micromonosporaceae</em> (post-QC filtering, excluding WMMA1363, WMMB482, WMMB486, and WMMC500) in Cytoscape.<br><strong>Data S10</strong>: A folder containing:<br>The pre-processed dataset of 1.2 million BGCs from BiG-SLiCE.<br>All report folders generated by BiG-SLiCE for the 779 <em>Micromonosporaceae </em>BGCs queried against the 1.2 million BGCs.<br>The results data.db and associated folders for the pre-processed dataset of 1.2 million BGCs.<br><strong>Data S11</strong>: A folder containing the scripts necessary to regenerate the figures and perform independent analyses, and the relevant data used for the analyses.<br><strong>Data S12: </strong>A folder containing the results of the nucleotide blast of WMMA1947.region12's siderophore contig against WMMD1120.region14's siderophore contig.</p> <p><strong>Supplementary Information: </strong>Supplementary Table S1 and Supplementary Figures S1-S181.</p>
ShareScore
36/100
Overall dataset sharing score
Score breakdown
These five areas show where the dataset supports — or may limit — practical reuse.
- Stewardship
- 8
- Harmonization
- 4
- Access
- 16
- Reuse readiness
- 8
- Engagement
- 0