Emergence and evolution of heterocyte glycolipid biosynthesis enabled specialized nitrogen fixation in cyanobacteria
<h1><strong>Abstract </strong></h1> <p>Paleontological and phylogenomic observations have shed light on the evolution of cyanobacteria. Nevertheless, the emergence of heterocytes, specialized cells for nitrogen fixation, remains unclear. Heterocytes are surrounded by heterocyte glycolipids (HGs), which contribute to protection of the nitrogenase enzyme from oxygen. Here, by comprehensive HG identification and screening of HG biosynthesis genes throughout cyanobacteria, we identify HG analogs produced by specific and distantly related non-heterocytous cyanobacteria. These structurally less complex molecules probably acted as precursors of HGs, suggesting that HGs arose after a genomic reorganization and expansion of ancestral biosynthetic machinery, enabling the rise of cyanobacterial heterocytes in an increasingly oxygenated atmosphere. Subsequently, HG chemical structure evolved convergently in response to environmental pressures. Our results open a new chapter in the potential use of diagenetic products of HGs and HG analogs as fossils for reconstructing the evolution of multicellularity and division of labor in cyanobacteria.</p> <h2><strong>Here we supply:</strong></h2> <div> <ul> <li><strong><span>Supplementary Data 1. Selected cyanobacterial genomes from the PATRIC genome database (now part of the BV-BRC database).</span></strong><span> Files called ‘selected_Cyanogenomes.genome_*.20220430.txt’ are sourced from the PATRIC File Transfer Protocol server (ftp.patricbrc.org). ‘gtdbtk.bac120.summary.tsv’ is the GTDB-Tk output file, and ‘qa.summary_extended.txt’ the CheckM output file.</span></li> <li><strong><span>Supplementary Data 2. HG biosynthetic gene clusters in selected PATRIC genomes and 14 newly sequenced genomes.</span></strong><span> The file ‘islands_on_contigs.3_ORFs_in_between.expanded_island_with_nucleotide_positions.txt’ contains the location of all hits to <em>Anabaena</em> sp. PCC 7120 HG biosynthesis genes. ORFs were predicted with Prodigal. The structure of a contig is as follows: “genome | contig”. The structure of a hit is as follows: “ORF number on contig | query (<em>e</em>-value; bit-score; start of alignment in query; end of alignment in query; query coverage per subject; start of alignment in subject; end of alignment in subject; subject coverage) [nucleotide position on contig start; nucleotide position on contig end; direction]”. Non-overlapping hits on the same ORF (see Online Methods) are connected with ‘&&&’ characters. An asterisk (‘*’) indicates that the hit is located at most three ORFs from a contig edge. Clusters of hits that are at most three open reading frames (ORFs) apart are connected with ‘~~~’ characters. The file ‘Supplementary_table.script_1.txt’ contains a summary of all identified <em>hgl</em> islands (i.e. clusters containing at least seven unique HG biosynthesis gene hits).</span></li> <li><strong><span>Supplementary Data 3. HG biosynthetic gene clusters in 255,388 prokaryotic genomes from the PATRIC genome database (now part of the BV-BRC database).</span></strong><span> The file called ‘PATRIC_20230120.selection_c50_c10.txt’ contains information on the selected PATRIC genomes based on data sourced from the PATRIC File Transfer Protocol server (<a>ftp.patricbrc.org</a>). The file ‘all_tree_of_life_genomes.islands_on_contigs.3_ORFs_in_between.expanded_island_with_nucleotide_positions.txt’ contains the location of all hits to <em>Anabaena</em> sp. PCC 7120 HG biosynthesis genes. ORFs were predicted with Prodigal. The structure of a contig is as follows: “genome | contig”. The structure of a hit is as follows: “ORF number on contig | query (<em>e</em>-value; bit-score; start of alignment in query; end of alignment in query; query coverage per subject; start of alignment in subject; end of alignment in subject; subject coverage) [nucleotide position on contig start; nucleotide position on contig end; direction]”. Non-overlapping hits on the same ORF (see Online Methods) are connected with ‘&&&’ characters. An asterisk (‘*’) indicates that the hit is located at most three ORFs from a contig edge. Clusters of hits that are at most three open reading frames (ORFs) apart are connected with ‘~~~’ characters.</span></li> <li><strong><span>Supplementary Data 4. Phylogeny of representative cyanobacterial genomes based on a core gene superalignment.</span></strong><span> The folder contains the files used to generate Fig. 2a. The directory ‘IQ-TREE’ contains the tree file and iTOL annotation files. The file ‘dRep.representative_to_cluster.txt’ contains the dRep clusters. Note that the manually defined subclades in the iTOL annotation file ‘iTOL_annotation.manually_defined_clades.DATASET_STYLE.txt’ have a different numbering from the paper: subclades 0 and 1 are the ‘heterocytous sister clades’, and subclades 2-10 in the annotation file are heterocytous subclades 1-9 in the paper, respectively.</span></li> <li><strong><span>Supplementary Data 5. Lipid data files.</span></strong><span> The folder contains all the UHPLC-HRMS<em><sup>n</sup></em> (Orbitrap) datafiles used in this study. The directory ‘CCY strains’ includes 24 heterocytous cyanobacterial cultures corresponding to 23 strains grown in nitrogen-deficient media, the resulting data are shown in Supplementary Table 10. Directory ‘HglT mutant’ contains the datafiles used to generate Supplementary Table 15. The directory ‘LEGE strains’ includes the UHPLC-HRMS<em><sup>n</sup></em> (Orbitrap) and GC-MS datafiles corresponding to eight cultures of two non-heterocytous strains grown in media with and without nitrogen for 38 to 77 days, the resulting data are shown in Supplementary Tables 10, 17 and 18.</span></li> <li><strong><span>Supplementary Data 6. Plasmid maps. </span></strong><span>GenBank and FASTA files of plasmids generated in this study. ‘HglT deletion’ directory contains the genomic region surrounding <em>hglT</em> in the <em>wild-type </em>strain and after deletion used to generate Supplementary Fig. 15. pAM5404 is shown in Supplementary Fig. 16 and p(A)RP0XX are shown in Supplementary Fig. 17.<span><span></span></span></span></li> <li><strong><span>Supplementary Data 7. Phylogenies of seven <em>hgl</em> island genes and of a concatenated alignment of these genes.</span></strong><span> The folder contains the files used to generate Supplementary Fig. 18 (in the directory ‘gene_trees_hgl_islands’), and Fig. 4 and related figures (in the directory ‘gene_trees_hgl_islands_4’. The directories contain the alignments and trimmed alignments, IQ-TREE output files, and iTOL annotation files. The file ‘gene_trees_hgl_islands/analysis_individual_gene_trees/explore_individual_gene_clusters.ipynb’ contains the code to identify the five <em>hgl</em> islands that contain genes with incongruent evolutionary histories.</span></li> <li><strong><span>Supplementary Data 8. Phylogeny of <em>hglE<sub>A</sub></em> homologs.</span></strong><span> The folder contains the files used to generate Supplementary Fig. 24 and related figures. The file ‘selected_hglE_hits.txt’ contains the selected <em>hglE<sub>A</sub></em> hits and the genomic cluster on which they are located. The folder contains the alignment and trimmed alignment, IQ-TREE output files, and iTOL annotation files.</span></li> <li><strong><span>All the code used in this publication </span></strong><span>including scripts used for: genome assemblies, download of genomes from public repositories, quality and contamination checks, genome analysis, construction of the phylogenetic trees, <em>hgl</em> island identification, etc. The shell script 'commands.sh' within each directory contains all the code used to generate the content in the directory.</span></li> <li><strong><span>All the figures used in this publication </span></strong><span>including the figures in the Supplementary Information file.</span></li> </ul> <div> <div> <div></div> </div> </div> </div>
ShareScore
32/100
Overall dataset sharing score
Score breakdown
These five areas show where the dataset supports — or may limit — practical reuse.
- Stewardship
- 8
- Harmonization
- 4
- Access
- 16
- Reuse readiness
- 0
- Engagement
- 4