Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

2

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

2 results for “heterocytous cyanobacteria”

Learn how ShareScore rates datasets ↗
zenodo32/100

Emergence and evolution of heterocyte glycolipid biosynthesis enabled specialized nitrogen fixation in cyanobacteria

<h1><strong>Abstract&nbsp;</strong></h1> <p>Paleontological and phylogenomic observations have shed light on the evolution of cyanobacteria. Nevertheless, the emergence of heterocytes, specialized cells for nitrogen fixation, remains unclear. Heterocytes are surrounded by heterocyte glycolipids (HGs), which contribute to protection of the nitrogenase enzyme from oxygen. Here, by comprehensive HG identification and screening of HG biosynthesis genes throughout cyanobacteria, we identify HG analogs produced by specific and distantly related non-heterocytous cyanobacteria. These structurally less complex molecules probably acted as precursors of HGs, suggesting that HGs arose after a genomic reorganization and expansion of ancestral biosynthetic machinery, enabling the rise of cyanobacterial heterocytes in an increasingly oxygenated atmosphere. Subsequently, HG chemical structure evolved convergently in response to environmental pressures. Our results open a new chapter in the potential use of diagenetic products of HGs and HG analogs as fossils for reconstructing the evolution of multicellularity and division of labor in cyanobacteria.</p> <h2><strong>Here we supply:</strong></h2> <div> <ul> <li><strong><span>Supplementary Data 1. Selected cyanobacterial genomes from the PATRIC genome database (now part of the BV-BRC database).</span></strong><span> Files called &lsquo;selected_Cyanogenomes.genome_*.20220430.txt&rsquo; are sourced from the PATRIC File Transfer Protocol server (ftp.patricbrc.org). &lsquo;gtdbtk.bac120.summary.tsv&rsquo; is the GTDB-Tk output file, and &lsquo;qa.summary_extended.txt&rsquo; the CheckM output file.</span></li> <li><strong><span>Supplementary Data 2. HG biosynthetic gene clusters in selected PATRIC genomes and 14 newly sequenced genomes.</span></strong><span> The file &lsquo;islands_on_contigs.3_ORFs_in_between.expanded_island_with_nucleotide_positions.txt&rsquo; contains the location of all hits to <em>Anabaena</em> sp. PCC 7120 HG biosynthesis genes. ORFs were predicted with Prodigal. The structure of a contig is as follows: &ldquo;genome | contig&rdquo;. The structure of a hit is as follows: &ldquo;ORF number on contig | query (<em>e</em>-value; bit-score; start of alignment in query; end of alignment in query; query coverage per subject; start of alignment in subject; end of alignment in subject; subject coverage) [nucleotide position on contig start; nucleotide position on contig end; direction]&rdquo;. Non-overlapping hits on the same ORF (see Online Methods) are connected with &lsquo;&amp;&amp;&amp;&rsquo; characters. An asterisk (&lsquo;*&rsquo;) indicates that the hit is located at most three ORFs from a contig edge. Clusters of hits that are at most three open reading frames (ORFs) apart are connected with &lsquo;~~~&rsquo; characters. The file &lsquo;Supplementary_table.script_1.txt&rsquo; contains a summary of all identified <em>hgl</em> islands (i.e. clusters containing at least seven unique HG biosynthesis gene hits).</span></li> <li><strong><span>Supplementary Data 3. HG biosynthetic gene clusters in 255,388 prokaryotic genomes from the PATRIC genome database (now part of the BV-BRC database).</span></strong><span> The file called &lsquo;PATRIC_20230120.selection_c50_c10.txt&rsquo; contains information on the selected PATRIC genomes based on data sourced from the PATRIC File Transfer Protocol server (<a>ftp.patricbrc.org</a>). The file &lsquo;all_tree_of_life_genomes.islands_on_contigs.3_ORFs_in_between.expanded_island_with_nucleotide_positions.txt&rsquo; contains the location of all hits to <em>Anabaena</em> sp. PCC 7120 HG biosynthesis genes. ORFs were predicted with Prodigal. The structure of a contig is as follows: &ldquo;genome | contig&rdquo;. The structure of a hit is as follows: &ldquo;ORF number on contig | query (<em>e</em>-value; bit-score; start of alignment in query; end of alignment in query; query coverage per subject; start of alignment in subject; end of alignment in subject; subject coverage) [nucleotide position on contig start; nucleotide position on contig end; direction]&rdquo;. Non-overlapping hits on the same ORF (see Online Methods) are connected with &lsquo;&amp;&amp;&amp;&rsquo; characters. An asterisk (&lsquo;*&rsquo;) indicates that the hit is located at most three ORFs from a contig edge. Clusters of hits that are at most three open reading frames (ORFs) apart are connected with &lsquo;~~~&rsquo; characters.</span></li> <li><strong><span>Supplementary Data 4. Phylogeny of representative cyanobacterial genomes based on a core gene superalignment.</span></strong><span> The folder contains the files used to generate Fig. 2a. The directory &lsquo;IQ-TREE&rsquo; contains the tree file and iTOL annotation files. The file &lsquo;dRep.representative_to_cluster.txt&rsquo; contains the dRep clusters. Note that the manually defined subclades in the iTOL annotation file &lsquo;iTOL_annotation.manually_defined_clades.DATASET_STYLE.txt&rsquo; have a different numbering from the paper: subclades 0 and 1 are the &lsquo;heterocytous sister clades&rsquo;, and subclades 2-10 in the annotation file are heterocytous subclades 1-9 in the paper, respectively.</span></li> <li><strong><span>Supplementary Data 5. Lipid data files.</span></strong><span> The folder contains all the UHPLC-HRMS<em><sup>n</sup></em> (Orbitrap) datafiles used in this study. The directory &lsquo;CCY strains&rsquo; includes 24 heterocytous cyanobacterial cultures corresponding to 23 strains grown in nitrogen-deficient media, the resulting data are shown in Supplementary Table 10. Directory &lsquo;HglT mutant&rsquo; contains the datafiles used to generate Supplementary Table 15. The directory &lsquo;LEGE strains&rsquo; includes the UHPLC-HRMS<em><sup>n</sup></em> (Orbitrap) and GC-MS datafiles corresponding to eight cultures of two non-heterocytous strains grown in media with and without nitrogen for 38 to 77 days, the resulting data are shown in Supplementary Tables 10, 17 and 18.</span></li> <li><strong><span>Supplementary Data 6. Plasmid maps.&nbsp;</span></strong><span>GenBank and FASTA files of plasmids generated in this study. &lsquo;HglT deletion&rsquo; directory contains the genomic region surrounding <em>hglT</em> in the <em>wild-type </em>strain and after deletion used to generate Supplementary Fig. 15. pAM5404 is shown in Supplementary Fig. 16 and p(A)RP0XX are shown in Supplementary Fig. 17.<span><span></span></span></span></li> <li><strong><span>Supplementary Data 7. Phylogenies of seven <em>hgl</em> island genes and of a concatenated alignment of these genes.</span></strong><span> The folder contains the files used to generate Supplementary Fig. 18 (in the directory &lsquo;gene_trees_hgl_islands&rsquo;), and Fig. 4 and related figures (in the directory &lsquo;gene_trees_hgl_islands_4&rsquo;. The directories contain the alignments and trimmed alignments, IQ-TREE output files, and iTOL annotation files. The file &lsquo;gene_trees_hgl_islands/analysis_individual_gene_trees/explore_individual_gene_clusters.ipynb&rsquo; contains the code to identify the five&nbsp;<em>hgl</em> islands that contain genes with incongruent evolutionary histories.</span></li> <li><strong><span>Supplementary Data 8. Phylogeny of <em>hglE<sub>A</sub></em> homologs.</span></strong><span> The folder contains the files used to generate Supplementary Fig. 24 and related figures. The file &lsquo;selected_hglE_hits.txt&rsquo; contains the selected <em>hglE<sub>A</sub></em> hits and the genomic cluster on which they are located. The folder contains the alignment and trimmed alignment, IQ-TREE output files, and iTOL annotation files.</span></li> <li><strong><span>All the code used in this publication&nbsp;</span></strong><span>including scripts used for: genome assemblies, download of genomes from public repositories, quality and contamination checks, genome analysis, construction of the phylogenetic trees, <em>hgl</em> island identification, etc. The shell script 'commands.sh' within each directory contains all the code used to generate the content in the directory.</span></li> <li><strong><span>All the figures used in this publication </span></strong><span>including the figures in the Supplementary Information file.</span></li> </ul> <div> <div> <div></div> </div> </div> </div>

opencc-by-4.0May 2024View details →
zenodo32/100

HglT fragments of heterocytous cyanobacteria

<p>Assembled <em>hglT</em> gene fragments obtained via PCR and direct sequencing using primer set 2 (Fw1 mix B + Rv1, Supplementary table 2) on the heterocytous cyanobacterial cultures&nbsp;listed in Table 2, Supplementary Table 8&nbsp;and shown in Figure 5 (sequences in blue) in P&eacute;rez Gallego et al. 2023&nbsp;</p> <p>Nucleotide sequences were obtained via Sanger sequencing and were processed using Geneious Prime (v 2023.0.4). Sequences were trimmed using an error probability limit of 0.01. Potential heterozygous bases in single reads were identified using a 50% peak similarity cutoff, peak detection height was set at 10%. When available, the consensus sequence was obtained by aligning forward and reverse reads with Geneious assembler using the highest sensitivity settings. When appropriate, sequences belonging to the pCR&trade;4-TOPO&trade; vector (Invitrogen, Carlsbad, CA, USA) were identified and removed.</p> <p>To generate the phylogenetic trees sequences were aligned using MAFFT (v7.407) with L-INS-i iterative refinement method (Katoh and Standley, 2013) and poorly aligned regions were removed using trimAl (Capella-Guti&eacute;rrez et al., 2009). Phylogenetic trees were built using IQ-tree (v1.6.7) and its in-built nucleotide substitution model finder (Kalyaanamoorthy et al., 2017), using 1000 replicates to perform SH-like approximate likelihood ratio test (SH-aLRT) (Guindon et al., 2010) and 1000 bootstrap replicates.&nbsp;</p>

opencc-by-4.0Aug 2023View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record