Skip to main content
zenodoopen

Protistan metabolism across the western North Atlantic Ocean revealed through autonomous underwater profiling

<p>Metatranscriptomic assembly, predicted open reading frames, counts, and annotation files from seawater samples obtained in the western North Atlantic Ocean. GitHub notebooks are located here:&nbsp;<a href="https://github.com/cnatalie/BATS">https://github.com/cnatalie/BATS</a>. Assembly was created using the <em>eukrhythmic</em> pipeline:&nbsp;<a href="https://github.com/AlexanderLabWHOI/eukrhythmic">https://github.com/AlexanderLabWHOI/eukrhythmic</a></p> <ul> <li>merged_merged.fasta.gz = Final assembly, merged across 44 metatranscriptomes using&nbsp;4 different assemblers</li> <li>merged.fasta.transdecoder.pep.zip = Open reading frames of final assembly, predicted by Transdecoder&nbsp;</li> <li>merged.fasta.transdecoder-estimated-taxonomy.out.zip&nbsp;= EUKulele-derived taxonomic&nbsp;annotations of ORFs using a combined EukProt, PhyloDB, and RefSeq reference database</li> <li>newtaxa.eukprot.merged.fasta.transdecoder-estimated-taxonomy.out.zip = similar to above, but manually curated mid-level taxonomy for supergroups of interest</li> <li>eggnog.emapper.annotations.zip = eggnog-mapper annotations of ORFs</li> <li>table.tab.zip =&nbsp;counts associated with ORFs (merged.fasta.transdecoder.pep) generated with Salmon</li> <li>TPM_table.tab.zip = community-wide TPM (normalized) counts associated with ORFs (merged.fasta.transdecoder.pep) generated with Salmon</li> <li>copiesperL_ORFs_FactorIncluded.csv.zip = raw counts associated with ORFs (merged.fasta.transdecoder.pep) converted to copies per L taking into account spiked-in RNA standard concentration (copies), standard reads mapped, volume of seawater filtered, and dilution factor used in library preparation</li> <li>assembly.table.tab.zip =&nbsp;counts associated with final assembly (merged_merged.fasta) generated with Salmon</li> <li>SamplesViewReportCLIO_AE1913merged_trans210506_updated220606exclusive.zip = Exclusive spectral counts associated with ORFs (merged.fasta.transdecoder.pep). Peptide-spectrum&nbsp;matches were performed using Sequest algorithm within IseNode Proteome Discoverer 2.2.0.388 (Thermo Fisher Scientific). Scaffold 5.1.2 (Proteome Software)&nbsp; was used for protein grouping and exclusive spectral counting. Note, the (+x) data has been removed from protein names, which indicates whether (and how many) proteins sharing peptides were designated into the same protein&nbsp;group.&nbsp;</li> <li>cds.length2.tab.zip = Length of proteins (ORFs)&nbsp;in nucleotide base pairs</li> <li>CTD.zip = CTD files from cruise AE1913</li> </ul>

ShareScore

28/100

Overall dataset sharing score

Score breakdown

These five areas show where the dataset supports — or may limit — practical reuse.

Stewardship
8
Harmonization
4
Access
16
Reuse readiness
0
Engagement
0

Topics