Protistan metabolism across the western North Atlantic Ocean revealed through autonomous underwater profiling
<p>Metatranscriptomic assembly, predicted open reading frames, counts, and annotation files from seawater samples obtained in the western North Atlantic Ocean. GitHub notebooks are located here: <a href="https://github.com/cnatalie/BATS">https://github.com/cnatalie/BATS</a>. Assembly was created using the <em>eukrhythmic</em> pipeline: <a href="https://github.com/AlexanderLabWHOI/eukrhythmic">https://github.com/AlexanderLabWHOI/eukrhythmic</a></p> <ul> <li>merged_merged.fasta.gz = Final assembly, merged across 44 metatranscriptomes using 4 different assemblers</li> <li>merged.fasta.transdecoder.pep.zip = Open reading frames of final assembly, predicted by Transdecoder </li> <li>merged.fasta.transdecoder-estimated-taxonomy.out.zip = EUKulele-derived taxonomic annotations of ORFs using a combined EukProt, PhyloDB, and RefSeq reference database</li> <li>newtaxa.eukprot.merged.fasta.transdecoder-estimated-taxonomy.out.zip = similar to above, but manually curated mid-level taxonomy for supergroups of interest</li> <li>eggnog.emapper.annotations.zip = eggnog-mapper annotations of ORFs</li> <li>table.tab.zip = counts associated with ORFs (merged.fasta.transdecoder.pep) generated with Salmon</li> <li>TPM_table.tab.zip = community-wide TPM (normalized) counts associated with ORFs (merged.fasta.transdecoder.pep) generated with Salmon</li> <li>copiesperL_ORFs_FactorIncluded.csv.zip = raw counts associated with ORFs (merged.fasta.transdecoder.pep) converted to copies per L taking into account spiked-in RNA standard concentration (copies), standard reads mapped, volume of seawater filtered, and dilution factor used in library preparation</li> <li>assembly.table.tab.zip = counts associated with final assembly (merged_merged.fasta) generated with Salmon</li> <li>SamplesViewReportCLIO_AE1913merged_trans210506_updated220606exclusive.zip = Exclusive spectral counts associated with ORFs (merged.fasta.transdecoder.pep). Peptide-spectrum matches were performed using Sequest algorithm within IseNode Proteome Discoverer 2.2.0.388 (Thermo Fisher Scientific). Scaffold 5.1.2 (Proteome Software) was used for protein grouping and exclusive spectral counting. Note, the (+x) data has been removed from protein names, which indicates whether (and how many) proteins sharing peptides were designated into the same protein group. </li> <li>cds.length2.tab.zip = Length of proteins (ORFs) in nucleotide base pairs</li> <li>CTD.zip = CTD files from cruise AE1913</li> </ul>
ShareScore
28/100
Overall dataset sharing score
Score breakdown
These five areas show where the dataset supports — or may limit — practical reuse.
- Stewardship
- 8
- Harmonization
- 4
- Access
- 16
- Reuse readiness
- 0
- Engagement
- 0