Skip to main content
zenodoopen

Machine Learning based identification of putative coral pathogens in endangered Caribbean staghorn coral

<h1>Supplementary Files</h1> <p>SupplementaryFile1.csv.gz &ndash; Metadata for field collected samples with columns:</p> <ul> <li>&ldquo;sample_id&rdquo; &ndash; individual sample names.</li> <li>&ldquo;health&rdquo; &ndash; &ldquo;H&rdquo; healthy and &ldquo;D&rdquo; diseased fragments.</li> <li>&ldquo;year&rdquo; &ndash; year fragment collected.</li> <li>&nbsp;&ldquo;season&rdquo; &ndash; season fragment collected (&ldquo;S&rdquo; July and &ldquo;W&rdquo; January)</li> <li>&ldquo;site&rdquo; &ndash; location fragment collected from</li> <li>&ldquo;lib.size&rdquo; &ndash; total number of sequenced reads</li> <li>&ldquo;norm.factors&rdquo; &ndash; factor used to normalize read counts of ASVs</li> </ul> <p>SupplementaryFile2.csv.gz &ndash; Metadata for tank collected samples with columns:</p> <ul> <li>&nbsp;&ldquo;sample_id&rdquo;<a name="_Hlk163818341"></a> &ndash; individual sample names.</li> <li>&ldquo;geno&rdquo; &ndash; fragment genotype</li> <li>&nbsp;&ldquo;fragment_id&rdquo; &ndash; fragment identification tracked through repeated sampling</li> <li>&nbsp;&ldquo;tank_id&rdquo; &ndash; tank identification</li> <li>&ldquo;time_treat&rdquo; &ndash; concatenated metric for sampling time, exposure, and disease outcome separated by &ldquo;_&rdquo; <ul> <li>&nbsp;Time &ndash; 0, 2, 8</li> <li>Exposure &ndash; &ldquo;D&rdquo; Diseased, &ldquo;N&rdquo; Healthy</li> <li>Disease Outcome - &ldquo;D&rdquo; Diseased, &ldquo;H&rdquo; Healthy</li> </ul> </li> <li>&nbsp;&ldquo;lib.size&rdquo; &ndash; total number of sequenced reads</li> <li>&ldquo;norm.factors&rdquo; &ndash; factor used to normalize read counts of ASVs</li> </ul> <p>SupplementaryFile3.fasta &ndash; FASTA file including complete 16s sequences named with ASV identifier and taxonomy.</p> <p>SupplementaryFile4.csv.gz &ndash; Matrix of the number of reads of each ASV sequenced in each sample. Combined both field and tank samples.</p> <p>SupplementaryFile5.csv.gz &ndash; Matrix of the log2 CPM of each ASV sequenced in each sample. Combined both field and tank samples.</p> <p>SupplementaryFile6.csv.gz &ndash; Complete results for each ASV association.</p> <ul> <li>&nbsp;&ldquo;top_classification&rdquo; &ndash; lowest taxonomic classification with more than 80% confidence.</li> <li>&ldquo;taxonomy&rdquo; &ndash; Full taxonomy including confidence in each taxonomic level.</li> <li>&ldquo;passedFilter&rdquo; &ndash; indicates taxa filtered from analysis due to rarity and/or lack of observations across sample times.</li> <li>&nbsp;&ldquo;rank_*&rdquo; &ndash; machine learning model rankings, median ranking, and model estimated ranking along with standard error, confidence interval, and FDR adjusted p-value used to identify important ASVs.</li> <li>&ldquo;ml_retained&rdquo; &ndash; Indicates if the ASV was of above average importance to ML models. NA values indicate ASVs which were filtered prior to ML modelling.</li> <li>&ldquo;fieldModel_*&rdquo; &ndash; ANOVA table results for each ASV testing the effects of health, year, season and all possible interactions indicating: <ul> <li>Sums of squares, mean squares, numerator and denominator degrees of freedom, F statistic, p-value, and FDR corrected p-value.</li> <li>NA values are filled for ASVs filtered prior to differential abundance analysis.</li> </ul> </li> <li>&ldquo;diffAbundance_healthAssociation&rdquo; &ndash; Marks the health association of ASVs from differential abundance analysis of field samples: &ldquo;H&rdquo; health, &ldquo;D&rdquo; diseased, &ldquo;N&rdquo; none, NA &ndash; filtered prior to differential abundance analysis.</li> <li>&ldquo;fieldLogFC_*&rdquo; &ndash; Post-hoc contrasts for ML retained ASVs testing the significance of the log2 fold-change between disease and healthy fragments within each sampling time (year: 2016, 2017 &amp; season: &ldquo;S&rdquo; July, &ldquo;W&rdquo; January) showing: <ul> <li>Mean estimate, standard error, degrees of freedom, lower and upper 95% confidence interval, t-statistic, p-value, FDR adjusted p-value.</li> <li>NA values are filled for ASVs which were not marked as important by ML models.</li> </ul> </li> <li>&nbsp;&ldquo;field_consistent&rdquo; &ndash; Indicates if the ASV was consistently healthy or disease associated across sampling times. NA values are filled for ASVs which were not marked as important by ML models.</li> <li>&ldquo;tankModel_*&rdquo; &ndash; ANOVA table results for each ASV testing the effect of the combination of time, disease exposure, and disease outcome, indicating: <ul> <li>Sums of squares, mean squares, numerator and denominator degrees of freedom, F statistic, p-value, and FDR corrected p-value.</li> <li>NA values are filled for ASVs filtered prior to tank experimental analysis.</li> </ul> </li> <li>&nbsp;&ldquo;tankLogFC_*&rdquo; &ndash; Post-hoc contrasts for ASVs tested in tank exposure experiments. <ul> <li>Contrasts include: <ul> <li>Post-exposure diseased vs healthy outcome regardless of exposure (DvH)</li> <li>Post-exposure diseased vs healthy exposure regardless of outcome (DvN)</li> <li>Post-exposure disease exposed corals with disease symptoms compared to disease exposed but still healthy corals (DDvDH)</li> <li>Post-exposure disease exposed corals with disease symptoms compared to healthy exposed and still healthy corals (DDvNH)</li> <li>Post-exposure disease exposed corals which stay healthy compared to healthy exposed and still healthy corals (DHvNH)</li> <li>Pre-exposure compared to Post-exposure in corals with the disease regardless of exposure (PostvPreD)</li> <li>Pre-exposure compared to Post-exposure in corals without the disease regardless of exposure (PostvPreH)</li> </ul> </li> <li>Mean estimate, standard error, degrees of freedom, lower and upper 95% confidence interval, t-statistic, p-value, FDR adjusted p-value.</li> <li>NA values are filled for ASVs which were not consistently associated with healthy or diseased corals in the field experiment.</li> </ul> </li> <li>&ldquo;pathogen_classification&rdquo; &ndash; Indicates the predicted microbial classification based on the tank results. Pathogen, Opportunist, Commensal <ul> <li>NA values are filled for ASVs which were not consistently associated with healthy or diseased corals in the field experiment.</li> </ul> </li> </ul>

ShareScore

40/100

Overall dataset sharing score

Score breakdown

These five areas show where the dataset supports — or may limit — practical reuse.

Stewardship
4
Harmonization
4
Access
20
Reuse readiness
8
Engagement
4