Machine Learning based identification of putative coral pathogens in endangered Caribbean staghorn coral
<h1>Supplementary Files</h1> <p>SupplementaryFile1.csv.gz – Metadata for field collected samples with columns:</p> <ul> <li>“sample_id” – individual sample names.</li> <li>“health” – “H” healthy and “D” diseased fragments.</li> <li>“year” – year fragment collected.</li> <li> “season” – season fragment collected (“S” July and “W” January)</li> <li>“site” – location fragment collected from</li> <li>“lib.size” – total number of sequenced reads</li> <li>“norm.factors” – factor used to normalize read counts of ASVs</li> </ul> <p>SupplementaryFile2.csv.gz – Metadata for tank collected samples with columns:</p> <ul> <li> “sample_id”<a name="_Hlk163818341"></a> – individual sample names.</li> <li>“geno” – fragment genotype</li> <li> “fragment_id” – fragment identification tracked through repeated sampling</li> <li> “tank_id” – tank identification</li> <li>“time_treat” – concatenated metric for sampling time, exposure, and disease outcome separated by “_” <ul> <li> Time – 0, 2, 8</li> <li>Exposure – “D” Diseased, “N” Healthy</li> <li>Disease Outcome - “D” Diseased, “H” Healthy</li> </ul> </li> <li> “lib.size” – total number of sequenced reads</li> <li>“norm.factors” – factor used to normalize read counts of ASVs</li> </ul> <p>SupplementaryFile3.fasta – FASTA file including complete 16s sequences named with ASV identifier and taxonomy.</p> <p>SupplementaryFile4.csv.gz – Matrix of the number of reads of each ASV sequenced in each sample. Combined both field and tank samples.</p> <p>SupplementaryFile5.csv.gz – Matrix of the log2 CPM of each ASV sequenced in each sample. Combined both field and tank samples.</p> <p>SupplementaryFile6.csv.gz – Complete results for each ASV association.</p> <ul> <li> “top_classification” – lowest taxonomic classification with more than 80% confidence.</li> <li>“taxonomy” – Full taxonomy including confidence in each taxonomic level.</li> <li>“passedFilter” – indicates taxa filtered from analysis due to rarity and/or lack of observations across sample times.</li> <li> “rank_*” – machine learning model rankings, median ranking, and model estimated ranking along with standard error, confidence interval, and FDR adjusted p-value used to identify important ASVs.</li> <li>“ml_retained” – Indicates if the ASV was of above average importance to ML models. NA values indicate ASVs which were filtered prior to ML modelling.</li> <li>“fieldModel_*” – ANOVA table results for each ASV testing the effects of health, year, season and all possible interactions indicating: <ul> <li>Sums of squares, mean squares, numerator and denominator degrees of freedom, F statistic, p-value, and FDR corrected p-value.</li> <li>NA values are filled for ASVs filtered prior to differential abundance analysis.</li> </ul> </li> <li>“diffAbundance_healthAssociation” – Marks the health association of ASVs from differential abundance analysis of field samples: “H” health, “D” diseased, “N” none, NA – filtered prior to differential abundance analysis.</li> <li>“fieldLogFC_*” – Post-hoc contrasts for ML retained ASVs testing the significance of the log2 fold-change between disease and healthy fragments within each sampling time (year: 2016, 2017 & season: “S” July, “W” January) showing: <ul> <li>Mean estimate, standard error, degrees of freedom, lower and upper 95% confidence interval, t-statistic, p-value, FDR adjusted p-value.</li> <li>NA values are filled for ASVs which were not marked as important by ML models.</li> </ul> </li> <li> “field_consistent” – Indicates if the ASV was consistently healthy or disease associated across sampling times. NA values are filled for ASVs which were not marked as important by ML models.</li> <li>“tankModel_*” – ANOVA table results for each ASV testing the effect of the combination of time, disease exposure, and disease outcome, indicating: <ul> <li>Sums of squares, mean squares, numerator and denominator degrees of freedom, F statistic, p-value, and FDR corrected p-value.</li> <li>NA values are filled for ASVs filtered prior to tank experimental analysis.</li> </ul> </li> <li> “tankLogFC_*” – Post-hoc contrasts for ASVs tested in tank exposure experiments. <ul> <li>Contrasts include: <ul> <li>Post-exposure diseased vs healthy outcome regardless of exposure (DvH)</li> <li>Post-exposure diseased vs healthy exposure regardless of outcome (DvN)</li> <li>Post-exposure disease exposed corals with disease symptoms compared to disease exposed but still healthy corals (DDvDH)</li> <li>Post-exposure disease exposed corals with disease symptoms compared to healthy exposed and still healthy corals (DDvNH)</li> <li>Post-exposure disease exposed corals which stay healthy compared to healthy exposed and still healthy corals (DHvNH)</li> <li>Pre-exposure compared to Post-exposure in corals with the disease regardless of exposure (PostvPreD)</li> <li>Pre-exposure compared to Post-exposure in corals without the disease regardless of exposure (PostvPreH)</li> </ul> </li> <li>Mean estimate, standard error, degrees of freedom, lower and upper 95% confidence interval, t-statistic, p-value, FDR adjusted p-value.</li> <li>NA values are filled for ASVs which were not consistently associated with healthy or diseased corals in the field experiment.</li> </ul> </li> <li>“pathogen_classification” – Indicates the predicted microbial classification based on the tank results. Pathogen, Opportunist, Commensal <ul> <li>NA values are filled for ASVs which were not consistently associated with healthy or diseased corals in the field experiment.</li> </ul> </li> </ul>
ShareScore
40/100
Overall dataset sharing score
Score breakdown
These five areas show where the dataset supports — or may limit — practical reuse.
- Stewardship
- 4
- Harmonization
- 4
- Access
- 20
- Reuse readiness
- 8
- Engagement
- 4