Prospects for a sequence-based taxonomy of influenza A virus subtypes
<p>This dataset comprises the multiple sequence alignments (*.fasta) and maximum likelihood phylogenies (Newick tree strings, *.nwk) for all available protein sequences corresponding to the eight genome segments of influenza A virus from the NCBI Genbank database. </p> <p>Each sequence is labelled with the Genbank accession number (e.g., "CY103884"), WHO strain identifier ("A/little yellow-shouldered bat/Guatemala/164/2009"), subtype label ("H17N10"), host species ("Sturnira lilium; gender M"), sampling location ("Guatemala: El Jobo"), and sample collection date ("May-2009"). These fields are separated by underscore characters.</p> <p>These data are provided under a Creative Commons license in support of a manuscript in progress, "Prospects for a sequence-based taxonomy of influenza A virus subtypes".</p>
ShareScore
40/100
Overall dataset sharing score
Score breakdown
These five areas show where the dataset supports — or may limit — practical reuse.
- Stewardship
- 8
- Harmonization
- 4
- Access
- 20
- Reuse readiness
- 8
- Engagement
- 0