Skip to main content
zenodoopen

Prospects for a sequence-based taxonomy of influenza A virus subtypes

<p>This dataset comprises the multiple sequence alignments (*.fasta) and maximum likelihood phylogenies (Newick tree strings, *.nwk) for all available protein sequences corresponding to the&nbsp;eight genome segments of influenza A virus from the NCBI Genbank database.&nbsp;</p> <p>Each sequence is labelled with the Genbank accession number (e.g., &quot;CY103884&quot;), WHO strain identifier (&quot;A/little yellow-shouldered bat/Guatemala/164/2009&quot;), subtype label (&quot;H17N10&quot;), host species (&quot;Sturnira lilium; gender M&quot;), sampling location (&quot;Guatemala: El Jobo&quot;), and sample collection date (&quot;May-2009&quot;).&nbsp; These fields are separated by underscore characters.</p> <p>These data are provided under a Creative Commons license&nbsp;in support of a manuscript in progress, &quot;Prospects for a sequence-based taxonomy of influenza A virus subtypes&quot;.</p>

ShareScore

40/100

Overall dataset sharing score

Score breakdown

These five areas show where the dataset supports — or may limit — practical reuse.

Stewardship
8
Harmonization
4
Access
20
Reuse readiness
8
Engagement
0

Topics