Skip to main content
zenodoopen

Simulated Illumina metagenomic reads

<p>We simulated metagenomic Illumina sequencing reads to a mixture ratio that approximates that found in patient sputa, albeit with a slightly higher mycobacterial component. In total, 0.9 gigabases were generated, at proportions: 46% each for bacteria and human, 6\% <em>Mycobacterium tuberculosis </em>complex (MTBC), and 1% each for virus and non-tuberculous mycobacteria (NTM).</p> <p>The reference genomes that reads were simulated from for these groups were gathered as follows. The references for the virus group were obtained using kraken&#39;s (v2.1.2) --download-library functionality. The viral library was downloaded on June 15 2023. The human genome from which the reads were simulated was KOREF_S1v2.1 (RefSeq accession GCA_020497085.1), with contigs shorter than 10kbp removed. The bacterial references were obtained by first downloading the bacteria library through kraken, followed by a subsampling due to the size (166Gb) of the resulting FASTA file. We subsampled the file by first removing sequences with a length &lt;50kbp. We then extracted each sequence into its own FASTA file under a directory for the genus of the sequence - excluding the <em>Mycobacterium</em> genus. Genera were randomly subsampled to contain a maximum of 1000 assemblies. Each genus was then reduced to a representative subset using Assembly Dereplicator (commit 2dfcb14; https://github.com/rrwick/Assembly-Dereplicator) by keeping only 10% of the assemblies for each genus (-f 0.1). The NTM references selected were <em>M. abscessus</em> (accession GCF_017190695.1), <em>M. avium</em> (GCF_020735285.1), <em>M. kansasii</em> (GCA_014701265.1), <em>M. ulcerans</em> (GCF_000013925.1), <em>M. intracellulare</em> (GCF_016756075.1), <em>M. terrae</em> (GCF_010727125.1), and <em>M. fortuitum</em> (GCF_001307545.1). The MTBC reference is a lineage 1 assembly (GCF_932530395.1).</p> <p>Illumina reads were simulated with ART (v2016.06.05). We simulated paired reads from a MiSeq v3 system (-ss MSv3) with a read length of 150, a mean fragment length of 250 and fragment length standard deviation 10 (-l 150 -m 250 -s 10).</p> <p>We removed simulated Illumina reads with any ambiguous base.</p>

ShareScore

36/100

Overall dataset sharing score

Score breakdown

These five areas show where the dataset supports — or may limit — practical reuse.

Stewardship
8
Harmonization
4
Access
16
Reuse readiness
8
Engagement
0

Topics