Skip to main content
zenodoopen

Nominal FAST5/FASTQ Evaluation Data Set

<p>FAST5/FASTQ data used for accuracy characterization of decoding techniques applied to the HEDGEs DNA-information storage code. FASTQ data is used to evaluate the hard-decoding algorithm as explained by Press et al. in (<a href="https://doi.org/10.1073/pnas.2004821117">https://doi.org/10.1073/pnas.2004821117</a>). FAST5 data is used in evaluation for both our novel Alignment Matrix soft decoder (<a href="https://doi.org/10.5281/zenodo.11454877">https://doi.org/10.5281/zenodo.11454877</a>), and the soft decoder developed by Chandak et al. in the publication (<a href="https://doi-org.prox.lib.ncsu.edu/10.1109/ICASSP40776.2020.9053441">10.1109/ICASSP40776.2020.9053441</a>). Our code repository at <a href="https://doi.org/10.5281/zenodo.11454877">https://doi.org/10.5281/zenodo.11454877</a> includes a GPU accelerated adaptation of Chandak et al.&rsquo;s algorithm in order to scale analysis on the submitted FAST5 data, and this is the version of code used to evaluate the algorithm&rsquo;s accuracy and runtime overhead.</p> <p>&nbsp;</p> <p>Within the archive there are several sub-archives. Explanations for each sub-archive can be found for the corresponding archive name within the README.md file.</p> <p>&nbsp;</p>

ShareScore

32/100

Overall dataset sharing score

Score breakdown

These five areas show where the dataset supports — or may limit — practical reuse.

Stewardship
4
Harmonization
4
Access
16
Reuse readiness
8
Engagement
0