Nominal FAST5/FASTQ Evaluation Data Set
<p>FAST5/FASTQ data used for accuracy characterization of decoding techniques applied to the HEDGEs DNA-information storage code. FASTQ data is used to evaluate the hard-decoding algorithm as explained by Press et al. in (<a href="https://doi.org/10.1073/pnas.2004821117">https://doi.org/10.1073/pnas.2004821117</a>). FAST5 data is used in evaluation for both our novel Alignment Matrix soft decoder (<a href="https://doi.org/10.5281/zenodo.11454877">https://doi.org/10.5281/zenodo.11454877</a>), and the soft decoder developed by Chandak et al. in the publication (<a href="https://doi-org.prox.lib.ncsu.edu/10.1109/ICASSP40776.2020.9053441">10.1109/ICASSP40776.2020.9053441</a>). Our code repository at <a href="https://doi.org/10.5281/zenodo.11454877">https://doi.org/10.5281/zenodo.11454877</a> includes a GPU accelerated adaptation of Chandak et al.’s algorithm in order to scale analysis on the submitted FAST5 data, and this is the version of code used to evaluate the algorithm’s accuracy and runtime overhead.</p> <p> </p> <p>Within the archive there are several sub-archives. Explanations for each sub-archive can be found for the corresponding archive name within the README.md file.</p> <p> </p>
ShareScore
32/100
Overall dataset sharing score
Score breakdown
These five areas show where the dataset supports — or may limit — practical reuse.
- Stewardship
- 4
- Harmonization
- 4
- Access
- 16
- Reuse readiness
- 8
- Engagement
- 0