Skip to main content
zenodoopen

K-mer collision statistics (BLEND: A Fast, Memory-Efficient, and Accurate Mechanism to Find Fuzzy Seed Matches in Genome Analysis)

<p>This dataset&nbsp;contains 1,077 FASTA files and CSV files. Each FASTA file includes 25-character long sequences similar to each other.</p> <p>We have a CSV file for each tool (i.e., minimap2 and BLEND)&nbsp;and configuration (i.e., different number of neighbors in BLEND).&nbsp;CSV files include the&nbsp;non-identical&nbsp;k-mer pairs (16-mers)&nbsp;that generate the same hash value (i.e., collisions). These k-mers are extracted from sequences that are similar to each other. In each line, we show the hash value of the k-mers, the actual sequene pairs&nbsp;that the k-mers are extracted from, k-mer pairs that generate the same hash value, and the edit distance between these k-mers.</p> <p>&nbsp;</p>

ShareScore

32/100

Overall dataset sharing score

Score breakdown

These five areas show where the dataset supports — or may limit — practical reuse.

Stewardship
4
Harmonization
4
Access
16
Reuse readiness
8
Engagement
0