Skip to main content
zenodoopen

Vantage-point search (vpsearch) dataset and index

<p>This is a dataset for use with the <a href="https://github.com/enthought/vpsearch">vpsearch</a> tool.</p> <p>This dataset contains the following files:</p> <ol> <li>A compressed sequence file,&nbsp;bac120_ssu_reps_r207-sliced-dedup.fa.gz, which was obtained from the <a href="https://gtdb.ecogenomic.org/">GTDB</a> dataset of bacterial sequences (v207) by extracting the v3-v4 hypervariable region and removing duplicate sequences.</li> <li>A vantage-point tree, built with version 0.1.2 of the vpsearch software. Note: the vantage-point tree needs to be decompressed (tar xzvf&nbsp;bac120_ssu_reps_r207-sliced-dedup.db.tar.gz) before it can be used for querying.</li> <li>A sample query file, query.fa, containing the sequence&nbsp;NR_126253.1 from RefSeq.</li> </ol> <p>The scripts used to prepare the dataset can be found in the <a href="https://github.com/enthought/vpsearch">vpsearch</a> GitHub repository. Also included in the repository is a description of how the primary data was obtained.</p>

ShareScore

36/100

Overall dataset sharing score

Score breakdown

These five areas show where the dataset supports — or may limit — practical reuse.

Stewardship
4
Harmonization
4
Access
16
Reuse readiness
8
Engagement
4