Vantage-point search (vpsearch) dataset and index
<p>This is a dataset for use with the <a href="https://github.com/enthought/vpsearch">vpsearch</a> tool.</p> <p>This dataset contains the following files:</p> <ol> <li>A compressed sequence file, bac120_ssu_reps_r207-sliced-dedup.fa.gz, which was obtained from the <a href="https://gtdb.ecogenomic.org/">GTDB</a> dataset of bacterial sequences (v207) by extracting the v3-v4 hypervariable region and removing duplicate sequences.</li> <li>A vantage-point tree, built with version 0.1.2 of the vpsearch software. Note: the vantage-point tree needs to be decompressed (tar xzvf bac120_ssu_reps_r207-sliced-dedup.db.tar.gz) before it can be used for querying.</li> <li>A sample query file, query.fa, containing the sequence NR_126253.1 from RefSeq.</li> </ol> <p>The scripts used to prepare the dataset can be found in the <a href="https://github.com/enthought/vpsearch">vpsearch</a> GitHub repository. Also included in the repository is a description of how the primary data was obtained.</p>
ShareScore
36/100
Overall dataset sharing score
Score breakdown
These five areas show where the dataset supports — or may limit — practical reuse.
- Stewardship
- 4
- Harmonization
- 4
- Access
- 16
- Reuse readiness
- 8
- Engagement
- 4