FOR-species20K dataset and code for review
<h1>Description</h1> <p>Data and code (code.zip) corresponding to the manuscript entitled "Benchmarking tree species classification from proximally-sensed laser scanning data: introducing the FOR-species20K dataset"</p> <h1>Code</h1> <p>The code folder contains all of the code to train and predict using the methods benchmarked in the manuscript</p> <h1>Data split and usage</h1> <p>The data is split into:</p> <ul> <li><strong>Development data (dev)</strong>: these includes 90% of the trees in the dataset and consists of individual tree point clouds (*.laz) named according to the <em>treeID </em>column available in the tree_metadata_dev.csv file, from which <em>tree_species </em>labels are available. These data are meant to be used for model development and can thus be further split into training and validation datasets.</li> <li><strong>Test data (test)</strong>: these are 10% of the trees (balanced sample) and include individual tree point clouds (*.laz) but, for benchmarking purposes, the species labels are witheld for benchmarking purposes. Thus to make use of the test data the users should predict species on the test trees, and output a table (.csv file) with a row per predicted tree and two columns (<em>treeID </em>and <em>predicted_species</em>). This table can then be used to create a new submission in the FOR-species20K Codabench benchmarking platform and obtain the evaluation metrics corresponding to the test data.</li> </ul>
ShareScore
8/100
Overall dataset sharing score
Score breakdown
These five areas show where the dataset supports — or may limit — practical reuse.
- Stewardship
- 4
- Harmonization
- 4
- Access
- 0
- Reuse readiness
- 0
- Engagement
- 0