DeepCanola: Phenotyping Brassica Pods Using Semi-Synthetic Data and Active Learning
<p>Dataset accomapnying the publication: <em>DeepCanola: Phenotyping Brassica Pods Using Semi-Synthetic Data and Active Learning </em>by Van Vliet, Atkins et al.</p> <p>We provide model weights, training and vlidation datasets, as well as phenotype outputs. Each file/folder is outlined below:</p> <ul> <li><em>deepcanola.pth - </em>Model weights for the final Model 4, named DeepCanola</li> <li><em>generated_datasets</em> - Datasets generated at each stage of the active learning process, datasets 1-4 <ul> <li>Each folder is an iteration of the active learning process, inside each folder are the generated images and associated annotations stored in the COCO format.</li> </ul> </li> <li><em>real_world_datasets - </em>Real-world datasets used for either creation of the pod pools or validation. Datasets include: <ul> <li><em>br9</em> - Ordered and disordered dataset of images with generated pod length data of the ordered images stored in the `br9_gt_lengths.csv` file</li> <li><em>br11</em> - Ordered and disordered dataset of images only</li> <li><em>br17</em> - Ordered dataset with ground-truth of images with length annotations collected in ImageJ and stored in the `BR017 POD SCAN DATA.csv` file.</li> <li><em>misc</em> - Dataset of m<span>iscellaneous images including Brassica napus from Rothamstead and brassica relatives</span></li> </ul> </li> <li><em>data_generation_pools</em> - Pools used to generate semi-synthetic data at each step of the active learning process. Pools include: <ul> <li><em>background_pool</em> - Created background images to be selected at random by the semi-synthetic data generation script</li> <li><em>pod_pools - </em>Pools of pods used in the semi-synthetic data generation process. Pod pools include: <ul> <li><em>br9 - </em>673 pods with associated masks</li> <li><em>br9 and br17 - </em>673 + 332 pods with associated masks</li> </ul> </li> </ul> </li> <li><em>deepcanola_outputs - </em>Phenotype data outputs generated by DeepCanola. Each output is stored as a .csv file of both length measurements of each pod (with <em>_objects.csv</em> suffix), and average length measurements per image (with <em>_averages.csv</em> suffix). Outputs include: <ul> <li><em>br9_ordered</em></li> <li><em>br9_disordered</em></li> <li><em>br17</em></li> </ul> </li> </ul>
ShareScore
32/100
Overall dataset sharing score
Score breakdown
These five areas show where the dataset supports — or may limit — practical reuse.
- Stewardship
- 4
- Harmonization
- 4
- Access
- 16
- Reuse readiness
- 8
- Engagement
- 0