Skip to main content
zenodoopen

DeepCanola: Phenotyping Brassica Pods Using Semi-Synthetic Data and Active Learning

<p>Dataset accomapnying the publication: <em>DeepCanola: Phenotyping Brassica Pods Using Semi-Synthetic Data and Active Learning&nbsp;</em>by Van Vliet, Atkins et al.</p> <p>We provide model weights, training and vlidation datasets, as well as phenotype outputs. Each file/folder is outlined below:</p> <ul> <li><em>deepcanola.pth -&nbsp;</em>Model weights for the final Model 4, named DeepCanola</li> <li><em>generated_datasets</em> - Datasets generated at each stage of the active learning process, datasets 1-4 <ul> <li>Each folder is an iteration of the active learning process, inside each folder are the generated images and associated annotations stored in the COCO format.</li> </ul> </li> <li><em>real_world_datasets -&nbsp;</em>Real-world datasets used for either creation of the pod pools or validation. Datasets include: <ul> <li><em>br9</em> - Ordered and disordered dataset of images with generated pod length data of the ordered images stored in the `br9_gt_lengths.csv` file</li> <li><em>br11</em> - Ordered and disordered dataset of images only</li> <li><em>br17</em> - Ordered dataset with ground-truth of images with length annotations collected in ImageJ and stored in the `BR017 POD SCAN DATA.csv` file.</li> <li><em>misc</em> - Dataset of m<span>iscellaneous images including Brassica napus from Rothamstead and brassica relatives</span></li> </ul> </li> <li><em>data_generation_pools</em> - Pools used to generate semi-synthetic data at each step of the active learning process. Pools include: <ul> <li><em>background_pool</em> - Created background images to be selected at random by the semi-synthetic data generation script</li> <li><em>pod_pools -&nbsp;</em>Pools of pods used in the semi-synthetic data generation process. Pod pools include: <ul> <li><em>br9 -&nbsp;</em>673 pods with associated masks</li> <li><em>br9 and br17 -&nbsp;</em>673 + 332 pods with associated masks</li> </ul> </li> </ul> </li> <li><em>deepcanola_outputs - </em>Phenotype data outputs generated by DeepCanola. Each output is stored as a .csv file of both length measurements of each pod (with <em>_objects.csv</em> suffix), and average length measurements per image (with <em>_averages.csv</em> suffix). Outputs include: <ul> <li><em>br9_ordered</em></li> <li><em>br9_disordered</em></li> <li><em>br17</em></li> </ul> </li> </ul>

ShareScore

32/100

Overall dataset sharing score

Score breakdown

These five areas show where the dataset supports — or may limit — practical reuse.

Stewardship
4
Harmonization
4
Access
16
Reuse readiness
8
Engagement
0