CAD 120 affordance dataset
<p>% ==============================================================================<br> % CAD 120 Affordance Dataset<br> % Version 1.0<br> % ------------------------------------------------------------------------------<br> % If you use the dataset please cite:<br> %<br> % Johann Sawatzky, Abhilash Srikantha, Juergen Gall.<br> % Weakly Supervised Affordance Detection.<br> % IEEE Conference on Computer Vision and Pattern Recognition (CVPR'17)<br> %<br> % and<br> %<br> % H. S. Koppula and A. Saxena.<br> % Physically grounded spatio-temporal object affordances.<br> % European Conference on Computer Vision (ECCV'14)<br> %<br> % Any bugs or questions, please email sawatzky AT iai DOT uni-bonn DOT de.<br> % ==============================================================================</p> <p>This is the CAD 120 Affordance Segmentation Dataset based on the Cornell Activity<br> Dataset CAD 120 (see http://pr.cs.cornell.edu/humanactivities/data.php).</p> <p>Content</p> <p>frames/*.png:<br> RGB frames selected from Cornell Activity Dataset. To find out the location of the frame<br> in the original videos, see video_info.txt.</p> <p>object_crop_images/*.png<br> image crops taken from the selected frames and resized to 321*321. Each crop is a padded<br> bounding box of an object the human interacts with in the video. Due to the padding,<br> the crops may contain background and other objects.<br> In each selected frame, each bounding box was processed. The bounding boxes are already<br> given in the Cornell Activity Dataset.<br> The 5-digit number gives the frame number, the second number gives the bounding box number<br> within the frame.</p> <p>segmentation_mat/*.mat<br> 321*321*6 segmentation masks for the image crops. Each channel corresponds to an<br> affordance (openabe, cuttable, pourable, containable, supportable, holdable, in this order).<br> All pixels belonging to a particular affordance are labeled 1 in the respective channel,<br> otherwise 0. </p> <p>segmentation_png/*.png<br> 321*321 png images, each containing the binary mask for one of the affordances.</p> <p>lists/*.txt<br> Lists containing the train and test sets for two splits. The actor split ensures that<br> train and test images stem from different videos with different actors while the object split ensures<br> that train and test data have no (central) object classes in common.<br> The train sets are additionally subdivided into 3 subsets A,B and C. For the actor split,<br> the subsets stem from different videos. For the object split, each subset contains<br> every third crop of the train set.</p> <p>crop_coordinate_info.txt<br> Maps image crops to their coordinates in the frames.</p> <p>hpose_info.txt<br> Maps frames to 2d human pose coordinates. Hand annotated by us.</p> <p>object_info.txt<br> Maps image crops to the (central) object it contains.</p> <p>visible_affordance_info.txt<br> Maps image crops to affordances visible in this crop</p> <p> </p> <p>%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%55<br> The crops contain the following object classes:<br> 1.table<br> 2.kettle<br> 3.plate<br> 4.bottle<br> 5.thermal cup<br> 6.knife<br> 7.medicine box<br> 8.can<br> 9.microwave<br> 10.paper box<br> 11.bowl<br> 12.mug</p> <p>Affordances in our set:<br> 1.openable<br> 2.cuttable<br> 3.pourable<br> 4.containable<br> 5.supportable<br> 6.holdable</p> <p>Note that our object affordance labeling differs from the Cornell Activity Dataset:<br> E.g. the cap of a pizza box is considered to be supportable.</p> <p> </p>
ShareScore
44/100
Overall dataset sharing score
Score breakdown
These five areas show where the dataset supports — or may limit — practical reuse.
- Stewardship
- 8
- Harmonization
- 8
- Access
- 16
- Reuse readiness
- 8
- Engagement
- 4