Annotations
Structured, layered annotation
Raw footage is just the starting point. Every project can be layered with narration, action labels, object state, hand-object interaction, and more — as deep as the use case demands.
Every tier adds a layer on top of the last
Tier 1
Raw narrated video — head-mounted POV, contributor narration, QA-passed
Tier 2
+ Action segments + Object state change
Tier 3
+ HOI (Hand Object Interaction) — frame-by-frame hand contact log
Tier 4
+ Phase decomposition, teleoperation demonstrations
Tier 5
+ Gaze data (hardware-dependent)
Narration
Timestamped contributor description synchronized to video.
Action Segments
Verb + noun + start/end timestamps for each action.
Object State Change
Before/after states of objects (bolt: loose to tight).
HOI
Frame-by-frame: which hand, what object, grip type, contact state.
Gaze
Fixation and saccade data, x/y coordinates per frame.
Bounding Boxes
Spatial object localization per frame.
Keypoints
Hand and body landmark coordinates (21-point hand model).