Execution-level annotation for VLA demonstration datasets
VLA data teams should add a short execution annotation pass to manipulation demonstrations where the goal label hides important choices. The useful fields are operational: active arm, target object, approach direction, contact region, motion path, orientation, final configuration, and recovery behavior. FineVLA shows a practical version of the workflow: convert heterogeneous robot datasets into a shared format, cluster redundant demonstrations with dynamic time warping, then annotate a representative subset rather than every trajectory.
The payoff is better control over how a robot completes the same task. FineVLA reports that mixed fine-grained and raw goal labels performed best, with AlohaMix-OFT reaching 86.8% Easy and 82.5% Hard on RoboTwin. In real dual-arm manipulation, a 1:1 fine-grained-to-raw mix scored 62.7/100, compared with 49.9 for raw-only training. A cheap internal test is to pick one high-volume task with multiple valid executions, label a clustered subset, and measure compliance on the fields that operators can see, such as pose, approach direction, and contact point.