Steerable VLA policies
FineVLA targets a specific weakness in robot datasets: many trajectories say what task was completed, but leave out how it was done. The work adds human-verified execution instructions for arm choice, approach direction, contact region, final pose, and other action details across 47,159 selected trajectories. Training with a mix of fine-grained and raw labels gives the best reported results, with AlohaMix-OFT reaching 86.8% Easy and 82.5% Hard on RoboTwin. In real dual-arm manipulation, the same line of evidence reports 62.7/100 for a 1:1 fine-grained-to-raw mix, compared with 49.9 for raw-only training.