Explicit spatial grounding for manipulation
Several robot papers add intermediate target signals between language and motor control. AVP makes the vision-language model output visual primitives before action prediction, then feeds those tokens to a flow-matching action expert. On Chinese chess manipulation, it reports 90.28% average success versus 62.67% for π₀.₅, with 0.27 seconds per instruction.
GesVLA handles a different source of ambiguity: human pointing. It turns wrist and index-finger keypoints into gesture tokens and combines them with language and scene perception. Across three real-robot tasks, success reaches 83.3%, compared with 31.7% for a text-only VLA. SOMA adds persistent spatial memory for objects outside the current camera view, using head-camera scans and memory tokens with semantic and 3D position data. Its gains are smaller in success rate, but it cuts target search time and grasp attempts across out-of-vision tasks.