3D grounding for VLA manipulation
Spatial alignment is a main performance lever. Dex-BEV puts visual geometry, proprioception, and output actions into a shared bird’s-eye-view coordinate frame. It reports 97.8% average success on official LIBERO, 76.0% on RoboTwin 2.0 Clean, and 89.9% on modified LIBERO camera and pose settings where listed 2D baselines fall below 10%.
GeoAlign adds RGB-derived geometry features queried by robot state. The gains are clearest on geometry-sensitive tasks: real ALOHA average success is 78.8%, versus 65.0% for the RGB-only baseline, with transparent-bottle success at 75.0% versus 35.0%. 3DThinkVLA takes a lighter deployment route. It trains latent 3D perception and reasoning adapters, then runs on 2D images at inference while reaching 98.7% on LIBERO and 81.0% on LIBERO-PLUS.