Source note
Co-VLA: Coordination-Aware Structured Action Modeling for Dual-Arm Vision-Language-Action Systems
Vision Language ActionDual Arm ManipulationRobot Foundation ModelStructured Action ModelingBimanual Coordination
Summary
Co-VLA adds explicit dual-arm coordination structure to a VLA action head and uses that structure during execution. It reports the largest gains on tightly coupled bimanual tasks, with smaller gains under heavy simulation randomization.
Problem
- Standard VLA policies often output one concatenated dual-arm action vector, so timing, role split, and safety-related motion smoothing are learned implicitly.
- This matters for bimanual manipulation because tasks such as handover, lifting, and joint transport need synchronized motion and asymmetric arm roles.
- Implicit coordination is hard to inspect or adjust at deployment time when actions jitter, desynchronize, or collide.
Approach
- Co-VLA keeps a pretrained VLA backbone, based on in the excerpt, and replaces the monolithic action head with a Structured Action Expert.
- The Structured Action Expert predicts one shared latent for task-level coordination and two residual latents for left-arm and right-arm adjustments.
- Final 7-DoF joint velocity commands for each arm are the sum of shared and residual action components.
- Task-adaptive auxiliary losses shape the decomposition: sparse residual loss for near-symmetric motion, shared mean velocity loss for asymmetric roles, and temporal synchronization loss for coupled timing. The auxiliary weight is .
- A Latent-Aware Controller reads shared and residual action energies at deployment, then low-pass filters joint commands with adaptive stiffness to preserve coordinated micro-adjustments and suppress jitter. It does not require force sensing or impedance control.
Results
- On RoboTwin 2.0 Easy settings across 8 selected bimanual tasks, average success increased to 82% for Co-VLA, compared with 76% for and 73% for .
- On RoboTwin 2.0 Hard settings, average success was 22% for Co-VLA, compared with 21% for and 21.9% for , so the hard-setting aggregate gain is small.
- On Handover Block Easy, Co-VLA reached 91% success, compared with 64% for and 44% for , a +27 point gain over .
- The abstract reports a 27% success-rate gain in tight-coordination tasks and more than 2x OOD real-world improvement, from 13% to 27%.
- The abstract reports task completion time reductions of up to 25%.
- The training setup used 1,000 successful demonstrations per simulated task, 100 evaluation rollouts per setting, 1,000 SAE warm-up steps, 30,000 full fine-tuning steps, batch size 32, and 4 GPUs with FSDP.