World models for staged manipulation
HarmoWAM treats manipulation as a phase-dependent control problem. A video world model predicts future frames, while two action experts handle different parts of the task: a reactive expert for transit and a predictive expert for precise interaction. A learned gate chooses the expert during inference.
The paper reports tests on six real-world tasks with background, position, and object variation. In out-of-distribution settings, HarmoWAM claims average gains of 33 percentage points over prior VLA models and 29 points over prior World Action Models. The key detail is the separation between reaching and contact control, which the paper measures directly in its motivation study.