World-model training and adaptation
World models are now used as training spaces and control priors, with clear pressure to reduce task-specific robot data. HarmoWAM combines a video world model with predictive and reactive action experts, then routes between transit and interaction phases. Its reported out-of-domain gains are 33 percentage points over prior VLA models and 29 points over prior World Action Models (WAMs) across six real-world tasks.
RAW-Dream gives the week’s clearest data-efficiency claim. It trains OpenVLA-OFT with reinforcement learning inside a task-agnostic video world model and uses Qwen3-VL for reward judgments. On LIBERO, it raises a 1-shot supervised fine-tuning baseline from 43.4% to 52.3% with 10 target demonstrations and no target rollouts for world-model training; with in-domain world-model tuning, it reaches 66.0%.