Compact and plannable world models
World-model work centers on smaller latent state and better planning signals. OneWM-VLA compresses each camera view and frame into one semantic token, then generates future latent tokens and action chunks together. It reports MetaWorld average success of 61.3% against 47.9% for π0, 98.1% average success on LIBERO, and 71.7% real Piper-arm success under clean conditions against 50.0% for π0.
RLA-WM uses Residual Latent Action (RLA), a compact code for DINO feature changes. It predicts future visual features with 3.5T FLOPs per inference and beats listed feature and flow baselines on ManiSkill and IWS prediction metrics. RC-aux adds reachability supervision to latent world models, showing that accurate short-horizon prediction can still mislead a planner. On Wall, it raises success to 83.6 ± 3.6 compared with 50.4 ± 6.5 for the LeWM control.