Control quality now depends on the action interface as much as the perception stack
Robot control papers are getting more explicit about where action quality is lost and how to recover it. The clearest evidence comes from two directions. One paper shows that better visual encoders do not reliably help a vision-language-action policy when actions are squeezed through discrete tokens. On LIBERO-10, Diffusion Policy climbed from 36.4% to 57.6% with a ResNet-18 to SigLIP upgrade at size M, while OAT moved from 53.8% to 57.4%. Another paper improves long-horizon world-model control by splitting planning into latent macro-actions and low-level actions. HWM raised real Franka pick-and-place success from 0% to 70% and drawer tasks from 30% to 70%, with lower planning cost claims in the same report.