Physical feedback becomes a practical VLA input
MoSS makes the day's strongest empirical case for adding physical feedback directly into vision-language-action models. It keeps tactile and torque inputs in separate streams, then lets them interact with the action model through shared attention. On four real robot tasks, the full model lifts GR00T N1.5 from 20.8% average success to 49.0%, and pi_0 from 26.1% to 45.9%. The reported overhead is small at 1.11x for the dual-signal setup. The task mix matters here: cup unstacking, egg pick-and-place, board erase, and plug insertion all depend on contact cues that vision alone can miss.