Tactile and torque adapter retrofits for contact-heavy VLA tasks
Adding tactile and torque inputs to an existing VLA now looks like a practical upgrade for teams working on contact-heavy manipulation. The clearest target is a retrofit around tasks where camera views miss the deciding event: grasp stability on fragile items, contact onset during board wiping, or misalignment during plug insertion. MoSS keeps the physical signals in separate streams and joins them to the action model with shared attention, which matters because the gain did not come from a generic sensor concatenation story. On four real-robot tasks, the full setup lifted GR00T N1.5 from 20.8% average success to 49.0%, and pi_0 from 26.1% to 45.9%, with reported inference cost rising only to 1.11x for the dual-signal version.
The near-term build is a sensor add-on path for one or two failure-prone skills, not a full policy rewrite. A team already running GR00T, pi_0, or a similar diffusion-style VLA could attach fingertip tactile sensing and joint torque logging, keep those streams separate in the adapter, and test on a narrow set of contact-driven tasks. The cheap check is simple: compare vision-only, tactile-only, torque-only, and dual-signal variants on the same task mix. If the pattern matches the paper, the combined model should beat each single-signal version on the tasks where contact timing and force correction matter most.