Robust VLA training becomes a concrete evaluation target
STRONG-VLA centers this day on failure tolerance in embodied control. The paper argues that robustness training should be split into two steps: first learn under perturbations, then recover clean-task fidelity. The evidence is concrete. On LIBERO, gains reach +12.60% seen and +7.77% unseen for OpenVLA, +14.48% and +13.81% for OpenVLA-OFT, and +16.49% and +5.58% for pi0. Clean performance stays close to baseline. The benchmark also matters. It covers 28 perturbation types across text and vision, including held-out tests such as semantic drift and dynamic visual artifacts. This keeps the paper tied to deployment problems like occlusion, instruction corruption, and sensor noise, not just synthetic stress tests.