Pre-deployment smoke-test routing for frozen VLA experts
Robot teams that already run smoke tests before deployment can keep those rollouts as selection data for each task and perturbation. RouterVLA shows a simple probe-success rule selecting among frozen experts reached 0.6149 held-out success on LIBERO-Plus, compared with 0.4686 for the global best expert. The result is useful because learned scorers added little over the simple rule, while same-trial reuse inflated measured gains.
A practical version would store each candidate checkpoint’s probe outcomes, rollout length, duration, termination behavior, and missing-statistic flags, then choose a policy for the target condition with a trial-disjoint validation rule. The cheap check is to replay the lab’s existing commissioning logs through this rule and compare against the single-checkpoint choice on held-out rollouts.