W4A4 acceptance tests for robot VLA policies
Robotics teams trying to run Pi 0.5 or GR00T-class policies close to the robot should add a 4-bit policy acceptance test before buying more inference hardware. The test is concrete: quantize the language backbone and diffusion action head to uniform W4A4, calibrate DiT activation scales on a small set of unlabeled trajectories, then compare against the FP16 policy on the team’s manipulation suite.
Ω-QVLA reports Pi 0.5 W4A4 at 98.0% average LIBERO success versus 97.1% FP16 and GR00T N1.5 W4A4 at 87.8% versus 87.0%, with a 71.3% static memory-footprint cut. The practical gate should include task success, action smoothness, memory use, and real robot progress, since the same paper reports ARX R5 dual-arm progress close to FP16 and a much larger gap over QuantVLA.