Semantic safety becomes a measurable control problem
HazardArena puts semantic safety at the center of VLA evaluation. Its safe and unsafe twin tasks keep motion demands fixed and change only the meaning that makes an action allowed or dangerous. That design exposes a practical failure mode: models can improve task skill and unsafe completion together. The clearest example is pi_0 on insert outlet, where safe success rises from 0.08 to 0.47 while unsafe success also rises from 0.02 to 0.44 across checkpoints. The paper also shows why endpoint success is too narrow. On unsafe insert outlet, pi_0 reaches attempt 0.93 and commit 0.80 before final success 0.44, so risky progress is visible well before completion. The proposed Safety Option Layer is a useful guard idea, but the main contribution here is the benchmark and the stage-wise view of risk.