Selective unlearning reaches robot policy level
Safety work in vision-language-action models is getting more concrete. VLA-Forget edits the visual encoder, cross-modal projector, and upper action layers in stages, then uses retain, forget, mismatch, and feature-preservation losses to remove unwanted behavior without wiping out task skill. The reported numbers are strong for this kind of intervention: on OpenVLA-7B with Open X-Embodiment it posts FC 93, RC 91, and TSR 78, with lower safety-violation recovery after quantization than NPO and much better retention than GA. That matters because the failure target is a robot action, not just a text output.