Semantic preservation during VLA fine-tuning
Two independent methods identify representation loss during behavior-cloning fine-tuning as a generalization failure. Semantic Anchoring aligns a shared action channel to a frozen text manifold and finds that alignment tracks out-of-distribution success with Spearman ρ=0.964. Anchor-Align instead distills frozen vision-language hidden states and adds motion-direction prediction on robot observations. It reaches 71.9% on LIBERO-PRO, versus 61.0% for its VLA-Adapter baseline. Together, the studies support preserving pretrained concepts while leaving room for execution-specific features.