Sample-efficient adaptation with compact control signals
Real-robot adaptation is getting more targeted. RL Token keeps a pretrained vision-language-action model frozen, exposes a compact state for reinforcement learning, and updates only a small actor-critic online. The payoff is concrete on precision work: up to 3× faster execution on the hardest phase and a screw insertion gain from 20% to 65% after minutes to a few hours of practice. GazeVLA attacks the same data bottleneck from another side. It uses human gaze as an intention signal, pretrains on more than 150M egocentric frames, and reports stronger few-shot transfer with only 10 robot trajectories and 50 human trajectories per task, including 85% success on simple pick-and-place and a reported 2× gain over pi0.5 on screw tightening.