Efficient policy post-training
Learning from Hindsight relabels failed robot rollouts with the behavior they actually completed, then scores those trajectories against the new instruction. For vision-language-action (VLA) post-training, this keeps 70%–80% of trajectory groups usable, versus 20%–40% with standard methods. On out-of-distribution LIBERO-PRO tasks, it reaches standard training’s final performance in about five steps instead of nearly 30. With 160 physical-robot rollouts, success reaches 56%, compared with 22% for standard group-relative policy optimization.