State tracking in vision-language navigation
Dual-Anchoring makes long-horizon vision-language navigation more explicit about task state. The core idea is simple: force the model to state which instruction sub-goals are done, and force it to retain a landmark-level memory of where it has been. That supervision is large-scale, with 3.6M progress samples and 937K grounded landmark samples. The reported gains are strong on continuous-environment VLN benchmarks: success rate reaches 65.6 on R2R-CE and 61.7 on RxR-CE, with about +8.7 and +8.8 points over StreamVLN. In this small period, that makes memory and progress tracking the clearest algorithmic result.