Persistent state for closed-loop manipulation
Current images often omit the information needed to finish a physical task. POT-VLA keeps role-indexed 3D object records across walking, contact, occlusion, and recovery, reaching 71/80 real-world successes versus 39/80 for its matched baseline. FM-VLA instead compresses wrist force history into eight memory tokens; it averages 83.3% across three contact-rich tasks, compared with 27.8% for a memoryless policy. Together, the studies show that useful memory should preserve physical events and entities, not merely add more past images.