Visual action representations
RoboInter1.5 and Masked Visual Actions both place spatially explicit signals between intent and predicted outcomes. RoboInter1.5 supplies object grounding, affordances, contact points, and motion traces across more than 230,000 episodes. Masked Visual Actions instead exposes pixel-space entity trajectories to a pretrained video model; one checkpoint can predict scene responses or infer robot motion from desired object movement. On DROID, it reports LPIPS of 0.0945 versus 0.362 for Ctrl-World. Together, the papers support visual structure as an embodiment-flexible control interface, although RoboInter1.5’s inspected excerpt does not provide downstream comparison metrics.