Predictive supervision
Future scene changes are becoming direct training targets for control. Xiaomi-Robotics-U0 generates embodied observations and manipulation videos; adding its synthetic data raises out-of-distribution policy success from 36.9% to 63.2%. WALA learns semantic and geometric future deltas from action-free video, then connects those latent changes to executable actions. It reaches 75.2% average success on RoboCasa, versus 54.2% for its base policy. Lumo-2 follows the same design logic by predicting action-relevant latent dynamics before producing action chunks, though its excerpt does not provide numerical benchmark margins.