Open VLA models and robot datasets
MolmoAct2 is the clearest deployment-oriented release in the period. The paper describes an open VLA system with released weights, code, and training data. Its backbone, Molmo2-ER, is a 4B vision-language model (VLM) trained on a 3.3M-sample embodied-reasoning corpus, then connected to robot actions through an action tokenizer and a continuous action expert.
The data release is a large part of the claim. The authors report 720 hours of bimanual YAM data, a filtered SO-100/101 community dataset with 38,059 episodes, and a filtered DROID Franka subset with 74,604 successful episodes. They also report 63.8% average performance across 13 embodied-reasoning benchmarks for Molmo2-ER, with a 17-point gain over Molmo2. The provided excerpt says MolmoAct2 beats strong baselines across simulation and real-world benchmarks, but it does not include the underlying task success rates.