Data curation for VLA pretraining
EmbodiedMidtrain makes data selection a first-class part of VLA training. The paper measures a real mismatch between generic vision-language model data and robot trajectories, then mid-trains on samples scored as closer to robot data. The gains are large for small backbones: InternVL3.5-1B rises from 36.5 to 56.3 success on SimplerEnv-Bridge and from 39.0 to 54.2 on Libero-10. Qwen3VL-2B also improves across Calvin, SimplerEnv-Bridge, and Libero-10. This gives the day a concrete message: better robot policies are coming from better pre-action data alignment, not only larger action heads or more robot fine-tuning.