Generalist VLA policies
Qwen-VLA is the scale-and-unification anchor. It uses Qwen3.5-4B for vision-language understanding and a DiT flow-matching action decoder for continuous actions. The model reads robot descriptions, images, and instructions, then predicts action or trajectory chunks across manipulation, navigation, and trajectory prediction. Reported results span LIBERO, Simpler-WidowX, RoboTwin, R2R, RxR, ALOHA out-of-distribution trials, and DOMINO dynamic manipulation.
VLA-Pro takes a modular route to transfer. It stores task-specific LoRA adapters with structured procedural states, retrieves related memories at inference, and fuses the weights for the current action chunk. The gains are large in the reported settings: π0.5 real-world success on six held-out UR7e tasks rises from 5.8% to 65.0%, and RoboTwin results improve across X-VLA, RDT, and π0.5 backbones.