Simple VLA baselines hold up under controlled evaluation
StarVLA-α makes the clearest period-level claim: a plain vision-language model (VLM) plus a small MLP action head can match or beat heavier vision-language-action systems when the recipe is controlled. Its specialist model reports 98.8 average on LIBERO, 88.2 on RoboTwin 2.0 clean*, and 53.8 on RoboCasa-GR1, with the simple MLP head matching or beating more complex heads in the same setup. The ablations are as important as the headline numbers. Extra robot pretraining helps some benchmarks and hurts others, and common pipeline additions give only small gains once task data is large. This day’s strongest concrete result is not a new stack. It is a cleaner accounting of what actually changes outcomes.