Structured action generation
Several papers add control structure inside the policy output rather than treating robot action as a flat vector. EquiVLA builds rotation equivariance into a frozen vision-language backbone plus a flow-matching action head, reaching 92.6% average success on LIBERO relative control and 72% average success on five Mobile ALOHA real-robot tasks. Co-VLA separates bimanual action into shared coordination and per-arm residual latents, with its largest reported gain on Handover Block Easy: 91% success versus 64% for the π0 baseline. FAFM predicts frequency coefficients for continuous trajectories, which targets mixed control rates and smoother motion without adding network parameters.