VLA manipulation fine-tuning
VLA manipulation papers put more weight on real execution metrics. EXPO-FT keeps a pretrained π0.5 policy and trains a lightweight edit policy with off-policy reinforcement learning. It reports 30/30 final successes on each of 8 real-world manipulation tasks after an average of 19.1 minutes of online robot data.
OASIS attacks the action decoding problem directly. It predicts an 8-step SE(3) end-effector trajectory, meaning 3D position and rotation, before producing 6-DoF actions and gripper commands. The paper reports 97.6% average success on LIBERO and 89.2% average success in real-world tests on Franka Research 3 and Kinova Gen3 robots.