Action representations and critical timesteps
Several papers attack the action stream itself. RotVLA encodes frame transitions as continuous SO(n) latent actions and composes them during training. It reports 98.2% average success on LIBERO and 89.6% / 88.5% on RoboTwin2.0 clean and randomized settings after pretraining on more than 1700 hours of robot and human video.
FrameSkip and AttenA+ make a related claim at the data and loss levels. FrameSkip keeps 20% of unique trajectory frames, with priority for alignment, contact, grasp closure, and release. Its macro-average success across RoboCasa-GR1, SimplerEnv, and LIBERO rises from 66.50% to 76.15%. AttenA+ gives more loss weight to slow, precision-heavy actions and lifts OpenVLA-OFT on LIBERO from 97.10% to 98.60%.