Action-labeled synthetic data for robot transfer
Robot learning papers today focus on better action grounding, not bigger generic models. CRAFT is the clearest example. It uses simulator rollouts plus a Canny-guided video diffusion pipeline to generate photorealistic demonstrations that keep paired action labels. That matters in bimanual tasks, where contact and coordination errors break transfer quickly. The gains are large in cross-embodiment tests: from UR5 to Franka, CRAFT reports 82.6% on Lift Pot, 89.3% on Place Cans, and 86.0% on Stack Bowls with no target-robot demos. In real tests from xArm7 to Franka, it reaches 17/20, 15/20, and 16/20 successes on three tasks. The method also supports viewpoint, lighting, background, and embodiment changes in one pipeline.