Trend

Robotics papers put structure into control and data collection

Day · 2026-04-16 · Embodied AI

This day’s robotics set is strongest on one point: researchers are putting task structure into the data path and decision path. The clearest examples are π0.7\pi_{0.7}, WAV, and ShapeGen. One line adds richer prompts and latent planning to VLA control. Another line builds narrower, geometry-aware data generation and capture systems so policies see the right variation in the first place.

Structured VLA control

Generalist robot models now add more structure to the prompt and the rollout, not just more data. π0.7\pi_{0.7} conditions action generation on task text, subtask text, episode metadata, control mode, and optional subgoal images from a world model. The paper ties that prompt design to zero-shot behavior on long dexterous tasks such as espresso making, folding laundry, and trash handling, plus cross-embodiment transfer. WAV pushes the same theme on the planning side. It combines a future predictor, a value model, and an action decoder, then searches in latent space rather than raw actions. The evidence here is stronger on mechanism than on exact margins, because the excerpt does not include full task tables for either paper.

  • Physical Intelligence, Bo Ai, Ali Amin, Raichelle Aniceto, Ashwin Balakrishna, Greg Balke, Kevin Black, George Bokinsky, Shihao Cao, Thomas Charbonnier, Vedant Choudhary, Foster Collins, Ken Conley, Grace Connors, James Darpinian, Karan Dhabalia, Maitrayee Dhaka, Jared DiCarlo, Danny Driess, Michael Equi, Adnan Esmail, Yunhao Fang, Chelsea Finn, Catherine Glossop, Thomas Godden, Ivan Goryachev, Lachlan Groom, Haroun Habeeb, Hunter Hancock, Karol Hausman, Gashon Hussein, Victor Hwang, Brian Ichter, Connor Jacobsen, Szymon Jakubczak, Rowan Jen, Tim Jones, Gregg Kammerer, Ben Katz, Liyiming Ke, Mairbek Khadikov, Chandra Kuchi, Marinda Lamb, Devin LeBlanc, Brendon LeCount, Sergey Levine, Xinyu Li, Adrian Li-Bell, Vladislav Lialin, Zhonglin Liang, Wallace Lim, Yao Lu, Enyu Luo, Vishnu Mano, Nandan Marwaha, Aikys Mongush, Liam Murphy, Suraj Nair, Tyler Patterson, Karl Pertsch, Allen Z. Ren, Gavin Schelske, Charvi Sharma, Baifeng Shi, Lucy Xiaoyang Shi, Laura Smith, Jost Tobias Springenberg, Kyle Stachowicz, Will Stoeckle, Jiaming Tang, Jimmy Tanner, Shalom Tekeste, Marcel Torne, Kyle Vedder, Quan Vuong, Anna Walling, Haohuan Wang, Jason Wang, XuDong Wang, Chris Whalen, Samuel Whitmore, Blake Williams, Charles Xu, Sukwon Yoo, Lili Yu, Wuming Zhang, Zhuoyang Zhang, Ury Zhilinsky
  • Runze Li, Hongyin Zhang, Junxi Jin, Qixin Zeng, Zifeng Zhuang, Yiqi Tang, Shangke Lyu, Donglin Wang

Data generation targets the gap directly

Several papers attack generalization by generating better robot data around the failure mode they care about. ShapeGen creates new real-to-real demonstrations by swapping objects within a category using learned 3D warps. The gains are concrete: hang_mug rises from 5% to 45%, hang_mug_hard from 5% to 50%, and serve_kettle from 35% to 75% on unseen instances. DockAnywhere does the same for mobile manipulation under docking error. It reuses the contact-rich part of a demo, replans only the approach segments, and edits RGB-D observations in 3D. In ManiSkill with five docking points, it reaches 78.9% overall success, versus 17.8% for plain DP3 and 74.2% for DP3+DemoGen. The common pattern is narrow augmentation with task geometry kept intact.

Dexterous data collection gets more practical

Dexterous manipulation work in this window puts weight on how data is captured, not only on model design. DEX-Mouse is a portable hand-held interface built from off-the-shelf parts for under USD 150. In its attached forearm setup, it reaches 86.67% overall success and 10.05 s average completion time, ahead of two glove baselines in the reported study. HRDexDB tackles the same bottleneck at dataset scale. It pairs human and robot grasps on the same 100 objects, with 1.4K trials, 12.8M frames, 23 video views, object pose, tactile signals, and success labels across four embodiments. That gives the field more grounded material for cross-embodiment learning and grasp analysis.

Abstract simulators get a grounded transfer recipe

Sim2real work here focuses on the case where the simulator is missing real state, not just mis-tuned physics. ASTRA treats that as a history-dependent grounding problem. It learns a recurrent latent state from abstracted real trajectories, then corrects the simulator with transition, reward, and next-state losses. The reported evidence is selective but useful: in a morphology-shift AntMaze test with 1.25× leg length, success reaches 65% on U-Maze versus 21% for direct transfer. This matters because many practical simulators stay coarse on purpose, and the paper gives a concrete recipe for training in that setting with limited real data.

NewerCoding-agent progress is coming from tighter control over trajectories, costs, and environmentsOlderCoding-agent progress is coming from tighter feedback loops and harder evidence