Trend

Robot policies and world models are being judged by memory, contact, and action fidelity

Day · 2026-05-05 · Embodied AI

The day’s clearest signal is evaluation pressure on embodied AI. After several days of Vision-Language-Action (VLA) deployment work, the current papers make success depend on memory, contact sensing, action-conditioned prediction, and spatial consistency. RLDX-1, RoboAlign-R1, and iWorld-Bench give the strongest evidence.

Dexterous VLA policies

RLDX-1 treats dexterous manipulation as a multi-signal control problem. The policy adds video motion, a small memory store, and tactile or torque inputs to a Qwen3-VL-based VLA. Its Multi-Stream Action Transformer keeps cognition, proprioception, and physics streams separate before cross-stream attention combines them for action prediction.

The reported gains are largest on tasks where current-image policies struggle. RLDX-1 reports 86.8% success on ALLEX humanoid tasks, while π₀.₅ and GR00T N1.6 are around 40%. On ALLEX Object-in-Box Selection, a memory-heavy task, it reports 91.7% success while the two baselines are in the 30% range. The report also ties deployment to latency: inference optimization cuts per-step latency on an RTX 5090 from 71.2 ms to 43.7 ms.

  • Dongyoung Kim, Huiwon Jang, Myungkyu Koo, Suhyeok Jang, Taeyoung Kim, Beomjun Kim, Byungjun Yoon, Changsung Jang, Daewon Choi, Dongsu Han, Donguk Lee, Heeseung Kwon, Hojin Jeon, Jaehyun Kang, Jaekyoung Bae, Jihyuk Lee, Jimin Lee, John Won, Joonwoo Ahn, Junhyeong Park, Junyoung Sung, Kyungmin Lee, Minseong Han, Minsung Yoon, Sejune Joo, Seonil Son, Seungcheol Park, Seunggeun Cho, Seungjun Moon, Seungku Kim, Yonghoon Dong, Yongjin Cho, Youngchan Kim, Chang Hwan Kim, Dohyeon Kim, Heecheol Kim, Heewon Lee, Hensen Ahn, Hyungkyu Ryu, Hyunsoo Choi, Hyunsoo Shin, Jaeheon Jung, Jaewoo Kim, Jinwook Kim, Joochul Chang, Joonsoo Kim, Junghun Park, Jungwoo Park, Junho Cho, Junhyeok Park, Junwon Lee, Kangwook Lee, Kwanghoon Kim, Kyoungwhan Choe, Manoj Bhadu, Nayoung Oh, Sangjun Kim, Sangwoo Kim, Seunghoon Shim, Seunghyun Kim, Seungjun Lee, Seungyup Ka, Sungryol Yang, Wook Jung, Yashu Shukla, Yeonjae Lee, Yeonwoo Bae, Jinwoo Shin

Robot video world models

RoboAlign-R1 makes robot video prediction answer to task-level behavior, not only pixel loss. It fine-tunes an 8B multimodal judge, distills it into a 98M reward model, and uses that reward for post-training. The judge scores instruction following, manipulation success, action-outcome consistency, temporal consistency, contact realism, and physics adherence.

The paper reports a RobotWorldBench score of 8.52±0.15, compared with 7.74±0.62 for iVideoGPT. Its Sliding Window Re-encoding method refreshes rollout context during long predictions and reports better SSIM, PSNR, LPIPS, and ROI-LPIPS with about 1% added latency. The practical point is clear: long-horizon robot video models need both behavioral rewards and rollout maintenance.

Interactive world-model benchmarks

iWorld-Bench targets a measurement gap for interactive world models: whether generated futures follow actions and preserve memory. It standardizes action inputs across text commands, one-hot controls, and camera parameters, allowing models with different control formats to face comparable tasks.

The benchmark is substantial in coverage. It contains 330,000 video clips, selects 2,100 evaluation videos, and defines 4,900 test tasks across action-control difficulty, memory, and camera following. Its data spans four viewpoints, nine outdoor weather types, five indoor lighting types, and 18 simulator environments. The supplied text does not give a model leaderboard, so the grounded contribution is the benchmark scale and action mapping.

Spatial reasoning for unified visual models

JoyAI-Image connects image understanding, generation, and editing around spatial supervision. The system uses Qwen3-VL-8B-Instruct for multimodal understanding and instruction parsing, then conditions a 16B diffusion transformer for generation and editing. Its OpenSpatial data engine builds spatial QA and editing data from 3D boxes, masks, visibility checks, and multi-view consistency checks.

The reported result is strongest on spatial understanding. JoyAI-Image-Und reaches a 64.4 average across nine spatial benchmarks, up 5.3 points over Qwen3-VL-8B-Instruct and matching Gemini-2.5-Pro in the supplied summary. General benchmark scores stay close to the base model, which matters because the added spatial training does not erase broad visual skills in the reported tests.

NewerExecutable checks are setting the bar for software agentsOlderCoding agents are being judged by interfaces, token cost, and repository evidence