Unified VLA control and spatial action
Vision-Language-Action (VLA) models are being judged on whether their internal state helps action, prediction, and spatial placement at the same time. Pelican-Unified trains language reasoning, future video generation, and robot action chunks through a shared latent state. It reports 93.5% average success on the 50-task RoboTwin dual-arm suite and an EWM Score of 66.03 on WorldArena.
Evo-Depth attacks a narrower deployment bottleneck: spatial precision without extra depth hardware. Its RGB-derived depth features feed a 0.9B-parameter VLA model. The reported real-world result is 90% average success across three tasks, using 3.2 GB of GPU memory and running at 12.3 Hz.