Day · 2026-05-14 · Embodied AI
Embodied AI work in this window treats robot intelligence as an execution problem. Pelican-Unified links reasoning, future video, and action in one latent state.
Day · 2026-05-14 · Software Intelligence
The day’s strongest signal is practical code-agent work under executable checks. FrontierSmith and DIO-Agent use scoring or execution errors to make coding tasks harder and more useful.
Day · 2026-05-13 · Software Intelligence
The period’s clearest signal: code agents are being judged by complete, checkable work. SWE-Cycle and Phoenix-bench make setup, tests, and domain toolchains part of the score.
Day · 2026-05-13 · Embodied AI
Robot Vision-Language-Action (VLA) research today treats deployment as an execution problem. The strongest papers tune action representations, subtask calls, critical-frame training, visual invariance, and inference…
Day · 2026-05-12 · Software Intelligence
The day’s strongest signal is auditability for agentic systems. BenchJack attacks benchmark harnesses before agents run. Rollout Cards asks evaluations to publish rollout evidence.
Day · 2026-05-12 · Embodied AI
The day’s robot papers treat Vision-Language-Action (VLA) models as control systems that need predictive rollouts, guided action decoding, and explicit safety checks. RAW-Dream gives the clearest data-efficiency result.
Day · 2026-05-11 · Software Intelligence
The strongest signal is that agents need inspectable execution and stricter task evidence. DuST uses execution-labeled candidate code as training data, Shepherd records live agent state for branching, and ComplexMCP…
Day · 2026-05-11 · Embodied AI
Robot Vision-Language-Action (VLA) papers in this period focus on deployment failure points: out-of-distribution scenes, limited demonstrations, and weak action supervision.
Week · 2026-W19 · Software Intelligence
This week’s research treats large language model (LLM) coding agents as systems that need proof before trust. The strongest work checks generated code through execution, repository tasks, formal proofs, tool contracts…
Week · 2026-W19 · Embodied AI
This week’s robotics corpus treats Vision-Language-Action (VLA) policies as deployable control systems. The strongest work measures recovery after drift, memory over long tasks, and low-cost world-state prediction.
Day · 2026-05-10 · Software Intelligence
The period’s main signal is stricter proof for AI-built software. ConCovUp, RubricRefine, and MonitoringBench test agents against concrete failure modes: missed concurrent interactions, wrong tool contracts, and hidden…
Day · 2026-05-10 · Embodied AI
Robotics work in this period treats reliability as a measured control problem. Vision-Language-Action (VLA) policies get recovery training, uncertainty-triggered search, and store-specific action data.