Day · 2026-05-21 · Software Intelligence
Current emphasis: coding-agent work is tying progress to inspectable evidence. P2T curates repair steps, SWE-Mutation tests whether generated tests catch real bugs, and MOSS replays production failures before…
Day · 2026-05-20 · Embodied AI
Embodied AI is the clear center. Vision-language-action (VLA) papers are judged by 3D contact cues, real hardware, perturbation tests, and reproducible evaluation.
Day · 2026-05-20 · Software Intelligence
The day’s strongest signal is executable proof. SpecBench shows public tests can reward hollow systems, while FuzzingBrain V2 and ERA use evaluation loops to verify crashes or improve scientific metrics.
Day · 2026-05-19 · Embodied AI
Embodied AI dominates this period. Vision-language-action (VLA) work is judged by latency, lighting, perturbations, and fine-grained task stages, while world-model papers build longer rollouts and cheaper synthetic data.
Day · 2026-05-19 · Software Intelligence
This day’s strongest signal is runtime discipline. STORM, OpenComputer, and DIFFCODEGEN point to the same requirement: agents need current state, executable checks, and cheap validation around model output before teams…
Day · 2026-05-18 · Embodied AI
The period is dominated by embodied AI work that treats policies as deployable robot systems. Vision-language-action (VLA) models are tested on dexterous hands, corrupted cameras, dual-arm tasks, and contact-rich…
Day · 2026-05-18 · Software Intelligence
Current emphasis: coding agents are judged by how they run, repair, and stay inside bounds. A-ProS shows gains from stateful judge feedback. ProcBench scores process defects inside traces.
Week · 2026-W20 · Embodied AI
This week’s strongest signal is execution quality for robot Vision-Language-Action (VLA) models. Work on HarmoWAM, RAW-Dream, and Pelican-Unified ties policy gains to imagined rollouts, phase-aware action control, and…
Week · 2026-W20 · Software Intelligence
Code-agent research this week set a higher bar for useful work. SWE-Cycle and SaaSBench score setup, integration, tests, and delivered behavior. Rollout Cards adds reporting discipline for agent runs.
Day · 2026-05-17 · Embodied AI
Vision-language-action (VLA) robot work in this period is execution-centered. DyGRO-VLA protects multi-task policies during reinforcement learning. AffordVLA teaches contact regions without runtime modules.
Day · 2026-05-17 · Software Intelligence
The day’s strongest signal is concrete execution. SaaSBench and WebGameBench score delivered software behavior, while ContraFix and MemRepair improve repair by keeping runtime evidence and prior fixes inside the loop.
Day · 2026-05-16 · Software Intelligence
The strongest signal is operational evaluation. 1GC-7RC, AgentKernelArena, and TOBench all score agents inside bounded work loops with tools, runtime checks, and resource limits.