<?xml version='1.0' encoding='utf-8'?>
<feed xmlns="http://www.w3.org/2005/Atom" xml:lang="zh-CN"><title>Recoleta Research Radar</title><author><name>Recoleta</name></author><id>https://neapolitanicecream.github.io/recoleta/zh-cn/index.html</id><updated>2026-07-27T08:30:33.773828Z</updated><link rel="alternate" href="https://neapolitanicecream.github.io/recoleta/zh-cn/index.html" type="text/html" /><link rel="self" href="https://neapolitanicecream.github.io/recoleta/zh-cn/feed.xml" type="application/atom+xml" /><entry><title>智能体评测触及模糊项目，可靠性机制转移到 harness · Recoleta Trends</title><id>https://neapolitanicecream.github.io/recoleta/zh-cn/trends/software-intelligence--day--2026-07-23--trend--2079.html</id><updated>2026-07-23T00:00:00Z</updated><link href="https://neapolitanicecream.github.io/recoleta/zh-cn/trends/software-intelligence--day--2026-07-23--trend--2079.html" type="text/html" /><summary>在过去几天聚焦于编码循环中的可执行反馈之后，今天的证据拓宽了控制面。新的基准测试智能体处理不完整的产品意图和混合型办公任务；可靠性机制则在预先定义的检查点交付记忆、逻辑推理和审查。结果仍处于早期阶段：几项研究缺乏广泛的量化比较，其中一个工作流虽然提升了可审计性，但成本显著增加。</summary></entry><entry><title>面向模糊智能体工作的评估与审查控制 · Recoleta Ideas</title><id>https://neapolitanicecream.github.io/recoleta/zh-cn/ideas/software-intelligence--day--2026-07-23--ideas.html</id><updated>2026-07-23T00:00:00Z</updated><link href="https://neapolitanicecream.github.io/recoleta/zh-cn/ideas/software-intelligence--day--2026-07-23--ideas.html" type="text/html" /><summary>将澄清、形式化验证和专家审查安排在错误变得难以逆转的决策点，能够更精确地测试智能体的可靠性。最有价值的改进涉及智能体何时提问、策略结论是否具有可执行的推导，以及哪些中间产物确实会触发审查。</summary></entry><entry><title>可执行接口正成为提升机器人可靠性的共同抓手 · Recoleta Trends</title><id>https://neapolitanicecream.github.io/recoleta/zh-cn/trends/embodied-ai--day--2026-07-22--trend--946.html</id><updated>2026-07-22T00:00:00Z</updated><link href="https://neapolitanicecream.github.io/recoleta/zh-cn/trends/embodied-ai--day--2026-07-22--trend--946.html" type="text/html" /><summary>前两个有内容的日期强调了与行动相关的状态和结构化接口。今天的证据在部署、评估和训练中延伸了这一信号：当学习模型的输出被收窄为明确目标、任务相关场景、稳定动力学或可执行轨迹时，表现会更好。结果涵盖实体机器人和仿真环境，但其中几项研究仍局限于特定任务，或缺少受控的硬件对比。</summary></entry><entry><title>可执行反馈优于仅依赖提示的编码工作流 · Recoleta Trends</title><id>https://neapolitanicecream.github.io/recoleta/zh-cn/trends/software-intelligence--day--2026-07-22--trend--2064.html</id><updated>2026-07-22T00:00:00Z</updated><link href="https://neapolitanicecream.github.io/recoleta/zh-cn/trends/software-intelligence--day--2026-07-22--trend--2064.html" type="text/html" /><summary>近期关于编码代理控制机制的研究仍在推进，但今天的证据表明，控制信号正变得更加针对具体任务。性能分析器、变异补丁、静态分析和仓库上下文在循环中引导生成并验证结果。报告的收益幅度较大，但它们来自不同基准，不能据此确定某一种架构总体上更优。</summary></entry><entry><title>机器人学习正在改变显式执行接口 · Recoleta Ideas</title><id>https://neapolitanicecream.github.io/recoleta/zh-cn/ideas/embodied-ai--day--2026-07-22--ideas.html</id><updated>2026-07-22T00:00:00Z</updated><link href="https://neapolitanicecream.github.io/recoleta/zh-cn/ideas/embodied-ai--day--2026-07-22--ideas.html" type="text/html" /><summary>机器人部署团队可以通过保留可执行、可检查、可纠正的明确决策，让想象练习和真实世界恢复数据更有用。现有证据支持：在演练阶段使用轨迹级可行性过滤器，并在接口层进行标注，以诊断失败究竟始于目标选择、场景抽象还是控制环节。</summary></entry><entry><title>代码优化与生成式测试的可执行控制 · Recoleta Ideas</title><id>https://neapolitanicecream.github.io/recoleta/zh-cn/ideas/software-intelligence--day--2026-07-22--ideas.html</id><updated>2026-07-22T00:00:00Z</updated><link href="https://neapolitanicecream.github.io/recoleta/zh-cn/ideas/software-intelligence--day--2026-07-22--ideas.html" type="text/html" /><summary>通过结合互补的可执行信号，性能优化和测试生成工作流可以提高模型输出的可靠性：使用运行时性能分析来优先处理静态优化匹配，使用语义变异来检验验收测试，并在行为级验证之前采用确定性的项目脚手架。</summary></entry><entry><title>结构化动作接口支撑具身世界模型 · Recoleta Trends</title><id>https://neapolitanicecream.github.io/recoleta/zh-cn/trends/embodied-ai--day--2026-07-21--trend--935.html</id><updated>2026-07-21T00:00:00Z</updated><link href="https://neapolitanicecream.github.io/recoleta/zh-cn/trends/embodied-ai--day--2026-07-21--trend--935.html" type="text/html" /><summary>围绕动作相关状态的前一日信号仍在延续，但今天的五篇论文更直接地将结构引入世界建模。视觉轨迹、物理分解和可模拟的回放记录，将动作与预测后果连接起来。证据来自异质的预印本，且评估大多彼此独立，因此它表明了一种共同的设计方向，而不是已经确定的最优架构。</summary></entry><entry><title>结构化上下文和执行反馈减少编码代理的浪费 · Recoleta Trends</title><id>https://neapolitanicecream.github.io/recoleta/zh-cn/trends/software-intelligence--day--2026-07-21--trend--2050.html</id><updated>2026-07-21T00:00:00Z</updated><link href="https://neapolitanicecream.github.io/recoleta/zh-cn/trends/software-intelligence--day--2026-07-21--trend--2050.html" type="text/html" /><summary>近期围绕编码代理控制机制的关注仍在持续，但目前最有力的证据已转向代理循环内部的工作。语义化的代码仓库结构减少了重复探索和脆弱编辑，执行反馈则引导更低成本的恢复和更有力的功能检查。大多数收益仍来自作者报告或特定任务，因此尚未确立其对广泛生产应用的影响。</summary></entry><entry><title>面向机器人世界模型的反事实与轨迹测试 · Recoleta Ideas</title><id>https://neapolitanicecream.github.io/recoleta/zh-cn/ideas/embodied-ai--day--2026-07-21--ideas.html</id><updated>2026-07-21T00:00:00Z</updated><link href="https://neapolitanicecream.github.io/recoleta/zh-cn/ideas/embodied-ai--day--2026-07-21--ideas.html" type="text/html" /><summary>可回放的成对事件副本能够提供录制机器人视频所缺少的反事实监督，而空间轨迹可以将整段事件回放失败转化为可修复的对齐、接触和动力学错误。双向视觉动作模型还需要循环测试，以验证推断出的机器人运动是否确实产生了所要求的物体结果。</summary></entry><entry><title>面向代码代理的仓库感知编辑、恢复与评估 · Recoleta Ideas</title><id>https://neapolitanicecream.github.io/recoleta/zh-cn/ideas/software-intelligence--day--2026-07-21--ideas.html</id><updated>2026-07-21T00:00:00Z</updated><link href="https://neapolitanicecream.github.io/recoleta/zh-cn/ideas/software-intelligence--day--2026-07-21--ideas.html" type="text/html" /><summary>代码代理的控制机制可以更贴近工作的语义：需求链接能够约束跨文件编辑，不变量违规可以改进恢复决策，而受控的代码变换则能揭示仓库检索究竟何时节省了工作量。现有证据支持有针对性的评估，但不足以支持广泛的生产环境结论。</summary></entry><entry><title>具身策略通过保留与行动相关的状态得到改进 · Recoleta Trends</title><id>https://neapolitanicecream.github.io/recoleta/zh-cn/trends/embodied-ai--day--2026-07-20--trend--928.html</id><updated>2026-07-20T00:00:00Z</updated><link href="https://neapolitanicecream.github.io/recoleta/zh-cn/trends/embodied-ai--day--2026-07-20--trend--928.html" type="text/html" /><summary>当天最有力的证据强化了最近一次有内容的日度信号：可靠的具身控制依赖于能够在执行过程中持续保留的状态。持久化的三维物体信息、力历史、密集视觉图像块和结构化的未来引导，都能改善操作或规划。结果很有前景，但大多局限于单个机器人和基准；一项鲁棒性研究还表明，加入推理并不能可靠地让策略更安全。</summary></entry><entry><title>编码代理生成的输出正在合并前被精简和检查 · Recoleta Trends</title><id>https://neapolitanicecream.github.io/recoleta/zh-cn/trends/software-intelligence--day--2026-07-20--trend--2031.html</id><updated>2026-07-20T00:00:00Z</updated><link href="https://neapolitanicecream.github.io/recoleta/zh-cn/trends/software-intelligence--day--2026-07-20--trend--2031.html" type="text/html" /><summary>围绕编码代理控制措施的近期工作仍在继续，但今天的证据主要集中在代理留下的产物上。轨迹感知的清理会移除冗余编辑，而覆盖率检查和明确的需求则能揭示仅凭测试通过可能遗漏的缺口。大多数结果来自单项研究或供应商数据，因此其对生产环境的广泛影响仍不确定。</summary></entry><entry><title>具身控制的状态表示与检查 · Recoleta Ideas</title><id>https://neapolitanicecream.github.io/recoleta/zh-cn/ideas/embodied-ai--day--2026-07-20--ideas.html</id><updated>2026-07-20T00:00:00Z</updated><link href="https://neapolitanicecream.github.io/recoleta/zh-cn/ideas/embodied-ai--day--2026-07-20--ideas.html" type="text/html" /><summary>具身控制团队应以不同速率保留不同信息：为当前场景保留密集的空间细节，跨时间压缩为紧凑的物理记录，并使用独立刷新的状态检查执行结果。现有证据还支持通过对动作敏感的干预来评估世界模型，而不能只看生成轨迹是否合理。</summary></entry></feed>