要点速览:我们的编程 Agent 在 Terminal Bench 2.0 上从前 30 名跃升至前 5 名。我们只改动了运行框架。下面介绍我们的运行框架工程方法,先透露一点:自我验证与执行轨迹记录非常有帮助。
TLDR: Our coding agent went from Top 30 to Top 5 on Terminal Bench 2.0. We only changed the harness. Here’s our approach to harness engineering (teaser: self-verification & tracing help a lot).
运行框架工程的目标
The Goal of Harness Engineering
运行框架的目标,是将模型先天参差不齐的智能塑造成适合目标任务的能力。运行框架工程关注的是系统:围绕模型构建工具与配套机制,优化任务表现、token 效率、时延等目标。设计决策包括系统提示词、工具选择和执行流程。
The goal of a harness is to mold the inherently spiky intelligence of a model for tasks we care about. Harness Engineering is about systems, you’re building tooling around the model to optimize goals like task performance, token efficiency, latency, etc. Design decisions include the system prompt, tool choice, and execution flow.
但应当怎样改动运行框架,才能改进你的 Agent?
But how should you change the harness to improve your agent?
在 LangChain,我们利用执行轨迹大规模分析 Agent 的失败模式。当今的模型在很大程度上仍是黑箱,内部机制难以解释。不过,我们能够在文本层面看到它们的输入和输出,并将这些信息用于改进循环。
At LangChain, we use Traces to understand agent failure modes at scale. Models today are largely black-boxes, their inner mechanisms are hard to interpret. But we can see their inputs and outputs in text space which we then use in our improvement loops.
我们采用了一套简单的方法,迭代改进 deepagents-cli(我们的编程 Agent),使其在 Terminal Bench 2.0 上提升了 13.7 points,得分从 52.8 提高到 66.5。我们只调整了运行框架,模型始终保持为 gpt-5.2-codex。
We used a simple recipe to iteratively improve deepagents-cli (our coding agent) 13.7 points from 52.8 to 66.5 on Terminal Bench 2.0. We only tweaked the harness and kept the model fixed, gpt-5.2-codex.
实验设置与运行框架中的可调项
Experiment Setup & The Knobs on a Harness
我们使用了 Terminal Bench 2.0,它如今已是评测 Agent 编程能力的常用基准。它包含 89 项任务,覆盖机器学习、调试和生物学等领域。我们使用 Harbor 编排各次运行。它负责启动沙箱(Daytona)、与我们的 Agent 循环交互,以及执行验证和评分。
We used Terminal Bench 2.0, a now standard benchmark to evaluate agentic coding. It has 89 tasks across domains like machine learning, debugging, and biology. We use Harbor to orchestrate the runs. It spins up sandboxes (Daytona), interacts with our agent loop, and runs verification + scoring.
Agent 的每项操作都会保存在 LangSmith 中,其中也包括时延、token 数量和成本等指标。
Every agent action is stored in LangSmith. It also includes metrics like latency, token counts, and costs.
我们可以调整什么
The Knobs we can Turn
Agent 运行框架有很多可调项:系统提示词、工具、钩子/中间件、技能、子 Agent 委派、记忆系统,等等。我们有意缩小优化空间,聚焦于三项:系统提示词、工具和中间件(我们用这个词指围绕模型调用与工具调用设置的钩子)。
An agent harness has a lot of knobs: system prompts, tools, hooks/middleware, skills, sub-agent delegation, memory systems, and more. We deliberately compress the optimization space and focus on three: System Prompt, Tools, and Middleware (our term for hooks around model and tool calls).
我们从默认提示词和标准工具加中间件的配置开始。使用 GPT-5.2-Codex 时,这套配置的得分为 52.8%。这是一个不错的成绩,略低于当时排行榜前 30 名的水平,但仍有提升空间。
We start with a default prompt and standard tools+middleware. This scores 52.8% with GPT-5.2-Codex. A solid score, just outside the Top 30 of the leaderboard today, but room to grow.
执行轨迹分析技能
The Trace Analyzer Skill
我们希望执行轨迹分析能够重复开展,因此将其做成了一个 Agent Skill。它提供了一套方法,用于分析多次运行中的错误,并改进运行框架。流程如下:
We wanted trace analysis to be repeatable so we made it into an Agent Skill. This serves as our recipe to analyze errors across runs and make improvements to the harness. The flow is:
- 从 LangSmith 获取实验的执行轨迹。
- 启动并行的错误分析 Agent → 由主 Agent 汇总发现与建议。
- 汇总反馈,对运行框架进行有针对性的修改。
- Fetch experiment traces from LangSmith
- Spawn parallel error analysis agents → main agent synthesizes findings + suggestions
- Aggregate feedback and make targeted changes to the harness.
这种方法类似于 boosting,重点关注此前运行中出现的错误。第 3 步中,人工参与可以很有帮助(但并非必需),用于核验和讨论提出的改动。对某项任务过拟合的改动会损害泛化能力,也可能使其他任务的表现退化。
This works similarly to boosting which focuses on mistakes from previous runs. A human can be pretty helpful in Step 3 (though not required) to verify and discuss proposed changes. Changes that overfit to a task are bad for generalization and can lead to regressions in other Tasks.
自动化执行轨迹分析节省了数小时的时间,也让我们能够方便、快速地尝试实验。我们将很快发布这项技能,目前正在测试它在通用提示词优化中的表现。
Automated trace analysis saves hours of time and made it easy to quickly try experiments. We’ll be publishing this skill soon, we’re currently testing it for prompt optimization generally.
哪些改动真正改善了 Agent 的表现
What Actually Improved Agent Performance
自动化执行轨迹分析让我们能够排查 Agent 在哪里出了问题。这些问题包括推理错误、不遵循任务指令、缺少测试与验证、耗尽时间,等等。下面几节将更详细地介绍相应的改进。
Automated Trace analysis allowed us to debug where agents were going wrong. Issues included reasoning errors, not following task instructions, missing testing and verification, running out of time, etc. We go into these improvements in more details in the sections below.
构建与自我验证
Build & Self-Verify
当今的模型是出色的自我改进机器。
Today’s models are exceptional self-improvement machines.
自我验证让 Agent 能够在一次运行过程中,通过反馈改进自身表现。不过,它们并不会自然而然地进入这种构建—验证循环。
Self-verification allows agents to self-improve via feedback within a run. However, they don’t have a natural tendency to enter this build-verify loop.
最常见的失败模式是:Agent 写出一个方案,重新读一遍自己的代码,确认看起来没问题,然后就停止了。测试是自主 Agent 编程的关键环节。它既有助于检验整体正确性,也为 Agent 提供逐步优化所依据的反馈信号。
The most common failure pattern was that the agent wrote a solution, re-read its own code, confirmed it looks ok, and stopped. Testing is a key part of autonomous agentic coding. It helps test for overall correctness and simultaneously gives agents signal to hill-climb against.
我们在系统提示词中加入了解决问题的方法指导。
We added guidance to the system prompt on how to approach problem solving.
- 规划与探索:阅读任务、浏览代码库,并根据任务规格以及如何验证解决方案,制定初步计划。
- 构建:在实施计划时就考虑验证。如果还没有测试,就编写测试,同时覆盖正常流程和边界情况。
- 验证:运行测试,阅读完整输出,并对照用户要求检查结果,而不是对照你自己的代码。
- 修复:分析出现的错误,重新查看原始规格,并修复问题。
- Planning & Discovery: Read the task, scan the codebase, and build an initial plan based on the task specification and how to verify the solution.
- Build: Implement the plan with verification in mind. Build tests, if they don’t exist and test both happy paths and edge cases.
- Verify: Run tests, read the full output, compare against what was asked (not against your own code).
- Fix: Analyze any errors, revisit the original spec, and fix issues.
我们非常重视测试,因为它为每一轮迭代中的改动提供依据。我们发现,除了提示词,确定性的上下文注入也能帮助 Agent 验证自己的工作。我们使用 PreCompletionChecklistMiddleware,在 Agent 退出前拦截它,提醒它对照任务规格进行一轮验证。这类似于 Ralph Wiggum Loop:通过钩子在 Agent 退出时强制它继续执行;我们将这一机制用于验证。
We really focus on testing because it powers the changes in every iteration. We found that alongside prompting, deterministic context injection helps agents verify their work. We use a PreCompletionChecklistMiddleware that intercepts the agent before it exits and reminds it to run a verification pass against the Task spec. This is similar to a Ralph Wiggum Loop where a hook forces the agent to continue executing on exit, we use this for verification.
为 Agent 提供关于其环境的上下文
Giving Agents Context about their Environment
运行框架工程的一部分,是为上下文工程构建良好的信息交付机制。Terminal Bench 任务自带目录结构、内置工具,以及严格的超时限制。
Part of harness engineering is building a good delivery mechanism for context engineering. Terminal Bench tasks come with directory structures, built-in tooling, and strict timeouts.
- 目录上下文与工具:
LocalContextMiddleware会在 Agent 启动时运行,梳理cwd以及其他父目录和子目录。我们运行bash命令,查找已安装的Python等工具。上下文发现与搜索很容易出错,因此注入上下文能够减少这类出错机会,帮助Agent 熟悉其工作环境。 - 教 Agent 编写可测试的代码:Agent 并不知道代码需要具备怎样的可测试性。我们在提示词中说明,它的工作会接受程序化测试的检验,类似于提交代码时的检查。例如,任务规格提到的文件路径应当严格遵守,这样解决方案才能在自动评分环节正常工作。在提示词中强调边界情况,可以帮助 Agent 避免只检查“正常流程”。要求模型遵循测试标准,是防止低质量产物随时间不断堆积的一项有力策略。
- 时间预算:我们注入时间预算提醒,促使 Agent 完成实现并转向验证。众所周知,Agent 不擅长估算时间,因此这一启发式方法在这种环境中很有帮助。现实中的编程通常没有严格的时间限制,但如果完全不给 Agent 提供约束信息,它们就不会在时间边界内开展工作。
- Directory Context & Tooling: A
LocalContextMiddlewareruns on agent start to map thecwdand other parent+children directories. We runbashcommands to find tools likePythoninstallations. Context discovery and search are error prone, so injecting context reduces this error surface and helps onboard the agent into its environment. - Teaching Agents to Write Testable Code: Agents don’t know how their code needs to be testable. We add prompting say their work will be measured against programatic tests, similar to when committing code. For example, Task specs that mention file paths should be followed exactly so the solutions works in an automated scoring step. Prompting that stresses edge-cases helps the agent avoid only checking “happy path” cases. Forcing models to conform to testing standards is a powerful strategy to avoid “slop buildup” over time.
- Time Budgeting: We inject time budget warnings to nudge the agent to finish work and shift to verification. Agents are famously bad at time estimation so this heuristic helps in this environment. Real world coding usually doesn’t have strict time limits, but without adding any knowledge of constraints, agents won’t work within time bounds.
Agent 越了解自己的环境、约束和评测标准,就越能自主安排工作。
The more that agents know about their environment, constraints, and evaluation criteria, the better they can autonomously self-direct their work.
运行框架工程师的职责:准备并提供上下文,让 Agent 能够自主完成工作。
The purpose of the harness engineer: prepare and deliver context so agents can autonomously complete work.
鼓励 Agent 暂停推进,重新审视计划
Encouraging Agents to Step Back & Reconsider Plans
Agent 一旦确定了计划,就可能陷入狭隘视角,进入“死循环”,只对同一种行不通的方法做微小变动(某些执行轨迹中会重复 10 多次)。
Agents can be myopic once they’ve decided on a plan which results in “doom loops” that make small variations to the same broken approach (10+ times in some traces).
我们使用 LoopDetectionMiddleware,通过工具调用钩子跟踪每个文件的编辑次数。在同一个文件被编辑 N 次后,它会加入类似“……考虑重新审视你的方法”的上下文。这可以帮助 Agent 摆脱死循环,不过,如果模型认为自己的方法正确,它仍可以沿着原来的路径继续。
We use a LoopDetectionMiddleware that tracks per-file edit counts via tool call hooks. It adds context like “…consider reconsidering your approach” after N edits to the same file. This can help agents recover from doom loops, though the model can continue down the same path if it thinks it’s correct.
需要特别说明,这是一种针对目前观察到的模型问题所设计的启发式方法。随着模型进步,这些防护机制很可能不再必要;但在当下,它们有助于 Agent 正确、自主地执行任务。
Important note. This is a design heuristic that engineers around today’s perceived model issues. As models improve, these guardrails will likely be unnecessary, but today helps agents execute correctly and autonomously.
决定在推理上投入多少计算资源
Choosing How Much Compute to Spend on Reasoning
推理模型可以自主运行数小时,因此我们必须决定每项子任务应投入多少计算资源。你可以对每项任务都使用最高推理预算,但对大多数工作而言,优化推理计算资源的投入会带来收益。
Reasoning models can run autonomously for hours so we have to decide how much compute to spend on every subtask. You can use the max reasoning budget on every task, but most work can benefit from optimizing reasoning compute spend.
Terminal Bench 的超时限制带来了取舍。更多推理有助于 Agent 评估每一步,但 token 和时间消耗可能达到 2x 以上。gpt-5.2-codex 有 4 种推理模式:low、medium、high 和 xhigh。
Terminal Bench timeout limits create a tradeoff. More reasoning helps agents evaluate each step, but can burn over 2x more tokens/time. gpt-5.2-codex has 4 reasoning modes, low, medium, high, and xhigh.
我们发现,推理有助于在规划阶段充分理解问题,而一些 Terminal Bench 任务非常困难。好的计划能够帮助模型更快地找到可行方案。
We found that reasoning helps with planning to fully understand the problem, some Terminal Bench tasks are very difficult. A good plan helps get to a working solution more quickly.
在后期验证阶段,投入更多推理同样有助于发现错误,并提交解决方案。作为一种启发式方法,我们选择 xhigh-high-xhigh 的“推理三明治”作为基线。
Later stage verification also benefits from more reasoning to catch mistakes and get a solution submitted. As a heuristic, we choose a xhigh-high-xhigh "reasoning sandwich" as a baseline.
全程只使用 xhigh 时,Agent 会因超时而表现不佳,得分仅为 53.9%;相比之下,使用 high 时得分为 63.6%。不同推理预算分配方式在试运行中的差异并不大,因此我们保留了自己的方案,将得分提升至 66.5%。
Running only at xhigh scored poorly at 53.9% due to agent timeouts compared to 63.6% at high. There weren’t large differences in trial runs across reasoning budget splits so we stuck with our approach which pushed the score to 66.5%.
在支持多模型的运行框架中,平衡推理预算可以体现为:使用大模型制定计划,再交接给较小的模型负责实现。
In a multi-model harness, balancing reasoning budgets could play out as using a large model for planning and handing off to a smaller model for implementation.
构建 Agent 运行框架的实践要点
Practical Takeaways for Building Agent Harnesses
Agent 的设计空间很大。下面是我们从这些实验以及整体构建 deepagents 的过程中总结出的一些通用原则。
The design space of agents is big. Here are some general principles from our experiments and building deepagents overall.
- 为 Agent 做好上下文工程。对当今的 Agent 而言,组织上下文仍然困难,尤其是在未见过的环境中。提供目录结构、可用工具、编程最佳实践和问题解决策略等上下文,帮助模型熟悉环境,可以减少搜索不当和可避免的规划错误所带来的出错机会。
- 帮助 Agent 验证自己的工作。模型倾向于接受第一个看似可行的方案。应在提示词中强烈要求它通过运行测试、改进方案来验证工作。这对没有人工参与的自主编程系统尤其重要。
- 将执行轨迹作为反馈信号。执行轨迹让 Agent 能够自我评估和自我调试。将工具与推理放在一起排查很重要,例如,模型可能因为缺少工具,或缺少如何完成某件事的指令,而走上错误的路径。
- 在短期内检测并纠正不良模式。当今的模型并不完美。运行框架设计者的工作,是针对当前缺陷设计应对机制,同时为未来更聪明的模型做好准备。盲目重试、不验证工作,都是很好的例子。这些防护机制几乎肯定会随时间推移而消失,但要在当下构建稳健的 Agent 应用,它们仍是值得尝试的工具。
- 针对模型定制运行框架。Codex 和 Claude 的提示词指南表明,不同模型需要不同的提示方式。在较早版本运行框架上的一次试运行中,Claude Opus 4.6 得分为
59.6%,有竞争力,但低于 Codex,因为我们没有针对 Claude 运行同样的改进循环。许多原则具有通用性,例如充分准备上下文、重视验证;但针对自己的任务进行几轮运行框架迭代,有助于最大程度地提升 Agent 在各项任务中的表现。
- Context Engineering on Behalf of Agents. Context assembly is still difficult for agents today, especially in unseen environments. Onboarding models with context like directory structures, available tools, coding best practices, and problem solving strategies helps reduce the error surface for poor search and avoidable errors in planning.
- Help agents self-verify their work. Models are biased towards their first plausible solution. Prompt them aggressively to verify their work by running tests and refining solutions. This is especially important in autonomous coding systems that don’t have humans in the loop.
- Tracing as a feedback signal. Traces allow agents to self-evaluate and debug themselves. It’s important to debug tooling and reasoning together (ex: models go down wrong paths because they lack a tool or instructions how to do something).
- Detect and fix bad patterns in the short term. Models today aren’t perfect. The job of the harness designer is to design around today’s shortcomings while planning for smarter models in the future. Blind retries and not verifying work are good examples. These guardrails will almost surely dissolve over time, but to build robust agent applications today, they’re useful tools to experiment with.
- Tailor Harnesses to Models. The Codex and Claude prompting guides show that models require different prompting. A test run with Claude Opus 4.6 scored
59.6%with an earlier harness version, competitive but worse than Codex because we didn’t run the same Improvement Loop with Claude. Many principles generalize like good context preparation and a focus on verification, but running a few rounds of harness iterations for your task helps maximize agent performance across tasks.
运行框架设计还有更多值得开放研究的方向。有意思的探索包括:将 Codex、Gemini 和 Claude 结合使用的多模型系统;用于持续学习的记忆基础机制,让 Agent 能够自主改进任务表现;以及衡量运行框架改动对不同模型的影响。
There’s more open research to do in harness design. Interesting avenues include multi-model systems (Codex, Gemini, and Claude together), memory primitives for continual learning so agents can autonomously improve on tasks, and measuring harness changes across models.
对于改进 Agent 的外环,我们正在探索 RLM 等方法,以更高效地挖掘执行轨迹。我们会继续改进运行框架,并公开分享研究成果。
For the outer loop of improving agents, we’re looking at methods like RLMs to more efficiently mine traces. We’ll be continuing work to improve the harness and openly share our research.
我们创建了一个执行轨迹数据集,与社区分享。
We created a dataset of our Traces to share with the community.
Deep Agents 是开源的,提供 Python 和 Javascript 版本。
Deep Agents is open source. Python and Javascript.
期待更多逐步优化与开放研究。
To more hill climbing and open research.
— 全文完 —
原文来自 LangChain,中文为非官方学习译文。
查看原始出处 ↗




