SWE-bench 是一项 AI 评测基准,用于衡量模型完成真实软件工程任务的能力。
SWE-bench is an AI evaluation benchmark that assesses a model's ability to complete real-world software engineering tasks.
我们的最新模型——升级版 Claude 3.5 Sonnet——在软件工程评测 SWE-bench Verified 上取得了 49% 的成绩,超过了此前最先进模型的 45%。本文将介绍我们围绕该模型构建的“Agent”,旨在帮助开发者尽可能充分地发挥 Claude 3.5 Sonnet 的能力。
Our latest model, the upgraded Claude 3.5 Sonnet, achieved 49% on SWE-bench Verified, a software engineering evaluation, beating the previous state-of-the-art model's 45%. This post explains the "agent" we built around the model, and is intended to help developers get the best possible performance out of Claude 3.5 Sonnet.
SWE-bench 是一项 AI 评测基准,用于衡量模型完成真实软件工程任务的能力。具体来说,它测试模型解决热门开源 Python 仓库中 GitHub issue 的能力。对于基准中的每项任务,AI 模型都会获得一个配置好的 Python 环境,以及该仓库在相应 issue 得到解决之前的检出版本(即本地工作副本)。随后,模型需要理解、修改并测试代码,再提交它提出的解决方案。
SWE-bench is an AI evaluation benchmark that assesses a model's ability to complete real-world software engineering tasks. Specifically, it tests how the model can resolve GitHub issues from popular open-source Python repositories. For each task in the benchmark, the AI model is given a set up Python environment and the checkout (a local working copy) of the repository from just before the issue was resolved. The model then needs to understand, modify, and test the code before submitting its proposed solution.
每个解决方案都会使用关闭原始 GitHub issue 的拉取请求中的真实单元测试来评分。这是在检验:AI 模型能否实现与原始 PR 的人类作者相同的功能。
Each solution is graded against the real unit tests from the pull request that closed the original GitHub issue. This tests whether the AI model was able to achieve the same functionality as the original human author of the PR.
SWE-bench 评测的并非孤立的 AI 模型,而是一整套“Agent”系统。在这里,“Agent”指 AI 模型及其外围软件运行框架(scaffolding)的组合。运行框架负责生成输入模型的提示词,解析模型输出以执行操作,以及管理交互循环:将模型上一步操作的结果纳入下一条提示词。即使使用相同的底层 AI 模型,Agent 在 SWE-bench 上的表现也可能因运行框架不同而有显著差异。
SWE-bench doesn't just evaluate the AI model in isolation, but rather an entire "agent" system. In this context, an "agent" refers to the combination of an AI model and the software scaffolding around it. This scaffolding is responsible for generating the prompts that go into the model, parsing the model's output to take action, and managing the interaction loop where the result of the model's previous action is incorporated into its next prompt. The performance of an agent on SWE-bench can vary significantly based on this scaffolding, even when using the same underlying AI model.
衡量大语言模型编程能力的基准还有很多,但 SWE-bench 因以下几个原因而日益受到关注:
There are many other benchmarks for the coding abilities of Large Language Models, but SWE-bench has gained in popularity for several reasons:
- 它使用真实项目中的实际工程任务,而不是竞赛题或面试题;
- 它尚未达到性能饱和,仍有很大的提升空间。目前还没有模型在 SWE-bench Verified 上的任务完成率超过 50%(不过,截至本文撰写时,更新版 Claude 3.5 Sonnet 已达到 49%);
- 它衡量的是整个“Agent”,而非孤立的模型。开源开发者和初创公司在优化运行框架方面取得了显著成果,能够在使用同一个模型的情况下,大幅提升系统表现。
- It uses real engineering tasks from actual projects, rather than competition- or interview-style questions;
- It is not yet saturated—there’s plenty of room for improvement. No model has yet crossed 50% completion on SWE-bench Verified (though the updated Claude 3.5 Sonnet is, at the time of writing, at 49%);
- It measures an entire "agent", rather than a model in isolation. Open-source developers and startups have had great success in optimizing scaffoldings to greatly improve the performance around the same model.
需要注意,原始 SWE-bench 数据集中有些任务,如果没有 GitHub issue 之外的额外上下文,就无法解决(例如,关于应返回哪些特定错误消息的信息)。SWE-bench-Verified 是 SWE-bench 的一个包含 500 道问题的子集,经过人工审核,以确保这些问题可以解决,因此能够最清晰地衡量编程 Agent 的表现。本文所指的就是这一基准。
Note that the original SWE-bench dataset contains some tasks that are impossible to solve without additional context outside of the GitHub issue (for example, about specific error messages to return). SWE-bench-Verified is a 500 problem subset of SWE-bench that has been reviewed by humans to make sure they are solvable, and thus provides the most clear measure of coding agents' performance. This is the benchmark to which we’ll refer in this post.
达到最先进水平
Achieving state-of-the-art
使用工具的 Agent
Tool Using Agent
在为更新版 Claude 3.5 Sonnet 创建经过优化的 Agent 运行框架时,我们的设计理念是尽可能把控制权交给语言模型本身,并让运行框架保持精简。这个 Agent 配有一条提示词、一个用于执行 bash 命令的 Bash 工具,以及一个用于查看和编辑文件及目录的 Edit 工具。我们持续从模型采样,直到它自行决定任务已经完成,或者超出其 200k 的上下文长度。这个运行框架让模型能够自行判断如何着手解决问题,而不是通过硬编码将其限定在某种特定模式或工作流中。
Our design philosophy when creating the agent scaffold optimized for updated Claude 3.5 Sonnet was to give as much control as possible to the language model itself, and keep the scaffolding minimal. The agent has a prompt, a Bash Tool for executing bash commands, and an Edit Tool, for viewing and editing files and directories. We continue to sample until the model decides that it is finished, or exceeds its 200k context length. This scaffold allows the model to use its own judgment of how to pursue the problem, rather than be hardcoded into a particular pattern or workflow.
提示词为模型勾勒了一个建议采用的方法,但对于这项任务而言,它既不过长,也不过于详细。模型可以自由选择如何从一个步骤转向另一个步骤,而不必遵循严格、彼此分立的阶段切换。如果你对 token 用量不敏感,明确鼓励模型生成较长的回复可能会有所帮助。
The prompt outlines a suggested approach for the model, but it’s not overly long or too detailed for this task. The model is free to choose how it moves from step to step, rather than having strict and discrete transitions. If you are not token-sensitive, it can help to explicitly encourage the model to produce a long response.
下面的代码展示了我们 Agent 运行框架中的提示词:
The following code shows the prompt from our agent scaffold:
模型的第一个工具用于执行 Bash 命令。它的 schema 很简单,只接受要在环境中运行的命令。不过,工具描述发挥着更重要的作用。它为模型提供了更详细的指令,包括如何转义输入、环境无法访问互联网,以及如何在后台运行命令。
The model's first tool executes Bash commands. The schema is simple, taking only the command to be run in the environment. However, the description of the tool carries more weight. It includes more detailed instructions for the model, including escaping inputs, lack of internet access, and how to run commands in the background.
接下来展示的是 Bash 工具的规格说明:
Next, we show the spec for the Bash Tool:
模型的第二个工具(Edit 工具)要复杂得多,包含了模型查看、创建和编辑文件所需的全部功能。同样,我们在工具描述中提供了详细信息,告诉模型如何使用这个工具。
The model's second tool (the Edit Tool) is much more complex, and contains everything the model needs for viewing, creating, and editing files. Again, our tool description contains detailed information for the model about how to use the tool.
针对各种各样的 Agent 任务,我们在这些工具的描述和规格说明上投入了大量精力。我们通过测试找出模型可能误解规格说明的地方,以及使用工具时可能遇到的陷阱,然后修改描述,提前避免这些问题。我们认为,应当在为模型设计工具接口时投入更多精力,就像人们在为人类设计工具界面时投入大量精力一样。
We put a lot of effort into the descriptions and specs for these tools across a wide variety of agentic tasks. We tested them to uncover any ways that the model might misunderstand the spec, or the possible pitfalls of using the tools, then edited the descriptions to preempt these problems. We believe that much more attention should go into designing tool interfaces for models, in the same way that a large amount of attention goes into designing tool interfaces for humans.
下面的代码展示了我们为 Edit 工具编写的描述:
The following code shows the description for our Edit Tool:
我们提升表现的一种方法,是为工具做“防错”设计。例如,当 Agent 离开根目录后,模型有时会弄错相对文件路径。为了避免这种情况,我们直接让工具始终要求使用绝对路径。
One way we improved performance was to "error-proof" our tools. For instance, sometimes models could mess up relative file paths after the agent had moved out of the root directory. To prevent this, we simply made the tool always require an absolute path.
我们尝试了几种不同的策略,让模型指定如何编辑现有文件,其中可靠性最高的是字符串替换:模型指定将给定文件中的 `old_str` 替换为 `new_str`。只有在 `old_str` 恰好匹配一处时,替换才会执行。如果匹配数量更多或更少,就会向模型显示相应的错误消息,让它重试。
We experimented with several different strategies for specifying edits to existing files and had the highest reliability with string replacement, where the model specifies `old_str` to replace with `new_str` in the given file. The replacement will only occur if there is exactly one match of `old_str`. If there are more or fewer matches, the model is shown an appropriate error message for it to retry.
下面展示的是 Edit 工具的规格说明:
The spec for our Edit Tool is shown below:
结果
Results
总体而言,升级版 Claude 3.5 Sonnet 在推理、编程和数学能力上,都优于我们此前的模型,以及此前最先进的模型。它的 Agent 能力也有所提升:工具与运行框架有助于让这些增强后的能力得到充分发挥。
In general, the upgraded Claude 3.5 Sonnet demonstrates higher reasoning, coding, and mathematical abilities than our prior models, and the previous state-of-the-art model. It also demonstrates improved agentic capabilities: the tools and scaffolding help put those improved abilities to their best use.
Agent 行为示例
Examples of agent behavior
运行这项基准时,我们以 SWE-Agent 框架作为 Agent 代码的基础。在下面的日志中,我们将 Agent 的文本输出、工具调用和工具响应分别呈现为 THOUGHT、ACTION 和 OBSERVATION,不过我们并未将模型限定为必须按固定顺序执行这些步骤。
For running the benchmark, we used the SWE-Agent framework as a foundation for our agent code. In our logs below, we render the agent's text output, tool calls, and tool responses as THOUGHT, ACTION, and OBSERVATION, even though we don’t constrain the model to a fixed ordering.
下面的代码块将逐步展示 Sonnet 3.5 解决一道 SWE-bench 问题的典型过程。
The code blocks below will walk through a typical case of the Sonnet 3.5 solving a SWE-bench problem.
在第一个代码块中,你可以看到提供给模型的初始提示词的一部分,其中 `{pr_description}` 已填入某项 SWE-bench 任务的实际内容。关键在于,这项任务包含复现问题的步骤,为模型开展调查提供了一个有价值的起点。
In this first block, you can see part of the initial prompt given to the model, with `{pr_description}` filled in with the real value from a SWE-bench task. Importantly, this task contains steps to reproduce the issue, which will give the model a valuable starting point to investigate.
模型作出回应,首先使用 Edit 工具查看仓库结构。你可以在 THOUGHT 和 ACTION 下看到模型的文本输出与工具调用参数,在 OBSERVATION 下看到工具输出的一部分:
The model responds and first uses the Edit Tool to view the repository structure. You can see the model's text output and tool call arguments under THOUGHT and ACTION, and part of the tool's output under OBSERVATION:
现在,模型对仓库结构有了更充分的了解,于是使用 Edit 工具创建一个新脚本,用来复现问题并测试它的修复:
Now that the model has a better understanding of the repository structure, it uses the Edit Tool to create a new script that it will use to reproduce the issue and test its fix:
随后,模型使用 Bash 工具执行它编写的脚本,成功复现了任务中的问题:
The model then uses the Bash Tool to execute the script it wrote, and successfully reproduces the issue from the task:
接下来,模型使用 Edit 工具修改仓库中的源代码,并重新运行脚本,验证改动是否解决了问题:
From here on, the model uses the Edit Tool to change the source code in the repository and reruns its script to verify whether the change has resolved the issue:
在这个具体示例中,模型执行了 12 个步骤,随后决定已经可以提交解决方案。之后,任务的测试成功运行,证实模型的方案解决了问题。有些任务中,模型经过 100 多轮才提交方案;另一些任务中,模型则一直尝试,直到耗尽上下文。
In this particular example, the model worked for 12 steps before deciding that it was ready to submit. The task's tests then ran successfully, verifying that the model's solution addressed the problem. Some tasks took more than 100 turns before the model submitted its solution; in others, the model kept trying until it ran out of context.
通过将更新版 Claude 3.5 Sonnet 的解题尝试与旧版模型进行对比,我们发现更新版 3.5 Sonnet 更经常自我纠正。它还表现出尝试多种不同解决方案的能力,而不是陷入反复犯同一个错误的循环。
From reviewing attempts from the updated Claude 3.5 Sonnet compared to older models, updated 3.5 Sonnet self-corrects more often. It also shows an ability to try several different solutions, rather than getting stuck making the same mistake over and over.
挑战
Challenges
SWE-bench Verified 是一项强有力的评测,但运行起来也比简单的单轮评测更复杂。下面是我们在使用过程中遇到的一些挑战,其他 AI 开发者也可能遇到这些问题。
SWE-bench Verified is a powerful evaluation, but it’s also more complex to run than simple, single-turn evals. These are some of the challenges that we faced in using it—challenges that other AI developers might also encounter.
- 耗时长,token 成本高。上面的示例展示了一个在 12 个步骤内成功完成的案例。不过,许多成功的运行都需要模型经过数百轮、消耗 >100k token 才能解决问题。更新版 Claude 3.5 Sonnet 很有韧性:只要给它足够的时间,它往往就能找到解决问题的办法,但这可能代价高昂;
- 评分。在检查失败任务时,我们发现,有些情况下模型的行为是正确的,但环境配置存在问题,或者安装补丁被应用了两次。解决这些系统层面的问题,对于准确了解 AI Agent 的表现至关重要。
- 隐藏测试。由于模型看不到用来给它评分的测试,它经常会“认为”自己已经成功,实际上任务却失败了。其中有些失败,是因为模型在错误的抽象层次上解决问题:只做了表面修补,而没有进行更深层次的重构。另一些失败则让人觉得不那么公平:模型解决了问题,却不符合原始任务的单元测试。
- 多模态。尽管更新版 Claude 3.5 Sonnet 具有出色的视觉和多模态能力,我们却没有实现让它查看保存在文件系统中或通过 URL 引用的文件的方法。这使得某些任务,尤其是来自 Matplotlib 的任务,格外难以调试,也容易出现模型幻觉。开发者在这里显然有一些容易实现的改进机会,而且 SWE-bench 已经推出了一项新的聚焦多模态任务的评测。我们期待在不久的将来,看到开发者借助 Claude 在这项评测中取得更高分数。
- Duration and high token costs. The examples above are from a case that was successfully completed in 12 steps. However, many successful runs took hundreds of turns for the model to resolve, and >100k tokens. The updated Claude 3.5 Sonnet is tenacious: it can often find its way around a problem given enough time, but that can be expensive;
- Grading. While inspecting failed tasks, we found cases where the model behaved correctly, but there were environment setup issues, or problems with install patches being applied twice. Resolving these systems issues is crucial for getting an accurate picture of an AI agent's performance.
- Hidden tests. Because the model cannot see the tests it's being graded against, it often “thinks” that it has succeeded when the task actually is a failure. Some of these failures are because the model solved the problem at the wrong level of abstraction (applying a bandaid instead of a deeper refactor). Other failures feel a little less fair: they solve the problem, but do not match the unit tests from the original task.
- Multimodal. Despite the updated Claude 3.5 Sonnet having excellent vision and multimodal capabilities, we did not implement a way for it to view files saved to the filesystem or referenced as URLs. This made debugging certain tasks (especially those from Matplotlib) especially difficult, and also prone to model hallucinations. There is definitely low-hanging fruit here for developers to improve upon—and SWE-bench has launched a new evaluation focused on multi-modal tasks. We look forward to seeing developers achieve higher scores on this eval with Claude in the near future.
升级版 Claude 3.5 Sonnet 仅凭一条简单提示词和两个通用工具,就在 SWE-bench Verified 上取得了 49% 的成绩,超过了此前的最先进水平(45%)。我们相信,使用新版 Claude 3.5 Sonnet 进行开发的开发者,很快就会找到新的、更好的方法,使 SWE-bench 得分进一步超过我们在这里展示的初步结果。
The upgraded Claude 3.5 Sonnet achieved 49% on SWE-bench Verified, beating the previous state-of-the-art (45%), with a simple prompt and two general purpose tools. We feel confident that developers building with the new Claude 3.5 Sonnet will quickly find new, better ways to improve SWE-bench scores over what we've initially demonstrated here.
致谢
Acknowledgements
Erik Schluntz 优化了 SWE-bench Agent,并撰写了这篇博客文章。Simon Biggs、Dawn Drain 和 Eric Christiansen 协助实现了这项基准。Shauna Kravec、Dawn Drain、Felipe Rosso、Nova DasSarma、Ven Chandrasekaran 以及许多其他人参与了 Claude 3.5 Sonnet 的训练,使其在 Agent 式编程方面表现出色。
Erik Schluntz optimized the SWE-bench agent and wrote this blog post. Simon Biggs, Dawn Drain, and Eric Christiansen helped implement the benchmark. Shauna Kravec, Dawn Drain, Felipe Rosso, Nova DasSarma, Ven Chandrasekaran, and many others contributed to training Claude 3.5 Sonnet to be excellent at agentic coding.
— 全文完 —
原文来自 Anthropic,中文为非官方学习译文。
查看原始出处 ↗