资料馆/工具与技能
Anthropic阅读档案 · 非官方中文译文

借助 Agent 为 Agent 编写高效工具Writing effective tools for agents — with agents

下载 PDF
中文 PDF ↓英文 PDF ↓
完整译文与原文逐段对应。图片、图注、表格和代码保留原文。A complete reading edition. Figures, captions, tables and code are preserved from the source.
中文译文ENGLISH ORIGINAL
This is an abstract illustration for the Eng Blog article, Writing effective tools for agents -- with agents.

Agent 的效能取决于我们为它提供的工具。本文将分享如何编写高质量工具和评测,以及如何让 Claude 为自己优化工具,从而提升表现。

Agents are only as effective as the tools we give them. We share how to write high-quality tools and evaluations, and how you can boost performance by using Claude to optimize its tools for itself.

模型上下文协议(Model Context Protocol,MCP)可以让大语言模型(LLM)Agent 使用多达数百种工具,解决现实世界中的任务。但我们怎样才能让这些工具发挥最大的效用?

The Model Context Protocol (MCP) can empower LLM agents with potentially hundreds of tools to solve real-world tasks. But how do we make those tools maximally effective?

本文介绍我们在多种 Agent 式 AI 系统中提升表现时最有效的一些方法1。

In this post, we describe our most effective techniques for improving performance in a variety of agentic AI systems1.

我们首先介绍如何:

We begin by covering how you can:

  • 构建并测试工具原型
  • 创建并运行全面的评测,检验 Agent 使用工具的表现
  • 与 Claude Code 等 Agent 协作,自动提升工具的表现
  • Build and test prototypes of your tools
  • Create and run comprehensive evaluations of your tools with agents
  • Collaborate with agents like Claude Code to automatically increase the performance of your tools

最后,我们总结在这一过程中发现的编写高质量工具的关键原则:

We conclude with key principles for writing high-quality tools we’ve identified along the way:

  • 选择合适的工具进行实现,并决定哪些工具不应实现
  • 通过命名空间为工具划定清晰的功能边界
  • 让工具向 Agent 返回有意义的上下文
  • 优化工具响应,提高 token 使用效率
  • 对工具描述和规格说明进行提示词工程优化
  • Choosing the right tools to implement (and not to implement)
  • Namespacing tools to define clear boundaries in functionality
  • Returning meaningful context from tools back to agents
  • Optimizing tool responses for token efficiency
  • Prompt-engineering tool descriptions and specs
This is an image depicting how an engineer might use Claude Code to evaluate the efficacy of agentic tools.
Building an evaluation allows you to systematically measure the performance of your tools. You can use Claude Code to automatically optimize your tools against this evaluation.

什么是工具?

What is a tool?

在计算领域,确定性系统在输入相同时,每次都会产生相同的输出;而 Agent 这样的非确定性系统,即使初始条件相同,也可能生成不同的响应。

In computing, deterministic systems produce the same output every time given identical inputs, while non-deterministic systems—like agents—can generate varied responses even with the same starting conditions.

在传统的软件开发中,我们是在确定性系统之间建立契约。例如,像 getWeather(“NYC”) 这样的函数调用,每次被调用时,都会以完全相同的方式获取纽约市的天气。

When we traditionally write software, we’re establishing a contract between deterministic systems. For instance, a function call like getWeather(“NYC”) will always fetch the weather in New York City in the exact same manner every time it is called.

工具是一种新型软件,体现的是确定性系统与非确定性 Agent 之间的契约。当用户问“今天我应该带伞吗?”时,Agent 可能调用天气工具,也可能凭一般知识回答,甚至可能先追问用户所在的位置。有时,Agent 可能产生幻觉,甚至无法理解该如何使用某个工具。

Tools are a new kind of software which reflects a contract between deterministic systems and non-deterministic agents. When a user asks "Should I bring an umbrella today?,” an agent might call the weather tool, answer from general knowledge, or even ask a clarifying question about location first. Occasionally, an agent might hallucinate or even fail to grasp how to use a tool.

这意味着,为 Agent 编写软件时,我们需要从根本上重新思考方法:不能再像为其他开发者或系统编写函数和 API 那样编写工具和 MCP 服务器,而应当面向 Agent 来设计它们。

This means fundamentally rethinking our approach when writing software for agents: instead of writing tools and MCP servers the way we’d write functions and APIs for other developers or systems, we need to design them for agents.

我们的目标是扩大 Agent 能够有效解决任务的范围,让它们通过工具采取多种可行策略,完成广泛的任务。幸运的是,根据我们的经验,对 Agent 而言最“顺手”的工具,往往也出乎意料地容易被人类直观理解。

Our goal is to increase the surface area over which agents can be effective in solving a wide range of tasks by using tools to pursue a variety of successful strategies. Fortunately, in our experience, the tools that are most “ergonomic” for agents also end up being surprisingly intuitive to grasp as humans.

如何编写工具

How to write tools

本节介绍如何与 Agent 协作,编写并改进供它们使用的工具。首先快速搭建工具原型,并在本地测试。接着,运行全面的评测,衡量后续修改的影响。通过与 Agent 协作,你可以反复评测和改进工具,直到 Agent 在现实任务中取得良好表现。

In this section, we describe how you can collaborate with agents both to write and to improve the tools you give them. Start by standing up a quick prototype of your tools and testing them locally. Next, run a comprehensive evaluation to measure subsequent changes. Working alongside agents, you can repeat the process of evaluating and improving your tools until your agents achieve strong performance on real-world tasks.

构建原型

Building a prototype

如果不亲自动手,很难预先判断哪些工具会让 Agent 用起来顺手,哪些不会。首先快速搭建工具原型。如果你使用 Claude Code 编写工具(甚至可能一次生成),为 Claude 提供工具所依赖的软件库、API 或 SDK 的文档会很有帮助,其中也可能包括 MCP SDK。适合 LLM 阅读的文档通常可以在官方文档网站的扁平化 llms.txt 文件中找到(这是我们 API 的文档)。

It can be difficult to anticipate which tools agents will find ergonomic and which tools they won’t without getting hands-on yourself. Start by standing up a quick prototype of your tools. If you’re using Claude Code to write your tools (potentially in one-shot), it helps to give Claude documentation for any software libraries, APIs, or SDKs (including potentially the MCP SDK) your tools will rely on. LLM-friendly documentation can commonly be found in flat llms.txt files on official documentation sites (here’s our API’s).

将工具封装为本地 MCP 服务器或桌面扩展(DXT),就可以在 Claude Code 或 Claude 桌面应用中连接并测试工具。

Wrapping your tools in a local MCP server or Desktop extension (DXT) will allow you to connect and test your tools in Claude Code or the Claude Desktop app.

要将本地 MCP 服务器连接到 Claude Code,请运行 claude mcp add <name> <command> [args...]。

To connect your local MCP server to Claude Code, run claude mcp add <name> <command> [args...].

要将本地 MCP 服务器或 DXT 连接到 Claude 桌面应用,请分别前往 Settings > Developer 或 Settings > Extensions。

To connect your local MCP server or DXT to the Claude Desktop app, navigate to Settings > Developer or Settings > Extensions, respectively.

也可以将工具直接传入 Anthropic API 调用,以编程方式进行测试。

Tools can also be passed directly into Anthropic API calls for programmatic testing.

亲自测试工具,找出不顺畅之处。收集用户反馈,逐步形成对预期使用场景和提示词的直觉,了解你的工具应当支持哪些需求。

Test the tools yourself to identify any rough edges. Collect feedback from your users to build an intuition around the use-cases and prompts you expect your tools to enable.

运行评测

Running an evaluation

接下来,需要通过评测衡量 Claude 使用工具的效果。首先,基于真实用途生成大量评测任务。我们建议与 Agent 协作,让它帮助分析结果,并判断如何改进工具。完整流程可参见我们的工具评测实践指南。

Next, you need to measure how well Claude uses your tools by running an evaluation. Start by generating lots of evaluation tasks, grounded in real world uses. We recommend collaborating with an agent to help analyze your results and determine how to improve your tools. See this process end-to-end in our tool evaluation cookbook.

This graph measures the test set accuracy of human-written vs. Claude-optimized Slack MCP servers.
Held-out test set performance of our internal Slack tools

生成评测任务

Generating evaluation tasks

有了早期原型,Claude Code 就可以快速探索你的工具,并创建数十组提示词与响应配对。提示词应来自真实使用场景,并基于贴近实际的数据源和服务,例如内部知识库和微服务。我们建议避免过于简单或流于表面的“沙箱”环境,因为它们的复杂度不足以对工具进行充分的压力测试。好的评测任务可能需要多次工具调用,甚至数十次。

With your early prototype, Claude Code can quickly explore your tools and create dozens of prompt and response pairs. Prompts should be inspired by real-world uses and be based on realistic data sources and services (for example, internal knowledge bases and microservices). We recommend you avoid overly simplistic or superficial “sandbox” environments that don’t stress-test your tools with sufficient complexity. Strong evaluation tasks might require multiple tool calls—potentially dozens.

以下是一些好的任务示例:

Here are some examples of strong tasks:

  • 安排下周与 Jane 开会,讨论我们最新的 Acme Corp 项目。附上上一次项目规划会议的记录,并预订一间会议室。
  • 客户 ID 为 9182 的客户反映,一次购买尝试被扣款三次。找出所有相关日志条目,并判断是否有其他客户受到同一问题的影响。
  • 客户 Sarah Chen 刚刚提交了取消服务的请求。请准备一份挽留优惠方案。确定:(1)客户为什么要离开;(2)什么挽留优惠最有吸引力;(3)提出优惠方案前,我们应了解哪些风险因素。
  • Schedule a meeting with Jane next week to discuss our latest Acme Corp project. Attach the notes from our last project planning meeting and reserve a conference room.
  • Customer ID 9182 reported that they were charged three times for a single purchase attempt. Find all relevant log entries and determine if any other customers were affected by the same issue.
  • Customer Sarah Chen just submitted a cancellation request. Prepare a retention offer. Determine: (1) why they're leaving, (2) what retention offer would be most compelling, and (3) any risk factors we should be aware of before making an offer.

以下则是一些较弱的任务:

And here are some weaker tasks:

  • 安排下周与 jane@acme.corp 开会。
  • 在支付日志中搜索 purchase_complete 和 customer_id=9182。
  • 查找客户 ID 为 45892 的取消请求。
  • Schedule a meeting with jane@acme.corp next week.
  • Search the payment logs for purchase_complete and customer_id=9182.
  • Find the cancellation request by Customer ID 45892.

每条评测提示词都应配有可验证的响应或结果。验证器可以很简单,例如将标准答案与采样生成的响应进行精确字符串比较;也可以更复杂,例如让 Claude 评判响应。应避免过于严格的验证器,以免仅因格式、标点或合理的替代表述等无关差异,就将正确响应判为错误。

Each evaluation prompt should be paired with a verifiable response or outcome. Your verifier can be as simple as an exact string comparison between ground truth and sampled responses, or as advanced as enlisting Claude to judge the response. Avoid overly strict verifiers that reject correct responses due to spurious differences like formatting, punctuation, or valid alternative phrasings.

对于每组提示词与响应,你还可以选择指定预期 Agent 为解决任务而调用的工具,从而在评测中衡量 Agent 是否成功理解了每个工具的用途。不过,正确完成任务可能有多条有效路径,因此应尽量避免对策略规定得过细,或对某种策略过拟合。

For each prompt-response pair, you can optionally also specify the tools you expect an agent to call in solving the task, to measure whether or not agents are successful in grasping each tool’s purpose during evaluation. However, because there might be multiple valid paths to solving tasks correctly, try to avoid overspecifying or overfitting to strategies.

运行评测

Running the evaluation

我们建议直接调用 LLM API,以编程方式运行评测。采用简单的 Agent 循环,即用 while 循环封装交替进行的 LLM API 调用与工具调用,每个评测任务对应一个循环。应向每个评测 Agent 提供一条任务提示词和你的工具。

We recommend running your evaluation programmatically with direct LLM API calls. Use simple agentic loops (while-loops wrapping alternating LLM API and tool calls): one loop for each evaluation task. Each evaluation agent should be given a single task prompt and your tools.

在评测 Agent 的系统提示词中,我们建议要求 Agent 不仅输出用于验证的结构化响应块,还输出推理块和反馈块。要求 Agent 在工具调用块和响应块之前输出这些内容,可能会通过触发思维链(CoT)行为,提高 LLM 实际发挥出的智能水平。

In your evaluation agents’ system prompts, we recommend instructing agents to output not just structured response blocks (for verification), but also reasoning and feedback blocks. Instructing agents to output these before tool call and response blocks may increase LLMs’ effective intelligence by triggering chain-of-thought (CoT) behaviors.

如果使用 Claude 运行评测,可以开启交错思考,直接获得类似的现成功能。这有助于探究 Agent 为什么调用或不调用某些工具,并找出工具描述和规格说明中需要改进的具体部分。

If you’re running your evaluation with Claude, you can turn on interleaved thinking for similar functionality “off-the-shelf”. This will help you probe why agents do or don’t call certain tools and highlight specific areas of improvement in tool descriptions and specs.

除了整体准确率,我们还建议收集其他指标,例如单次工具调用和任务的总运行时长、工具调用总次数、token 总消耗量,以及工具错误。跟踪工具调用有助于发现 Agent 经常采用的工作流程,并识别可以整合工具功能的机会。

As well as top-level accuracy, we recommend collecting other metrics like the total runtime of individual tool calls and tasks, the total number of tool calls, the total token consumption, and tool errors. Tracking tool calls can help reveal common workflows that agents pursue and offer some opportunities for tools to consolidate.

This graph measures the test set accuracy of human-written vs. Claude-optimized Asana MCP servers.
Held-out test set performance of our internal Asana tools

分析结果
Agent 可以成为有用的伙伴,帮助发现问题并提供反馈,涉及的范围包括相互矛盾的工具描述、低效的工具实现,以及令人困惑的工具 schema。不过,请记住:Agent 在反馈和响应中遗漏的内容,往往可能比它写出的内容更重要。LLM 并不总能如实表达其实际思考。

Analyzing results
Agents are your helpful partners in spotting issues and providing feedback on everything from contradictory tool descriptions to inefficient tool implementations and confusing tool schemas. However, keep in mind that what agents omit in their feedback and responses can often be more important than what they include. LLMs don’t always say what they mean.

观察 Agent 在哪里卡住或感到困惑。通读评测 Agent 的推理和反馈(或思维链),找出不顺畅之处。检查原始交互记录,包括工具调用和工具响应,捕捉 Agent 未在思维链中明确描述的行为。要读出字里行间的信息;记住,评测 Agent 并不一定知道正确答案和正确策略。

Observe where your agents get stumped or confused. Read through your evaluation agents’ reasoning and feedback (or CoT) to identify rough edges. Review the raw transcripts (including tool calls and tool responses) to catch any behavior not explicitly described in the agent’s CoT. Read between the lines; remember that your evaluation agents don’t necessarily know the correct answers and strategies.

分析工具调用指标。大量冗余工具调用可能说明应适当调整分页或 token 上限参数;大量因参数无效产生的工具错误,可能说明工具需要更清晰的描述或更好的示例。我们推出 Claude 的网页搜索工具时,发现 Claude 会不必要地在工具的 query 参数中追加 2025,导致搜索结果产生偏差,降低表现。我们通过改进工具描述,将 Claude 引导到了正确方向。

Analyze your tool calling metrics. Lots of redundant tool calls might suggest some rightsizing of pagination or token limit parameters is warranted; lots of tool errors for invalid parameters might suggest tools could use clearer descriptions or better examples. When we launched Claude’s web search tool, we identified that Claude was needlessly appending 2025 to the tool’s query parameter, biasing search results and degrading performance (we steered Claude in the right direction by improving the tool description).

与 Agent 协作

Collaborating with agents

你甚至可以让 Agent 分析结果,并替你改进工具。只需将各个评测 Agent 的交互记录拼接起来,粘贴到 Claude Code 中即可。Claude 擅长分析交互记录,并同时重构大量工具,例如在引入新修改时,确保工具实现与描述保持一致。

You can even let agents analyze your results and improve your tools for you. Simply concatenate the transcripts from your evaluation agents and paste them into Claude Code. Claude is an expert at analyzing transcripts and refactoring lots of tools all at once—for example, to ensure tool implementations and descriptions remain self-consistent when new changes are made.

事实上,本文大部分建议都来自我们使用 Claude Code 反复优化内部工具实现的过程。我们的评测基于内部工作空间创建,复现了内部工作流程的复杂性,其中包含真实的项目、文档和消息。

In fact, most of the advice in this post came from repeatedly optimizing our internal tool implementations with Claude Code. Our evaluations were created on top of our internal workspace, mirroring the complexity of our internal workflows, including real projects, documents, and messages.

我们依靠预留测试集,确保没有对“训练”评测过拟合。这些测试集表明,即使工具已经由“专家”实现,我们仍然能够在其基础上进一步提升表现,无论这些工具是由研究人员手动编写,还是由 Claude 自己生成。

We relied on held-out test sets to ensure we did not overfit to our “training” evaluations. These test sets revealed that we could extract additional performance improvements even beyond what we achieved with "expert" tool implementations—whether those tools were manually written by our researchers or generated by Claude itself.

下一节将分享我们从这一过程中学到的一些经验。

In the next section, we’ll share some of what we learned from this process.

编写有效工具的原则

Principles for writing effective tools

本节将我们的经验提炼为几条指导原则,帮助编写有效的工具。

In this section, we distill our learnings into a few guiding principles for writing effective tools.

为 Agent 选择合适的工具

Choosing the right tools for agents

工具更多,并不总能带来更好的结果。我们观察到的一个常见错误是,工具只是对现有软件功能或 API 端点做了一层封装,却没有考虑这些工具是否适合 Agent。原因在于,Agent 与传统软件有着不同的“可供性(交互能力)”,也就是说,它们感知自己能够通过这些工具采取哪些行动的方式不同。

More tools don’t always lead to better outcomes. A common error we’ve observed is tools that merely wrap existing software functionality or API endpoints—whether or not the tools are appropriate for agents. This is because agents have distinct “affordances” to traditional software—that is, they have different ways of perceiving the potential actions they can take with those tools

LLM Agent 的“上下文”有限,也就是它们一次能够处理的信息量有限,而计算机内存则便宜且充裕。以在通讯录中查找联系人为例,传统软件程序可以高效地存储联系人列表,逐个处理联系人,检查完一个再继续下一个。

LLM agents have limited "context" (that is, there are limits to how much information they can process at once), whereas computer memory is cheap and abundant. Consider the task of searching for a contact in an address book. Traditional software programs can efficiently store and process a list of contacts one at a time, checking each one before moving on.

但如果 LLM Agent 使用的工具返回了全部联系人,它就必须逐个 token 阅读每个联系人的信息,把有限的上下文空间浪费在无关信息上。想象一下,为了在通讯录中找到某个联系人,你从头到尾逐页阅读,也就是采用暴力搜索。对 Agent 和人类而言,更好、更自然的做法都是先跳到相关页面,例如按字母顺序定位。

However, if an LLM agent uses a tool that returns ALL contacts and then has to read through each one token-by-token, it's wasting its limited context space on irrelevant information (imagine searching for a contact in your address book by reading each page from top-to-bottom—that is, via brute-force search). The better and more natural approach (for agents and humans alike) is to skip to the relevant page first (perhaps finding it alphabetically).

我们建议先针对特定且影响较大的工作流程,精心设计少量工具,使其与你的评测任务相匹配,再在此基础上扩展。以通讯录为例,可以选择实现 search_contacts 或 message_contact 工具,而不是 list_contacts 工具。

We recommend building a few thoughtful tools targeting specific high-impact workflows, which match your evaluation tasks and scaling up from there. In the address book case, you might choose to implement a search_contacts or message_contact tool instead of a list_contacts tool.

工具可以整合功能,在内部处理多个独立操作或 API 调用。例如,工具可以在响应中补充相关元数据,或用一次工具调用处理经常串联执行的多步骤任务。

Tools can consolidate functionality, handling potentially multiple discrete operations (or API calls) under the hood. For example, tools can enrich tool responses with related metadata or handle frequently chained, multi-step tasks in a single tool call.

以下是一些例子:

Here are some examples:

  • 与其分别实现 list_users、list_events 和 create_event 工具,不如考虑实现一个 schedule_event 工具,查找空闲时间并安排活动。
  • 与其实现 read_logs 工具,不如考虑实现一个 search_logs 工具,仅返回相关日志行及其周围的一些上下文。
  • 与其分别实现 get_customer_by_id、list_transactions 和 list_notes 工具,不如实现一个 get_customer_context 工具,一次性汇总某位客户近期的所有相关信息。
  • Instead of implementing a list_users, list_events, and create_event tools, consider implementing a schedule_event tool which finds availability and schedules an event.
  • Instead of implementing a read_logs tool, consider implementing a search_logs tool which only returns relevant log lines and some surrounding context.
  • Instead of implementing get_customer_by_id, list_transactions, and list_notes tools, implement a get_customer_context tool which compiles all of a customer’s recent & relevant information all at once.

确保构建的每个工具都有明确且独立的用途。工具应让 Agent 能够像拥有相同底层资源的人类一样拆解和解决任务,同时减少原本会被中间输出占用的上下文。

Make sure each tool you build has a clear, distinct purpose. Tools should enable agents to subdivide and solve tasks in much the same way that a human would, given access to the same underlying resources, and simultaneously reduce the context that would have otherwise been consumed by intermediate outputs.

工具过多或功能重叠,也可能干扰 Agent 采用高效策略。认真、有选择地规划要构建哪些工具、不要构建哪些工具,确实能够带来显著回报。

Too many tools or overlapping tools can also distract agents from pursuing efficient strategies. Careful, selective planning of the tools you build (or don’t build) can really pay off.

为工具设置命名空间

Namespacing your tools

你的 AI Agent 可能会接入数十个 MCP 服务器和数百种不同工具,其中包括其他开发者提供的工具。当工具功能重叠或用途模糊时,Agent 可能会困惑,不知道该使用哪一个。

Your AI agents will potentially gain access to dozens of MCP servers and hundreds of different tools–including those by other developers. When tools overlap in function or have a vague purpose, agents can get confused about which ones to use.

命名空间,即通过共同前缀对相关工具进行分组,有助于划定大量工具之间的边界;MCP 客户端有时会默认这样做。例如,按服务设置工具命名空间,如 asana_search、jira_search,以及按资源设置命名空间,如 asana_projects_search、asana_users_search,可以帮助 Agent 在合适的时机选择正确的工具。

Namespacing (grouping related tools under common prefixes) can help delineate boundaries between lots of tools; MCP clients sometimes do this by default. For example, namespacing tools by service (e.g., asana_search, jira_search) and by resource (e.g., asana_projects_search, asana_users_search), can help agents select the right tools at the right time.

我们发现,在采用前缀还是后缀来设置命名空间之间作出选择,会对工具使用评测产生不可忽视的影响。影响因 LLM 而异,因此我们鼓励你根据自己的评测选择命名方案。

We have found selecting between prefix- and suffix-based namespacing to have non-trivial effects on our tool-use evaluations. Effects vary by LLM and we encourage you to choose a naming scheme according to your own evaluations.

Agent 可能调用错误的工具,给正确的工具传入错误的参数,调用的工具过少,或错误地处理工具响应。通过有选择地实现那些名称能反映任务自然拆分方式的工具,你既能减少加载到 Agent 上下文中的工具及工具描述数量,又能把原本由 Agent 在上下文中完成的计算转移到工具调用本身。这会降低 Agent 整体犯错的风险。

Agents might call the wrong tools, call the right tools with the wrong parameters, call too few tools, or process tool responses incorrectly. By selectively implementing tools whose names reflect natural subdivisions of tasks, you simultaneously reduce the number of tools and tool descriptions loaded into the agent’s context and offload agentic computation from the agent’s context back into the tool calls themselves. This reduces an agent’s overall risk of making mistakes.

让工具返回有意义的上下文

Returning meaningful context from your tools

同样,工具实现应注意只向 Agent 返回高信息量的内容。与灵活性相比,应优先考虑内容与当前上下文的相关性,并避免使用底层技术标识符,例如 uuid、256px_image_url、mime_type。相比之下,name、image_url 和 file_type 等字段,更有可能直接为 Agent 后续的行动和响应提供依据。

In the same vein, tool implementations should take care to return only high signal information back to agents. They should prioritize contextual relevance over flexibility, and eschew low-level technical identifiers (for example: uuid, 256px_image_url, mime_type). Fields like name, image_url, and file_type are much more likely to directly inform agents’ downstream actions and responses.

与晦涩的标识符相比,Agent 通常也更擅长处理自然语言名称、术语或标识符。我们发现,仅仅把任意字母数字组合的 UUID 转换成语义更明确、更容易解释的语言,甚至只是改成从 0 开始编号的 ID 方案,就能通过减少幻觉,显著提高 Claude 在检索任务中的精确率。

Agents also tend to grapple with natural language names, terms, or identifiers significantly more successfully than they do with cryptic identifiers. We’ve found that merely resolving arbitrary alphanumeric UUIDs to more semantically meaningful and interpretable language (or even a 0-indexed ID scheme) significantly improves Claude’s precision in retrieval tasks by reducing hallucinations.

在某些情况下,Agent 可能需要灵活处理同时含有自然语言和技术标识符的输出,哪怕只是为了发起后续工具调用,例如 search_user(name=’jane’) → send_message(id=12345)。你可以在工具中提供一个简单的 response_format 枚举参数,同时支持两种形式,让 Agent 控制工具返回 “concise”(简洁)还是 “detailed”(详细)响应,见下图。

In some instances, agents may require the flexibility to interact with both natural language and technical identifiers outputs, if only to trigger downstream tool calls (for example, search_user(name=’jane’) → send_message(id=12345)). You can enable both by exposing a simple response_format enum parameter in your tool, allowing your agent to control whether tools return “concise” or “detailed” responses (images below).

还可以增加更多格式,以获得更大的灵活性,类似于 GraphQL 中可以精确选择想接收哪些信息。下面是一个用于控制工具响应详细程度的 ResponseFormat 枚举示例:

You can add more formats for even greater flexibility, similar to GraphQL where you can choose exactly which pieces of information you want to receive. Here is an example ResponseFormat enum to control tool response verbosity:

enum ResponseFormat {
   DETAILED = "detailed",
   CONCISE = "concise"
}

下面是一个详细工具响应的示例(206 个 token):

Here’s an example of a detailed tool response (206 tokens):

This code snippet depicts an example of a detailed tool response.

下面是一个简洁工具响应的示例(72 个 token):

Here’s an example of a concise tool response (72 tokens):

This code snippet depicts a concise tool response.
Slack threads and thread replies are identified by unique thread_ts which are required to fetch thread replies. thread_ts and other IDs (channel_id, user_id) can be retrieved from a “detailed” tool response to enable further tool calls that require these. “concise” tool responses return only thread content and exclude IDs. In this example, we use ~⅓ of the tokens with “concise” tool responses.

甚至工具响应的结构,例如 XML、JSON 或 Markdown,也会影响评测表现;不存在适用于所有情况的解决方案。这是因为 LLM 通过预测下一个 token 进行训练,通常在格式与训练数据相符时表现更好。最佳响应结构会随任务和 Agent 的不同而存在很大差异。我们鼓励你根据自己的评测选择最佳响应结构。

Even your tool response structure—for example XML, JSON, or Markdown—can have an impact on evaluation performance: there is no one-size-fits-all solution. This is because LLMs are trained on next-token prediction and tend to perform better with formats that match their training data. The optimal response structure will vary widely by task and agent. We encourage you to select the best response structure based on your own evaluation.

优化工具响应,提高 token 使用效率

Optimizing tool responses for token efficiency

优化上下文质量很重要。但优化工具响应返回给 Agent 的上下文数量,同样重要。

Optimizing the quality of context is important. But so is optimizing the quantity of context returned back to agents in tool responses.

对于任何可能占用大量上下文的工具响应,我们建议组合使用分页、范围选择、过滤和/或截断,并设置合理的默认参数值。在 Claude Code 中,我们默认将工具响应限制在 25,000 个 token 以内。我们预计 Agent 的有效上下文长度会随时间增长,但对高效使用上下文的工具的需求仍将持续存在。

We suggest implementing some combination of pagination, range selection, filtering, and/or truncation with sensible default parameter values for any tool responses that could use up lots of context. For Claude Code, we restrict tool responses to 25,000 tokens by default. We expect the effective context length of agents to grow over time, but the need for context-efficient tools to remain.

如果选择截断响应,一定要提供有帮助的指引。可以直接鼓励 Agent 采用更节省 token 的策略,例如在知识检索任务中,多次进行小范围、有针对性的搜索,而不是只进行一次宽泛搜索。同样,如果工具调用报错,例如输入验证失败,可以通过提示词工程优化错误响应,清楚传达具体且可执行的改进方法,而不是返回不透明的错误码或堆栈追踪。

If you choose to truncate responses, be sure to steer agents with helpful instructions. You can directly encourage agents to pursue more token-efficient strategies, like making many small and targeted searches instead of a single, broad search for a knowledge retrieval task. Similarly, if a tool call raises an error (for example, during input validation), you can prompt-engineer your error responses to clearly communicate specific and actionable improvements, rather than opaque error codes or tracebacks.

下面是一个工具响应被截断的示例:

Here’s an example of a truncated tool response:

This image depicts an example of a truncated tool response.

下面是一个没有帮助的错误响应示例:

Here’s an example of an unhelpful error response:

This image depicts an example of an unhelpful tool response.

下面是一个有帮助的错误响应示例:

Here’s an example of a helpful error response:

This image depicts an example of a helpful error response.
Tool truncation and error responses can steer agents towards more token-efficient tool-use behaviors (using filters or pagination) or give examples of correctly formatted tool inputs.

对工具描述进行提示词工程优化

Prompt-engineering your tool descriptions

现在介绍改进工具最有效的方法之一:对工具描述和规格说明进行提示词工程优化。由于这些内容会加载到 Agent 的上下文中,它们可以共同引导 Agent 形成有效的工具调用行为。

We now come to one of the most effective methods for improving tools: prompt-engineering your tool descriptions and specs. Because these are loaded into your agents’ context, they can collectively steer agents toward effective tool-calling behaviors.

编写工具描述和规格说明时,可以设想你会如何向团队的一位新同事介绍工具。考虑那些你可能默认对方已经知道的背景,例如专用查询格式、小众术语的定义、底层资源之间的关系,并把它们明确写出来。清楚描述预期的输入和输出,并用严格的数据模型加以约束,以避免歧义。尤其是输入参数的名称,应当明确无误:例如,不要将参数命名为 user,可以改用 user_id。

When writing tool descriptions and specs, think of how you would describe your tool to a new hire on your team. Consider the context that you might implicitly bring—specialized query formats, definitions of niche terminology, relationships between underlying resources—and make it explicit. Avoid ambiguity by clearly describing (and enforcing with strict data models) expected inputs and outputs. In particular, input parameters should be unambiguously named: instead of a parameter named user, try a parameter named user_id.

借助评测,你可以更有把握地衡量提示词工程的影响。即便只是微调工具描述,也可能带来大幅提升。我们对工具描述进行精准调整后,Claude Sonnet 3.5 在 SWE-bench Verified 评测中取得了当时最先进的表现,错误率大幅降低,任务完成情况得到改善。

With your evaluation you can measure the impact of your prompt engineering with greater confidence. Even small refinements to tool descriptions can yield dramatic improvements. Claude Sonnet 3.5 achieved state-of-the-art performance on the SWE-bench Verified evaluation after we made precise refinements to tool descriptions, dramatically reducing error rates and improving task completion.

其他工具定义的最佳实践可参见我们的开发者指南。如果你正在为 Claude 构建工具,我们还建议了解工具如何动态加载到 Claude 的系统提示词中。最后,如果你正在为 MCP 服务器编写工具,工具注解有助于说明哪些工具需要访问外部开放环境,或会造成破坏性修改。

You can find other best practices for tool definitions in our Developer Guide. If you’re building tools for Claude, we also recommend reading about how tools are dynamically loaded into Claude’s system prompt. Lastly, if you’re writing tools for an MCP server, tool annotations help disclose which tools require open-world access or make destructive changes.

展望未来

Looking ahead

要为 Agent 构建有效工具,我们需要调整软件开发实践,将重心从可预测的确定性模式转向非确定性模式。

To build effective tools for agents, we need to re-orient our software development practices from predictable, deterministic patterns to non-deterministic ones.

通过本文描述的迭代式、评测驱动的过程,我们发现了让工具成功发挥作用的一些共同规律:有效的工具经过有意识且清晰的定义,审慎使用 Agent 的上下文,能够组合成多种工作流程,并让 Agent 以直观的方式解决现实任务。

Through the iterative, evaluation-driven process we’ve described in this post, we've identified consistent patterns in what makes tools successful: Effective tools are intentionally and clearly defined, use agent context judiciously, can be combined together in diverse workflows, and enable agents to intuitively solve real-world tasks.

未来,我们预计 Agent 与世界交互的具体机制还会持续演进,从 MCP 协议的更新,到底层 LLM 本身的升级,都是如此。通过系统化、评测驱动的方法改进 Agent 的工具,我们可以确保:随着 Agent 能力增强,它们使用的工具也会同步演进。

In the future, we expect the specific mechanisms through which agents interact with the world to evolve—from updates to the MCP protocol to upgrades to the underlying LLMs themselves. With a systematic, evaluation-driven approach to improving tools for agents, we can ensure that as agents become more capable, the tools they use will evolve alongside them.

致谢

Acknowledgements

本文由 Ken Aizawa 撰写,研究团队(Barry Zhang、Zachary Witten、Daniel Jiang、Sami Al-Sheikh、Matt Bell、Maggie Vo)、MCP 团队(Theodora Chu、John Welsh、David Soria Parra、Adam Jones)、产品工程团队(Santiago Seira)、市场团队(Molly Vorwerck)、设计团队(Drew Roper)和应用 AI 团队(Christian Ryan、Alexander Bricken)的同事作出了宝贵贡献。

Written by Ken Aizawa with valuable contributions from colleagues across Research (Barry Zhang, Zachary Witten, Daniel Jiang, Sami Al-Sheikh, Matt Bell, Maggie Vo), MCP (Theodora Chu, John Welsh, David Soria Parra, Adam Jones), Product Engineering (Santiago Seira), Marketing (Molly Vorwerck), Design (Drew Roper), and Applied AI (Christian Ryan, Alexander Bricken).

1本文讨论的是训练底层 LLM 本身之外的方法。

1Beyond training the underlying LLMs themselves.

Interlocking puzzle piece with complex geometric shape and detailed surface texture

想进一步了解?

Looking to learn more?

探索课程

Explore courses

— 全文完 —

原文来自 Anthropic,中文为非官方学习译文。
查看原始出处 ↗

点击空白处或按 Esc 关闭