资料馆/基础与架构
OpenAI阅读档案 · 非官方中文译文

构建 Agent 的实用指南A practical guide to building agents

下载 PDF
原版 PDF ↓
完整译文与原文逐段对应。图片、图注、表格和代码保留原文。A complete reading edition. Figures, captions, tables and code are preserved from the source.
中文译文ENGLISH ORIGINAL
cover art (original PDF, page 1)

目录

Contents

什么是 Agent? 4
何时应该构建 Agent? 5
Agent 设计基础 7
安全护栏 24
结语 32

What is an agent? 4
When should you build an agent? 5
Agent design foundations 7
Guardrails 24
Conclusion 32

引言

Introduction

大语言模型处理复杂、多步骤任务的能力正日益增强。推理、多模态和工具使用方面的进步,催生了一类由 LLM 驱动的新系统,称为 Agent。

Large language models are becoming increasingly capable of handling complex, multi-step tasks. Advances in reasoning, multimodality, and tool use have unlocked a new category of LLM-powered systems known as agents.

本指南面向探索如何构建首个 Agent 的产品与工程团队,将大量客户部署实践中的经验提炼为实用、可操作的最佳实践。内容包括识别有前景的使用场景的框架、设计 Agent 逻辑与编排的清晰模式,以及确保 Agent 安全、可预测且有效运行的最佳实践。

This guide is designed for product and engineering teams exploring how to build their first agents, distilling insights from numerous customer deployments into practical and actionable best practices. It includes frameworks for identifying promising use cases, clear patterns for designing agent logic and orchestration, and best practices to ensure your agents run safely, predictably, and effectively.

读完本指南,你将掌握所需的基础知识,有信心开始构建自己的第一个 Agent。

After reading this guide, you’ll have the foundational knowledge you need to confidently start building your first agent.

什么是 Agent?

What is an agent?

传统软件使用户能够简化并自动化工作流,而 Agent 能以高度独立的方式,代表用户执行同样的工作流。

While conventional software enables users to streamline and automate workflows, agents are able to perform the same workflows on the users’ behalf with a high degree of independence.

Agent 是能够代表你独立完成任务的系统。

Agents are systems that independently accomplish tasks on your behalf.

工作流是为实现用户目标而必须执行的一系列步骤,无论目标是解决客服问题、预订餐厅、提交代码变更,还是生成报告。

A workflow is a sequence of steps that must be executed to meet the user’s goal, whether that's resolving a customer service issue, booking a restaurant reservation, committing a code change, or generating a report.

集成了 LLM,却不使用 LLM 控制工作流执行的应用,例如简单聊天机器人、单轮 LLM 应用或情感分类器,都不是 Agent。

Applications that integrate LLMs but don’t use them to control workflow execution—think simple chatbots, single-turn LLMs, or sentiment classifiers—are not agents.

更具体地说,Agent 具备一些核心特征,使其能够可靠、一致地代表用户行动:

More concretely, an agent possesses core characteristics that allow it to act reliably and consistently on behalf of a user:

  1. 它利用 LLM 管理工作流执行并作出决策。它能识别工作流何时完成,并在需要时主动纠正自身行动。如果失败,它可以停止执行,将控制权交还用户。
  2. 它能够访问多种工具,与外部系统交互,既用于获取上下文,也用于采取行动;并根据工作流当前状态动态选择合适的工具,始终在明确定义的安全护栏内运行。
  1. It leverages an LLM to manage workflow execution and make decisions. It recognizes when a workflow is complete and can proactively correct its actions if needed. In case of failure, it can halt execution and transfer control back to the user.
  2. It has access to various tools to interact with external systems—both to gather context and to take actions—and dynamically selects the appropriate tools depending on the workflow’s current state, always operating within clearly defined guardrails.

何时应该构建 Agent?

When should you build an agent?

构建 Agent,需要重新思考系统如何作出决策和处理复杂性。不同于传统自动化,Agent 尤其适合传统确定性方法和基于规则的方法难以胜任的工作流。

Building agents requires rethinking how your systems make decisions and handle complexity. Unlike conventional automation, agents are uniquely suited to workflows where traditional deterministic and rule-based approaches fall short.

以支付欺诈分析为例。传统规则引擎像一份检查清单,根据预设条件标记交易。相比之下,LLM Agent 更像经验丰富的调查员:评估上下文、考虑细微模式,即使没有违反明确规则,也能识别可疑活动。正是这种细致的推理能力,使 Agent 能够有效应对复杂、模糊的情况。

Consider the example of payment fraud analysis. A traditional rules engine works like a checklist, flagging transactions based on preset criteria. In contrast, an LLM agent functions more like a seasoned investigator, evaluating context, considering subtle patterns, and identifying suspicious activity even when clear-cut rules aren’t violated. This nuanced reasoning capability is exactly what enables agents to manage complex, ambiguous situations effectively.

评估 Agent 能在哪些地方创造价值时,应优先考虑过去难以自动化的工作流,尤其是传统方法遇到阻碍的场景:

As you evaluate where agents can add value, prioritize workflows that have previously resisted automation, especially where traditional methods encounter friction:

Use case selection table (original PDF, page 6)

在决定投入构建 Agent 之前,请验证你的使用场景是否明确符合这些标准。否则,确定性方案可能已经足够。

Before committing to building an agent, validate that your use case can meet these criteria clearly. Otherwise, a deterministic solution may suffice.

Agent 设计基础

Agent design foundations

在最基本的形式下,Agent 由三个核心组件构成:

In its most fundamental form, an agent consists of three core components:

Agent components table (original PDF, page 7)

使用 OpenAI 的 Agents SDK 时,对应代码如下。你也可以使用偏好的库,或直接从零构建,实现同样的概念。

Here’s what this looks like in code when using OpenAI’s Agents SDK. You can also implement the same concepts using your preferred library or building directly from scratch.

Python example (original PDF, page 7)

选择模型

Selecting your models

不同模型在任务复杂度、时延和成本方面,有着不同的优势与取舍。正如下文“编排”一节将介绍的,你可能需要考虑,为工作流中的不同任务使用不同模型。

Different models have different strengths and tradeoffs related to task complexity, latency, and cost. As we’ll see in the next section on Orchestration, you might want to consider using a variety of models for different tasks in the workflow.

并非每项任务都需要最聪明的模型:简单检索或意图分类任务,可能由更小、更快的模型就能完成;而决定是否批准退款等更困难的任务,则可能受益于能力更强的模型。

Not every task requires the smartest model—a simple retrieval or intent classification task may be handled by a smaller, faster model, while harder tasks like deciding whether to approve a refund may benefit from a more capable model.

一种有效方法是,在构建 Agent 原型时,为每项任务都使用能力最强的模型,以建立性能基线。然后,尝试替换成更小的模型,观察是否仍能获得可接受的结果。这样既不会过早限制 Agent 的能力,也能诊断较小模型在哪些地方成功、在哪些地方失败。

An approach that works well is to build your agent prototype with the most capable model for every task to establish a performance baseline. From there, try swapping in smaller models to see if they still achieve acceptable results. This way, you don’t prematurely limit the agent’s abilities, and you can diagnose where smaller models succeed or fail.

总的来说,选择模型的原则很简单:

In summary, the principles for choosing a model are simple:

  1. 建立评测,以确立性能基线。
  2. 首先使用可用的最佳模型,达到准确率目标。
  3. 在可行的情况下,用较小模型替换较大模型,优化成本和时延。
  1. Set up evals to establish a performance baseline
  2. Focus on meeting your accuracy target with the best models available
  3. Optimize for cost and latency by replacing larger models with smaller ones where possible

你可以在这里找到关于选择 OpenAI 模型的完整指南。

You can find a comprehensive guide to selecting OpenAI models here.

定义工具

Defining tools

工具通过调用底层应用或系统的 API,扩展 Agent 的能力。对于没有 API 的旧系统,Agent 可以依靠计算机使用模型,通过网页和应用界面直接与这些应用和系统交互,就像人类一样。

Tools extend your agent’s capabilities by using APIs from underlying applications or systems. For legacy systems without APIs, agents can rely on computer-use models to interact directly with those applications and systems through web and application UIs—just as a human would.

每个工具都应有标准化定义,使工具与 Agent 之间能够建立灵活的多对多关系。文档完善、经过充分测试且可复用的工具,能够提高可发现性、简化版本管理,并避免重复定义。

Each tool should have a standardized definition, enabling flexible, many-to-many relationships between tools and agents. Well-documented, thoroughly tested, and reusable tools improve discoverability, simplify version management, and prevent redundant definitions.

概括来说,Agent 需要三类工具:

Broadly speaking, agents need three types of tools:

Tool categories table (original PDF, page 9)

例如,使用 Agents SDK 时,可以按如下方式,为上面定义的 Agent 配备一组工具:

For example, here’s how you would equip the agent defined above with a series of tools when using the Agents SDK:

Python example (original PDF, page 10)

随着所需工具数量增加,可以考虑将任务拆分给多个 Agent(参见“编排”)。

As the number of required tools increases, consider splitting tasks across multiple agents (see Orchestration).

配置指令

Configuring instructions

高质量指令对任何由 LLM 驱动的应用都至关重要,对 Agent 尤其如此。清晰的指令能减少歧义,改善 Agent 决策,从而让工作流执行更顺畅、错误更少。

High-quality instructions are essential for any LLM-powered app, but especially critical for agents. Clear instructions reduce ambiguity and improve agent decision-making, resulting in smoother workflow execution and fewer errors.

Agent 指令的最佳实践

Best practices for agent instructions

Agent instruction best practices table (original PDF, page 11)

你可以使用 o1 或 o3-mini 等先进模型,根据现有文档自动生成指令。下面的提示词示例展示了这种做法:

You can use advanced models, like o1 or o3-mini, to automatically generate instructions from existing documents. Here’s a sample prompt illustrating this approach:

Instruction generation prompt (original PDF, page 12)

编排

Orchestration

具备基础组件后,你就可以考虑采用编排模式,让 Agent 有效执行工作流。

With the foundational components in place, you can consider orchestration patterns to enable your agent to execute workflows effectively.

虽然人们很容易想要立即构建架构复杂、完全自主的 Agent,但客户通常采用渐进式方法会取得更好的成果。

While it’s tempting to immediately build a fully autonomous agent with complex architecture, customers typically achieve greater success with an incremental approach.

一般而言,编排模式分为两类:

In general, orchestration patterns fall into two categories:

  1. 单 Agent 系统:一个配备适当工具和指令的模型,通过循环执行工作流。
  2. 多 Agent 系统:工作流执行分布在多个相互协调的 Agent 之间。
  1. Single-agent systems, where a single model equipped with appropriate tools and instructions executes workflows in a loop
  2. Multi-agent systems, where workflow execution is distributed across multiple coordinated agents

下面详细探讨这两种模式。

Let’s explore each pattern in detail.

单 Agent 系统

Single-agent systems

通过逐步添加工具,单个 Agent 就能处理许多任务,同时将复杂性控制在可管理范围内,并简化评测与维护。每个新工具都会扩展它的能力,而不会过早迫使你编排多个 Agent。

A single agent can handle many tasks by incrementally adding tools, keeping complexity manageable and simplifying evaluation and maintenance. Each new tool expands its capabilities without prematurely forcing you to orchestrate multiple agents.

Single-agent execution diagram (original PDF, page 14)

任何编排方式都需要“运行”(run)这一概念,通常通过循环实现,让 Agent 持续运行,直到满足退出条件。常见退出条件包括工具调用、某种结构化输出、错误,或达到最大轮数。

Every orchestration approach needs the concept of a ‘run’, typically implemented as a loop that lets agents operate until an exit condition is reached. Common exit conditions include tool calls, a certain structured output, errors, or reaching a maximum number of turns.

例如,在 Agents SDK 中,通过 Runner.run() 方法启动 Agent,该方法循环调用 LLM,直到出现以下任一情况:

For example, in the Agents SDK, agents are started using the Runner.run() method, which loops over the LLM until either:

  1. 调用了由特定输出类型定义的最终输出工具。
  2. 模型返回不包含任何工具调用的响应(例如直接面向用户的消息)。
  1. A final-output tool is invoked, defined by a specific output type
  2. The model returns a response without any tool calls (e.g., a direct user message)

使用示例:

Example usage:

Python example (original PDF, page 15)

这种 while 循环的概念,是 Agent 运作的核心。正如下文将介绍的,在多 Agent 系统中,可以存在一系列工具调用以及 Agent 之间的移交,同时允许模型连续执行多个步骤,直到满足退出条件。

This concept of a while loop is central to the functioning of an agent. In multi-agent systems, as you’ll see next, you can have a sequence of tool calls and handoffs between agents but allow the model to run multiple steps until an exit condition is met.

使用提示词模板,是一种无需转向多 Agent 框架就能管理复杂性的有效策略。不必为不同使用场景维护大量独立提示词,而是使用一个灵活的基础提示词,接收策略变量。这种模板方法容易适应各种上下文,能显著简化维护与评测。出现新的使用场景时,只需更新变量,而不必重写整个工作流。

An effective strategy for managing complexity without switching to a multi-agent framework is to use prompt templates. Rather than maintaining numerous individual prompts for distinct use cases, use a single flexible base prompt that accepts policy variables. This template approach adapts easily to various contexts, significantly simplifying maintenance and evaluation. As new use cases arise, you can update variables rather than rewriting entire workflows.

Prompt template example (original PDF, page 15)

何时考虑创建多个 Agent

When to consider creating multiple agents

我们的一般建议是,先充分发挥单个 Agent 的能力。更多 Agent 可以直观地分离不同概念,但也会引入额外复杂性和开销,因此,很多时候一个配备工具的 Agent 就已经足够。

Our general recommendation is to maximize a single agent’s capabilities first. More agents can provide intuitive separation of concepts, but can introduce additional complexity and overhead, so often a single agent with tools is sufficient.

对于许多复杂工作流,将提示词和工具分配给多个 Agent,可以提高性能与可扩展性。当 Agent 无法遵循复杂指令,或持续选错工具时,你可能需要进一步拆分系统,引入职责更明确的 Agent。

For many complex workflows, splitting up prompts and tools across multiple agents allows for improved performance and scalability. When your agents fail to follow complicated instructions or consistently select incorrect tools, you may need to further divide your system and introduce more distinct agents.

拆分 Agent 的实用指导原则包括:

Practical guidelines for splitting agents include:

Multi-agent selection table (original PDF, page 16)

多 Agent 系统

Multi-agent systems

虽然针对具体工作流和需求,可以用多种方式设计多 Agent 系统,但根据我们与客户合作的经验,有两类模式具有广泛适用性:

While multi-agent systems can be designed in numerous ways for specific workflows and requirements, our experience with customers highlights two broadly applicable categories:

Multi-agent patterns table (original PDF, page 17)

多 Agent 系统可以建模为图,其中 Agent 表示为节点。在管理者模式中,边表示工具调用;在去中心化模式中,边表示在 Agent 之间转移执行权的移交。

Multi-agent systems can be modeled as graphs, with agents represented as nodes. In the manager pattern, edges represent tool calls whereas in the decentralized pattern, edges represent handoffs that transfer execution between agents.

无论采用哪种编排模式,同样的原则都适用:保持组件灵活、可组合,并由清晰、结构良好的提示词驱动。

Regardless of the orchestration pattern, the same principles apply: keep components flexible, composable, and driven by clear, well-structured prompts.

管理者模式

Manager pattern

管理者模式让一个中心 LLM——“管理者”——通过工具调用,无缝编排一个专门化 Agent 网络。管理者不会丢失上下文或控制权,而是智能地在合适时间,将任务委派给合适的 Agent,再轻松地将结果整合为连贯的交互。这能确保流畅、统一的用户体验,同时让各种专门能力始终可以按需调用。

The manager pattern empowers a central LLM—the “manager”—to orchestrate a network of specialized agents seamlessly through tool calls. Instead of losing context or control, the manager intelligently delegates tasks to the right agent at the right time, effortlessly synthesizing the results into a cohesive interaction. This ensures a smooth, unified user experience, with specialized capabilities always available on-demand.

如果你只希望由一个 Agent 控制工作流执行并与用户接触,这种模式就非常适合。

This pattern is ideal for workflows where you only want one agent to control workflow execution and have access to the user.

Manager pattern diagram (original PDF, page 18)

例如,可以在 Agents SDK 中按如下方式实现这种模式:

For example, here’s how you could implement this pattern in the Agents SDK:

Manager pattern Python example — part 1 (original PDF, page 19)
Manager pattern Python example — part 2 (original PDF, page 20)

声明式图与非声明式图

Declarative vs non-declarative graphs

有些框架采用声明式方式,要求开发者事先通过由节点(Agent)和边(确定性或动态移交)组成的图,显式定义工作流中的每个分支、循环和条件。这种方式有助于直观展示结构,但随着工作流变得更动态、更复杂,它很快就会变得繁琐且难以处理,通常还需要学习专门的领域特定语言。

Some frameworks are declarative, requiring developers to explicitly define every branch, loop, and conditional in the workflow upfront through graphs consisting of nodes (agents) and edges (deterministic or dynamic handoffs). While beneficial for visual clarity, this approach can quickly become cumbersome and challenging as workflows grow more dynamic and complex, often necessitating the learning of specialized domain-specific languages.

相比之下,Agents SDK 采用更灵活、以代码为先的方式。开发者可以直接用熟悉的编程结构表达工作流逻辑,无需事先定义整张图,从而实现更动态、更具适应性的 Agent 编排。

In contrast, the Agents SDK adopts a more flexible, code-first approach. Developers can directly express workflow logic using familiar programming constructs without needing to pre-define the entire graph upfront, enabling more dynamic and adaptable agent orchestration.

去中心化模式

Decentralized pattern

在去中心化模式中,Agent 可以将工作流执行“移交”(handoff)给彼此。移交是一种单向转移,允许一个 Agent 将任务委派给另一个 Agent。在 Agents SDK 中,移交是一种工具或函数。如果 Agent 调用了移交函数,我们就会立即开始执行接收移交的新 Agent,同时转移最新的对话状态。

In a decentralized pattern, agents can ‘handoff’ workflow execution to one another. Handoffs are a one way transfer that allow an agent to delegate to another agent. In the Agents SDK, a handoff is a type of tool, or function. If an agent calls a handoff function, we immediately start execution on that new agent that was handed off to while also transferring the latest conversation state.

这种模式使用多个地位平等的 Agent,一个 Agent 可以直接把工作流控制权交给另一个 Agent。当你不需要单个 Agent 维持集中控制或汇总结果,而是希望每个 Agent 按需接管执行并与用户交互时,这种模式最为合适。

This pattern involves using many agents on equal footing, where one agent can directly hand off control of the workflow to another agent. This is optimal when you don’t need a single agent maintaining central control or synthesis—instead allowing each agent to take over execution and interact with the user as needed.

Decentralized pattern diagram (original PDF, page 21)

例如,对于同时处理销售和支持的客服工作流,可以使用 Agents SDK 按如下方式实现去中心化模式:

For example, here’s how you’d implement the decentralized pattern using the Agents SDK for a customer service workflow that handles both sales and support:

Decentralized pattern Python example — part 1 (original PDF, page 22)
Decentralized pattern Python example — part 2 (original PDF, page 23)

在上面的示例中,最初的用户消息会发送给 triage_agent。triage_agent 识别出输入与近期购买有关后,会调用移交,将控制权转给 order_management_agent。

In the above example, the initial user message is sent to triage_agent. Recognizing that the input concerns a recent purchase, the triage_agent would invoke a handoff to the order_management_agent, transferring control to it.

这种模式特别适合对话分流等场景,或任何希望专门化 Agent 完全接管某些任务、而不再需要原 Agent 继续参与的情况。你也可以为第二个 Agent 配备移交回原 Agent 的能力,让它在必要时再次转移控制权。

This pattern is especially effective for scenarios like conversation triage, or whenever you prefer specialized agents to fully take over certain tasks without the original agent needing to remain involved. Optionally, you can equip the second agent with a handoff back to the original agent, allowing it to transfer control again if necessary.

安全护栏

Guardrails

设计良好的安全护栏可以帮助你管理数据隐私风险(例如防止系统提示词泄露)或声誉风险(例如确保模型行为符合品牌要求)。你可以针对使用场景中已识别的风险设置护栏,并在发现新漏洞时逐层添加更多护栏。护栏是任何基于 LLM 的部署的关键组成部分,但还应配合健全的身份验证与授权协议、严格的访问控制,以及标准的软件安全措施。

Well-designed guardrails help you manage data privacy risks (for example, preventing system prompt leaks) or reputational risks (for example, enforcing brand aligned model behavior). You can set up guardrails that address risks you’ve already identified for your use case and layer in additional ones as you uncover new vulnerabilities. Guardrails are a critical component of any LLM-based deployment, but should be coupled with robust authentication and authorization protocols, strict access controls, and standard software security measures.

可以将护栏视为分层防御机制。单个护栏通常不足以提供充分保护,而组合使用多个专门化护栏,可以构建更稳健的 Agent。

Think of guardrails as a layered defense mechanism. While a single one is unlikely to provide sufficient protection, using multiple, specialized guardrails together creates more resilient agents.

在下图中,我们结合了基于 LLM 的护栏、正则表达式等基于规则的护栏,以及 OpenAI 内容审核 API,来检查用户输入。

In the diagram below, we combine LLM-based guardrails, rules-based guardrails such as regex, and the OpenAI moderation API to vet our user inputs.

Layered guardrails diagram (original PDF, page 25)

护栏的类型

Types of guardrails

Types of guardrails table — part 1 (original PDF, page 26)
Types of guardrails table — part 2 (original PDF, page 27)

构建护栏

Building guardrails

针对使用场景中已经识别出的风险设置护栏,并在发现新漏洞时逐层添加更多护栏。

Set up guardrails that address the risks you’ve already identified for your use case and layer in additional ones as you uncover new vulnerabilities.

我们发现,以下经验原则很有效:

We’ve found the following heuristic to be effective:

  1. 关注数据隐私与内容安全。
  2. 根据实际遇到的边界情况和失败,添加新护栏。
  3. 同时优化安全性和用户体验,随着 Agent 演进,不断调整护栏。
  1. Focus on data privacy and content safety
  2. Add new guardrails based on real-world edge cases and failures you encounter
  3. Optimize for both security and user experience, tweaking your guardrails as your agent evolves.

例如,使用 Agents SDK 时,可以按如下方式设置护栏:

For example, here’s how you would set up guardrails when using the Agents SDK:

Guardrail Python example — part 1 (original PDF, page 28)
Guardrail Python example — part 2 (original PDF, page 29)
Guardrail Python example — part 3 (original PDF, page 30)

Agents SDK 将护栏视为一等概念,默认采用乐观执行。在这种方式下,主 Agent 主动生成输出,护栏同时并发运行;如果违反约束,就触发异常。

The Agents SDK treats guardrails as first-class concepts, relying on optimistic execution by default. Under this approach, the primary agent proactively generates outputs while guardrails run concurrently, triggering exceptions if constraints are breached.

护栏可以实现为函数或 Agent,用来执行防止越狱、相关性验证、关键词过滤、阻止列表约束或安全分类等策略。例如,上面的 Agent 会先以乐观方式处理数学问题输入,直到 math_homework_tripwire 护栏识别出违规并抛出异常。

Guardrails can be implemented as functions or agents that enforce policies such as jailbreak prevention, relevance validation, keyword filtering, blocklist enforcement, or safety classification. For example, the agent above processes a math question input optimistically until the math_homework_tripwire guardrail identifies a violation and raises an exception.

为人工介入做好准备

Plan for human intervention

人工介入是一项关键保障,使你能够在不损害用户体验的情况下,改善 Agent 在现实环境中的表现。它在部署初期尤其重要,有助于发现失败、识别边界情况,并建立健全的评测循环。

Human intervention is a critical safeguard enabling you to improve an agent’s real-world performance without compromising user experience. It’s especially important early in deployment, helping identify failures, uncover edge cases, and establish a robust evaluation cycle.

实现人工介入机制,可以让 Agent 在无法完成任务时,平稳地转移控制权。在客服场景中,这意味着将问题升级给人工客服;对于编程 Agent,则意味着把控制权交还用户。

Implementing a human intervention mechanism allows the agent to gracefully transfer control when it can’t complete a task. In customer service, this means escalating the issue to a human agent. For a coding agent, this means handing control back to the user.

通常有两类主要触发条件需要人工介入:

Two primary triggers typically warrant human intervention:

超过失败阈值:为 Agent 的重试次数或行动次数设置上限。如果 Agent 超过这些限制(例如多次尝试后仍无法理解客户意图),就升级到人工介入。

Exceeding failure thresholds: Set limits on agent retries or actions. If the agent exceeds these limits (e.g., fails to understand customer intent after multiple attempts), escalate to human intervention.

高风险行动:对于敏感、不可逆或影响重大的行动,在对 Agent 的可靠性建立足够信心之前,应触发人工监督。例如取消用户订单、批准大额退款或付款。

High-risk actions: Actions that are sensitive, irreversible, or have high stakes should trigger human oversight until confidence in the agent’s reliability grows. Examples include canceling user orders, authorizing large refunds, or making payments.

结语

Conclusion

Agent 标志着工作流自动化进入新时代:系统能够在模糊情境中推理、跨工具采取行动,并以高度自主的方式处理多步骤任务。与较简单的 LLM 应用不同,Agent 能端到端地执行工作流,因此非常适合涉及复杂决策、非结构化数据或脆弱规则系统的使用场景。

Agents mark a new era in workflow automation, where systems can reason through ambiguity, take action across tools, and handle multi-step tasks with a high degree of autonomy. Unlike simpler LLM applications, agents execute workflows end-to-end, making them well-suited for use cases that involve complex decisions, unstructured data, or brittle rule-based systems.

要构建可靠的 Agent,应从扎实的基础开始:将能力强的模型,与定义完善的工具、清晰且结构化的指令配合使用。采用与复杂程度相匹配的编排模式,从单个 Agent 起步,仅在需要时演进为多 Agent 系统。从输入过滤和工具使用,到有人参与的介入机制,护栏在每个阶段都至关重要,有助于确保 Agent 在生产环境中安全、可预测地运行。

To build reliable agents, start with strong foundations: pair capable models with well-defined tools and clear, structured instructions. Use orchestration patterns that match your complexity level, starting with a single agent and evolving to multi-agent systems only when needed. Guardrails are critical at every stage, from input filtering and tool use to human-in-the-loop intervention, helping ensure agents operate safely and predictably in production.

成功部署的道路并非非此即彼。从小规模开始,用真实用户验证,再逐步扩展能力。有了正确的基础和迭代方法,Agent 就能创造真实的业务价值:不仅将单个任务自动化,还能以智能和适应能力,将整个工作流自动化。

The path to successful deployment isn’t all-or-nothing. Start small, validate with real users, and grow capabilities over time. With the right foundations and an iterative approach, agents can deliver real business value—automating not just tasks, but entire workflows with intelligence and adaptability.

如果你正在为组织探索 Agent,或准备首次部署,欢迎联系我们。我们的团队可以提供专业知识、指导和实际支持,帮助你取得成功。

If you’re exploring agents for your organization or preparing for your first deployment, feel free to reach out. Our team can provide the expertise, guidance, and hands-on support to ensure your success.

更多资源

More resources

OpenAI 是一家 AI 研究与部署公司。我们的使命是确保通用人工智能造福全人类。

OpenAI is an AI research and deployment company. Our mission is to ensure that artificial general intelligence benefits all of humanity.

— 全文完 —

原文来自 OpenAI,中文为非官方学习译文。
查看原始出处 ↗

点击空白处或按 Esc 关闭