资料馆/基础与架构
LangChain阅读档案 · 非官方中文译文

Agent 工程:一门新学科Agent Engineering: A New Discipline

下载 PDF
中文 PDF ↓英文 PDF ↓
完整译文与原文逐段对应。图片、图注、表格和代码保留原文。A complete reading edition. Figures, captions, tables and code are preserved from the source.
中文译文ENGLISH ORIGINAL

如果你构建过 Agent,就会知道,“在我的机器上能用”与“在生产环境中能用”之间,可能存在巨大的差距。传统软件假设你大体了解输入,并且能够定义输出。Agent 却两者都无法保证:用户真的可以说任何话,而可能出现的行为几乎没有边界。这正是它们强大的原因,也正是它们可能以你未曾预料的方式偏离正轨的原因。

If you’ve built an agent, you know that the delta between “it works on my machine” and “it works in production” can be huge. Traditional software assumes you mostly know the inputs and can define the outputs. Agents give you neither: users can say literally anything, and the space of possible behaviors is wide open. That’s why they’re powerful — and why they can also go a little sideways in ways you didn’t see coming.

过去 3 年,我们见过数千个团队努力应对这一现实。那些成功将可靠产品交付到生产环境的团队——例如 Clay、Vanta、LinkedIn 和 Cloudflare——并没有照搬传统软件开发的方法。它们正在开创一套新实践:Agent 工程。

Over the past 3 years, we’ve watched thousands of teams struggle with this reality. The ones who’ve succeeded in shipping something reliable to production — companies like Clay, Vanta, LinkedIn, and Cloudflare — aren’t following the traditional software playbook. They’re pioneering something new: agent engineering.

什么是 Agent 工程?

What is agent engineering?

Agent 工程,是通过迭代改进,将非确定性的 LLM 系统打造成可靠生产体验的过程。这是一个循环过程:构建、测试、发布、观察、改进,再次循环。

Agent engineering is the iterative process of refining non-deterministic LLM systems into reliable production experiences. It is a cyclical process: build, test, ship, observe, refine, repeat.

关键在于,发布并不是最终目标。它只是让你持续获得新认识、改进 Agent 的方式。要做出有意义的改进,就需要理解生产环境中正在发生什么。完成这个循环的速度越快,Agent 就越可靠。

The key here is that shipping isn't the end goal. It’s just the way you keep moving to get new insights and improve your agent. To make improvements that matter, you need to understand what’s happening in production. The faster you move through this cycle, the more reliable your agent becomes.

我们把 Agent 工程视为一门新学科,它将 3 类技能结合起来,让它们协同发挥作用:

We see agent engineering as a new discipline that combines 3 skillsets working together:

  • 产品思维定义任务范围,塑造 Agent 行为。这包括:
    • 编写驱动 Agent 行为的提示词(通常有数百行,甚至数千行)。良好的沟通和写作能力在这里至关重要。
    • 深入理解 Agent 所复现的“待完成的工作”。
    • 定义评测,检验 Agent 的表现是否符合这项“待完成的工作”的要求。
  • 工程能力构建让 Agent 能够投入生产的基础设施。这包括:
    • 编写供 Agent 使用的工具。
    • 开发 Agent 交互的 UI/UX,包括流式输出、中断处理等。
    • 构建稳健的运行时,处理持久执行、需要人类介入时的暂停,以及记忆管理。
  • 数据科学持续衡量并改进 Agent 的表现。这包括:
    • 构建评测、A/B 测试、监控等系统,衡量 Agent 的表现与可靠性。
    • 分析使用模式并开展错误分析,因为相较于传统软件,用户使用 Agent 的方式更加广泛。
  • Product thinking defines the scope and shapes agent behavior. This involves:
    • Writing prompts that drive agent behavior (often hundreds or thousands of lines). Good communication and writing skills are key here.
    • Deeply understanding the "job to be done" that the agent replicates
    • Defining evaluations that test whether the agent performs as intended by the “job to be done”
  • Engineering builds the infrastructure that makes agents production-ready. This involves:
    • Writing tools for agents to use
    • Developing UI/UX for agent interactions (with streaming, interrupt handling, etc.)
    • Creating robust runtimes that handle durable execution, human-in-the-loop pauses, and memory management.
  • Data science measures and improves agent performance over time. This involves:
    • Building systems (evals, A/B testing, monitoring etc.) to measure agent performance and reliability
    • Analyzing usage patterns and error analysis (since agents have a broader scope of how users use them than traditional software)

Agent 工程体现在哪里

Where agent engineering shows up

Agent 工程并不是一种新的岗位名称。它是一组职责:当现有团队开始构建能够推理、适应变化且行为难以预测的系统时,就需要承担这些职责。如今能交付可靠 Agent 的组织,正在拓展工程、产品与数据团队的技能,以满足非确定性系统的要求。

Agent engineering isn’t a new job title. Instead, it’s a set of responsibilities that existing teams take on when they’re building systems that reason, adapt, and behave unpredictably. The organizations shipping reliable agents today are extending the skills of engineering, product, and data teams to meet the demands of non-deterministic systems.

这套实践通常体现在以下工作中:

Here’s where the practice typically shows up:

  • 软件工程师和机器学习工程师编写提示词、构建供 Agent 使用的工具,追踪 Agent 为什么发起特定的工具调用,并改进底层模型。
  • 平台工程师构建 Agent 基础设施,处理持久执行与人类参与的工作流。
  • 产品经理编写提示词、定义 Agent 的任务范围,并确保 Agent 解决的是正确的问题。
  • 数据科学家衡量 Agent 的可靠性,识别改进机会。
  • Software engineers and ML engineers writing prompts and building tools for agents to use, tracing why an agent made specific tool calls, and refining the underlying models
  • Platform engineers building agent infrastructure that handles durable execution and human-in-the-loop workflows
  • Product managers writing prompts, defining agent scope, and ensuring the agent solves the right problem
  • Data scientists measuring agent reliability and identifying opportunities for improvement

这些团队拥抱快速迭代。你经常会看到,软件工程师追踪错误后,把发现交给产品经理据此调整提示词;也可能是产品经理发现任务范围存在问题,需要工程师提供新工具。大家都认识到:让 Agent 变得稳健,真正依靠的是这样的循环——观察生产环境中的行为,再根据所学到的东西系统地改进。

These teams embrace rapid iteration, and you'll often see software engineers tracing errors and handing off to PMs to tweak prompts based on those insights, or PMs identifying scope issues that require new tools from engineers. Each recognizes that the real work of hardening an agent happens through this cycle of observing production behavior and systematically refining based on what they learn.

为什么需要 Agent 工程?为什么是现在?

Why agent engineering, and why now?

两个根本性的变化,使 Agent 工程成为必要。

Two fundamental shifts have made agent engineering necessary.

首先,LLM 已经足够强大,能够处理复杂的多步骤工作流。我们已经看到,Agent 开始承担整项工作,而不只是单个任务。Clay 用 Agent 处理从潜在客户研究到个性化联系、再到 CRM 更新的全部流程。LinkedIn 用 Agent 为招聘筛查海量人才库,对候选人排序,并即时呈现最合适的人选。我们正在跨越一个门槛:Agent 开始在生产环境中交付实质性的商业价值。

First, LLMs are powerful enough to handle complex, multi-step workflows. We’ve been seeing it with agents taking on whole jobs, not just tasks. Clay uses agents to handle everything from prospect research to personalized outreach and CRM updates. LinkedIn uses agents to scan massive talent pools for recruiting, ranking candidates and surfacing the strongest matches instantly. We’re starting to cross the threshold where agents are delivering meaningful business value in production.

其次,这种能力伴随着真实存在的不可预测性。简单的 LLM 应用虽然也具有非确定性,但其行为范围通常更受限制。Agent 则不同。它们会进行多步推理、调用工具,并根据上下文调整行为。让 Agent 变得有用的那些特性,也使其行为有别于传统软件。这通常意味着:

Second, that power comes with real unpredictability. Simple LLM apps, though non-deterministic, tend to have more contained behavior. Agents are different. They reason across multiple steps, call tools, and adapt based on context. The same things that make agents useful also make them behave differently than traditional software. This usually means that:

  • 每一种输入都可能是边缘情况。当用户可以用自然语言询问任何事情时,就不存在所谓“正常”的输入。当你输入“让它更出彩”,或者“按上次那样做,但换一种方式”时,Agent 和人一样,可能对提示词作出不同解读。
  • 你不能再用老办法调试。由于大量逻辑存在于模型内部,你必须检查每一次决策与工具调用。提示词或配置中的细微调整,都可能造成行为上的巨大变化。
  • “能用”并不是一个非黑即白的判断。Agent 的可用率即使达到 99.99%,仍可能偏离正轨、无法正常完成工作。对于真正重要的问题,答案未必只是简单的是或否:Agent 作出的决策是否正确?使用工具的方式是否恰当?是否遵循了指令背后的意图?
  • Every input is an edge case. There's no such thing as a "normal" input when users can ask anything in natural language. When you type in “make it pop” or “do what you did last time but differently”, the agent (just like a human) can interpret the prompts in different ways.
  • You can’t debug the old way. Because so much logic lives inside the model, you have to inspect each decision and tool call. Small prompt or config tweaks can create huge shifts in behavior.
  • “Working” isn’t binary. An agent can have 99.99% uptime while still being off the rails and broken. There aren’t always simple yes/no answers to the questions that matter, like: is the agent making the right calls? Using tools the right way? Following the intent behind your instructions?

把这一切放在一起看:Agent 正在运行真实且影响重大的工作流,却又展现出传统软件方法难以应对的行为。这为一门新学科创造了机会,也产生了需求。Agent 工程让你能够发挥 LLM 的能力,同时构建在生产环境中真正值得信赖的系统。

When you put this all together — agents running real, high impact workflows yet behaving in ways that traditional software can’t solve — there’s an opportunity and the need for a new discipline. Agent engineering lets you harness the power of LLMs while building systems you can actually trust in production.

Agent 工程在实践中是什么样的?

What does agent engineering look like in practice?

Agent 工程遵循的原则与传统软件开发不同。要获得可靠的 Agent 系统,发布本身就是学习的方式,而不是学完以后才做的事。

Agent engineering operates on a different principle than traditional software development. To achieve a reliable agent system, shipping is how you learn, not what you do after learning.

我们看到,成功的工程团队通常按类似下面的节奏开发 Agent:

We’ve seen successful eng teams follow a cadence for agent development that looks something like this:

  • 构建 Agent 的基础。从设计 Agent 的基础结构开始,无论它是一次带工具的简单 LLM 调用,还是复杂的多 Agent 系统。架构取决于你需要多少工作流(确定性的逐步流程),以及多少自主性(由 LLM 驱动的决策)。
  • 针对你能想到的场景进行测试。用示例场景测试 Agent,发现提示词、工具定义和工作流中的明显问题。传统软件可以预先规划用户流程,而对于自然语言输入,你无法预见用户交互的每一种方式。因此,要将思路从“穷尽测试,然后发布”转变为“合理测试,通过发布弄清什么真正重要”。
  • 通过发布观察真实世界的行为。一旦发布,你就会立刻看到未曾考虑过的输入;每一条生产环境执行轨迹,都展示了 Agent 实际需要处理什么。
  • 观察。追踪每一次交互,查看完整对话、每一个被调用的工具,以及 Agent 每次决策所依据的确切上下文。针对生产数据运行评测,衡量 Agent 的质量,无论你关注的是准确性、时延、用户满意度,还是其他标准。
  • 改进。识别出失败模式后,通过编辑提示词、修改工具定义来改进。这个过程始终持续进行:你可以把有问题的案例重新加入示例场景集,用于回归测试。
  • 重复。发布改进,观察生产环境发生了什么变化。每一次循环都会让你对用户如何与 Agent 交互,以及在你的场景中可靠性究竟意味着什么,获得新的认识。
  • Build your agent’s foundation. Start with designing your agent's foundation, whether it's a simple LLM call with tools or a complex multi-agent system. Your architecture depends on how much workflow (deterministic step-by-step processes) versus agency (LLM-driven decisions) you need.
  • Test based on scenario you can imagine. Test your agent against example scenarios to catch obvious issues with prompts, tool definitions, and workflows. Unlike traditional software where you can map out user flows, you can't anticipate every way users will interact with natural language input. Shift your mindset from "test exhaustively, then ship" to "test reasonably, ship to learn what actually matters.”
  • Ship to see real-world behavior. Once you ship, you’ll immediately start seeing inputs you hadn’t considered and every production trace shows what your agent actually needs to handle.
  • Observe. Trace every every interaction to see the full conversation, every tool called, and the exact context that informed each decision the agent made. Run evals over your production data to measure agent quality, whether you care about accuracy, latency, user satisfaction, or other criteria.
  • Refine. Once you’ve identified patterns in what's failing, refine by editing prompts and modifying tool definitions. It’s all continuous, as you can add problematic cases back to your set of example scenarios for regression testing.
  • Repeat. Ship your improvements and watch what’s changing in production. Each cycle teaches you something new about how users are interacting with your agent and what reliability actually means in your context.

工程实践的新标准

A new standard for engineering

如今能交付可靠 Agent 的团队有一个共同点:它们不再试图在发布前就让 Agent 完美无缺,而是开始把生产环境当作最主要的老师。换句话说,就是追踪每一次决策、大规模开展评测,并以天为单位发布改进,而不是按季度推进。

The teams shipping reliable agents today share one thing: they've stopped trying to perfect agents before launch and started treating production as their primary teacher. In other words, tracing every decision, evaluating at scale, and shipping improvements in days instead of quarters.

Agent 工程正在兴起,是因为新的机会要求我们这样做。Agent 已经能够处理过去需要人类判断的工作流,但前提是你能够让它足够可靠、值得信赖。这里没有捷径,只有系统性的迭代工作。问题不在于 Agent 工程会不会成为标准实践,而在于你的团队能多快采用它,从而释放 Agent 的能力。

Agent engineering is emerging because the opportunity demands it. Agents can now handle workflows that previously required human judgment, but only if you can make them reliable enough to trust. There is no shortcut, just the systematic work of iteration. The question isn't whether agent engineering will become standard practice. It's how quickly your team can adopt it to unlock what agents can do.

— 全文完 —

原文来自 LangChain,中文为非官方学习译文。
查看原始出处 ↗

点击空白处或按 Esc 关闭