Agent 在循环中运行:LLM 决定做什么,工具执行操作,模型评估结果,然后继续循环,直到任务完成。
Agents run in a loop: an LLM decides what to do, a tool executes, a model evaluates the results, and then continues in that loop until the task is complete.
起初,Agent 和 LLM 很难集成到软件应用中,因为软件应用依赖结构化数据与可预测的接口。后来出现了两项基础能力,让集成容易了许多:
Agents and LLMs were initially difficult to integrate into software applications, which depend on structured data and predictable interfaces. Two primitives emerged that made this much easier:
- Tool calling let models make structured requests and receive structured results.
- Structured outputs let models return restructured results.
但即使具备了这些能力,Agent 循环仍然缓慢而且昂贵:每一次决策,都需要再调用一次模型。
But even with those in place, the agent loop is still slow and costly: every decision requires another model call.
这时,Jev 出现了。Jev 是 TypeSafe AI 发布的一个新模型。该公司报告称,在分类任务上,相比同类 LLM,Jev 的推理速度最高可达 200 倍,成本可低至 1/400。
Enter, Jev. Jev is a new model released from TypeSafe AI. The company reports up to 200x faster inference and 400x lower cost than comparable LLMs on classification tasks.
本文将介绍 Jev 如何工作、它适合放在 Agent 循环的哪个位置,以及如何与 LangChain 一起使用。
This post covers how Jev works, where it fits into the agent loop, and how to use it with LangChain.
了解 Jev
All about Jev
实际上,Jev 不是传统意义上的 LLM,它不会生成文本。TypeSafe AI 团队将其称为 System One 模型:
Jev is actually not a traditional LLM, it doesn’t generate text. It’s what the TypeSafe AI team calls a System One model:
📖 System One 模型是一类 AI 模型,专门用于快速做出结构化决策,供软件直接使用。System One 模型评估一个状态,并返回带类型的答案及概率。
📖 System One models are a class of AI models built to make fast, structured decisions that software can use directly. A System One model evaluates a state and returns typed answers and probabilities.
它使用面向校准决策的强化学习(RLCD)进行训练。你的代码可以利用这些结果,指导 Agent 接下来做什么,而不必为每次决策都进行一次完整的聊天 LLM 调用。
It’s trained using reinforcement learning for calibrated decisions (RLCD). Your code uses those results to guide what an agent does next, without a full chat LLM call for each decision.
调用 Jev 模型时,你需要向它发送一个状态(即上下文),以及有关该状态的问题。下面是其文档中客服工单示例的单问题版本:
To invoke a Jev model, you send it a state (the context) and questions about that state. Here’s a single-question version of the support-ticket example in their docs:
文档示例给出了下面这个关于紧急程度的答案。这里省略了响应中的其他部分:
The docs’ example gives this urgency answer, shown here without the rest of the response:
这表示该消息有 99.9% 的概率属于紧急事项,你的应用可以据此优先处理该工单。
That’s a 99.9% probability that the message is urgent, which your application can use to prioritize the ticket.
它支持三种问题类型:
There are three types of supported questions:
- Choice:从一组选项中选择。返回每个选项的概率,以及一个总体置信度分数。
- Score:按照低、中、高等有序等级对输入评分。返回一个连续分数、其底层分布,以及一个置信度值。
- Noul:回答是非问题。返回某个陈述为真的概率。
- Choice: Pick from a set of options. Returns a probability for each option and an overall confidence score.
- Score: Rate an input against ordered levels, such as low, medium, and high. Returns a continuous score, the underlying distribution, and a confidence value.
- Noul: Answer a yes-or-no question. Returns the probability that a statement is true.
这里有一个关键特性:你可以在一次请求中,针对同一个状态提出多个问题。
One key feature here is that you can ask multiple questions about the same state in one request.
💡 System One 模型会并行评估请求中的所有问题。增加问题几乎不会改变响应时间,只会增加额外问题所消耗的 token 成本,而这部分成本很低。
💡 System One models evaluate every question in a request in parallel. Adding questions barely changes the response time and costs only the tokens for the extra questions, which are cheap.
如需查看针对一张客服工单提出多个问题的示例,请参阅 TypeSafe 快速入门。
For an example of asking multiple questions about a support ticket, see the TypeSafe Quickstart.
总之,与传统 LLM 不同,Jev 既不受文本生成过程的限制,也不受顺序决策的限制!
In sum, unlike traditional LLMs, Jev is neither constrained by text generation or sequential decision making!
如何结合 LangChain 使用 Jev
How to Use Jev with LangChain
LangChain 不依赖特定供应商的设计,非常适合在支持数千种其他集成和模型供应商的同时,支持 Jev。
LangChain's provider agnostic model is well suited for supporting Jev alongside thousands of other integrations and model providers.
LangChain 集成通过 TypeSafeClassifier 提供 Jev。你将状态与问题传给 .invoke(),得到的会是分类结果,而不是聊天响应。
The LangChain integration exposes Jev through TypeSafeClassifier. You pass your state and questions to .invoke(), and get classification results rather than a chat response.
安装 langchain-typesafe,设置 TYPESAFE_API_KEY,然后发起调用:
Install langchain-typesafe and set your TYPESAFE_API_KEY, then make a call:
状态可以是文本、结构化数据,或者 LangChain 消息。因此,你可以在节点或中间件钩子中,直接利用 Agent 已有的上下文调用 Jev。
The state can be text, structured data, or LangChain messages. That makes it straightforward to call Jev from a node or middleware hook using the context your agent already has.
你可以将这种能力加入自定义中间件或工具中!
You can build this into custom middleware or tools!
使用场景
Use Cases
Jev 并不能直接替代 LLM。它不会生成文本,但能以更低的时延与成本,处理我们如今经常交给 LLM 的分类任务。这使它有望成为驱动 Agent 的模型的有益补充:用 LLM 进行开放式推理和生成,用 Jev 在过程中快速做出结构化决策。
Jev isn’t a drop-in replacement for an LLM. It doesn’t generate text, but it can handle classification tasks we often use LLMs for today, without the same latency and cost. That makes it a promising complement to the model driving your agent: use an LLM for open-ended reasoning and generation, and Jev for fast, structured decisions along the way.
模型路由
Model routing
简单查询与困难的调试任务,并不需要使用同一个模型。模型路由中间件让 Jev 评估请求,并根据你定义的标准选择模型:简单任务使用快速、便宜的模型,复杂任务则使用能力更强的模型。
A simple lookup doesn’t need the same model as a difficult debugging task. Model-routing middleware lets Jev assess the request and choose a model based on criteria you define, so fast and inexpensive for straightforward tasks, more capable for complex ones.
路由器根据最新的用户消息选择模型,并在整个运行过程中使用它。相应的概率与置信度也会保留在 Agent 状态中,供后续使用。
The router selects a model from the latest user message and uses it throughout the run. The probabilities and confidence remain available in agent state, too.
自动模式
Auto Mode
Agent 本质上仍然不值得完全信任。Agent 可能接收到不当指令,无论这些指令是自然出现的,还是由足够有动机的攻击者构造的;这些指令都可能诱使它采取我们不希望发生的行动。
Agents are still inherently untrustworthy. An agent can receive bad instructions (either naturally or from a motivated enough attacker) which can persuade it into taking actions we didn’t want it to.
Claude、Codex、Cursor 等编程运行框架,已经提供了某种机制,在危险操作执行之前对其进行分类。这逐渐帮助人们建立起对 Agent 的信任。但直到现在,这个分类步骤一直封装在运行框架的闭源部分中。
Coding harnesses like claude, codex, cursor have shipped some kind of way to classify dangerous actions before they’re taken which has slowly helped to build trust in agents. Up until now, this classifier step has been locked away in the closed source parts of the harness.
如今,既便宜又高效的分类模型已经出现,我们就可以采用同样的模式,并将其应用到所有 Agent 上!
Now that a cheap and performant classifier model exists, we can take the same pattern and adopt it to all agents!
AutoModeMiddleware 使用 Jev 检查工具调用中可能包含的高风险决策,并在工具执行前阻止这些调用。
AutoModeMiddleware uses Jev to check tool calls for risky decisions it may take, and block calls before the tool executes.
开始使用!
Get Started!
我们对 Jev 及其带来的可能性十分兴奋。我们已经看到了一些很有意思的项目:Browserbase 的 Kyle Jeong 正在以不到一美分的成本驱动浏览器操作 Agent;Jarrod Watts 构建了一个实时交易 Agent;Ryan Vogel 则在大规模进行邮件分类处理。
We're pretty thrilled about Jev and the possibilities that come with it. A few cool projects that we’ve seen already: Kyle Jeong from Browserbase is powering browser use agents for fractions of a cent, Jarrod Watts built a live trading agent, and Ryan Vogel is doing email triage at scale.
如今几乎每周都有新模型发布,但这个模型引起的反响格外强烈。我们很期待看到你用 LangChain 和 Jev 构建出什么。
New models drop every week at this point, but this one had a pretty outsized response. We’re excited to see what you build with LangChain and Jev.
欢迎在论坛上分享你的想法,在 X 上提及我们并展示你正在构建的东西,或者参与 LangChain issues 的讨论!
Let us know what you think on the forum, tag us on X and share what you’re building, or engage with LangChain issues!
— 全文完 —
原文来自 LangChain,中文为非官方学习译文。
查看原始出处 ↗
