<?xml version="1.0" encoding="UTF-8"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>6yuan博客 · 更新订阅</title><link>https://blog.wanglvyuan.com/</link><description>6yuan 的个人博客。收藏关于 AI Agent、上下文工程与智能体运行框架的原始资料和中文译文。</description><language>zh-CN</language><atom:link href="https://blog.wanglvyuan.com/feed.xml" rel="self" type="application/rss+xml"/><item><title>构建高效的 Agent</title><link>https://blog.wanglvyuan.com/library/building/</link><guid isPermaLink="true">https://blog.wanglvyuan.com/library/building/</guid><description>Anthropic 工程文章的中文译文，区分预定义工作流与模型自主决策的 Agent，介绍提示链、路由、并行和编排等常见模式，并通过客服与编程案例讨论工具接口设计、复杂性、成本和适用场景。</description><pubDate>Sat, 03 Oct 2026 19:38:28 +0800</pubDate><category>基础与架构</category></item><item><title>详解 AI Agent 评测</title><link>https://blog.wanglvyuan.com/library/evals/</link><guid isPermaLink="true">https://blog.wanglvyuan.com/library/evals/</guid><description>Anthropic Agent 评测指南的中文译文，解释任务、试验、执行轨迹与评分器，比较代码、模型和人工评分，讨论非确定性、能力评测与回归评测，并给出从真实失败案例建立和长期维护评测体系的方法。</description><pubDate>Sat, 03 Oct 2026 19:38:28 +0800</pubDate><category>评测与改进</category></item><item><title>面向 AI Agent 的高效上下文工程</title><link>https://blog.wanglvyuan.com/library/context/</link><guid isPermaLink="true">https://blog.wanglvyuan.com/library/context/</guid><description>Anthropic 上下文工程文章的中文译文，解释上下文与提示词工程的关系，讨论系统提示、工具、检索和自主搜索的设计，以及通过压缩、结构化笔记和子 Agent 支持长程任务的具体方法与取舍。</description><pubDate>Sat, 03 Oct 2026 19:38:28 +0800</pubDate><category>上下文工程</category></item><item><title>传统软件由代码描述，而 Agent 由执行轨迹描述</title><link>https://blog.wanglvyuan.com/library/traces-documentation/</link><guid isPermaLink="true">https://blog.wanglvyuan.com/library/traces-documentation/</guid><description>Harrison Chase 文章的中文译文，讨论 Agent 的决策为何难以仅靠代码理解，以及执行轨迹如何记录实际行为。内容涵盖调试、评测、性能优化、质量监控、产品分析和跨团队协作所需关键信息的变化。</description><pubDate>Sat, 03 Oct 2026 19:38:28 +0800</pubDate><category>评测与改进</category></item><item><title>扩展 Managed Agents：将大脑与双手解耦</title><link>https://blog.wanglvyuan.com/library/managed/</link><guid isPermaLink="true">https://blog.wanglvyuan.com/library/managed/</guid><description>Anthropic Managed Agents 架构文章的中文译文，介绍如何将模型推理、工具执行环境与持久会话解耦，使运行框架能够独立演进，区分持久会话与模型当前上下文，并讨论状态恢复、扩展执行和稳定接口的设计。</description><pubDate>Sat, 03 Oct 2026 19:38:28 +0800</pubDate><category>基础与架构</category></item><item><title>基于执行轨迹数据才能建立 Agent 迭代机制</title><link>https://blog.wanglvyuan.com/library/improvement/</link><guid isPermaLink="true">https://blog.wanglvyuan.com/library/improvement/</guid><description>LangChain 指南的中文译文，说明如何收集执行轨迹，通过在线评分与人工反馈发现失败，再转化为离线评测和经过验证的改动。结合 LangSmith 介绍持续迭代流程，并讨论自动化与人工判断的边界。</description><pubDate>Sat, 03 Oct 2026 19:38:28 +0800</pubDate><category>评测与改进</category></item><item><title>借助 Agent 为 Agent 编写高效工具</title><link>https://blog.wanglvyuan.com/library/tools/</link><guid isPermaLink="true">https://blog.wanglvyuan.com/library/tools/</guid><description>Anthropic 工具工程文章的中文译文，介绍如何与 Agent 协作构建、评测和优化工具，讨论工具选择、命名空间、响应信息与 token 成本，并说明如何通过清晰描述和留出测试验证改进，避免仅凭调用成功判断质量。</description><pubDate>Sat, 03 Oct 2026 19:38:28 +0800</pubDate><category>工具与技能</category></item><item><title>我们如何构建多 Agent 研究系统</title><link>https://blog.wanglvyuan.com/library/multi-agent/</link><guid isPermaLink="true">https://blog.wanglvyuan.com/library/multi-agent/</guid><description>Anthropic Research 系统的工程经验中文译文，介绍主 Agent 与并行子 Agent 的分工、搜索和信息汇总，讨论任务委派、提示词、工具设计、评测与生产可靠性，并说明任务并行程度和额外成本如何影响架构选择。</description><pubDate>Sat, 03 Oct 2026 19:38:28 +0800</pubDate><category>基础与架构</category></item><item><title>Claude Code 最佳实践</title><link>https://blog.wanglvyuan.com/library/best-practices/</link><guid isPermaLink="true">https://blog.wanglvyuan.com/library/best-practices/</guid><description>Claude Code 官方实践文档的中文译文，围绕验证结果、探索规划和上下文管理，说明如何配置项目指令、权限及外部工具，借助技能、子 Agent 与自动化机制组织编程任务，并识别会话中常见的失败模式。</description><pubDate>Sat, 03 Oct 2026 19:38:28 +0800</pubDate><category>基础与架构</category></item><item><title>think 工具：让 Claude 在复杂工具使用场景中停下来思考</title><link>https://blog.wanglvyuan.com/library/think/</link><guid isPermaLink="true">https://blog.wanglvyuan.com/library/think/</guid><description>Anthropic think 工具研究文章的中文译文，围绕 Claude 3.7 Sonnet 的实验，说明如何在工具调用过程中加入专门的思考步骤，比较其与扩展思考的区别，并讨论基准结果、提示示例、业务规则遵循和适用边界。</description><pubDate>Sat, 03 Oct 2026 19:38:28 +0800</pubDate><category>工具与技能</category></item><item><title>为长时间运行的 Agent 构建高效运行框架</title><link>https://blog.wanglvyuan.com/library/long-running/</link><guid isPermaLink="true">https://blog.wanglvyuan.com/library/long-running/</guid><description>Anthropic 长程 Agent 工程文章的中文译文，针对跨上下文窗口丢失进展的问题，介绍初始化与编程 Agent 的分工，以及功能清单、进度记录、Git 提交和端到端测试如何帮助任务持续推进、避免过早宣布完成。</description><pubDate>Sat, 03 Oct 2026 19:38:28 +0800</pubDate><category>基础与架构</category></item><item><title>通过 MCP 执行代码：构建更高效的 Agent</title><link>https://blog.wanglvyuan.com/library/mcp-code/</link><guid isPermaLink="true">https://blog.wanglvyuan.com/library/mcp-code/</guid><description>Anthropic MCP 工程文章的中文译文，解释预加载工具定义与中间结果带来的上下文开销，介绍通过代码按需发现、组合和调用工具的方法，并讨论结果过滤、状态持久化、隐私及安全执行环境的成本。</description><pubDate>Sat, 03 Oct 2026 19:38:28 +0800</pubDate><category>工具与技能</category></item><item><title>借助 Agent Skills，让 Agent 胜任现实世界的任务</title><link>https://blog.wanglvyuan.com/library/skills/</link><guid isPermaLink="true">https://blog.wanglvyuan.com/library/skills/</guid><description>Anthropic Agent Skills 设计文章的中文译文，介绍如何用包含指令、脚本和资源的文件夹提供领域知识，解释如何按需发现、渐进式加载与代码执行，并讨论技能的开发、评测、安全审查及跨平台复用。</description><pubDate>Sat, 03 Oct 2026 19:38:28 +0800</pubDate><category>工具与技能</category></item><item><title>超越权限提示：让 Claude Code 更安全、更自主</title><link>https://blog.wanglvyuan.com/library/sandboxing/</link><guid isPermaLink="true">https://blog.wanglvyuan.com/library/sandboxing/</guid><description>Anthropic 沙箱工程文章的中文译文，介绍 Claude Code 如何结合文件系统与网络隔离减少权限提示和审批疲劳，并说明沙箱化 bash、云端执行和 Git 凭据代理的设计，以及安全边界与自主执行之间的关系。</description><pubDate>Sat, 03 Oct 2026 19:38:28 +0800</pubDate><category>安全与可靠性</category></item><item><title>我们如何构建 Claude Code 自动模式：更安全地跳过权限确认</title><link>https://blog.wanglvyuan.com/library/auto-mode/</link><guid isPermaLink="true">https://blog.wanglvyuan.com/library/auto-mode/</guid><description>Anthropic Claude Code 自动模式文章的中文译文，介绍权限分类器如何判断操作是否获得用户授权，讨论两阶段分类、提示注入防御、子 Agent 交接与拒绝后的恢复，并披露误报、漏报及需要人工判断的剩余风险。</description><pubDate>Sat, 03 Oct 2026 19:38:28 +0800</pubDate><category>安全与可靠性</category></item><item><title>面向长时间应用开发的运行框架设计</title><link>https://blog.wanglvyuan.com/library/long-apps/</link><guid isPermaLink="true">https://blog.wanglvyuan.com/library/long-apps/</guid><description>Anthropic 长时间应用开发实验的中文译文，介绍如何把前端审美转化为评分标准，并用规划器、生成器和评估器组织全栈开发。文章比较运行框架迭代的结果，讨论模型更新后应保留或移除哪些机制。</description><pubDate>Sat, 03 Oct 2026 19:38:28 +0800</pubDate><category>基础与架构</category></item><item><title>我们如何在各产品中限制 Claude 的影响范围</title><link>https://blog.wanglvyuan.com/library/containment/</link><guid isPermaLink="true">https://blog.wanglvyuan.com/library/containment/</guid><description>Anthropic Agent 隔离实践的中文译文，对比不同产品使用的临时容器、人工参与沙箱和本地虚拟机，结合权限信任、提示注入与获准域名外传数据等实际问题，讨论如何为故障影响范围设定上限。</description><pubDate>Sat, 03 Oct 2026 19:38:28 +0800</pubDate><category>安全与可靠性</category></item><item><title>用并行运行的 Claude 团队构建 C 编译器</title><link>https://blog.wanglvyuan.com/library/c-compiler/</link><guid isPermaLink="true">https://blog.wanglvyuan.com/library/c-compiler/</guid><description>Nicholas Carlini 编译器实验的中文译文，记录 16 个 Claude Agent 并行用 Rust 构建 C 编译器的过程，讨论长期运行、测试质量、共享代码库协作与角色分工，并说明测试通过仍不等于软件可靠，并分析自主开发的真实风险。</description><pubDate>Sat, 03 Oct 2026 19:38:28 +0800</pubDate><category>评测与改进</category></item><item><title>Claude 开发者平台的高级工具使用能力</title><link>https://blog.wanglvyuan.com/library/advanced-tools/</link><guid isPermaLink="true">https://blog.wanglvyuan.com/library/advanced-tools/</guid><description>Anthropic 高级工具使用文章的中文译文，介绍工具搜索、程序化工具调用和工具使用示例三项能力，解释如何减少预加载与中间结果开销、明确控制流和参数用法，并讨论组合使用时的适用场景。</description><pubDate>Sat, 03 Oct 2026 19:38:28 +0800</pubDate><category>工具与技能</category></item><item><title>剖析 Agent 运行框架</title><link>https://blog.wanglvyuan.com/library/harness-anatomy/</link><guid isPermaLink="true">https://blog.wanglvyuan.com/library/harness-anatomy/</guid><description>Vivek Trivedy 文章的中文译文，从模型能力与期望行为的差距出发，梳理 Agent 运行框架的组成，涵盖工具执行、文件系统、沙箱、记忆、上下文管理与验证循环，并讨论各类组件与模型训练的关系。</description><pubDate>Sat, 03 Oct 2026 19:38:28 +0800</pubDate><category>基础与架构</category></item><item><title>Agent 工程：一门新学科</title><link>https://blog.wanglvyuan.com/library/agent-engineering/</link><guid isPermaLink="true">https://blog.wanglvyuan.com/library/agent-engineering/</guid><description>LangChain 团队文章的中文译文，将 Agent 工程描述为构建、测试、发布、观察与改进的持续循环，讨论非确定性系统的生产可靠性，以及评测、可观测性和真实使用反馈在日常团队实践中的作用。</description><pubDate>Sat, 03 Oct 2026 19:38:28 +0800</pubDate><category>基础与架构</category></item><item><title>如何通过中间件定制 Agent 运行框架</title><link>https://blog.wanglvyuan.com/library/middleware/</link><guid isPermaLink="true">https://blog.wanglvyuan.com/library/middleware/</guid><description>LangChain 中间件指南的中文译文，解释如何在 Agent 循环中接入自定义逻辑，修改提示词、工具和模型行为，并以 Deep Agents 为例讨论上下文管理、策略执行及可复用运行框架组件的设计与职责分离。</description><pubDate>Sat, 03 Oct 2026 19:38:28 +0800</pubDate><category>基础与架构</category></item><item><title>如何在运行框架中构建模型路由器</title><link>https://blog.wanglvyuan.com/library/model-router/</link><guid isPermaLink="true">https://blog.wanglvyuan.com/library/model-router/</guid><description>LangChain 模型路由实验的中文译文，介绍如何结合任务特征、模型档位和结果反馈，在运行框架中选择模型。结合 Open SWE 案例讨论成本、质量、缓存与路由粒度，并交代实验结果的统计与适用局限。</description><pubDate>Sat, 03 Oct 2026 19:38:28 +0800</pubDate><category>基础与架构</category></item><item><title>借助 Jev 构建运行框架</title><link>https://blog.wanglvyuan.com/library/jev/</link><guid isPermaLink="true">https://blog.wanglvyuan.com/library/jev/</guid><description>LangChain 介绍 Jev 的中文译文，解释这种分类模型如何输出决策及概率，并用于模型路由和工具权限检查。文章展示中间件集成方式，区分分类决策与文本生成，并转述供应商报告的速度和成本表现。</description><pubDate>Sat, 03 Oct 2026 19:38:28 +0800</pubDate><category>基础与架构</category></item><item><title>编程 Agent 是黑箱：如何看清它们的内部运行过程</title><link>https://blog.wanglvyuan.com/library/coding-blackbox/</link><guid isPermaLink="true">https://blog.wanglvyuan.com/library/coding-blackbox/</guid><description>LangChain 编程 Agent 可观测性文章的中文译文，从旧辅助函数导致分页错误的案例出发，说明如何通过执行轨迹检查上下文、工具调用和子 Agent 交接，并用统一记录比较不同编程工具的行为。</description><pubDate>Sat, 03 Oct 2026 19:38:28 +0800</pubDate><category>评测与改进</category></item><item><title>构建 Agent 的实用指南</title><link>https://blog.wanglvyuan.com/library/practical-guide/</link><guid isPermaLink="true">https://blog.wanglvyuan.com/library/practical-guide/</guid><description>OpenAI Agent 实用指南的中文译文，面向首次构建系统的产品与工程团队，介绍场景选择、模型、工具与指令设计，比较单体与多智能体协作的编排模式，并讨论安全护栏、人工介入和通过真实用户逐步验证部署。</description><pubDate>Sat, 03 Oct 2026 19:38:28 +0800</pubDate><category>基础与架构</category></item><item><title>构建 Agent 的新工具</title><link>https://blog.wanglvyuan.com/library/agent-tools/</link><guid isPermaLink="true">https://blog.wanglvyuan.com/library/agent-tools/</guid><description>OpenAI 开发工具发布文章的中文译文，介绍 Responses API 及网页搜索、文件检索和计算机操作工具，说明如何组织多步骤任务、交接控制权与追踪执行过程，并保留原发布时的接口迁移说明和安全限制。</description><pubDate>Sat, 03 Oct 2026 19:38:28 +0800</pubDate><category>工具与技能</category></item><item><title>面向组织、合作伙伴与生态系统的 Skills</title><link>https://blog.wanglvyuan.com/library/skills-ecosystem/</link><guid isPermaLink="true">https://blog.wanglvyuan.com/library/skills-ecosystem/</guid><description>Anthropic Skills 生态发布文章的中文译文，介绍组织管理员如何集中配置工作流、用户如何创建和编辑技能，以及合作伙伴目录与开放标准的作用，说明可复用知识与 MCP 工具之间的互补关系。</description><pubDate>Sat, 03 Oct 2026 19:38:28 +0800</pubDate><category>工具与技能</category></item><item><title>使用 Mods 定制 Claude Code</title><link>https://blog.wanglvyuan.com/library/code-mods/</link><guid isPermaLink="true">https://blog.wanglvyuan.com/library/code-mods/</guid><description>Anthropic Mods 介绍文章的中文译文，说明如何用 TypeScript 改写 Claude Code 事件、扩展界面或替换内置功能，介绍插件分发、团队管理、加载顺序与默认安全设置，并明确这些扩展不在沙箱中运行，强调只应安装可信来源的扩展。</description><pubDate>Sat, 03 Oct 2026 19:38:28 +0800</pubDate><category>工具与技能</category></item><item><title>上下文工程的兴起</title><link>https://blog.wanglvyuan.com/library/context-rise/</link><guid isPermaLink="true">https://blog.wanglvyuan.com/library/context-rise/</guid><description>Harrison Chase 上下文工程文章的中文译文，讨论为何 Agent 的成败不仅取决于提示词措辞，还取决于运行时提供的信息，梳理指令、历史、记忆、工具和检索内容，并解释上下文组织与产品可靠性的关系。</description><pubDate>Sat, 03 Oct 2026 19:38:28 +0800</pubDate><category>上下文工程</category></item><item><title>面向 Agent 的上下文工程</title><link>https://blog.wanglvyuan.com/library/context-agents/</link><guid isPermaLink="true">https://blog.wanglvyuan.com/library/context-agents/</guid><description>LangChain 上下文工程综述的中文译文，将常见方法归纳为写入、选择、压缩和隔离，结合记忆、检索、工具输出与多 Agent 案例，讨论如何管理有限的上下文资源，并介绍相关框架与可观测性支持。</description><pubDate>Sat, 03 Oct 2026 19:38:28 +0800</pubDate><category>上下文工程</category></item><item><title>上下文检索</title><link>https://blog.wanglvyuan.com/library/context-retrieval/</link><guid isPermaLink="true">https://blog.wanglvyuan.com/library/context-retrieval/</guid><description>Anthropic 上下文检索文章的中文译文，介绍为文档片段补充上下文后构建嵌入和 BM25 索引的方法，比较重排序及不同检索配置的实验结果，并讨论提示词缓存、实现成本和评测中的取舍。</description><pubDate>Sat, 03 Oct 2026 19:38:28 +0800</pubDate><category>上下文工程</category></item><item><title>上下文工程实践：构建 Manus 的经验教训</title><link>https://blog.wanglvyuan.com/library/manus-context/</link><guid isPermaLink="true">https://blog.wanglvyuan.com/library/manus-context/</guid><description>Manus 上下文工程经验的中文译文，介绍 KV 缓存、工具屏蔽、文件系统、任务复述、错误记录和示例多样性等设计选择，解释这些机制如何影响长程 Agent 的成本、注意力与行为，并保留实践中的限制。</description><pubDate>Sat, 03 Oct 2026 19:38:28 +0800</pubDate><category>上下文工程</category></item><item><title>动态发现上下文</title><link>https://blog.wanglvyuan.com/library/cursor-context/</link><guid isPermaLink="true">https://blog.wanglvyuan.com/library/cursor-context/</guid><description>Cursor 动态上下文发现文章的中文译文，介绍将长工具结果、聊天历史和终端会话转为可按需读取的文件，以及渐进式加载 Skills 与 MCP 工具的方法，讨论如何减少上下文占用并保留可追溯信息。</description><pubDate>Sat, 03 Oct 2026 19:38:28 +0800</pubDate><category>上下文工程</category></item><item><title>大语言模型驱动的自主 Agent</title><link>https://blog.wanglvyuan.com/library/autonomous-agents/</link><guid isPermaLink="true">https://blog.wanglvyuan.com/library/autonomous-agents/</guid><description>Lilian Weng 自主 Agent 综述的中文译文，围绕规划、记忆和工具使用介绍任务分解、自我反思与检索方法，并结合科学发现、生成式模拟和早期项目案例，讨论长期规划、上下文及可靠性的挑战。</description><pubDate>Sat, 03 Oct 2026 19:38:28 +0800</pubDate><category>基础与架构</category></item><item><title>能力问题：编程 Agent 的运行框架工程</title><link>https://blog.wanglvyuan.com/library/skill-issue/</link><guid isPermaLink="true">https://blog.wanglvyuan.com/library/skill-issue/</guid><description>HumanLayer 编程 Agent 工程文章的中文译文，讨论如何用指令文件、外部工具、技能、子 Agent 和钩子定制运行框架，结合具体失败经验解释渐进式披露、上下文隔离、成本控制与持续验证反馈的作用。</description><pubDate>Sat, 03 Oct 2026 19:38:28 +0800</pubDate><category>基础与架构</category></item><item><title>AutoHarness：通过自动合成代码运行框架改进 LLM Agent（论文摘要）</title><link>https://blog.wanglvyuan.com/library/autoharness/</link><guid isPermaLink="true">https://blog.wanglvyuan.com/library/autoharness/</guid><description>AutoHarness 论文摘要的中文译文与文献信息，介绍根据游戏环境反馈自动合成代码运行框架、约束非法行动的研究，并概述其在 TextArena 上的实验结论。本页不含论文全文，保留原文与全文入口。</description><pubDate>Sat, 03 Oct 2026 19:38:28 +0800</pubDate><category>评测与改进</category></item><item><title>持续改进我们的 Agent 运行框架</title><link>https://blog.wanglvyuan.com/library/cursor-harness/</link><guid isPermaLink="true">https://blog.wanglvyuan.com/library/cursor-harness/</guid><description>Cursor 运行框架迭代经验的中文译文，介绍如何结合离线评测、线上使用信号和回归跟踪验证改动，讨论上下文管理、针对不同模型的定制，以及对话中途切换模型带来的工程问题。</description><pubDate>Sat, 03 Oct 2026 19:38:28 +0800</pubDate><category>基础与架构</category></item><item><title>面向自我改进的运行框架工程</title><link>https://blog.wanglvyuan.com/library/self-improvement/</link><guid isPermaLink="true">https://blog.wanglvyuan.com/library/self-improvement/</guid><description>Lilian Weng 自我改进研究综述的中文译文，梳理上下文与工作流优化、运行框架演化、进化式搜索及模型权重联合更新，比较相关研究，并讨论评测、记忆、奖励投机、权限边界和人类监督的挑战。</description><pubDate>Sat, 03 Oct 2026 19:38:28 +0800</pubDate><category>评测与改进</category></item><item><title>Claude 3.5 Sonnet：刷新 SWE-bench Verified 成绩</title><link>https://blog.wanglvyuan.com/library/swe-bench-sonnet/</link><guid isPermaLink="true">https://blog.wanglvyuan.com/library/swe-bench-sonnet/</guid><description>Anthropic 介绍 Claude 3.5 Sonnet 在 SWE-bench Verified 上取得 49% 成绩的 Agent 设计：精简运行框架、Bash 与文件编辑工具、问题复现流程，以及隐藏测试和评测环境带来的挑战。本文为中文译文，保留原始代码、表格和英文对照。</description><pubDate>Sat, 03 Oct 2026 15:06:09 +0000</pubDate><category>评测与改进</category></item></channel></rss>