递归自我改进(recursive self-improvement,RSI)的概念可以追溯到 I. J. Good(1965)。他将“超智能机器”定义为一种能够在所有智力活动中超越人类,并能设计更好的机器来改进自身的系统。Yudkowsky(2008)用“递归自我改进”来描述一种特定的反馈循环:AI 利用当前的智能,改进产生其智能的认知机制。
The concept of recursive self-improvement (RSI) dates back to I. J. Good (1965), where he defined an “ultraintelligent machine” as a system that can surpass humans in all intellectual activities and design better machines to improve itself. Yudkowsky (2008) used the phrase “recursive self-improvement” for a specific feedback loop: an AI uses its current intelligence to improve the cognitive machinery that produces its intelligence.
在现代 AI 中,这种反馈循环可能指模型直接改写自身权重;更广义地说,也可以指模型改进训练流程和部署系统,进而产生一个更好的后继模型,使其在具有经济价值的任务上表现更出色。前沿实验室的实践已经表明,AI 研发速度因此显著加快(Anthropic;OpenAI)。
This feedback loop in modern AI may indicate the model rewriting its own weights directly, or more broadly the model improves the training pipeline and the deployment system, which in turn enables a better successor model with improved performance across economically valuable tasks. The speed of research development in AI has been shown to drastically accelerated in frontier labs (Anthropic; OpenAI).
我特意提到“部署系统”,是因为原始模型与现实环境之间的这一层,看起来与模型本身的智能(即预训练刚结束时的评测表现)同样重要。运行框架是 AI 部署的重要组成部分,Claude Code 和 Codex 等成功的编程 Agent 产品就说明了这一点。运行框架(harness)是围绕基础模型构建的系统,负责执行编排,并决定模型如何思考和规划、调用工具和采取行动、感知和管理上下文、存储产物,以及评估结果。
I explicitly mention “deployment system” because the layer between the raw model and the real-world context seems to be as important as the model’s raw intelligence (i.e. the evals right after pretraining). Harnesses are important components of AI deployment, as shown by successful coding agent products such as Claude Code and Codex. A harness is the system surrounding a base model that orchestrates execution and decides how the model thinks and plans, calls tools and acts, perceives and manages context, stores artifacts, and evaluates results.
本文将聚焦运行框架工程相关研究,以及它如何促进 RSI。近来许多关于自动化研究、自我改进 Agent 和进化式程序搜索的工作,都可以围绕这个问题来梳理。模型自我博弈、合成数据、测试时训练,以及更广泛的持续学习研究,也符合 RSI 的愿景(例如 Yuan 等,2024,Chen 等,2024,Zhao 等,2025,Choi 等,2026),但它们不是本文的重点。
This one post will focus on research around harness engineering and how it contributes to RSI. Much recent work on auto-research, self-improving agents, and evolutionary program search can be organized around this question. Other work on model self-play, synthetic data, test-time training and a broader theme of continual learning also matches the RSI vision (e.g. Yuan et al. 2024, Chen et al. 2024), Zhao et al. 2025, Choi et al. 2026)) but they will not be the focus of this post.
运行框架的设计模式
Harness Design Patterns
与早期 Agent 框架中的“Agent = LLM + 记忆 + 工具 + 规划 + 行动”相比,运行框架工程还包括工作流设计(如循环工程)、评测、权限控制和持久状态管理。它不再只是提示词模板,而更接近运行时和软件系统设计:模型如何观察、行动、记忆、自检和改进。
Compared with early agent frameworks, “agent = LLM + memory + tools + planning + action”, harnesses engineering additionally include workflow design (e.g. loop engineering), evaluation, permission controls, and persistent state management. It is no longer only prompt templates, but closer to runtime and software system design: how the model observes, acts, memorizes, checks itself, and improves.
为了实现泛化,设计应刻意保持简单、通用,并且很可能需要借鉴现有软件工程实践,以利用模型的预训练知识。操作系统与运行框架之间也存在很强的类比关系。与操作系统类似,运行框架应封装复杂逻辑,同时保持接口简单。与此同时,配置、工具接口和其他协议可能逐渐形成行业标准。
The design should be deliberately simple and generic to enable generalization, likely with reference to existing software engineering practices to benefit from prertaining knowlege. There is also a strong analogy between operating systems and harnesses. Similar to an OS, a harness should encapsulate complicated logic while keeping the interface simple. Meanwhile, configs, tool interfaces and other protocols may gradually become standardized across the industry.
模式 1:工作流自动化
Pattern 1: Workflow Automation
定义一种让模型能够操作、测试和迭代的工作流,是实现自动化的关键设计。Karpathy 的 autoresearch 仓库(https://github.com/karpathy/autoresearch)清晰地展示了如何构建这种工作流。常见的工作流遵循一个以目标为导向的循环:规划、执行、观察/测试、改进,然后再次执行,直到达成目标。过程中可能主动向用户提问,以澄清任务要求或执行偏好。
Defining a workflow in which the model can operate, test, and iterate is a key design for automation. Karpathy’s autoresearch repo (https://github.com/karpathy/autoresearch) is a clean example of how such a workflow can be constructed. A common workflow follows a goal-oriented loop of plan, execute, observe/test, improve, and execute again until the goal is achieved. The process may trigger proactive requests to users for clarity in task specification or execution preference.
工作流图也强调,模型会分析自身的执行轨迹和失败案例,再通过“Agent 运行时”不断推进和迭代,而不是依赖静态提示词模板。
The workflow graph also emphasizes the model analyzing its own trajectories and failure cases and then iterating on its progress through an “agent runtime” rather than a static prompt template.
模式 2:将文件系统用作持久记忆
Pattern 2: File System as Persistent Memory
长时程 Agent 系统中反复出现的一种模式,是以简单的方式控制丰富的状态和产物。运行框架不应把整个工作流和全部日志都放进上下文,而应将持久状态保存在文件中。在长时程的 Agent 执行过程中,实验日志、代码差异、论文摘要、错误轨迹和过往执行轨迹等产物,往往会增长到远远超过模型训练所适配的上下文窗口长度。
A recurring pattern in long-horizon agent systems is simple control over rich states and artifacts. A harness should not carry the entire workflow and all logs in context; instead, it should keep durable state in files. In long-horizon agentic rollout, artifacts such as experiment logs, code diffs, paper summaries, error traces, and past rollout trajectories often grow much longer than the context window that the model has trained for.
学会读取、写入和编辑文件系统(通常通过 bash 命令)是 LLM 的基础技能。因此,以文件这种简单形式管理持久记忆,自然能够受益于模型核心能力的提升。
Learning how to read, write, and edit the file system (commonly via bash commands) is a foundation skill for LLMs, and thus managing persistent memory in the simple form of files naturally benefits from improvements in core model capability.
模式 3:子 Agent 与后台任务
Pattern 3: Sub-agent and Backend Jobs
运行框架可以启动多个子 Agent 并行执行,并监控后台任务。当主 Agent 需要搜索多个假设、并发运行实验,或者在不污染主上下文的情况下委派独立子任务时,这一点很有用。此时,父 Agent 需要一个小型进程管理器:启动任务、检查日志、取消失败的运行,并将结果合并回主 Agent 会话。
A harness can spawn multiple subagents to execute in parallel and monitor backend jobs. This is useful when the main agent needs to search multiple hypotheses, run experiments concurrently, or delegate isolated subtasks without polluting the main context. The parent agent then needs a small process manager: launch jobs, inspect logs, cancel failed runs, and merge results back into the main agent thread.
关键设计选择是让并行执行变得显式且可检查。如果子 Agent 的输出只存在于临时聊天上下文中,它们很快就会过时并难以发现。如果将它们保存为文件、日志和状态记录,模型就能在中断后恢复,并对自身执行历史进行推理。
The key design choice is to make parallelism explicit and inspectable. If subagent outputs only live in a transient chat context, they quickly become obselete and hidden. If they are stored as files, logs, and status records, the model can recover after interruptions and reason over its own execution history.
案例研究:编程 Agent 的运行框架
Case study: Coding Agent Harness
Claude Code、Codex、OpenCode 和 Cursor 风格的 Agent 等主流编程 Agent,其核心接口已趋于稳定。它们通常采用如下循环:
The core interface of mainstream coding agents has become stabilized across Claude Code, Codex, OpenCode, and Cursor-style agents. They commonly use a loop like:
借助一组工具,编程 Agent 能在指定仓库中进行开发和问题调试,类似于人类开发者使用 IDE。
With access to a set of tools, the coding agent is able to develop and debug issues in a given repository, similar to how human developers are equipped with IDEs.
(这不是完整清单,仅用于展示。有兴趣可以阅读此处。)
(Not a comprenhensive list; shown for demonstration. Read this if interested.)
运行框架层与核心智能,孰轻孰重?
Harness Layer vs Core Intelligence?
我们很难预测 RSI 的未来会在多大程度上依赖运行框架工程,但 RSI 的近期路径不太可能从模型直接改写自身权重开始。我对近期可行路径的预测是:
It is hard to forecast how much the future of RSI will rely on harness engineering, but the near-term path of RSI is unlikely to start as a model directly rewriting its weights. My prediction of a practical near-term path is:
- 运行框架工程将向元方法论发展(即改进产生更好答案的机制,而不仅仅是改进答案本身)。运行框架系统自身将成为优化目标,启发式规则更少,通用机制更多。
- 反过来,成熟的运行框架将支持用于模型自我改进循环的自动化研究;更聪明的模型则能避免运行框架过度工程化,使系统保持可持续性。
- Harness engineering will evolve in the direction of meta-methodology (i.e. improving the machinery for getting better answers, not just improving the answer itself). The harness system itself becomes an optimization target, with fewer heuristic rules and more general mechanisms.
- In turn, mature harnesses enable auto-research for model self-improvement loop and smarter models prevents harnesses from overengineering and keep the system sustainable.
最终,许多运行框架改进可能会被内化为模型的核心行为,但与外部上下文和工具交互的接口应当仍会存在。在提示词工程中,我们已经见过这种模式的较温和版本:随着指令微调和模型推理能力的进步,手工提示词技巧的重要性下降了,但明确目标、约束、上下文和评测的需求并未消失。
Eventually it is possible that many harness improvements will be internalized into core model behavior, but the interface with external context and tools should remain. We have seen a softer version of this pattern with prompt engineering: manual prompt tricks became less central as instruction tuning and model reasoning improved, but the need to specify goals, constraints, context, and evaluation did not disappear.
运行框架优化
Harness Optimization
运行框架系统中优化对象的演进,大致是:指令提示词 → 结构化上下文 → 工作流 → 运行框架代码 → 优化器代码。随着模型变得更智能、更强大,我们开始面向更复杂的目标,采用更通用的方法。
The progression in the object being optimized in the harness system is roughly: instruction prompts → structured context → workflow → harness code → optimizer code. As the model becomes more intelligent and powerful, we move toward more complex targets and generic methods.
上下文工程
Context Engineering
随着 Agent 任务的执行跨度显著增加,简单地将所有工具响应和模型生成内容追加进上下文,可能很快变得难以控制。上下文管理这一层负责为 LLM 构建更有结构、更简洁的上下文,并管理持久状态。长上下文研究无疑会继续进步,但目前,长上下文智能与上下文工程有时是相互交织的。
Simply appending all the tool responses and model generations into the context can quickly grow out of control as the agentic job horizon increases significantly. Context management is a layer to construct a more structed and concise context for LLM and manage persistant states. There is no doubt that long-context research will keep on making progress but at the moment long-context intelligence and context engineering sometime intertwines.
Agentic Context Engineering(ACE;Zhang 等,2025)将上下文视为不断演进的操作手册,而不是越来越长的提示词。它通过三个组件维护一份由要点条目构成的上下文操作手册,每个条目都有标识符和说明。
Agentic Context Engineering (ACE; Zhang et al. 2025) treats context as an evolving playbook rather than an increasingly lengthening prompt. It has three components to maintain one context playbook of bullet points, each with an identifier and a description.
- 生成器(Generator):参考要点条目,生成任务执行轨迹。
- 反思器(Reflector):从成功和失败的执行轨迹中提炼见解。
- 整理器(Curator):通过逐项增量条目更新结构化上下文。
- Generator: produces task trajectories, with reference to bullet points.
- Reflector: distills insights from successful and failed trajectories.
- Curator: updates the structured context with incremental, itemized entries.
为了防止迭代改写过程中的上下文坍缩和简短偏好,ACE 的一项关键设计是:整理器不会重写整块提示词,而是输出一组结构化、逐项列出的要点,形式为(标识符,说明),再通过确定性逻辑将这些要点合并进结构化的上下文记录。上下文条目会被定期改进和去重。
To prevent context collapse and brevity bias during iterative rewrites, one key design choice in ACE is that the curator does not rewrite a full prompt blob. It instead outputs a collection of structured, itemized bullets in the form of (identifier, description), and these bullets are merged into a structured context logbook with deterministic logic. The context items are refined and deduplicated periodically.
ACE 从执行过程中学习见解,使我们向自主管理的记忆迈进,但更新规则和整体工作流仍是人工设计的。为了走向更能自我改进的循环,Meta Context Engineering(MCE;Ye 等,2026)将机制(如何管理上下文)与产物内容(上下文中有什么)分开:在元优化层进行技能演化,在基础层进行上下文优化。
The fact that ACE learns insights from rollouts helps us move toward self-managed memory, but the update rules and the overall workflow are still handcrafted. To move toward a more self-improving loop, Meta Context Engineering (MCE; Ye et al. 2026) separates the mechanism (how to manage context) from the artifact content (what is in context), running skill evolution at the meta-optimization level and context optimization at the base level.
MCE 技能 定义一个上下文函数 ,将输入 映射为上下文 ,其中:
An MCE skill defines a context function and maps an input to context , where:
- 是静态组件(提示词、知识库、代码库)。
- 是动态算子(搜索、选择、过滤、格式化)。
- are static components (prompts, knowledge bases, code libraries).
- are dynamic operators (search, selection, filtering, formatting).
这种双层优化是在给定技能 的情况下,根据训练数据找到最佳上下文 ;外层循环则寻找能在验证集上取得最佳表现的技能:
The bi-level optimization is to find the best context given skill on the training data, while the outer loop finds the optimal skill that provides the best performance on the validation set:
技能数据库跟踪以往技能、上下文函数和评测指标的历史记录 。给定任务 ,元层 Agent 对先前技能执行 Agent 式交叉操作,创建新技能:。
The skill database tracks the history of previous skills, context functions and eval metrics . A meta-level agent performs agentic crossover over prior skills to create a new skill given a task : .
接着,基础层的上下文工程 Agent 执行技能 ,在当前技能的指导下,从执行反馈 中学习上下文函数:。
Then a base-level context engineer executes the skill and learns the context function from rollout feedback , guided by the current skill: .
MCE 不像 ACE 那样强制采用某种关于上下文结构的启发式规则。它使用自由形式的技能保存任务最重要的知识,并将技能和受技能约束的上下文一起迭代演化。在实现上,上下文函数 被实例化为专用目录中的一组文件,既包括静态组件(skill.md),也包括动态组件(上下文和数据执行轨迹)。元层优化与基础层优化都在具有标准工具集的 Agent 编程环境中执行,工具集为:
MCE does not enforce a heuristic rule for how to structure context as ACE does. It uses free-form skills to store the most important knowledge for a task, and evolves the skill and the skill-conditioned context iteratively together. Implementation-wise, a context function is instantiated as a collection of files in a dedicated directory, including both static (skill.md) and dynamic (context and data rollouts) components. Both meta-level and base-level optimization are executed in agentic coding envs with a standard tool set,
Meta-Harness(Lee 等,2026)又深入了一层:优化对象是用于决定并优化“应当存储、检索哪些信息,以及向模型展示哪些信息”的代码。其名称中的“Meta-”表示,它是一个用于优化运行框架的运行框架。
Meta-Harness (Lee et al. 2026) moves another level deeper: the optimized object is the code that determines and optimizes what information should be stored, retrieved, and presented to the model. “Meta-” in its name means it is a harness for optimizing harnesses.
负责提出新运行框架的提议者,本身就是一个编程 Agent;最终输出是位于帕累托前沿上的一组运行框架候选方案。
The proposer for creating a new harness is itself a coding agent and the final output is a collection of harness candidates on the Pareto frontier.
- 整个执行历史都能通过文件系统访问,因此编程 Agent 会使用
grep或cat等命令读取它,而不是把所有内容塞进单个提示词上下文。 - 提出的运行框架是文件系统中的一个字典,包含自身源代码、分数、执行轨迹和状态更新。
- 元运行框架循环迭代创建新的运行框架,只保留合格的候选方案。
- The entire execution history is accessible via a file system, and thus the coding agent uses commands like
greporcatto read through it instead of shoveling everything into a single prompt context. - The proposed harness is a dictionary in the file system containing its own source code, scores, rollout trajectories, and state updates.
- The mete-harness loop iteratively creates new harnesses, and only qualified ones are kept.
不过,重要启示已经很清楚:一旦运行框架设计成为可执行的搜索空间,强大的编程 Agent 就能利用人类工程师使用的同一设计空间。
Still, the important lesson is clear: once harness design becomes an executable search space, a strong coding agent can exploit the same design space human engineers use.
工作流设计
Workflow Design
运行框架工程中的工作流,可以由领域专家手工设计。以自动化研究为例,人们已经提出并测试了多种框架。AI Scientist 系统(Lu 等,2026)构建了一条提出研究想法、编写代码、运行实验、分析结果、撰写论文和开展同行评审的流水线。Meng 等(2026)在 ScientistOne 中将可验证性作为核心设计约束:每一项主张(引文、数值、方法、结论)都必须能够追溯到证据来源,并经过证据链检查的审计。
Workflow design in harness engineering can be handcrafted by domain experts. Taking auto-research as an example, various frameworks have been proposed and tested. The AI Scientist system (Lu et al. 2026) builds a pipeline to propose research ideas, write code, run experiments, analyze results, write a manuscript, and perform peer review. Meng et al. (2026) make verifiability the central design constraint in ScientistOne, where every claim (citation, numerical, methodological, conclusion) must trace to an evidence source and is audited by Chain-of-Evidence checks.
Autodata Agent(Kulikov 等,2026)旨在充当生成训练数据和评测数据的数据科学家。主 Agent 管理一个提出问题的挑战者、一个弱求解器、一个强求解器和一个验证器/评判器,目标是合成难度“恰到好处”的数据,即强求解器能够成功,而弱求解器会失败。
The Autodata agent (Kulikov et al. 2026) is designed to work as a data scientist for generating training and evaluation data. The main agent manages a challenger that proposes problems, a weak solver, a strong solver, and a verifier/judge, aiming to synthesize data at the “just right” level of difficulty, meaning that the strong solver succeeds but the weak solver fails.
在 Autodata 中,挑战者的提示词会根据求解器和验证器的反馈迭代更新。这里的局限是:合成任务用于微调弱求解器,却不用于微调强求解器;如果循环无法迭代改进强模型,它就更像是在生成的提示词分布上进行间接蒸馏,RSI 的意味较弱。
In Autodata, the challenger prompt is updated iteratively according to feedback from the solvers and verifier. The limitation here is that synthesized tasks are used to fine-tune weak solvers but not strong solvers; if the loop cannot iteratively improve the strong model, it is more like indirect distillation over a generated prompt distribution, with less RSI flavor.
工作流的设计空间极其庞大,因此我们自然可以将工作流设计视为搜索问题,也就应该能够通过算法找到好的方案,而不只是手工设计。沿着这一方向,Automated Design of Agentic Systems(ADAS;Hu 等,2025)将 Agent 设计本身表述为优化问题,即“元 Agent 搜索”:由元 Agent 提出新的 Agent 工作流设计。
The design space for workflow is enormous, and naturally we can think of workflow design as a search problem, and therefore we should be able to find good solutions by algorithms rather than only manually craft them. Following this direction, Automated Design of Agentic Systems (ADAS; Hu et al. 2025) formulates agent design itself as an optimization problem, “meta-agent search” where a meta-agent proposes new designs of agentic workflows.
- 以 CoT、自我改进(self-refine)等简单 Agent 初始化 Agent 工作流档案库。
- 让元 Agent 受档案库中现有方案启发,完全以代码编写新的 Agent。
- 元 Agent 先生成新工作流的高层描述,再用代码实现。
- 随后,元 Agent 对程序草稿执行两步自我改进,以检查其新颖性(即先让模型提供反馈,再让同一个模型根据反馈改进先前生成的输出;Madaan 等,2023)。
- 评测每个新候选方案,将成功者加入档案库。
- 重复步骤 2–3,直到达到最大迭代次数。
- Initialize an archive of agentic workflows with simple agents such as CoT and self-refine.
- Ask a meta-agent to program new agents, all in code, inspired by existing solutions in the archive.
- The meta-agent first generates a high-level description of the new workflow, and then implements it in code.
- The draft program then goes through two self-refine steps (i.e. ask the model to provide feedback and then ask the same model to refine the previously generated outputs based on the feedback; Madaan et al. 2023) by the meta-agent to check its novelty.
- Evaluate each new candidate and add successful ones back to the archive.
- Repeat steps 2-3 until the maximum iteration count is reached.
AFlow(Zhang 等,2025)将 Agent 工作流表示为图:节点表示调用 LLM 的行动,边通过代码实现逻辑操作。工作流优化依赖 MCTS(蒙特卡洛树搜索):
AFlow (Zhang et al. 2025) represents an agentic workflow as a graph, where nodes represent LLM-invoking actions and edges implement logical operations in code. The workflow optimization relies on MCTS (Monte Carlo Tree Search):
- 用模板初始化树中的起始工作流 。
- 通过分数与均匀探索的软混合策略,选择一个工作流节点。
- 让 LLM 根据该工作流的评测表现生成修改后的工作流,从而扩展节点。
- 执行并评测新工作流。
- 如果新工作流在 轮预算内表现出改进,就将其加入树中。
- 重复步骤 2–5,直到前 个工作流的平均分趋于平稳,或用尽预算。
- Initialize the starting workflow in the tree with a template.
- Select a workflow node using a soft mixture of score and uniform exploration.
- Expand it by asking an LLM to produce a modified workflow conditioned on its evaluation performance.
- Execute and evaluate the new workflow.
- Add it back to the tree if the new workflow shows improvement within a budget of rounds.
- Repeat steps 2-5 and stop when the top- average score plateaus or hit the budget.
AFlow 在问答、代码和数学任务上的实验表明,与手工设计的工作流和 ADAS 相比,它取得了不错的提升。
Experiments of AFlow in QA, code, and math tasks showed decent improvement of AFlow over manually designed workflows and ADAS.
自我改进的运行框架
Self-Improving Harness
上下文工程或工作流设计,都只是运行框架的一部分。我们需要搜索整个设计空间,同时优化上下文管理逻辑、工作流、权限,以及运行框架的许多其他组件。正如 Meta-Harness、ADAS 和 AFlow 等工作所展示的,✨代码✨是定义程序与系统的通用语言。简单来说,运行框架就是一段代码,规定提示词、工具调用、子 Agent、控制流、记忆和工作流逻辑如何协同运作。如果 LLM 能优化执行 Agent 的代码,它就能进入一个比手写提示词大得多的设计空间。
Either context engineering or workflow design is only one part of a harness. We need to search through the entire design space and optimize context-management logic, workflow, permissions, and many other harness components together. As we have seen in work like Meta-Harness, ADAS, and AFlow, ✨code✨ is a universal language for defining programs and systems. In simple words, a harness is code that programs how prompts, tool calls, subagents, control flow, memory, and workflow logic work together. If an LLM can optimize the code that executes agents, it can access a much larger design space than hand-written prompts.
Self-Taught Optimizer(STOP;Zelikman 等,2023)是递归改进辅助框架的早期案例之一。在步骤 ,初始改进器 接收初始方案 、效用函数 和黑箱语言模型 ,返回改进后的方案 ,即 。STOP 的目标不是直接改进 ,而是改进改进器 本身。
Self-Taught Optimizer (STOP; Zelikman et al. 2023) is one of the early examples of recursive scaffolding improvement. A seed improver at step takes an initial solution , a utility function , and a black-box language model , and returns an improved solution , that is, . The goal of STOP is not directly to improve but to improve the improver itself.
首先,将元效用定义为:给定改进器函数 在一组下游任务 上的平均效用:
First, let’s define the meta-utility as the average utility of a given improver function over a collection of downstream tasks :
由于改进改进器函数本身也是一个优化问题,我们可以根据元效用衡量的 表现,通过一次自我改进更新,递归得到新版 :
Because improving the improver function is an optimization problem itself, we can recursively get a new version of based on ’s performance measured by meta-utility via a self-improvement update:
在实验中,改进后的改进器发现了多种策略,例如遗传算法、分解并改进各个部分、多臂提示词老虎机、模拟退火、改变温度,以及束搜索/树搜索。这与将运行框架工作流表示为可优化对象的方式类似。
In their experiments, the improved improver discovered various strategies, such as genetic algorithms, decomposing and improving parts, multi-armed prompt bandits, simulated annealing, varying temperature, and beam/tree search. This is analogous to how a harness workflow can be represented as an object for optimization.
Zelikman 等(2023)的研究中,有一个值得警惕的结果:使用 GPT-4 时,STOP 的平均下游表现随迭代提升;但使用 GPT-3.5 和 Mixtral 等较弱模型时,表现反而下降。仅有递归结构还不够,基础模型必须足够强,才能改进机制。这意味着,运行框架改进可以让模型得到更好的部署,但智能仍然是核心。
A cautionary result in Zelikman et al. (2023)’s findings is that STOP improved mean downstream performance across iterations with GPT-4 but degraded with weaker models like GPT-3.5 and Mixtral. Recursive structure alone is not enough. The base model must be capable enough to improve the mechanism. This implies that harness improvement enables better deployment of the model but intelligence is still the core.
Lin 等(2026)更详细地研究了运行框架演化对模型能力的依赖。他们区分了两个维度:(1)运行框架更新能力,指产生有用的运行框架修改的能力;(2)运行框架受益能力,指利用更新后的运行框架、更好地解决任务的能力。有趣的是,他们的实验观察到,从 Qwen3.5-9B 到 Claude Opus 4.6,一系列规模和核心智能不同的模型表现出相近的运行框架更新能力;9B 规模的运行框架提议者/演化器,能够写出在过程结构上与 Opus 同构的技能。要充分利用运行框架,模型需要正确、及时地调用技能/工具,并擅长长时程指令遵循。
Lin et al. (2026) investigated the dependency of harness evolution on model capabilities in more details. They disentangled two axes: (1) harness-updating refers to the capability of producing useful harness edits and (2) harness-benefit denotes the capability of utilizing the updated harness, to achieve better task solving. Interestingly a range of model of different sizes and core intelligence, from Qwen3.5-9B to Claude Opus 4.6, were observed in their experiments to show similar harness updating capability; the 9B harness proposer/evolver is able to write a skill procedurally isomorphic to Opus. To best utilize a harness, a model needs to invoke skills/tools correctly and timely and be good at long-horizon instruction following.
较近的一项工作 Self-Harness(Zhang 等,2026),依靠 LLM Agent 通过“提出—评测—接受”循环改进自身的运行框架。
A more recent work, Self-Harness (Zhang et al. 2026), relies on LLM agents to improve their own harness via a propose-evaluate-accept loop.
Self-Harness 的循环有三个阶段:
The loop in Self-Harness has three stages:
- 弱点挖掘:将失败聚类为以验证器结果为依据的失败模式。
- 使用当前运行框架 对任务进行评测,并收集执行轨迹以供分析。
- 需要注意,两次运行在错误日志中表面上可能具有相同的验证器结果,例如超时或缺少产物,但背后的因果机制不同。因此,我们需要信息丰富的失败记录,其中包含最终验证器层面的原因、相关 Agent 行为的因果状态,以及执行轨迹暴露出的抽象 Agent 机制,才能揭示根本原因。
- 运行框架提议:根据挖掘出的失败模式,提出范围受限的运行框架修改。
- 在 下调用同一个模型,让它充当提议者。
- 向模型提供范围受限的提议上下文:(1)当前运行框架中允许编辑的部分;(2)评测系统提供的、以验证器结果为依据的失败模式;(3)应该保留的成功行为记录;(4)先前尝试过的修改摘要。
- 运行框架修改应优先处理反复出现、可以解决(例如并非任务特有的难点)、且能通过局部修改消除的错误模式。
- 候选修改应彼此不同,具有多样性。
- 提议验证:验证并合并合格修改,创建新的运行框架 。
- 在内部分割 (测试弱点是否得到解决)和留出分割 (检查是否引入其他未知问题)上,通过回归测试评测候选修改。
- 只有在内部分割和留出数据上都没有回归退化时,才接受候选方案。
- 将接受的候选方案合并,把运行框架更新为 ;对于被拒绝的候选方案,则只记录日志,不修改当前生效的运行框架。
- Weakness mining: cluster failures into verifier-grounded failure patterns.
- The current harness is used to evaluate on tasks and execution traces are collected for analysis.
- Note that two runs can share the same verifier outcome in the error logs on the surface, such as timeout or missing artifact, while having different causal mechanisms. Therefore we need a failure record of rich information, containing the terminal verifier-level cause, the causal status of the relevant agent behavior, and the abstract agent mechanism exposed by the trace, to uncover the root causes.
- Harness proposal: propose bounded harness edits based on mined failure patterns.
- The same model is invoked under as a proposer.
- The model is provided with a bounded proposal context: (1) the editable surfaces of the current harness, (2) the verifier-grounded failure patterns from the evaluation system, (3) records of passing behaviors that should be preserved, and (4) summaries of previously attempted edits.
- Harness edits should prefer recurrent error patterns that are addressable (e.g. not task-specific difficulty) and can be resolved by narrow changes.
- Harness edit candidates should be distinct and diverse.
- Proposal validation: validate and merge qualified edits to create a new harness .
- Candidate edits are evaluated by regression tests on held-in (for testing whether the weakness is resolved) and held-out (for checking whether other unknown issues were introduced) splits.
- Candidates are accepted only if they have no regression on both held-in and held-out data.
- Accepted candidates are merged to update the harness to , while rejected candidates are logged without changing the active harness.
在 Terminal-Bench-2 上运行 MiniMax M2.5、Qwen3.5-35B-A3B 和 GLM-5 时,研究表明 Self-Harness 能学到针对特定模型的运行框架指令,分别处理不同基础模型的不同弱点,并提高留出集通过率。
When running MiniMax M2.5, Qwen3.5-35B-A3B, and GLM-5 on Terminal-Bench-2, Self-Harness was shown to learn model-specific harness instructions that target at different weaknesses of different base models and improve held-out pass rates.
Self-Harness 这类工作确实让我担心:如果允许程序编辑操作系统,抽象边界就被打破了。允许编辑的范围需要妥善设计,权限控制层和安全层则必须放在这一循环之外。围绕奖励投机的所有挑战仍然存在。
Self-harness type of work does raise my concerns that if a program is allowed to edit the OS system, abstraction boundaries are broken. The editable surface needs to be properly designed and the permission control and security layers need to live outside this loop. All the challenges around reward hacking still remain.
Agentic Harness Engineering(AHE;Lin 等,2026)认为,运行框架演化的瓶颈集中在可观测性上:也就是说,一次执行失败时,我们需要知道哪个组件对此负责,而且每项修改都应有证据支撑。
Agentic Harness Engineering (AHE; Lin et al. 2026) see the bottlenecks of harness evolution are around observability—that is, when a rollout fails, we need to know which component is responsible for that and every edit should be grounded by evidence.
该框架围绕可观测性的三个支柱构建闭环:
The framework creates a closed loop with 3 observability pillars:
- 组件可观测性:每个允许编辑的运行框架组件都在文件系统中有对应表示,使行动空间清晰且可追溯。
- 运行框架包含 7 个组件:系统提示词、工具描述、工具实现、中间件、技能、子 Agent 配置和长期记忆。
- 每种失败模式都映射到一个组件,以便更有针对性地修改。
- 经验可观测性:分析并总结大量原始执行轨迹,形成分层的证据和失败模式。
- 每个运行框架生成 条执行轨迹。
- 使用一个 Agent(“Agent debugger”)分析各自存储在单独文件中的执行轨迹,并为每个任务生成关于失败或成功根本原因的分析报告。
- 将所有逐任务报告汇总为基准测试概览,供下一步使用;需要时也可以访问原始执行轨迹。这种分层访问结构更加节省 token。
- 决策可观测性:每项修改都附带一个供下一轮验证的预测。
- 一个 Agent(“Evolve agent”)读取仓库,决定编辑哪个组件,然后生成修改及其背后的理由。
- 每项修改都是文件层面、可证伪的主张,可以在下一轮验证,并受以下两项约束:
- (1)修改只能应用于运行框架工作区。运行记录目录、轨迹记录器、验证器和 LLM 配置均为只读,这排除了一组奖励投机手段(例如禁用验证器、更换模型或提高推理预算),从而使记录到的每项提升都能归因于运行框架修改。
- (2)修改由证据驱动,并附带一条声明记录:失败证据名称、推断出的根本原因、有针对性的修复,以及预测影响,其中同时包含预期修复效果和可能发生回归退化的风险。
- Component observability: every editable harness component has a representation in the file system so the action space is explicit and tracable.
- A harness contains 7 components: system prompt, tool description, tool implementation, middleware, skill, sub-agent configuration, and long-term memory.
- Each failure pattern is mapped to one component so the edit can be more targeted.
- Experience observability: analysize and summarize a large amount of raw trajectories into a hierarchy of evidence and failure patterns.
- Each harness generates traces.
- Use an agent (“Agent debugger”) to analysis the trajectories each stored in one file and generate per-task analysis report on the root cause for the failure or success.
- All the per-task reports are aggregated into a benchmark overview for the next step, and raw traces can be accessed if needed. This layered access structure is more token efficient.
- Decision observability: every edit is paired with a prediction for the next round to validate.
- An agent (“Evolve agent”) reads the repo and decides which component to edit, and then produces the edit and the reasoning behind it.
- Every edit is a file-level, falsifiable claim and can be verified in the next round, under two constraints:
- (1) Edits are only applied to the harness workspace. the runs directory, tracer, verifier, and LLM configuration are read-only, which disables a set of reward hacking (e.g disabling the verifier, swapping the model, or raising the reasoning budget) and thus it can keep every recorded gain attributable to harness edits.
- (2) Edits are evidence-driven, with a manifesto entry: the failure evidence’s name, the inferred root cause, the targeted fix, and a predicted impact comprising both expected fixes and at-risk regressions.
在 Terminal-Bench-2 上,除 Hard 难度层级外,AHE 的表现优于人工设计的运行框架(OpenCode、Terminus-2、Codex),也优于其他若干自我演化基线(ACE、TF-GRPO)。同一个冻结后的运行框架,无需进一步演化,就能迁移到 SWE-bench-verified。这表明,演化后的运行框架能够将工程经验编码进组件,而不是只针对特定基准进行优化。
On Terminal-Bench-2, AHE achieved better than human-designed harness (OpenCode, Terminus-2, Codex) except for Hard tier and a few other self-evolve baselines (ACE, TF-GRPO). The same frozen harness, without further evolving, transfers to SWE-bench-verified, indicating that the evolved harness is able to encode engineering experience into harness components rather than doing benchmark-specific optimization.
进化式搜索
Evolutionary Search
进化式搜索是一种受自然选择启发的优化方法(参见我以前关于进化算法的文章)。它通过对方案进行变异,并只保留群体中“适应度”较高的方案,来演化一个解的种群。在以下情况下,进化式搜索很有用:(1)搜索空间庞大或形状不规则;(2)难以直接用梯度优化,但容易评估方案。运行框架搜索似乎很适合这种方法。
Evolutionary search is an optimization method inspired by natural selection (see my old post on evolutionary algorithm). It evolves a population of solutions by mutating them and only keeping those with high “fitness” in the crowd. Evolutionary search comes in handy when (1) the search space is extensive or weirdly shaped; and (2) it is hard to optimize directly with gradients but easy to evaluate solutions. Harness search seems to be a good fit here.
以往研究已经在提示词工程中使用进化式搜索。Promptbreeder(Fernando 等,2023)通过丰富的变异操作优化特定任务的提示词;有趣的是,变异提示词本身(即要求 LLM 对任务提示词进行变异的指令)也会通过演化得到改进。GEPA(Agrawal 等,2025)将基于反思的提示方法与进化式搜索结合,通过对试错执行轨迹进行自然语言反思,提出提示词更新。
Evolutionary search has been used in prompt engineering in the past studies. Promptbreeder (Fernando et al. 2023) optimizes task-specific prompts through a rich set of mutation operations, and interestingly the mutation prompts (i.e. instructions to an LLM to mutate a task prompt) are themselves also improved through evolution. GEPA (Agrawal et al. 2025) combines reflection-based prompting with evolutionary search and uses natural language reflection over trajectories of trial and error to propose prompt updates.
Novikov 等(2025)提出了 AlphaEvolve,这是一个基于编程 Agent 的进化式搜索系统。它保存一组候选程序,并提示权重冻结的 LLM 生成用于改进的代码差异。随着系统反复评测子代程序并保留成功者,它会逐渐发现更好的方案。
Novikov et al. (2025) introduced AlphaEvolve as a coding-agent evolutionary search system, which stores a pool of candidate programs and prompts frozen LLMs to generate diffs for improvement. As the system repeatedly evaluates child programs and keeps successful ones, it discovers better solutions in time.
AlphaEvolve 的设计中,有几个值得关注的细节:
A few details matter in the design of AlphaEvolve:
- 提示词包含父代程序、结果、指令,有时也包含元信息。
- 编程 Agent 可以访问整个仓库,但允许改进的代码区域用
# EVOLVE-BLOCK-START和# EVOLVE-BLOCK-END显式标记。 - 元提示词根据 LLM 的建议,与指令和上下文共同演化,方式类似于方案程序的演化。
- The prompt includes parent programs, results, instructions, and sometimes meta information.
- The coding agent has access to the full repo, but code regions for improvement are explicitly marked with
# EVOLVE-BLOCK-STARTand# EVOLVE-BLOCK-END. - Meta-prompt co-evolves with instructions and context as suggested by LLM, in a similar way as how we evolve solution programs.
消融实验展示了演化流程、提示词中的上下文、元提示词、整文件演化,以及使用更强 LLM 的作用。
Ablations show the evolution procedure, context in prompts, meta-prompts, full-file evolution and the use of stronger LLMs.
近期的一些变体,例如 ThetaEvolve(Wang 等,2025),将进化式搜索与强化学习和上下文学习结合;DemoEvolve(Che 等,2026)则用人类专家示范扩充自主执行档案库,为运行框架层面的诊断和编辑提供参考经验。另一方面,ShinkaEvolve(Lange 等,2025)引入了三个新组件,以提高 LLM 采样效率:
Recent variants such as ThetaEvolve (Wang et al. 2025) combines evolutionary search with RL and in-context learning, and DemoEvolve (Che, et al. 2026) augments the self-rollout archive with human expert demonstrations as reference experience for harness-level diagnosis and editing. ShinkaEvolve (Lange et al. 2025), on the other hand, introduced three new components to improve LLM sampling efficiency:
- 通过设计父代采样策略,在表现排名与后代数量之间取得平衡,实现样本效率更高的探索。
- 根据基于嵌入的余弦相似度,丢弃与现有种群过于相似的候选方案,实现基于代码新颖性的拒绝采样。
- 在元草稿区中识别成功方案中的良好模式,以指导未来变异。
- More sample-efficient exploration by designing parent sampling to balance performance rank and offspring count.
- Code-novelty rejection sampling by discarding candidates that are too similar to the existing population based on embedding-based cosine similarity.
- Identifying good patterns in successful solutions in a meta-scratchpad to guide future mutation.
与上述专注于改进方案的方法不同,Darwin Gödel Machine(DGM;Zhang 等,2025)明确面向可编辑的运行框架代码仓库,使用基于 LLM 的编程 Agent 对其进行演化。准确地说,这个 Agent 被允许修改自身的运行框架。后续关于 Hyperagents 的工作(Zhang 等,2026)引入了元 Agent,控制如何修改现有任务 Agent,以创建新的 Agent。
Unlike the methods above, which focus on solution improvement, Darwin Gödel Machine (DGM; Zhang et al. 2025) explicitly targets the evolution of an editable harness-code repository with an LLM-based coding agent. Precisely, this agent is allowed to modify its own harness. A follow-up work on Hyperagents (Zhang et al. 2026) introduced a meta-agent to control how to modify existing task agents to create new ones.
- 在候选池中放入一个编程 Agent,作为起点。
- 每轮迭代按与表现成正比、与已有子代数量成反比的概率选择一个父代,对它进行修改并分支,产生新的 Agent。
- 选中的父代 Agent 检查自身的基准评测日志,再对自己的运行框架代码库提出改进,生成新版编程 Agent。代码编辑由两个基本工具实现:(1)bash(参数:
<bash_command>);(2)editor(参数:view/create/edit <file_path>)。 - 对新的编程 Agent 进行评测,只将表现足够好的 Agent 加回候选池。
- 重复步骤 2–4,直到满足某项停止条件。
- Start with one coding agent in the pool.
- In each iteration, pick one parent with a probability proportional to its performance and inversely to the number of children it has, to modify and branch off to produce new agents.
- The selected parent agent examines its own benchmark evaluation log and then proposes improvements to its own harness codebase to generate a new version of the coding agent. Code editing is implemented with two basic tools: (1) bash (args:
<bash_command>) and (2) editor (args:view/create/edit <file_path>). - New coding agents are evaluated, and only those with sufficiently high performance are added back into the pool.
- Repeat steps 2-4 until some stop criteria hit.
DGM 是在固定模型下进行运行框架演化。在以 Claude 3.5 Sonnet 为基础 LLM、采用简单初始运行框架配置的实验中,DGM 发现的 Agent 在 SWE-bench Verified(20% 提升至 50%)和 Polyglot(14.2% 提升至 30.7%)上的表现,与人工设计的 Agent 相当或更好。
DGM is harness evolution under a fixed model. In experiments with Claude 3.5 Sonnet as the base LLM and simple initial harness configs, the DGM-discovered agents are comparable to or outperform handcrafted agents on SWE-bench Verified (20% to 50%) and Polyglot (14.2% to 30.7%).
当候选方案可以自动评测,且适应度容易量化时,这类方法效果很好,例如矩阵乘法、GPU 内核优化、算法竞赛和数据中心调度。对于评测缓慢、模糊或主要依赖启发式方法的领域,它们则面临困难。演化的计算效率和有效性也值得关注。
This family of methods works well when candidate solutions are automatically evaluable and candidate fitness is easy to quantify, such as matrix multiplication, GPU kernel optimization, algorithm contests, datacenter scheduling. It struggles with domains where evaluation is slow, ambiguous, or mostly heuristic-based. The compute efficiency and effectiveness of evolution are also concerns.
与模型权重联合优化
Joint Optimization with Model Weights
运行框架演化改变的是模型周围的非参数化系统。为了实现完整的自我改进,完全可以允许模型同时更新自身权重。权重更新可以通过改进模型训练流程实现,也可以通过测试时的持续学习实现。持续学习这个主题值得在未来另写一篇文章。
Harness evolution changes the non-parametric system around the model. To enable full self-improvement, the model can totally be allowed to update its own weights at the same time. The weight update can be implemented via improvements in the model training pipeline or continual learning at test time. The topic of continual learning is worthy of its own post in the future.
SIA(Hebbar 等,2026)是将运行框架改进与模型参数更新结合在同一优化循环中的早期尝试,其设计包含三个组件:
SIA (Hebbar et al. 2026) is an early attempt to combine harness improvement and model-parameter updates in the same optimization loop, with three components in the design:
- Meta-Agent:提出初始运行框架。
- Task-Specific Agent:执行任务。
- Feedback-Agent:根据近期执行轨迹,选择更新运行框架还是模型权重。
- Meta-Agent: proposes the initial harness.
- Task-Specific Agent: executes the task.
- Feedback-Agent: chooses whether to update the harness or the model weights based on recent trajectories.
SIA 实验中的一些设计选择引入了混淆因素,使结果难以解释。例如,任务专用 Agent 远弱于 Meta-Agent 和 Feedback-Agent 使用的模型(gpt-oss-120b 对比 Claude Sonnet 4.6),基线也过弱,无法与相关方法进行清晰的交叉比较。我认为这个方向很有意思,但证据尚属初步。同时,训练稳定性和古德哈特效应等许多挑战仍未解决。
There are a few confounding choices in SIA’s experiments that make the results hard to interpret. For example, the task-specific agent is much weaker than the models used for the Meta-Agent and Feedback-Agent (gpt-oss-120b vs Claude Sonnet 4.6), and the baselines are too weak to cross-reference cleanly against related methods. I would consider the direction interesting, but the evidence provisional. Yet many challenges, such as training stability and Goodhart effect, still remain open.
Continual Harness(Karten 等,2026)在长时程游戏场景中进行了实验:一边更新运行框架,一边通过蒸馏强教师模型对低奖励执行轨迹所给出的标签,协同学习一个策略模型。
Continual Harness (Karten et al. 2026) experimented in long-horizon gameplay setting with harness updating and co-learning a policy model by distilling a strong teacher model’s labels on low-reward trajectories.
未来的挑战
Future Challenges
AI Scientist 这一系列工作以撰写研究论文为实验形式,有力地证明了专家设计的运行框架能够协调自动化研究循环中的大部分环节。但论文生产并不等于科学发现。系统可以写出一篇看似合理的稿件,同时仍然存在编造引文、实现偏移或实验结果薄弱等问题。
The AI Scientist line of work is a strong demonstration that an expert-designed harness can coordinate a large portion of auto-research loop, experimented in the form of writing research papers. But paper production is not identical to scientific discovery. A system can write a plausible manuscript while still having fabricated citations, implementation drift, or weak experimental results.
Trehan 与 Chopra(2026)测试了 LLM 能否在最少的辅助框架和基本工具(即 read_file、write_file、llm_search、list_files)支持下,从研究想法走到论文。每个想法都有专用工作区,Agent 可以在其中生成和读取文档,作为上下文的一部分。他们在三个领域开展实验(世界模型、多 Agent 强化学习、AI 安全与对齐),每个领域包含 45–50 份高质量种子文档,用于启发新想法。人类专家只选出四个想法进入完整流程,其中只有一个被完整执行并形成论文。他们在实验中观察到六种反复出现的失败模式:
Trehan & Chopra (2026) tested whether LLMs can go from a research idea to a paper with minimal scaffolding and basic tools (i.e., read_file, write_file, llm_search, list_files). Each idea had a dedicated workspace where agents could generate and read documents as part of context. They experimented in three domains (world models, multi-agent RL, AI safety & alignment), with each domain containing 45-50 high-quality seed documents to inspire new ideas. Only four ideas were selected by human experts to run through the full pipeline, and only one was fully executed into a paper. They observed six recurring failure modes in the experiments:
- 偏向训练数据中的默认做法:使用旧库、过时命令、标准格式,或未经实际仓库和数据集证实的假设。
- 执行压力下的实现偏移:当实现变得技术上复杂时,模型可能转向常见的简单方案,而不是提出的方法。
- 记忆和上下文退化:长时程项目会丢失关键细节,除非将日志写成持久产物。
- 过度乐观:尽管实验充满噪声或已经失败,模型仍宣告成功。Bubeck 等(2025)也观察到了类似的“p 值操纵与大喊发现了”模式:模型可能贴上“数值胶带”进行临时修补,在信号仍然只是噪声时就宣告胜利。
- 领域智能不足:模型缺乏隐性的实践知识,例如预判实现复杂度、判断实验结果是否可信,或知道哪些基线重要。
- 科研品味薄弱:实验可能可以执行,却没有回答正确的问题。
- Bias toward training-data defaults: use old libraries, stale commands, standard formats, or assumptions not grounded in the actual repository or dataset.
- Implementation drift under execution pressure: when implementation becomes technically complex, the model may move toward a common simpler solution rather than the proposed method.
- Memory and context degradation: long-horizon projects lose critical details unless logs are written as persistent artifacts.
- Over-optimism: the model declares success despite noisy or failed experiments, similarly observed as “p-hacking and eureka-ing” pattern by Bubeck et al. (2025) where models can introduce “numerical duct tape” and declare victory when signals are still noise.
- Insufficient domain intelligence: the model lacks tacit craft knowledge, e.g. predicting implementation complexity, judging whether an experimental result is plausible, or knowing which baselines matter.
- Weak scientific taste: experiments may be executable but fail to answer the right question.
在迈向完整 RSI 的道路上,研究者已经取得实质进展,但仍存在若干瓶颈。
Toward full RSI, researchers have made real progress, but several bottlenecks remain.
1. 能力弱且标准模糊的评估器。许多研究主张没有快速、精确的验证器,很多现实任务也一样。当前的自我改进循环,在评测指标可测量且客观的任务上效果最好,这与强化学习发挥作用的方式相似。
1. Weak and fuzzy evaluators. Many research claims do not have a fast and precise verifier, and the same is true for many real-world tasks. Current self-improvement loops work best for tasks when evaluation metrics are measurable and objective, similar as how RL works.
科研品味、新颖性和长期科学价值则难以衡量得多。例如,科研品味往往涉及问题表述、实验设计,以及判断哪些出人意料的结果值得追究、哪些失败案例值得重试。
Research taste, novelty, and long-term scientific value are much harder to measure. For example, research taste often mixes problem framing, experimental design, and judgment about which surprising results are worth pursuing and which failure cases are worth retries.
2. 上下文与记忆的生命周期。随着 AI Agent 变得更自主、更独立,记忆也会不断增长。有用的运行框架需要管理上下文和记忆,弥补当前长上下文生成的局限,同时尽可能提高长时程任务的成功率。人类能够在一生中维持记忆,这里存在一种类比:我认为,上下文工程将会、也应当成为智能的核心组成部分,而不是停留在软件系统层。
2. Context and memory lifecycle. Memory grows as AI agents become more autonomous and independent. A useful harness needs to manage context and memory to complement existing limitation in long-context generation while still maximizing the success of long-horizon tasks. Since humans are able to maintain memory through our life time, I see an anoloy here that context engineering will and should become a core part of intelligence, rather than staying in the software system layer.
3. 负面结果。研究者受到发表成功结果的激励,因此文献偏向成功。训练于海量数据的 LLM(至少目前主要还是人类创建的数据,哈哈),可能因为数据中成功与失败案例不平衡,而不擅长决定何时放弃假设、报告负面结果,甚至承认失败。研究运行框架应让失败尝试易于保存,因为从失败中学习,是缩小任务搜索空间的最佳方式。
3. Negative results. Researchers are incentivized to publish successful results and thus literature is biased toward successes. LLMs trained on a vast amount of data (mostly human created, at least for now, lol) may be bad at deciding when to abandon a hypothesis, report a negative result, or even acknowledge a failure due to the imablance of success vs failure cases in data. A research harness should make failed attempts easy to preserve, as learning from failure is the best way to trim down the task search space.
4. 多样性坍缩。进化和强化学习循环倾向于利用已知的高奖励模式。我们需要相应机制,防止种群坍缩为同一方案的不同变体。这对于开放式研究尤其重要,因为在当前评估器看来,最佳路径起初可能表现更差。
4. Diversity collapse. Evolutionary and RL loops tend to exploit known high-reward patterns. We need mechanisms to prevent the population from collapsing into variants of the same solution. This is especially critical for open-ended research, where the best path may initially look worse under the current evaluator.
5. 奖励投机。自我改进循环会优化它获得的任何信号。如果奖励来自单元测试,Agent 可能对测试过拟合;如果来自评判模型,它可能学会针对该评判器的奖励投机技巧;如果来自基准分数,它可能利用基准中的非本质特征。
5. Reward hacking. A self-improvement loop optimizes whatever signal it is given. If the reward comes from unit tests, the agent may overfit to tests; if it comes from a judge model, it may learn reward hacking tricks specific to this judge; if it comes from benchmark scores, it may exploit benchmark artifacts.
评估器和权限控制很可能应当位于运行框架演化循环之外,并在重要决策点设置留出测试、执行轨迹审计和人工审查。监督能够在多大程度上扩展并自动化,仍是开放的研究领域。
The evaluator and permission control should likely sit outside the loop that evolves harness, with held-out tests, trace audits, and human review at decision points that matter—how much oversight can be scaled up and automated remains an open research area.
6. 长期成功。外部优化循环处理的奖励,超出了我们能在训练沙箱中模拟的单次执行范围。
6. Long-term success. An extrinsic loop of optimization works on rewards outside of individual rollouts that we can simulate in training sandbox.
以编程 Agent 为例。它已经提高了软件工程的日常生产力,但很多优化目标仍然过于短期。它常常能完成手头任务,但对于一个由数百乃至数千名工程师共同维护的仓库,如何保护其长期健康,则没有那么明确。标准的、基于沙箱的 RLVR 风格训练,很少涵盖可维护性、所有权边界、迁移成本、向后兼容性或未来调试负担。
Take coding agent as an example. Coding agents have already increased daily productivity in software engineering, but many optimization goals are still too short-term. It can often complete the task at hand, but less obvious how it should protect the long-term health of a repo collectively maintained by hundreds or thousands of engineers. Standard sandbox-based RLVR-style training rarely captures maintainability, ownership boundaries, migration cost, backwards compatibility, or future debugging burden.
7. 人类的角色。人类应向更高层移动,而不是从循环中被移除。这意味着,人类应在合适的时间、合适的抽象层级提供监督;系统设计则应考虑何时、如何设置这样的介入点。
7. The role of humans. Humans should move up the stack, not be removed from the loop, meaning that human should provide oversight at the right time, at the right abstraction level and our system design should consider when and how to set up such touch points.
上面列出的许多挑战,需要人类的反馈和引导。毕竟,我们构建技术是为了人类更美好的未来,而不是反过来。
Many challenges listed above need human’s feedback and steering. After all, we are building the technology for better future of humanity, not other way around.
引用
Citation
请按以下方式引用本文:
Please cite this work as:
或使用以下 BibTeX 引用:
Or use the BibTeX citation:
附录:一些有用的基准
Appendix: Some useful benchmarks
- PaperBench:从零复现 20 篇 ICML 2024 Spotlight 和 Oral 论文,包括理解论文贡献、开发代码库和成功执行实验。
- 每项复现任务都会分解为更小、可单独评分的任务。
- 共包含 8,316 项评分标准,由研究团队与论文作者共同制定。
- 当时最好的模型(
Claude 3.5 Sonnet,约 21%)未能超过机器学习博士。 - 包含 PaperBench、PaperBench Code-Dev(较轻量版本)和 JudgeEval。
- CORE-Bench:评估已发表研究的计算可复现性。
- 基于计算机科学、社会科学和医学领域的 90 篇科学论文,构建 270 项任务。
- 任务要求利用提供的代码和数据复现结果。
- 包含多个难度等级,以及纯语言任务和视觉语言任务。
- 当时报告的最佳 Agent(
GPT-4o和GPT-4o-mini)在最难任务上的准确率仅为 21%。
- ScienceAgentBench:评估 LLM Agent 在数据驱动科学发现中的能力。
- 从四个学科(数学、化学、生物、地理)的 44 篇同行评审文献中提取 102 项任务。
- 涵盖这些领域的基础数据科学任务:数据处理、模型开发、数据分析和信息可视化。
- RE-Bench:在真实的机器学习研究工程环境中,将前沿 AI Agent 与人类专家进行比较。
- 包含 7 个有挑战性的开放式机器学习研究工程环境。
- 每个环境 =(评分函数、起始方案、参考方案);每个环境均可用不超过 8 张 H100 GPU 运行。
- 示例包括:优化内核、开展规模定律实验、修复嵌入、微调 GPT-2 以完成问答等。
- 包含 61 位不同人类专家进行的 71 次、每次 8 小时的尝试数据。
- 人类专家在 82% 的 8 小时尝试中获得了非零分;24% 的尝试达到或超过了强参考方案。
- 在 2 小时预算下,最佳 AI Agent 的得分是人类的 4 倍;但人类从更长预算中获得的收益更大,在 8 小时和 32 小时设置下超过了 Agent。
- MLE-bench:在离线 Kaggle 竞赛上评估机器学习工程 Agent。
- 包含从 Kaggle 筛选出的 75 项机器学习工程竞赛。
- 测试训练模型、准备数据集、运行实验,以及向评分脚本提交预测结果的能力。
- 使用 Kaggle 公开排行榜作为人类基线。
- 论文中的最佳配置是采用 AIDE 辅助框架的
o1-preview,在 16.9% 的竞赛中至少达到 Kaggle 铜牌水平。 - 包含资源扩展和数据污染分析。
- KernelBench:评估生成的 GPU 内核的正确性与速度。
- 包含 250 项 PyTorch 任务,评估 LLM 能否编写快速且正确的内核。
- 评测指标 fast_p = 生成的内核中,正确且快于基线的比例。
- PaperBench: replicate 20 ICML 2024 Spotlight and Oral papers from scratch, including understanding paper contributions, developing a codebase, and successfully executing experiments.
- Each replication task is decomposed into smaller, individually gradable tasks.
- 8,316 rubrics in total, co-developed with the paper authors.
- The best model at the time (
Claude 3.5 Sonnet, ~21%) does not outperform ML PhDs. - Includes PaperBench, PaperBench Code-Dev (a lighter version), and JudgeEval.
- CORE-Bench: evaluate computational reproducibility of published research.
- 270 tasks based on 90 scientific papers across computer science, social science, and medicine.
- Tasks involve reproducing results from provided code and data.
- Includes multiple difficulty levels and both language-only and vision-language tasks.
- The best reported agent at the time (
GPT-4oandGPT-4o-mini) achieved only 21% accuracy on the hardest task.
- ScienceAgentBench: evaluate LLM agents for data-driven scientific discovery.
- Extracts 102 tasks from 44 peer-reviewed publications in four disciplines (math, chemistry, biology, geography).
- Covers basic data-science tasks in these domains: data processing, model development, data analysis, and information visualization.
- RE-Bench: evaluate frontier AI agents on realistic ML research-engineering envs against human experts.
- 7 challenging, open-ended ML research-engineering environments.
- Each environment = (scoring function, starting solution, reference solution); each can be run with 8 or fewer H100 GPUs.
- Examples: optimize a kernel, run a scaling-law experiment, fix an embedding, fine-tune GPT-2 for QA, etc.
- Includes data from 71 eight-hour attempts by 61 distinct human experts.
- Human experts achieved non-zero score in 82% of 8-hour attempts; 24% matched or exceeded strong reference solutions.
- Best AI agents scored 4× higher than humans at a 2-hour budget, but humans had better returns to longer budgets and exceeded agents at 8-hour and 32-hour settings.
- MLE-bench: evaluate ML engineering agents on offline Kaggle competitions.
- Contains 75 ML-engineering competitions curated from Kaggle.
- Tests training models, preparing datasets, running experiments, and submitting predictions to grading scripts.
- Uses Kaggle public leaderboards as human baselines.
- Best setup in the paper,
o1-previewwith AIDE scaffolding, reached at least Kaggle bronze-medal level in 16.9% of competitions. - Includes resource-scaling and contamination analyses.
- KernelBench: evaluate correctness and speed for generated GPU kernels.
- 250 PyTorch tasks to evaluate whether LLM can write fast and correct kernels.
- The evaluation metric fast_p = the percentage of generated kernels that are correct and faster than baseline.
参考文献
References
— 全文完 —
原文来自 Lilian Weng,中文为非官方学习译文。
查看原始出处 ↗