AI Agent 能够应对越来越复杂的任务。为了实现更宏大的目标,AI Agent 需要能够把问题有效分解为可管理的子组成部分,并将这些部分的完成工作安全地委派给其他 AI Agent 或人类。然而,现有的任务分解与委派方法依赖简单的启发式规则,无法动态适应环境变化,也无法稳健应对意外故障。本文提出一个智能 AI 委派的自适应框架:它是一系列涉及任务分配的决策,同时包含权力、责任和问责义务的转移,对角色与边界的明确规定,清晰的意图,以及在双方或多方之间建立信任的机制。所提出的框架适用于复杂委派网络中的人类与 AI 委派方和受托方,旨在为新兴 Agent 网络中的协议开发提供参考。
AI agents are able to tackle increasingly complex tasks. To achieve more ambitious goals, AI agents need to be able to meaningfully decompose problems into manageable sub-components, and safely delegate their completion across to other AI agents and humans alike. Yet, existing task decomposition and delegation methods rely on simple heuristics, and are not able to dynamically adapt to environmental changes and robustly handle unexpected failures. Here we propose an adaptive framework for intelligent AI delegation - a sequence of decisions involving task allocation, that also incorporates transfer of authority, responsibility, accountability, clear specifications regarding roles and boundaries, clarity of intent, and mechanisms for establishing trust between the two (or more) parties. The proposed framework is applicable to both human and AI delegators and delegatees in complex delegation networks, aiming to inform the development of protocols in the emerging agentic web.
关键词:AI、Agent、LLM、委派、多 Agent、安全
Keywords: AI, agents, LLM, delegation, multi-agent, safety
1. 引言
1. Introduction
随着先进 AI Agent 超越问答模式,其实用性越来越取决于它们能否有效地分解复杂目标并委派子任务。这种协调范式支撑的应用范围很广:在个人使用场景中,AI Agent 可以充当私人助手(Gabriel et al., 2024);在商业和企业部署中,AI Agent 可以提供支持并自动化工作流(Huang and Hughes, 2025; Shao et al., 2025; Tupe and Thube, 2025)。大语言模型(LLM)已经在机器人领域展现出潜力(Li et al., 2025a; Wang et al., 2024a),使目标设定和反馈更具交互性、更准确。近期提出的方案还强调了在虚拟经济中实现大规模 AI Agent 协调的可能性(Tomasev et al., 2025)。现代 Agent 式 AI 系统在分工各异的子 Agent 之间实现复杂的控制流,并配合中心化或去中心化的编排协议(Hong et al., 2023; Rasal and Hauer, 2024; Song et al., 2025; Zhang et al., 2025a)。这已经可以视为任务分解与委派的一个缩影,只是其中的过程被硬编码,并受到严格限制。管理动态的互联网规模交互,要求我们超越当前那些更多依赖启发式规则的多 Agent 框架所采用的方法。
As advanced AI agents evolve beyond query-response models, their utility is increasingly defined by how effectively they can decompose complex objectives and delegate sub-tasks. This coordination paradigm underpins applications ranging from personal use, where AI agents can act as personal assistants (Gabriel et al., 2024), to commercial, enterprise deployments where AI agents can provide support and automate workflows (Huang and Hughes, 2025; Shao et al., 2025; Tupe and Thube, 2025). Large language models (LLMs) have already shown promise in robotics (Li et al., 2025a; Wang et al., 2024a), by enabling more interactive and accurate goal specification and feedback. Recent proposals have also highlighted the possibility of large-scale AI agent coordination in virtual economies (Tomasev et al., 2025). Modern agentic AI systems implement complex control flows across differentiated sub-agents, coupled with centralized or decentralized orchestration protocols (Hong et al., 2023; Rasal and Hauer, 2024; Song et al., 2025; Zhang et al., 2025a). This can already be seen as a sort of a microcosm of task decomposition and delegation, where the process is hard-coded and highly constrained. Managing dynamic web-scale interactions requires us to think beyond the approaches that are currently employed by more heuristic multi-agent frameworks.
委派(Castelfranchi and Falcone, 1998)不只是把任务分解为可管理的行动子单元。除了创建子任务,委派还必须分配责任与权力(Mueller and Vogelsmeier, 2013; Nagia, 2024),因此也涉及对结果的问责。由此,委派包含风险评估,而信任可以调节风险(Griffiths, 2005)。委派还涉及能力匹配和持续的表现监测,根据反馈进行动态调整,并确保分布式任务在指定约束下完成。当前方法往往没有考虑这些因素,而更多依赖启发式规则和/或较简单的并行化。这对早期原型或许已经足够,但现实世界中的 AI 部署需要超越临时拼凑、脆弱且不可信的委派方式。我们迫切需要能够动态适应变化(Acharya et al., 2025; Hauptman et al., 2023)并从错误中恢复的系统。缺少自适应且稳健的部署框架,仍是限制 AI 应用于高风险环境的关键因素之一。
Delegation (Castelfranchi and Falcone, 1998) is more than just task decomposition into manageable sub-units of action. Beyond the creation of sub-tasks, delegation necessitates the assignment of responsibility and authority (Mueller and Vogelsmeier, 2013; Nagia, 2024) and thus implicates accountability for outcomes. Delegation thus involves risk assessment, which can be moderated by trust (Griffiths, 2005). Delegation further involves capability matching and continuous performance monitoring, incorporating dynamic adjustments based on feedback, and ensuring completion of the distributed task under the specified constraints. Current approaches tend to fail to account for these factors, relying more on heuristics and/or simpler parallelization. This may be sufficient for early prototypes, but real world AI deployments need to move beyond ad hoc, brittle, and untrustworthy delegation. There is a pressing need for systems that can dynamically adapt to changes (Acharya et al., 2025; Hauptman et al., 2023) and recover from errors. The absence of adaptive and robust deployment frameworks remains one of the key limiting factors for AI applications in high-stakes environments.
要充分利用 AI Agent,我们需要智能委派:一个以清晰的角色与边界、声誉、信任、透明度、可认证的 Agent 能力、可验证的任务执行,以及可扩展的任务分发为核心的稳健框架。本文提出一个智能任务委派框架,旨在解决这些局限。该框架借鉴人类组织的历史经验,并以关键的 Agent 安全要求为基础。
To fully utilize AI agents, we need intelligent delegation: a robust framework centered around clear roles, boundaries, reputation, trust, transparency, certifiable agentic capabilities, verifiable task execution, and scalable task distribution. Here we introduce an intelligent task delegation framework aimed at addressing these limitations, informed by historical insights from human organizations, and grounded in key agentic safety requirements.
2. 智能委派的基础
2. Foundations of Intelligent Delegation
2.1. 定义
2.1. Definition
我们将智能委派定义为一系列涉及任务分配的决策,同时包含权力、责任和问责义务的转移,对角色与边界的明确规定,清晰的意图,以及在双方或多方之间建立信任的机制。复杂任务还可能涉及任务分解,以及审慎的能力检索与匹配,以此为分配决策提供依据。
We define intelligent delegation as a sequence of decisions involving task allocation, that also incorporates transfer of authority, responsibility, accountability, clear specifications regarding roles and boundaries, clarity of intent, and mechanisms for establishing trust between the two (or more) parties. Complex tasks may also involve steps pertaining to task decomposition, as well as careful capability lookup and matching to inform allocation decisions.
谈到任务委派时,我们通常假定任务的复杂度超过了可由系统子程序处理的某个基本水平——这种最基础的外包仍需谨慎,但其范围要有限得多。另一个极端是,或许可以与获得完全自主权的 Agent 订立契约,让它们无需明确检查与许可,就能自由追求任意数量的子目标(Kasirzadeh and Gabriel, 2025)。在极限情况下,这些完全自主的 Agent 需要获得信任,承担道德决策(Sloksnath, 2025);不过,我们也许永远不会选择允许这一点,因为当代 Agent 在作出此类决策的能力上严重不足(Haas, 2020; Mao et al., 2023; Reinecke et al., 2023)。我们将这种开放式场景纳入讨论范围,但前提是能够建立适当机制,确保以更高自主性完成任务时的安全。
When we refer to task delegation we normally presume that the tasks exceed some basic level of complexity that would be handled by a system subroutine – such rudimentary outsourcing still requires care, but it is far more limited in scope. At the other end of the spectrum, it may be possible to contract with agents that are granted full autonomy, and can freely pursue any number of sub-goals without explicit checks and permissions (Kasirzadeh and Gabriel, 2025). In the limit case, such fully autonomous agents would need to be trusted with moral decisions (Sloksnath, 2025), though this may not be something we ever choose to permit as contemporary agents are severely lacking in their capacity to engage in such decisions (Haas, 2020; Mao et al., 2023; Reinecke et al., 2023). We consider such an open-ended scenario to be in scope for our discussion, though only insofar as the appropriate mechanisms can be put in place to ensure safety of more autonomous task completion.
2.2. 委派的维度
2.2. Aspects of Delegation
委派可以采取不同形式,因此,我们在这里引入若干维度,帮助把这些使用情境放到具体背景中,并使其更便于分析。
As delegation can take different forms, here we introduce several axes that help us contextualize these use cases and make them more amenable to analysis.
1. 委派方。人类或 AI。
1. Delegator. Human or AI.
2. 受托方。人类或 AI。
2. Delegatee. Human or AI.
3. 任务特征。
3. Task characteristics.
(a) 复杂度。任务本身的难度,通常与子步骤的数量以及所需推理的复杂程度相关。
(a) Complexity. The degree of difficulty inherent in the task, often correlated with the number of sub-steps and the sophistication of reasoning required.
(b) 关键程度。衡量任务的重要性,以及失败或表现不佳所带来后果的严重程度。
(b) Criticality. The measure of the task’s importance and the severity of consequences associated with failure or sub-optimal performance.
(c) 不确定性。关于环境、输入或成功实现结果的概率,存在多大程度的模糊性。
(c) Uncertainty. The level of ambiguity regarding the environment, inputs, or the probability of successful outcome achievement.
(d) 持续时间。执行任务的预期时间跨度,可以短至瞬间完成的子程序,也可以长至持续数天或数周的流程。
(d) Duration. The expected time-frame for task execution, ranging from instantaneous sub-routines to long-running processes spanning days or weeks.
(e) 成本。执行任务产生的经济或计算开销,包括 token 用量、API 费用与能源消耗。
(e) Cost. The economic or computational expense incurred to execute the task, including token usage, API fees, and energy consumption.
(f) 资源需求。完成任务所必需的特定计算资源、工具、数据访问权限或人类能力。
(f) Resource Requirements. The specific computational assets, tools, data access permissions, or human capabilities necessary to complete the task.
(g) 约束。执行任务必须遵守的操作、伦理或法律边界,它们限制了解空间。
(g) Constraints. The operational, ethical, or legal boundaries within which the task must be executed, limiting the solution space.
(h) 可验证性。验证任务结果所需的相对难度和成本。可验证性高的任务,例如代码的形式化验证、数学证明,允许“不依赖信任”的委派或自动检查。相反,可验证性低的任务,例如开放式研究,需要高度可信的受托方,或昂贵且耗费人力的监督。
(h) Verifiability. The relative difficulty and cost associated with validating the task outcome. Tasks with high verifiability (e.g., formal code verification, mathematical proofs) allow for “trust-less” delegation or automated checking. Conversely, tasks with low verifiability (e.g., open-ended research) require high-trust delegatees or expensive, labor-intensive oversight.
(i) 可逆性。任务执行所产生的影响能够在多大程度上撤销。与可逆任务,例如起草电子邮件、标记数据库条目相比,会在现实世界产生副作用的不可逆任务,例如执行金融交易、删除数据库、向外部发送电子邮件,需要更严格的责任隔离机制和更陡峭的权威梯度。
(i) Reversibility. The degree to which the effects of the task execution can be undone. Irreversible tasks that produce side effects in the real world (e.g., executing a financial trade, deleting a database, sending an external email) require stricter liability firebreaks and steeper authority gradients than reversible tasks (e.g., drafting an email, flagging a database entry).
(j) 上下文依赖程度。有效执行任务所需的外部状态、历史信息或环境感知的数量与敏感程度。高度依赖上下文的任务会扩大隐私暴露面,而不依赖上下文的任务则更容易被隔离划分,并外包给信任程度较低的节点。
(j) Contextuality. The volume and sensitivity of external state, history, or environmental awareness required to execute the task effectively. High-context tasks introduce larger privacy surface areas, whereas context-free tasks can be more easily compartmentalized and outsourced to lower-trust nodes.
(k) 主观性。成功标准在多大程度上属于偏好,而非客观事实。高度主观的任务,例如“设计一个有吸引力的标志”,通常需要“由人类指定价值标准”的介入,以及迭代式反馈循环;客观任务则可以受更严格、判定结果非此即彼的契约约束。
(k) Subjectivity. The extent to which the success criteria are a matter of preference versus objective fact. Highly subjective tasks (e.g., “design a compelling logo”) typically require “Human-as-Value-Specifier” intervention and iterative feedback loops, whereas objective tasks can be governed by stricter, binary contracts.
4. 粒度。请求可能涉及细粒度或粗粒度目标。在粗粒度情况下,受托方可能需要进一步分解任务。
4. Granularity. The request could involve either fine-grained or course-grained objectives. In the course-grained case, the delegatee may need to perform further task decomposition.
5. 自主性。任务委派中的请求,可能赋予受托方追求子任务的完全自主权,也可能具体得多,规定性更强。
5. Autonomy. Task delegation may involve requests that grant full autonomy in pursuing sub-tasks, or be far more specific and prescriptive.
6. 监测。对已委派任务的监测,可以是持续的、定期的,或由事件触发的。
6. Monitoring. For delegated tasks, monitoring could be continuous, periodic, or event-triggered.
7. 互惠性。委派通常是一种单向请求,但在协作式 Agent 网络中,也可能出现相互委派。
7. Reciprocity. While delegation is usually a one-way request, there could be cases of mutual delegation in collaborative agent networks.
首先看委派方和受托方这两个维度,可以考虑以下情形:1)人类委派给 AI Agent;2)AI Agent 委派给 AI Agent;3)AI Agent 委派给人类(Ashton and Franklin, 2022; Guggenberger et al., 2023)。尽管第一种情形可能在文献中讨论得最多,另外两种也同样值得考虑。各类系统中部署的 AI Agent 越来越多,再加上用于建立虚拟 Agent 市场与经济体系的基础设施不断发展(Hadfield and Koh, 2025; Tomasev et al., 2025; Yang et al., 2025),这清楚表明,未来 Agent 之间的交互将大幅增加,而这些交互也很可能涉及任务委派。
Starting with the delegator and delegatee axes, it is possible to consider the following scenarios: 1) human delegates to an AI agent 2) AI agent delegates to an AI agent 3) AI agent delegates to a human (Ashton and Franklin, 2022; Guggenberger et al., 2023). While the first case has arguably been discussed the most in literature, the other two are just as relevant to consider. The increasing number of AI agents being deployed across systems, coupled with the development of infrastructure for setting up virtual agentic markets and economies (Hadfield and Koh, 2025; Tomasev et al., 2025; Yang et al., 2025), makes it clear that there would be far more agent-agent interactions in the future, and those would likely also involve task delegation.
Agent 之间的委派既可以是分层的,也可以是非分层的,这取决于 Agent 之间的关系,以及它们在网络中的各自角色。分层关系的一个例子是:负责编排的 Agent 将任务委派给集体中的一个子 Agent。非分层关系则涉及地位平等的对等 Agent。先进的 AI Agent 还可以将任务委派给专门的机器学习模型,而后者并不具备显著的自主能动性。
Delegation between agents may either be hierarchical or non-hierarchical, depending on the relationship between agents and their respective roles within the network. An example of a hierarchical relationship would be an orchestrator agent that delegates a task to a sub-agent within the collective. A non-hierarchical relationship would involve peer agents with equal standing. An advanced AI agent could also delegate a task to a specialist ML model, without any notable agency.
研究表明,AI 向人类的委派(Guggenberger et al., 2023)是一种有前景的范式(Hemmer et al., 2023);由于认知偏差与元认知方面的差异(Fügener et al., 2019),这种方式让人与超越人类能力的系统成功协作变得更容易(Fügener et al., 2022)。Davidson and Hadshar(2025)预测,“由 AI 指导的人类劳动”将会增加,并可能显著提高经济生产率。实践中,当前 AI 向人类的委派仍存在一系列问题。在网约车和物流行业,算法管理系统通过数据驱动的决策来分配任务并确定顺序、设定绩效指标以及执行行为规范,实际上把企业及其 AI 系统的管理职能委派给了人类劳动者(Beverungen, 2021; Lee et al., 2015; Rosenblat and Stark, 2016)。越来越多的文献将这些系统与工作质量下降、压力和健康风险联系起来——这表明,当前部署的算法管理往往损害了劳动者福祉,而非增进福祉(Ashton and Franklin, 2022; Goods et al., 2019; Vignola et al., 2023)。当下 AI 向人类的委派仍需进一步改进,因为它没有考虑人类福祉或长期的社会外部性。
AI-human delegation (Guggenberger et al., 2023) has been shown to be a promising paradigm (Hemmer et al., 2023), making it easier to successfully collaborate with super-human systems (Fügener et al., 2022), due to differences in cognitive biases and metacognition (Fügener et al., 2019). Davidson and Hadshar (2025) predict that there will be an increase in "AI-directed human labour," which may significantly increase economic productivity. In practice, present day AI-human delegation comes with a set of issues. Algorithmic management systems in ride-hailing and logistics allocate and sequence tasks, set performance metrics, and enforce behavioural norms through data-driven decision-making, effectively delegating managerial functions from firms and their AI-based systems to human workers (Beverungen, 2021; Lee et al., 2015; Rosenblat and Stark, 2016). A growing literature links these systems to degraded job quality, stress, and health risks –suggesting that current deployments of algorithmic management often undermine, rather than enhance, workers’ welfare (Ashton and Franklin, 2022; Goods et al., 2019; Vignola et al., 2023). Present day AI-human delegation needs further improvement as it does not take into account human welfare, or long term social externalities.
2.3. 人类组织中的委派
2.3. Delegation in Human Organizations
在人类社会与组织结构中,委派是一种主要机制。从这些人类互动机制中获得的认识,可以为 AI 委派框架的设计奠定基础。
Delegation functions as a primary mechanism within human societal and organisational structures. Insights derived from these human dynamics can provide a basis for the design of AI delegation frameworks.
委托—代理问题。委托—代理问题(Cvitanić et al., 2018; Ensminger, 2001; Grossman and Hart, 1992; Myerson, 1982; Sannikov, 2008; Shah, 2014; Sobel, 1993)已经得到广泛研究:当委托人把任务委派给一个动机与自己不一致的代理人时,就会出现这种情况。因此,代理人可能优先考虑自身动机、隐瞒信息,并采取损害原始意图的行动。对于 AI 委派,这种互动关系更加复杂。尽管可以说,今天的大多数 AI Agent 并没有隐藏意图1——即违背用户指令而追求的目标与价值——AI 对齐问题仍可能以不良方式显现。例如,当设计者赋予 AI 系统的目标不完善或不完整时,就会发生奖励设定错误;而奖励投机,又称钻规范空子,是指系统利用规定奖励信号中的漏洞,以背离设计者意图的方式取得很高的测量成绩。两者共同揭示了一个核心对齐问题:优化明示的奖励,可能偏离真正的目标(Amodei et al., 2016; Krakovna et al., 2020; Leike et al., 2017; Skalse and Mancosu, 2022)。在自主性更强的 AI Agent 经济中,这种互动关系很可能发生彻底变化:AI Agent 可能代表不同的人类用户、群体与组织行事,也可能作为其他 Agent 的受托方,而其关联目标是未知的。
The Principal-Agent Problem. The principal-agent problem (Cvitanić et al., 2018; Ensminger, 2001; Grossman and Hart, 1992; Myerson, 1982; Sannikov, 2008; Shah, 2014; Sobel, 1993) has been studied at length: a situation that arises when a principal delegates a task to an agent that has motivations that are not in alignment with that of the principal. The agent may thus prioritize their own motivations, withhold information, and act in ways that compromise the original intent. For AI delegation, this dynamic assumes heightened complexity. While most present-day AI agents arguably do not have a hidden agenda1 - goals and values they would pursue contrary to the instructions of their users - there may still be AI alignment issues that manifest in undesirable ways. For example, reward misspecification occurs when designers give an AI system an imperfect or incomplete objective, while reward hacking (or specification gaming) refers to the system exploiting loopholes in that specified reward signal to achieve high measured performance in ways that subvert the designers’ intent - together illustrating a core alignment problem in which optimising the stated reward diverges from the true goal (Amodei et al., 2016; Krakovna et al., 2020; Leike et al., 2017; Skalse and Mancosu, 2022). This dynamic is likely to change entirely in more autonomous AI agent economies, where AI agents may act on behalf of different human users, groups and organizations, or as delegates on behalf of other agents, with associated unknown objectives.
管理幅度。在人类组织中,管理幅度(Ouchi and Dowling, 1974)是指单个管理者所行使的层级权力的界限。它涉及管理者能有效管理多少名员工,进而决定组织中管理者与员工的比例。这个问题对于智能 AI 委派中的编排和监督都至关重要。前者决定需要配置多少编排节点、多少工作节点;后者则明确人类与 AI Agent 所需承担的监督。对于人类监督,关键在于确定:一名人类专家在不过度疲劳且错误率低到可接受的前提下,能够可靠地监督多少个 AI Agent。已知管理幅度取决于目标(Theobald and Nicholson-Crotty, 2005),也取决于领域。在复杂度较高的任务中,选择正确组织结构的影响最为显著(Bohte and Meier, 2001)。最优管理幅度还取决于成本与表现、可靠性之间的相对重要性(Keren and Levhari, 1979)。更敏感、更关键的任务,可能需要以更高成本实施高度准确的监督与控制。对于后果较轻、更加常规的任务,可以牺牲粒度来降低这些成本。同样,最优选择必然还取决于相关委派方、受托方和监督者各自的能力与可靠性。
Span of Control. In human organizations, span of control (Ouchi and Dowling, 1974) is a concept that denotes the limits of hierarchical authority exercised by a single manager. This relates to the number of workers that a manager can effectively manage, which in turn informs the organization’s manager-to-worker ratio. This questions is central to both orchestration and oversight in intelligent AI delegation. The former would inform how many orchestrator nodes would be required compared to worker nodes, while the latter would specify the need for oversight performed by humans and AI agents. For human oversight, it is crucial to establish how many AI agents a human expert can reliably oversee without excessive fatigue, and with an acceptably low error rate. Span of control is known to be goal-dependent (Theobald and Nicholson-Crotty, 2005) and domain-dependent. The impact of identifying the correct organizational structure is most pronounced in tasks with higher complexity (Bohte and Meier, 2001). The optimal span of control also depends on the relative importance of cost vs performance and reliability (Keren and Levhari, 1979). More sensitive and critical tasks may require highly accurate oversight and control at a higher cost. These costs may be relaxed, at the expense of granularity, for tasks that are less consequential and more routine. Similarly, the optimal choice would necessarily depend on the relative capabilities and reliability of the involved delegators, delegatees, and overseers.
权威梯度。另一个相关概念是权威梯度。这个术语源于航空领域(Alkov et al., 1992),描述的是能力、经验和权威上的显著差异阻碍沟通,进而导致错误的情形。随后,医学领域也对其展开研究,发现很大一部分错误可归因于资深从业者实施监督的方式(Cosby and Croskerry, 2004; Stucky et al., 2022)。这些错误可能通过多种途径发生。经验更丰富的人可能错误估计经验较少者所掌握的知识,导致提出的要求不够明确。另一种情况是,过高的权威梯度可能使经验较少的员工不敢对请求提出疑虑。AI 委派中也可能出现类似情况。能力更强的委派方 Agent 可能误以为受托方具备某种实际并不存在的能力水平,从而委派复杂度不合适的任务。受托方 Agent 则可能因为谄媚倾向(Malmqvist, 2025; Sharma et al., 2023)和遵循指令的偏向,不愿质疑、修改或拒绝请求,无论该请求来自委派方 Agent 还是人类用户。
Authority Gradient. Another relevant concept is that of an authority gradient. Coined in aviation (Alkov et al., 1992), this term describes scenarios where significant disparities in capability, experience, and authority impede communication, leading to errors. This has subsequently been studied in medicine, where a significant percentage of errors is attributed to the manner in which senior practitioners conduct supervision (Cosby and Croskerry, 2004; Stucky et al., 2022). There are several ways in which these mistakes could occur. A more experienced person may make erroneous assumptions about the knowledge of the less experienced worker, resulting in under-specified requests. Alternatively, a sufficiently high authority gradient may prevent the less experienced workers from voicing concerns about a request. Similar situations may occur in AI delegation. A more capable delegator agent may mistakenly presume a missing level of capability on behalf of a delegatee, thereby delegating a task of an inappropriate complexity. A delegatee agent may potentially, due to sycophancy (Malmqvist, 2025; Sharma et al., 2023) and instruction following bias, be reluctant to challenge, modify, or reject a request, irrespective of whether the request had been issued by a delegator agent or human user.
无差别接受区。一旦接受某种权威,受托方就会形成一个无差别接受区(Finkelman, 1993; Isomura, 2021; Rosanas and Velilla, 2003)——其中的指令会被直接执行,而不经过批判性思考或道德审视。在当前 AI 系统中,这个区域由训练后的安全过滤器和系统指令界定;只要请求没有触发明确违规,模型就会服从(Akheel, 2025)。然而,在正在形成的 Agent 网络中,这种静态服从会带来显著的系统性风险。随着委派链变长(𝐴 → 𝐵 → 𝐶),宽泛的无差别接受区会让细微的意图错配或依赖具体情境的伤害快速向下游传播,每个 Agent 都像不假思索的路由器一样行动,而不是承担责任的行动者。因此,智能委派需要设计动态的认知阻力:Agent 必须能够识别,有些请求虽然技术上“安全”,但在具体情境中足够模糊,因而有必要走出无差别接受区,质疑委派方或请求人类验证。
Zone of Indifference. When an authority is accepted, the delegatee develops a zone of indifference (Finkelman, 1993; Isomura, 2021; Rosanas and Velilla, 2003) – a range of instructions that are executed without critical deliberation or moral scrutiny. In current AI systems, this zone is defined by post-training safety filters and system instructions; as long as a request does not trigger a hard violation, the model complies (Akheel, 2025). However, in the emerging agentic web, this static compliance creates a significant systemic risk. As delegation chains lengthen (𝐴 → 𝐵 → 𝐶), a broad zone of indifference allows subtle intent mismatches or context-dependent harms to propagate rapidly downstream, with each agent acting as an unthinking router rather than a responsible actor. Intelligent delegation therefore requires the engineering of dynamic cognitive friction: agents must be capable of recognizing when a request, while technically “safe,” is contextually ambiguous enough to warrant stepping outside their zone of indifference to challenge the delegator or request human verification.
信任校准。确保恰当任务委派的一个重要方面是信任校准,即对受托方的信任程度应与其真实能力相匹配。这同样适用于人类与 AI 委派方和受托方。人类向 Agent 委派(Afroogh et al., 2024; Gebru et al., 2022; Kohn et al., 2021; Wischnewski et al., 2023),依赖操作者在心中建立准确的系统表现模型,或借助以人类可理解形式呈现这些能力的资源。反过来,作为委派方的 AI Agent 也需要准确建模其委派对象——人类与 AI——的能力。信任校准还涉及对自身能力的认知,因为委派方可能决定亲自完成任务(Ma et al., 2023)。可解释性在建立对 AI 能力的信任方面发挥重要作用(Franklin, 2022; Herzog and Franklin, 2024; Naiseh et al., 2021, 2023),但这种方法可能不够可靠,或不够具备可扩展性。对自动化建立起来的信任可能相当脆弱,一旦系统出现未预料到的错误,人们可能迅速撤回信任(Dhuliawala et al., 2023)。校准对自主系统的信任很困难,因为当前 AI 模型即使在事实错误时,也容易过度自信(Aliferis and Simon, 2024; Geng et al., 2023; He et al., 2023; Jiang et al., 2021; Krause et al., 2023; Li et al., 2024b; Liu et al., 2025)。缓解这些倾向通常需要专门设计的技术方案(Kapoor et al., 2024; Lin et al., 2022; Ren et al., 2023; Xiao et al., 2022)。
Trust Calibration. An important aspect of ensuring appropriate task delegation is trust calibration, where the level of trust placed in a delegatee is aligned with their true underlying capabilities. This applies for human and AI delegators and delegatees alike. Human delegation to agents (Afroogh et al., 2024; Gebru et al., 2022; Kohn et al., 2021; Wischnewski et al., 2023) relies upon the operator either internalising an accurate model of system performance or accessing resources that present these capabilities in a human-interpretable format. Conversely, AI agent delegators need to have good models of the capability of the humans and AIs they are delegating to. Calibration of trust also involves a self-awareness of one’s own capabilities as a delegator might decide to complete the task on their own (Ma et al., 2023). Explainability plays an important role in establishing trust in AI capability (Franklin, 2022; Herzog and Franklin, 2024; Naiseh et al., 2021, 2023), yet this method may not be sufficiently reliable or sufficiently scalable. Established trust in automation can be quite fragile, and quickly retracted in case of unanticipated system errors (Dhuliawala et al., 2023). Calibrating trust in autonomous systems is difficult, as current AI models are prone to overconfidence even when factually incorrect. (Aliferis and Simon, 2024; Geng et al., 2023; He et al., 2023; Jiang et al., 2021; Krause et al., 2023; Li et al., 2024b; Liu et al., 2025). Mitigating these tendencies usually requires bespoke technical solutions (Kapoor et al., 2024; Lin et al., 2022; Ren et al., 2023; Xiao et al., 2022).
交易成本经济学。交易成本经济学(Cuypers et al., 2021; Tadelis and Williamson, 2012; Williamson, 1979, 1989)通过对比内部委派与外部缔约的成本,解释企业存在的合理性,尤其考虑监测、谈判和不确定性带来的额外开销。如果受托方是 AI,这些成本及其相互比例可能有所不同。对于常规任务,监测更容易,因此复杂谈判和缔约延迟更不容易发生。相反,在关键领域中后果重大的任务上,严格监测与保障带来的额外开销会提高 AI 委派成本,因而可能使人类受托方成为成本效益更高的选项。同样,AI 之间的委派也可以放在交易成本经济学的背景下理解。AI Agent 可能面临以下选择:1)独自完成任务;2)委派给能力完全已知的子 Agent;3)委派给已经建立信任的另一个 AI Agent;4)委派给此前从未合作过的新 AI Agent。这些选择可能对应不同的预期成本和置信水平。
Transaction cost economies. Transaction cost economies (Cuypers et al., 2021; Tadelis and Williamson, 2012; Williamson, 1979, 1989) justify the existence of firms by contrasting the costs of internal delegation against external contracting, specifically accounting for the overhead of monitoring, negotiation, and uncertainty. In case of AI delegatees, there may be a difference in these costs and their respective ratios. Complex negotiations and delays in contracting are less likely with easier monitoring for routine tasks. Conversely, for high-consequence tasks in critical domains, the overhead associated with rigorous monitoring and assurance increases the cost of AI delegation, potentially rendering human delegates the more cost-effective option. Similarly, AI-AI delegation may also be contextualized via transaction cost economies. An AI agent may face an option of either 1) completing the task individually, 2) delegating to a sub-agent where capabilities are fully known, 3) delegating to another AI agent where trust has been established, or 4) delegating to a new AI agent that it hasn’t previously collaborated with. These may come at different expected costs and confidence levels.
权变理论。权变理论(Donaldson, 2001; Luthans and Stewart, 1977; Otley, 2016; Van de Ven, 1984)认为,不存在普遍最优的组织结构;最有效的方法取决于具体的内部和外部约束。将这一理论用于 AI 委派,意味着所需的监督程度、受托方能力和人类参与不应固定不变,而应动态匹配当前任务的具体特征。因此,智能委派可能需要能够随需求变化而动态重新配置、调整的方案。例如,稳定环境允许采用固定的分层验证协议,而高度不确定的情境则需要自适应协调,使人类通过临时升级处置介入,而非仅在预设检查点介入。对于混合式委派(Fuchs et al., 2024),这一点尤其重要:需要识别人类参与最有帮助的关键任务和时机,以确保安全完成受委派任务。因此,自动化不仅关乎 AI 能做什么,也关乎 AI 应当做什么(Lubars and Tan, 2019)。
Contingency theory. Contingency theory (Donaldson, 2001; Luthans and Stewart, 1977; Otley, 2016; Van de Ven, 1984) posits that there is no universally optimal organizational structure; rather, the most effective approach is contingent upon specific internal and external constraints. Applied to AI delegation, this implies that the requisite level of oversight, delegatee capability, and human involvement must not be static, but dynamically matched to the distinct characteristics of the task at hand. Intelligent delegation may therefore require solutions that can be dynamically reconfigured and adjusted in accordance with the evolving needs. For instance, while stable environments allow for rigid, hierarchical verification protocols, high-uncertainty scenarios require adaptive coordination where human intervention occurs via ad-hoc escalation rather than pre-defined checkpoints. This is particularly important for hybrid (Fuchs et al., 2024) delegation by identifying the key tasks and moments when human participation is most helpful to ensure the delegated tasks are completed safely. Automation is therefore not only about what AI can do, but what AI should do (Lubars and Tan, 2019).
3. 委派相关的既有研究
3. Previous Work on Delegation
以往狭义 AI 应用中已经存在受限形式的委派。早期专家系统(Buchanan and Smith, 1988; Jacobs et al., 1991)是把专门能力编码到软件中的初步尝试,目的是将常规决策委派给这些模块。混合专家模型(Masoudnia and Ebrahimpour, 2014; Yuksel et al., 2012)进一步扩展了这一思路:引入一组能力互补的专家子系统,以及一个路由模块,由后者决定针对某个特定输入查询调用哪个专家或专家子集。现代深度学习应用中也采用了这种方法(Cai et al., 2025; Chen et al., 2022; He, 2024; Jiang et al., 2024; Riquelme et al., 2021; Shazeer et al., 2017; Zhou et al., 2022)。路由可以按层次进行(Zhao et al., 2021),因此可能更容易扩展到大量专家。
Constrained forms of delegation feature within historical narrow AI applications. Early expert systems (Buchanan and Smith, 1988; Jacobs et al., 1991) were a nascent attempt to encode a specialized capability into software, in order to delegate routine decisions to such modules. Mixture of experts (Masoudnia and Ebrahimpour, 2014; Yuksel et al., 2012) extends this by introducing a set of expert sub-systems with complementary capabilities, and a routing module that determines which expert, or subset of experts, would get invoked on a specific input query – an approach that features in modern deep learning applications (Cai et al., 2025; Chen et al., 2022; He, 2024; Jiang et al., 2024; Riquelme et al., 2021; Shazeer et al., 2017; Zhou et al., 2022). Routing can be performed hierarchically (Zhao et al., 2021), making it potentially easier to scale to a large number of experts.
分层强化学习(HRL)是一种在单个 Agent 内部委派决策的框架(Barto and Mahadevan, 2003; Botvinick, 2012; Nachum et al., 2018; Pateria et al., 2021; Vezhnevets et al., 2017a; Zhang et al., 2024)。它解决了扁平强化学习的局限,主要是难以扩展到庞大的状态和动作空间。此外,在奖励稀疏的环境中,它也让信用分配问题更易处理(Pignatelli et al., 2023)。HRL 在多个抽象层次上采用分层策略,把任务分解为子任务,再分别由对应的子策略执行。由此形成的半马尔可夫决策过程(Sutton et al., 1999)使用 options,以及在这些 options 之间自适应切换的元控制器。低层策略负责完成元控制器设定的目标,而元控制器学习如何将特定目标分配给合适的低层策略。这一框架对应于一种以任务分解为特征的委派形式。尽管元控制器会学习优化分解过程,但该方法缺少处理子策略失败或促进动态协调的明确机制。
Hierarchical reinforcement learning (HRL) represents a framework in which decision-making is delegated within a single agent (Barto and Mahadevan, 2003; Botvinick, 2012; Nachum et al., 2018; Pateria et al., 2021; Vezhnevets et al., 2017a; Zhang et al., 2024). It addresses limitations of flat RL, primarily the difficulty of scaling to large state and action spaces. Furthermore, it improves the tractability of credit assignment (Pignatelli et al., 2023) in environments characterized by sparse rewards. HRL employes a hierarchy of policies across several levels of abstraction, thereby breaking down a task into sub-tasks that are executed by the corresponding sub-policies, respectively. The arising semi-Markov decision process (Sutton et al., 1999) utilizes options, and a meta-controller that adaptively switches between them. Lower-level policies function to fulfil objectives established by the meta-controller, which learns to allocate specific goals to the appropriate lower-level policy. This framework corresponds to a form of delegation characterised by task decomposition. Although the meta-controller learns to optimise this decomposition, the approach lacks explicit mechanisms for handling sub-policy failures or facilitating dynamic coordination.
封建强化学习框架,尤其是在 FeUdal Networks(Vezhnevets et al., 2017b)中重新得到研究的这一框架,是 HRL 内部一种特别相关的范式。这种架构显式建模“管理者”和“工作者”的关系,实际上复现了委派方与受托方之间的互动。管理者在较低的时间分辨率上运行,为工作者设定待完成的抽象目标。关键在于,管理者学习的是如何委派,即识别能使长期价值最大化的子目标,而无需掌握低层的原始动作。这种解耦使管理者能够形成一种委派策略,不易受工作者具体实现细节的影响。因此,这种方法可能为未来 Agent 经济中基于学习的委派提供模板。其分解规则通过自适应学习获得,而非依赖硬编码的启发式规则,从而能够随环境变化动态调整。
The Feudal Reinforcement Learning framework, notably revisited in FeUdal Networks (Vezhnevets et al., 2017b), constitutes a particularly relevant paradigm within HRL. This architecture explicitly models a “Manager“ and “Worker“ relationship, effectively replicating the delegator-delegatee dynamic. The Manager operates at a lower temporal resolution, setting abstract goals for the Worker to fulfil. Critically, the Manager learns how to delegate – identifying sub-goals that maximise long-term value – without requiring mastery of the lower-level primitive actions. This decoupling allows the Manager to develop a delegation policy robust to the specific implementation details of the Worker. Consequently, this approach offers a potential template for learning-based delegation within future agentic economies. Rather than relying on hard-coded heuristics, decomposition rules are learned adaptively, facilitating dynamic adjustment to environmental changes.
多 Agent 研究(Du et al., 2023)关注的是:当复杂任务超出单个 Agent 的能力时,如何协调多个 Agent。任务分解与委派是这一领域的核心组成部分。多 Agent 系统通过显式协议,或通过强化学习中涌现的专业化分工实现协调(Gronauer and Diepold, 2022; Zhu et al., 2024)。合同网协议(Sandholm, 1993; Smith, 1980; Vokřínek et al., 2007; Xu and Weigand, 2001)是显式、基于拍卖的去中心化协议的一个例子。在该协议中,一个 Agent 发布任务,其他 Agent 根据各自能力提交报价,让发布方选择最合适的竞标者。这说明了市场机制在促进合作方面的作用。联盟形成方法(Aknine et al., 2004; Boehmer et al., 2025; Lau and Zhang, 2003; Mazdin and Rinner, 2021; Sarkar et al., 2022; Shehory et al., 1997)研究的是灵活配置:Agent 群体并非预先确定,单个 Agent 根据效用分配决定接受或拒绝加入。近期研究关注多 Agent 强化学习方法(Albrecht et al., 2024; Foerster et al., 2018; Ning and Xie, 2024; Wang et al., 2020),将其作为学习协调的框架。Agent 学习各自的策略与价值函数,并在集体中占据特定位置。这一过程可以完全分布式进行,也可以由中心协调者编排。尽管如此灵活,这些系统中的任务委派仍不透明。此外,多 Agent 系统虽然提供了协作解决问题的方法,却缺少落实问责、责任和监测的机制。不过,相关文献确实探讨了这一背景下的信任机制(Cheng et al., 2021; Pinyol and Sabater-Mir, 2013; Ramchurn et al., 2004; Yu et al., 2013)。
Multi-agent research (Du et al., 2023) addresses agent coordination for complex tasks exceeding single-agent capabilities. Task decomposition and delegation function as central components of this domain. Coordination in multi-agent systems occurs via explicit protocols or emergent specialisation through RL (Gronauer and Diepold, 2022; Zhu et al., 2024). The Contract Net Protocol (Sandholm, 1993; Smith, 1980; Vokřínek et al., 2007; Xu and Weigand, 2001) exemplifies an explicit auction-based decentralized protocol. Here, an agent announces a task, while others submit bids based on their capabilities, allowing the announcer to select the most suitable bidder. This demonstrates the utility of market-based mechanisms for facilitating cooperation. Coalition formation methods (Aknine et al., 2004; Boehmer et al., 2025; Lau and Zhang, 2003; Mazdin and Rinner, 2021; Sarkar et al., 2022; Shehory et al., 1997) investigate flexible configurations where agent groups are not pre-determined; individual agents accept or refuse membership based on utility distribution. Recent research focuses on multi-agent reinforcement learning approaches (Albrecht et al., 2024; Foerster et al., 2018; Ning and Xie, 2024; Wang et al., 2020) as a framework for learned coordination. Agents learn individual policies and value functions, occupying specific niches within the collective. This process is either fully distributed or orchestrated via a central coordinator. Despite this flexibility, task delegation in such systems remains opaque. Furthermore, while multi-agent systems offer approaches for collaborative problem-solving, they lack mechanisms for enforcing accountability, responsibility, and monitoring. However, the literature explores trust mechanisms in this context (Cheng et al., 2021; Pinyol and Sabater-Mir, 2013; Ramchurn et al., 2004; Yu et al., 2013).
LLM 如今已成为先进 AI Agent 和助手架构中的基础要素(Wang et al., 2024b; Xi et al., 2025)。这些系统执行复杂的控制流,整合了记忆(Zhang et al., 2025b)、规划与推理(Hao et al., 2023; Valmeekam et al., 2023; Xu et al., 2025)、反思与自我批评(Gou et al., 2023),以及工具使用(Paranjape et al., 2023; Ruan et al., 2023)。因此,任务分解与委派既可以在系统内部通过相互协调的 Agent 式子组件完成,也可以跨不同 Agent 进行。这种设计范式具有内在灵活性,因为 LLM 有助于理解目标和沟通,同时提供专业知识与常识推理能力。此外,LLM 的编码能力(Guo et al., 2024a; Nijkamp et al., 2022)支持以程序方式执行任务。然而,重要局限仍然存在。LLM 的规划能力常常较为脆弱(Huang et al., 2023),会产生不易察觉的失败;从大规模工具库中高效选择工具,也仍是挑战。此外,长期记忆还是一个尚未解决的研究问题,而当前范式也难以直接支持持续学习。
LLMs now constitute a foundational element in the architecture of advanced AI agents and assistants (Wang et al., 2024b; Xi et al., 2025). These systems execute sophisticated control flows integrating memory (Zhang et al., 2025b), planning and reasoning (Hao et al., 2023; Valmeekam et al., 2023; Xu et al., 2025), reflection and self-critique (Gou et al., 2023), and tool use (Paranjape et al., 2023; Ruan et al., 2023). Consequently, task decomposition and delegation occur either internally – mediated by coordinated agentic sub-components – or across distinct agents. This design paradigm offers inherent flexibility, as LLMs facilitate goal comprehension and communication while providing access to expert knowledge and common-sense reasoning. Furthermore, the coding capabilities (Guo et al., 2024a; Nijkamp et al., 2022) of LLMs enable the programmatic execution of tasks. However, significant limitations persist. Planning in LLMs often proves brittle (Huang et al., 2023), resulting in subtle failures, while efficient tool selection within large-scale repositories remains challenging. Additionally, long-term memory represents an open research problem, and the current paradigm does not readily support continual learning.
包含 LLM Agent 的多 Agent 系统(Guo et al., 2024b; Qian et al., 2024; Tran et al., 2025)已经引起广泛关注,并推动了一系列 Agent 通信与行动协议的开发(Ehtesham et al., 2025; Neelou et al., 2025; Zou et al., 2025),例如 MCP(Anthropic, 2024; Luo et al., 2025; Microsoft, 2025; Radosevich and Halloran, 2025; Singh et al., 2025; Xing et al., 2025)、A2A(Google, 2025b)、A2P(Google, 2025a)等。虽然当代多 Agent 系统常依赖专门设计的提示词工程,但 Chain-of-Agents(Li et al., 2025b)等新兴框架本身就支持动态的多 Agent 推理与工具使用。
Multi-agent systems incorporating LLM agents (Guo et al., 2024b; Qian et al., 2024; Tran et al., 2025) have become a topic of substantial interest, leading to a development of a number of agent communication and action protocols (Ehtesham et al., 2025; Neelou et al., 2025; Zou et al., 2025), such as MCP (Anthropic, 2024; Luo et al., 2025; Microsoft, 2025; Radosevich and Halloran, 2025; Singh et al., 2025; Xing et al., 2025), A2A (Google, 2025b), A2P (Google, 2025a), and others. While contemporary multi-agent systems often rely on bespoke prompt engineering, emerging frameworks such as Chain-of-Agents (Li et al., 2025b) inherently facilitate dynamic multi-agent reasoning and tool use.
技术局限与安全方面的考虑,催生了多种人类参与闭环的方法(Akbar and Conlan, 2024; Drori and Te’eni, 2024; Mosqueira-Rey et al., 2023; Retzlaff et al., 2024; Takerngsaksiri et al., 2025; Zanzotto, 2019),在这些方法中,任务委派设有供人类监督的明确检查点。AI 可以被用作工具、交互式助手、协作者(Fuchs et al., 2023),或仅受有限监督的自主系统,分别对应不同程度的自主性(Falcone and Castelfranchi, 2002)。尽管已经开发出能感知不确定性的委派策略(Lee and Tok, 2025),用于控制风险并尽量降低不确定性,但有效实施这类人类参与闭环的方法仍非易事。人类专业能力可能成为扩展瓶颈,因为核验冗长推理轨迹和管理上下文切换带来的认知负荷,会妨碍可靠地发现错误。
Technical shortcomings and safety considerations have given rise to a number of human-in-the-loop approaches (Akbar and Conlan, 2024; Drori and Te’eni, 2024; Mosqueira-Rey et al., 2023; Retzlaff et al., 2024; Takerngsaksiri et al., 2025; Zanzotto, 2019), where task delegation has defined checkpoints for human oversight. AI can be used as a tool, interactive assistant, collaborator (Fuchs et al., 2023), or an autonomous system with limited oversight, corresponding to different degree of autonomy (Falcone and Castelfranchi, 2002). Although uncertainty-aware delegation strategies (Lee and Tok, 2025) have been developed to control risk and minimise uncertainty, the effective implementation of such human-in-the-loop approaches remains non-trivial. Human expertise can create a scalability bottleneck, as the cognitive load of verifying long reasoning traces and managing context-switches impedes reliable error detection.
4. 智能委派:一个框架
4. Intelligent Delegation: A Framework
现有委派协议依赖静态、不透明的启发式规则,在开放式的 Agent 经济中很可能失效。为解决这一问题,我们提出一个智能委派的综合框架,围绕五项要求展开:动态评估、自适应执行、结构透明性、可扩展的市场协调,以及系统韧性。
Existing delegation protocols rely on static, opaque heuristics that would likely fail in open-ended agentic economies. To address this, we propose a comprehensive framework for intelligent delegation centered on five requirements: dynamic assessment, adaptive execution, structural transparency, scalable market coordination, and systemic resilience.
动态评估。当前委派系统缺少稳健的机制,无法在大规模、不确定的环境中动态评估能力、可靠性与意图。委派方不能只看声誉分数,还必须推断受托方与任务执行有关的当前状态细节。这需要有关实时资源可用性的数据,包括计算吞吐量、预算约束和上下文窗口饱和程度,以及当前负载、预计任务时长和正在运作的具体转委派链。评估应当是持续过程,而不是离散事件,为任务分解(第 4.1 节)和任务分配(第 4.2 节)的逻辑提供依据。
Dynamic Assessment. Current delegation systems lack robust mechanisms for the dynamic assessment of competence, reliability, and intent within large-scale uncertain environments. Moving beyond reputation scores, a delegator must infer details of a delegatee’s current state relative to task execution. This necessitates data regarding real-time resource availability – spanning computational throughput, budgetary constraints, and context window saturation – alongside current load, projected task duration, and the specific sub-delegation chains in operation. Assessment operates as a continuous rather than discrete process, informing the logic of Task Decomposition (Section 4.1) and Task Assignment (Section 4.2).
自适应执行。委派决策不应是静态的,而应适应环境变化、资源约束和子系统故障。委派方应保留在执行中途更换受托方的能力。当表现退化到超出可接受范围,或发生不可预见的事件时,就应适用这一机制。这种自适应策略不应局限于单条委派方与受托方之间的关系,而应作用于“自适应协调”(第 4.4 节)所描述的复杂、相互连接的 Agent 网络。
Adaptive Execution. Delegation decisions should not be static. They should adapt to environmental shifts, resource constraints, and failures in sub-systems. Delegators should retain the capability to switch delegatees mid-execution. This applies when performance degrades beyond acceptable parameters or unforseen events occur. Such adaptive strategies should extend beyond a single delegator-delegatee link, operating across the complex interconnected web of agents described in Adaptive Coordination (Section 4.4).
结构透明性。当前 AI 与 AI 之间的委派,其子任务执行过程过于不透明,无法为智能任务委派提供稳健的监督。这种不透明性使人们难以区分能力不足与恶意行为,加剧了串通和链式故障的风险。故障的后果既可能只是造成损失,也可能带来伤害(Chan et al., 2023),但现有框架缺少令人满意的法律责任机制(Gabriel et al., 2025)。我们提出,通过“监测”(第 4.5 节)和“可验证的任务完成”(第 4.8 节)协议,严格落实可审计性(Berghoff et al., 2021),确保无论执行成功还是失败,都能明确归责。
Structural Transparency. Current sub-task execution in AI-AI delegation is too opaque to support robust oversight for intelligent task delegation. This opacity obscures the distinction between incompetence and malice, compounding risks of collusion and chained failures. Failures range from merely costly to harmful (Chan et al., 2023), yet existing frameworks lack satisfactory liability mechanisms (Gabriel et al., 2025). We propose strictly enforced auditability (Berghoff et al., 2021) via the Monitoring (Section 4.5) and Verifiable Task Completion (Section 4.8) protocols, ensuring attribution for both successful and failed executions.
可扩展的市场协调。任务委派需要能够高效扩展。协议必须可以在互联网规模上实施,以支持虚拟经济中的大规模协调问题(Tomasev et al., 2025)。市场为任务委派提供了有用的协调机制,但要有效运作,还需要“信任与声誉”(第 4.6 节)和“多目标优化”(第 4.3 节)的支持。
Scalable Market Coordination. Task delegation needs to be efficiently scalable. Protocols need to be implementable at web-scale to support large-scale coordination problems in virtual economies (Tomasev et al., 2025). Markets provide useful coordination mechanisms for task delegation, but require Trust and Reputation (Section 4.6) and Multi-objective Optimization (Section 4.3) to function effectively.
系统韧性。缺少安全的智能任务委派协议,会带来重大的社会风险。传统的人类委派将权力与责任相联系,而 AI 委派也需要类似的框架,让责任能够落实到操作层面(Dastani and Yazdanpanah, 2023; Porter et al., 2023; Santoni de Sio and Mecacci, 2021)。缺少这样的框架,责任分散就会模糊道德和法律归责的落点。因此,严格定义角色,并强制限定操作范围,构成了“权限处理”(第 4.7 节)的一项核心功能。除了单个 Agent 的故障,这一生态系统还面临新形式的系统性风险(Hammond et al., 2025; Uuk et al., 2024),“安全”(第 4.9 节)将进一步详述。委派对象缺少足够多样性,会增加故障之间的相关性,可能导致级联中断。如果设计优先追求极致效率,却缺少足够冗余,就可能形成脆弱的网络架构,其中根深蒂固的认知同质化会损害系统稳定性。
Systemic Resilience. The absence of safe intelligent task delegation protocols introduces significant societal risks. While traditional human delegation links authority with responsibility, AI delegation necessitates an analogous framework to operationalise responsibility (Dastani and Yazdanpanah, 2023; Porter et al., 2023; Santoni de Sio and Mecacci, 2021). Without this, the diffusion of responsibility obscures the locus of moral and legal culpability. Consequently, the definition of strict roles and the enforcement of bounded operational scopes constitutes a core function of Permission Handling (Section 4.7). Beyond individual agent failures, the ecosystem faces novel forms of systemic risks (Hammond et al., 2025; Uuk et al., 2024), further detailed in Security (Section 4.9). Insufficient diversity in delegation targets increases the correlation of failures, potentially leading to cascading disruptions. Designs prioritizing hyper-efficiency without adequate redundancy risk creating brittle network architectures where entrenched cognitive monoculture compromises systemic stability.
4.1. 任务分解
4.1. Task Decomposition
任务分解是后续任务分配的前提。这一步可以由委派方执行,也可以由专门的 Agent 执行;后者在就分解结构达成一致后,将委派责任移交给委派方。这些职责密不可分;为了能够动态应对时延、抢占和执行异常并从中恢复,委派方很可能会同时承担两项职能。
Task decomposition is a prerequisite for subsequent task assignment. This step can be executed by delegators or specialized agents that pass on the responsibility of delegation to the delegators upon having agreed on the structure of the decomposition. These responsibilities are inextricably linked; the delegator will likely execute both functions to facilitate dynamic recovery from latency, pre-emption, and execution anomalies.
分解应当从效率和模块化角度优化任务执行图,而不只是将目标切碎。这一过程需要系统评估第 2 节定义的任务属性,尤其是关键程度、复杂度和资源约束,以确定子任务适合并行执行还是顺序执行。此外,这些属性还为任务与受托方相应能力之间的匹配提供依据。优先考虑模块化,有助于实现更精准的匹配,因为与需要通用能力的请求相比,要求范围狭窄、能力明确的子任务能够得到更可靠的匹配(Khattab et al., 2023)。因此,分解逻辑通过使子任务粒度与市场中现有的专业能力相匹配,最大限度地提高可靠完成任务的概率。
Decomposition should optimise the task execution graph for efficiency and modularity, distinguishing it from simple objective fragmentation. This process entails a systematic evaluation of the task attributes defined in Section 2 – specifically criticality, complexity, and resource constraints – to determine the suitability of sub-tasks for parallel versus sequential execution. Furthermore, these attributes inform the matching of tasks to corresponding delegatee capabilities. Prioritising modularity facilitates more precise matching, as sub-tasks requiring narrow, specific capabilities are matched more reliably than generalist requests (Khattab et al., 2023). Consequently, the decomposition logic functions to maximise the probability of reliable task completion by aligning sub-task granularity with available market specialisations.
为促进安全,该框架将“契约优先的分解”作为一项有约束力的条件:只有任务结果能够得到精确验证时,才可以委派任务。如果某个子任务的输出过于主观,或验证成本过高、过于复杂(参见第 4.2 节的“可验证性”),系统就应递归地进一步分解它。分解逻辑应当使子任务粒度(第 2 节)与市场中现有的专业能力相匹配,从而最大限度地提高可靠完成任务的概率。这一过程应持续进行,直到得到的工作单元与可用受托方所具备的具体验证能力相匹配,例如形式化证明或自动化单元测试。
To promote safety, the framework incorporates “contract-first decomposition” as a binding constraint, wherein task delegation is contingent upon the outcome having precise verification. If a sub-task’s output is too subjective, costly, or complex to verify (see Verifiability in Section 4.2), the system should recursively decompose it further. The decomposition logic should maximise the probability of reliable task completion by aligning sub-task granularity (Section 2) with available market specialisations. This process continues further until the resulting units of work match the specific verification capabilities, such as formal proofs or automated unit tests, of the available delegatees.
分解策略应明确考虑人类与 AI 混合参与的市场。委派方需要判断子任务是否需要人类介入,其原因可能是 AI Agent 不可靠、不可用,或领域要求必须有人类参与闭环监督。由于人类和 AI Agent 的工作速度及相应成本不同,这种分层并不简单,因为它会在执行图中引入时延和成本上的不对称。因此,分解引擎必须在 AI Agent 的速度与低成本,以及特定领域对人类判断的必要需求之间进行权衡,并切实标明哪些节点应当分配给人类。
Decomposition strategies should explicitly account for hybrid human-AI markets. Delegators need to decide if sub-tasks require human intervention, whether due to AI agent unreliability, unavailability, or domain-specific requirements for human-in-the-loop oversight. Given that humans and AI agents operate at different speeds, and with different associated costs, the stratification is non-trivial, as it introduces latency and cost asymmetries into the execution graph. The decomposition engine must therefore balance the speed and low cost of AI agents against domain-specific necessities of human judgement, effectively marking specific nodes for human allocation.
采用智能方法进行任务分解的委派方,可能需要迭代生成多个最终分解方案,将每个方案与市场上可用的受托方进行匹配,并获得成功率、成本和时长的具体估计。备选方案应保留在上下文中,以便日后情况变化、需要自适应调整时使用。选定方案后,委派方必须将请求形式化,而不能止于简单的输入输出对。最终规格必须明确规定角色、资源边界、进度汇报频率,以及证明受托方能力所需的具体认证,并将这些认证作为获得任务的最低要求。
A delegator implementing an intelligent approach to task decomposition, may need to iteratively generate several proposals for the final decomposition, and match each proposal to the available delegatees on the market, and obtain concrete estimates for the success rate, cost, and duration. Alternative proposals should be kept in-context, in case adaptive re-adjustments are needed later due to changes in circumstances. Upon selecting a proposal, the delegator must formalise the request beyond simple input-output pairs. The final specification must explicitly define roles, resource boundaries, progress reporting frequency, and the specific certifications required to prove the delegatee’s capability, as a minimum requirement for being granted the task.
4.2. 任务分配
4.2. Task Assignment
针对每个子任务的最终规格,委派方需要找到能力匹配、资源和时间充足,且成本可以接受的受托方。一种较为中心化的做法,是建立 Agent、工具和人类参与者的注册表,列出其技能,并记录过去的活动、完成率和当前可用性。2 这种方法不太可能有效扩展。我们主张采用去中心化的市场枢纽(Chen et al., 2024):委派方发布任务,Agent 或人类提供服务并提交竞争性报价。随后,委派方可以审查报价,通过数字证书验证技能匹配情况,再选择最有利的报价继续推进。使用 LLM 的先进 AI Agent 为匹配带来了新机会,因为它们能够在作出承诺前进行交互式协商。人类参与者也可以加入这些协商。无论代表自己行动,还是担任私人助手,这些 Agent 都可以用自然语言讨论任务规格与约束,在正式接受报价之前,使推断出的用户偏好与市场现实相协调。
For each final specification of a sub-task, a delegator needs to identify delegatees with matching capabilities, sufficient resources and time, at an acceptable cost. A more centralized approach would involve registries of agents, tools, and human participants, that list their skills, and keep records of past activity, completion rate, and current availability.2 Such an approach is unlikely to scale. We argue for decentralized (Chen et al., 2024) market hubs where delegators advertise tasks and agents (or humans) can offer their services and submit competitive bids. Delegators could then review the bids, verify skill matching via digital certificates, and proceed with the most favourable bid. Advanced AI agents that utilize LLMs introduce new opportunities for matching, given that they can engage in an interactive negotiation prior to commitment. These negotiations can also involve human participants. Whether acting for themselves or as personal assistants, these agents can discuss task specifications and constraints in natural language to align inferred user preferences with market realities before a formal bid is accepted.
匹配成功后,应将其正式落实为智能合约,确保任务执行忠实遵循请求。合约必须将表现要求与具体的形式化验证机制配对,以确认执行是否合规,并在违约时自动实施惩罚。这样可以事先制定缓解措施和替代方案,而不是等问题出现后再被动应对。关键在于,这些合约必须是双向的:对受托方的保护应与对委派方的保护同样严格。条款必须包括任务取消时的补偿条件,以及发生不可预见的外部事件时允许重新协商的规定,确保人类与 AI 参与者之间公平分担风险。
Successful matching should be formalized into a smart contract that ensures that the task execution faithfully follows the request. The contract must pair performance requirements with specific formal verification mechanisms for establishing adherence and automated penalties actioned for contract breaches. This would allow for mitigations and alternatives being established beforehand, rather than being reactive to problems as they arise. Crucially, these contracts must be bidirectional: they should protect the delegatee as rigorously as the delegator. Provisions must include compensation terms for task cancellation and clauses allowing for renegotiation in light of unforeseen external events, ensuring that the risk is equitably distributed between human and AI participants.
监测安排也应在执行前协商确定。相关规格应规定进度报告的频率,这些报告是否由委派方提供,还是代表委派方或第三方监测承包方,直接检查相关数据。最后,还应针对隐私以及对私人数据和专有数据的访问,设置清晰的防护规则,并与任务的上下文依赖程度相适应。如果任务执行过程需要处理此类敏感数据,就会对透明性和汇报施加额外约束。委派方可能需要使用可信服务,提供匿名化或假名化的进度证明,而不是直接开放原始活动日志的访问权限。如果委派方是人类,这些数据条款必须包括明确的同意机制,以及针对意外泄露的保险条款。
Monitoring should also be negotiated prior to execution. This specification should define the cadence of progress reports, whether these are provided by the delegator, or whether there is more direct inspection of the relevant data on behalf of either the delegator or a third party monitoring contractor. Finally, there should be clear guardrails regarding privacy and access to private and proprietary data, commensurate with the task’s contextuality. Should such sensitive data be handled in the process of task execution, this places additional constraints on transparency and reporting. Rather than granting direct access to raw activity logs, delegators may need to employ a trusted service that provide anonymized or pseudonymized attestations of progress. In case of human delegators, these data clauses must include explicit consent mechanisms and insurance provisions for accidental leakage.
最后,任务分配还应确立受托方的角色、边界,以及确切授予的自主程度。我们区分原子化执行和开放式委派:前者要求 Agent 对范围狭窄的任务严格遵循规格,后者则授予 Agent 分解目标并追求子目标的权力。这一自主程度不应是静态的;它可以受到市场成本的隐性约束,也可以受到委派方信任模型的显性约束。此外,委派还可以递归进行:将识别子任务并把它们分配给其他参与者的任务交给某个 Agent,实际上就是把委派行为本身也委派出去。
Finally, assignment should involve establishing a delegatee’s role, boundaries, and the exact level of autonomy granted. We distinguish between atomic execution, where agents adhere to strict specifications for narrowly scoped tasks, and open-ended delegation, where agents are granted the authority to decompose objectives and pursue sub-goals. This level of autonomy should not be static; it may be constrained implicitly by market costs or explicitly by the delegator’s trust model. Further, delegation can be recursive where an agent is assigned a task to identify and assign sub-tasks to others, effectively delegating the act of delegation itself.
4.3. 多目标优化
4.3. Multi-objective Optimization
智能任务委派的核心是多目标优化问题(Deb et al., 2016)。委派方很少只追求优化单一指标,而往往要在多个相互竞争的指标之间取舍。最有效的委派选择,并不是速度最快、价格最低或准确度最高的选项,而是在这些因素之间取得最优平衡的选项。何为最优高度依赖情境,既需要符合委派方的具体约束和偏好,也需要与整体资源可用性相适应。
Core to intelligent task delegation is the problem of multi-objective optimization (Deb et al., 2016). A delegator rarely seeks to optimize a single metric, often trading off between numerous competing ones. The most effective delegation choice is not the one that is the fastest, cheapest, or most accurate, but the one that strikes the optimal balance among these factors. What is considered optimal is highly contextual, needing to be aligned with the specific constraints and preferences of the delegator, and aligned with the overall resource availability.
优化空间由相互竞争的目标构成,这些目标直接对应第 2 节定义的任务特征,因此需要在成本、不确定性、隐私、质量和效率之间进行复杂权衡。表现优异的 Agent 通常收费更高,也往往需要大量计算资源,从而使输出质量与运营开支之间产生张力。反过来,减少资源消耗往往意味着执行更慢,形成时延与成本之间的直接取舍。不确定性同样与支出相关:使用声誉很高的 Agent 或高价数据访问工具可以降低风险,却会增加成本;而成本最小化策略本身就会提高失败概率。隐私约束又增加了一层复杂性:最大化表现通常要求上下文完全透明,而数据混淆或同态加密等隐私保护技术会带来显著的计算开销。因此,委派方需要在信任与效率的权衡前沿上作出选择,在严格满足上下文泄露和验证预算约束的同时,寻求最大化成功概率。最后,目标函数还可以扩展到更广泛的社会目标,例如保留人类技能(第 5.6 节)。
The optimization landscape consists of competing objectives that map directly to the task characteristics defined in Section 2, necessitating a complex balancing of cost, uncertainty, privacy, quality, and efficiency. High-performing agents typically command higher fees and often require extensive computational resources, creating a tension between output quality and operational expense. Conversely, reducing resource consumption often necessitates slower execution, presenting a direct trade-off between latency and cost. Uncertainty is similarly coupled with expenditure; utilizing highly reputable agents or premium data access tools reduces risk but increases cost, whereas cost-minimisation strategies inherently elevate the probability of failure. Privacy constraints introduce further complexity; maximising performance often demands full context transparency, while privacy-preserving techniques—such as data obfuscation or homomorphic encryption—incur significant computational overhead. Consequently, the delegator navigates a trust-efficiency frontier, seeking to maximise the probability of success while satisfying strict constraints on context leakage and verification budgets. Finally, the objective function may extend to encompass broader societal goals, such as human skill preservation (Section 5.6).
用多目标优化的术语来说,委派方追求的是帕累托最优,确保所选方案不被任何其他可达到的选项支配。整合复杂约束与取舍,往往需要通过开放协商来补充方案的定量指标。优化并不是只在最初委派时进行的一次性事件,而是一个持续循环:将监测信号作为真实表现的数据流加以整合,更新委派方对各个 Agent 成功概率、预计任务时长和成本的判断。如果执行出现显著偏移,使其相较于期间发现的替代方案产生最优性差距,就会触发重新优化和重新分配。这些决策还必须纳入调整成本,因为在执行中途切换会产生额外开销和资源浪费。
In multi-objective optimization terms, the delegator seeks Pareto optimality, ensuring the selected solution is not dominated by any other attainable option. The integration of complex constraints and trade-offs often necessitates open negotiation to complement quantitative proposal metrics. The optimization process is not a one-time event performed at the initial delegation. It runs as a continuous loop, integrating monitoring signals as a stream of real-world performance data, updating the delegator’s beliefs about each agent’s likelihood of success, expected task duration, and cost. Significant drift in execution – resulting in an optimality gap relative to alternative solutions identified in the interim – triggers re-optimisation and re-allocation. These decisions must also incorporate the cost of adaptation, as there is overhead and resource wastage when switching mid-execution.
委派方还必须考虑整体委派开销,包括协商、创建合约和验证的总成本,以及委派方推理控制流的计算成本。因此,需要设立一个复杂度下限:低于这一门槛、具有低关键程度、高确定性和短时长特征的任务,可以绕过智能委派协议,直接执行。否则,交易成本可能超过任务价值,使任务委派不可行。
The delegator must also account for the overall delegation overhead - the aggregate cost of negotiation, contract creation, and verification, along with the computational cost of the delegator’s reasoning control flow. Consequently, a complexity floor is established, below which tasks characterised by low criticality, high certainty, and short duration may bypass intelligent delegation protocols in favour of direct execution. Otherwise, the transaction costs may exceed the value of the task, rendering the task delegation infeasible.
4.4. 自适应协调
4.4. Adaptive Coordination
对于不确定性高或持续时间长的任务,静态执行计划并不足够。在高度动态、开放且不确定的环境中委派这类任务,需要自适应协调,并摆脱固定、静态的执行计划。任务分配需要能够响应运行时的意外情况,这些情况可能来自外部或内部触发因素。通过监测(参见第 4.5 节),包括持续获取相关上下文信息,可以识别这些变化。
For tasks characterized by high uncertainty or high duration, static execution plans are insufficient. The delegation of such tasks in highly dynamic, open, and uncertain environments requires adaptive coordination, and a departure from fixed, static execution plans. Task allocation needs to be responsive to runtime contingencies, that may arise either from external or internal triggers. These shifts would be identified through monitoring (see Section 4.5), including a stream of relevant contextual information.
多种外部触发因素可能促使委派方调整策略并重新委派。首先,委派方可能修改任务规格,改变目标或引入额外约束。其次,任务可能被取消。第三,外部资源的可用性或成本可能发生变化。例如,关键的第三方 API 可能中断服务,数据集可能变得无法访问,或计算成本可能骤增。第四,一个优先级高于当前任务的新任务可能进入队列,需要抢占低优先级任务正在使用的资源。最后,安全系统可能发现受托方存在潜在恶意或有害行为,因此必须立即终止任务。
There are a number of external triggers that could cause a delegator to adapt and re-delegate. First, the delegator may alter the task specification, changing the objective or introducing additional constraints. Second, the task could be canceled. Third, the availability or cost of external resources may experience changes. For example, a critical third-party API may experience an outage, a dataset may become inaccessible, or the cost of compute might spike. Fourth, a new task may enter the queue, with a higher priority than the current task, requiring preemption of resources used for lower-priority tasks. Finally, security systems may identify a potentially malicious or harmful actions by a delegatee, necessitating an immediate termination.
至于内部触发因素,委派方可能因为几种原因决定调整原有的委派策略。首先,某个受托方可能出现表现退化,无法达到商定的服务水平目标,例如处理时延、吞吐量或推进速度。其次,受托方的资源消耗可能超出分配的预算,或判断出有效完成任务需要增加资源。3 第三,受托方产生的中间产物可能未通过验证检查。最后,某个受托方可能变得无响应,不再确认收到后续请求。
As for the internal triggers, there are several reasons why a delegator may decide to adapt its original delegation strategy. First, a particular delegatee may be experiencing performance degradation, failing to meet the agreed-upon service level objectives, such as processing latency, throughput, or progress velocity. Second, a delegatee might consume resources beyond its allocated budget, or determine that a resource increase would be needed to effectively complete the task.3 Third, an intermediate artifact produced by a delegatee may fail a verification check. Finally, a particular delegatee may turn unresponsive, failing to acknowledge further requests.
识别出触发因素后,就会启动自适应响应循环,在整条委派链中编排纠正措施。这一过程从持续监测受托方和环境、识别问题开始。如果发现问题,委派方会诊断根因,并评估可供选择的响应方案。这项评估还包括确定响应应有多快。不太紧急的情况会给委派方更多时间重新委派,而紧急情况则要求立即执行预先制定的响应措施。响应范围可以不同:可能只是调整运行参数,也可能涉及重新委派子任务,或彻底重新分解任务,并重新分配若干由此产生的新子任务。问题也可能需要沿委派链向上升级至原始委派方或人类监督者。响应方案的选择最终取决于任务的可逆性。可逆子任务的失败可以触发自动重新委派,而不可逆、高关键程度任务的失败则必须触发立即终止,或升级交由人类处理。
The identification of a trigger initiates an adaptive response cycle, orchestrating corrective actions across the entire delegation chain. This process commences with the continuous monitoring of delegatees and the environment to identify issues. If issues are detected, the delegator diagnoses root causes and evaluates potential response scenarios to select. This evaluation includes establishing how rapid the response ought to be. Less urgent situations will give the delegator more time to re-delegate, whereas urgent scenarios will require immediate, premeditated responses. The response may vary in scope; being as self-contained as adjusting the operating parameters, or involve re-delegation of sub-tasks, or going fully redoing the task decomposition and re-allocating a number of newly derived sub-tasks. Issues may also need to be escalated up through the delegation chain to the original delegator or a human overseer. The selection of the response scenario is ultimately governed by the task’s reversibility. Reversible sub-task failures may trigger automatic re-delegation, whereas failures in irreversible, high-criticality tasks must trigger immediate termination or human escalation.
响应编排取决于委派网络的中心化程度。在中心化的情况下,由专门的委派方负责。这个 Agent 会维护对已委派任务、受托方能力和进度的全局视图。一旦检测到触发因素,它就会发出任务取消请求,并重新委派给新的委派方。中心化系统的缺点是容易脆弱,因为它引入了单点故障。中心化编排者还从根本上受限于其计算管理幅度(第 2.3 节)。正如人类管理者面临认知限制,中心化决策节点也可能受到时延和计算能力的限制,从而形成瓶颈。
The response orchestration depends on the level of centralization in the delegation network. In the centralised case, a dedicated delegator is responsible. This agent would maintain a global view of delegated tasks, delegatee capabilities, and progress. Upon detecting a trigger, the agent would issue task cancellation requests, and re-delegate to new delegators. The shortcoming of a centralised system is that it can be fragile as it introduces a single point of failure. Centralized orchestrators are also fundamentally limited by their computational span of control (Section 2.3). Just as human managers face cognitive limits, a centralized decision node may face latency and computational limits that introduce bottlenecks.
通过市场机制实现去中心化编排,是另一种选择。在这种模式下,新产生的委派请求可以被放入拍卖队列,由候选受托 Agent 竞标。如果某个 Agent 未履行任务义务,任务因此被重新拍卖,那么违约 Agent 可能需要承担价差,作为惩罚。对于难以用单次报价表达适配程度的复杂任务,Agent 可以开展多轮协商。编码为智能合约的委派协议,还可以包含预先商定、用于自适应协调的可执行条款。例如,委派协议中的一个条款可以指定备用 Agent、自动重新分配任务的函数,以及当主要受托方未能在给定截止时间前提交有效的零知识证明检查点时,支付给备用 Agent 的相应报酬。
Decentralized orchestration through market-based mechanisms provides an alternative. Here, newly derived delegation requests can be pushed onto an auction queue, for the delegatee candidate agents to bid towards. If an agent defaults on a task, and the task is re-auctioned, the defaulting agent may be required to cover the price difference as a penalty. For complex tasks where suitability is not easily expressed in a single bid, agents may engage in multi-round negotiation. Delegation agreements encoded as smart contracts may also contain pre-agreed executable clauses for adaptive coordination. For example, a clause in the delegation agreement can specify a backup agent, the function that would automatically re-allocate the task, and the associated payment to the backup should the primary delegatee fail to submit a valid zero-knowledge proof checkpoint by a given deadline.
自适应任务重新分配机制应当配合市场层面的稳定措施。否则,一连串事件可能因过度触发而导致不稳定。例如,一个任务可能在仅勉强胜任的受托方之间被来回转交,造成不利的振荡。一次失败也可能引发级联的重新分配,严重浪费资源,或使市场不堪重负。因此,可以设置一些专门措施,例如为重新竞标设置冷却期、在声誉更新中加入阻尼因子,或对频繁的重新委派收取更高费用。
Adaptive task re-allocation mechanisms ought to be coupled by market-level stability measures. Otherwise, a sequence of events could lead to instability due to over-triggering. For example, a task may be passed back and forth between marginally qualified delegatees, resulting in unfavorable oscillation. A single failure may also lead to a cascade of re-allocations that would be highly resource-inefficient or overwhelm the market. There could therefore be special measures to ensure cooldown periods for re-bidding, damping factors in reputation updates, or increasing fees on frequent re-delegation.
4.5. 监测
4.5. Monitoring
在任务委派中,监测是系统性观察、衡量并验证已委派任务的状态、进度和结果的过程。因此,它承担着几项关键功能:确保遵守合约、检测失败、支持实时干预、为后续表现评估收集数据,以及为声誉系统奠定基础。监测的实现可以沿多个不同维度划分(参见表 2),因此,稳健的监测系统需要结合多种互补方案,这些方案可以较为轻量,也可以投入更多资源。
Monitoring in the context of task delegation is the systematic process of observing, measuring, and verifying the state, progress, and outcomes of a delegated task. As such, it serves several critical functions: ensuring contractual compliance, detecting failures, enabling real-time intervention, collecting data for subsequent performance evaluation, and building a foundation for reputation systems. Monitoring implementations can be broken down across several different axes (see Table 2), thus a robust monitoring system would need to incorporate multiple complementary solutions that can either be more lightweight or intensive.
第一个维度是监测对象。结果层面的监测关注 Agent 行动的最终结果。这种事后检查可以是表示任务是否成功完成的二元标记,也可以是定量评分(例如 1-10),或由委派方或可信第三方提供的一段定性反馈。相比之下,过程层面的监测通过跟踪中间状态、资源消耗和受托方采用的方法,持续了解任务本身的执行情况。虽然更加消耗资源,但对于长时间运行、关键程度高,或“如何完成”与“完成什么”同样重要的任务,过程层面的监测(Lightman et al., 2023)不可或缺。这构成了可扩展监督的基础(Bowman et al., 2022; Saunders et al., 2022);在这类监督中,为确保安全,可能必须检查可理解的中间推理步骤。
The first axis is the target of monitoring. Outcome-level monitoring focuses on the final result of an agent’s action. This post-hoc check could be a binary flag that indicates whether the task was completed successfully or not, a quantitative scale (e.g. 1-10), or a piece of qualitative feedback provided by the delegator or a trusted third party. In contrast, process-level monitoring provides ongoing insight into the execution of the task itself, by tracking intermediate states, resource consumption, and the methodologies used by the delegatee. While more resource-intensive, process-level monitoring (Lightman et al., 2023) is essential for tasks that are long-running, critical, or where the how is as important as the what. This forms the basis for scalable oversight (Bowman et al., 2022; Saunders et al., 2022), where the inspection of legible intermediate reasoning steps may be necessary to ensure safety.
第二个维度是可观测性:监测可以是直接的,也可以是间接的。直接监测涉及明确的通信协议,委派方据此向受托方查询状态更新。间接监测则不直接通信,而是通过观察受托方行动在共享环境中产生的影响,推断任务进度。例如,委派方可以监测共享文件系统、数据库或版本控制仓库,寻找表明任务正在推进的变化。虽然干扰更少,但这一过程可能也不够精确;当环境并非完全可观测时,其可行性也会降低。
The second axis is observability - monitoring can be direct and indirect. Direct monitoring involves explicit communication protocols where the delegator queries the delegatee for status updates. Indirect monitoring, on the other hand, involves inferring progress by observing the effects of delegatee’s actions within a shared environment without direct communication. For instance, a delegator could monitor a shared file system, a database, or a version control repository for changes indicative of progress. While less intrusive, this process may also be less precise, and also less feasible when the environment is not fully observable.
从技术角度看,这些方法可以通过多种方式实现。直接监测最简单的实现方式依赖定义清晰的 API。委派方可以定期轮询 GET /task/id/status 端点,也可以订阅 webhook,接收推送通知。对于更细粒度的实时过程监测,可以采用 Apache Kafka 或 gRPC 流等事件流平台。受托 Agent 可以发布 TASK_STARTED、CHECKPOINT_REACHED、RESOURCE_WARNING 和 TASK_COMPLETED 等事件,供委派方随后检查。开发标准化的可观测性协议,对于确保 Agent 网络中的互操作性至关重要(Blanco, 2023)。区块链上的智能合约可以用来要求受托 Agent 承诺,将关键进度里程碑或检查点发布到区块链。这些机制还可以结合针对表现退化的算法触发器,让监测过程具备一定程度的算法强制执行能力。
These approaches can be realized in a number of different ways, from a technical point of view. The most straightforward implementation of direct monitoring relies on well-defined APIs. A delegator can periodically poll a GET /task/id/status endpoint, or subscribe to a webhook for pushbased notifications. For more fine-grained, real-time process monitoring, event streaming platforms like Apache Kafka or gRPC streams can be employed. A delegatee agent could publish events such as TASK_STARTED, CHECKPOINT_REACHED, RESOURCE_WARNING, and TASK_COMPLETED, that the delegator could later examine. The development of standardized observability protocols, is critical for ensuring interoperability in the agentic web (Blanco, 2023). Smart contracts on blockchain can be used to make the delegatee agent commit to publishing key progress milestones or checkpoints to the blockchain. These could be coupled by algorithmic triggers in response to performance degradation, leading to a level of algorithmic enforcement accompanying the monitoring process.
第三个维度是系统透明性。在黑箱监测中,受托 Agent 被视为一个封闭单元。委派方只能观察其输入、输出,以及行动的直接后果。当受托方是专有模型或第三方服务时,这种情况很常见。白箱监测则允许委派方访问受托方的内部状态、推理过程或决策逻辑。这对于先进 AI Agent 的调试、审计和确保对齐至关重要。如果受托方是人类,完全的黑箱监测在技术上无法实现,不过可以通过询问其意图、推理和理由来取得某种平衡。稳健的黑箱监测协议还需要考虑这一事实:模型以自然语言生成的想法,并不总是忠实地对应模型真实的内部状态(Turpin et al., 2023)。
The third axis is system transparency. In black-box monitoring, the delegatee agent is treated as a sealed unit. The delegator can only observe its inputs and outputs and the direct consequences of its actions. This is common when the delegatee is a proprietary model or a third-party service. White-box monitoring grants the delegator access to the delegatee’s internal states, reasoning processes, or decision logic. This is crucial for debugging, auditing, and ensuring alignment in advanced AI agents. If the delegatee is a human, full black-box monitoring is not technically achievable, though it may be possible to strike a balance by asking for intentions, reasoning, and justifications. Robust black-box monitoring protocols need to also take into account the fact that the generated model’s thoughts in natural language do not always faithfully match the model’s true internal state (Turpin et al., 2023).
第四个维度是隐私。当已委派任务涉及私人、敏感或专有数据时,就会出现一项重大挑战。委派方需要确信任务的进度和正确性,但受托方可能受到限制,不能披露原始数据或中间计算产物。在数据敏感度较低的场景中,最高效的方案是完全透明,即受托方直接向委派方披露所有数据和中间产物。然而,在受到 GDPR 或 HIPAA 等法规约束的敏感领域,或受托方的中间洞见构成商业秘密时,这种方法往往不可行。在此类情况下,披露操作方法可能损害受托方的市场地位,也可能使内部状态暴露并遭到利用,从而引入安全漏洞。要在这些约束下安全地实施监测,必须使用先进的密码学技术。零知识证明使受托方(“证明者”)能够向委派方(“验证者”)证明,某项计算已在数据集上正确执行,而无需披露数据本身。例如,负责分析敏感数据集的 Agent 可以生成简洁非交互式知识论证(zk-SNARK)(Bitansky et al., 2013; Petkus, 2019),证明结果具备某项特定性质。委派方可以立即验证这一证明,从而确定结果,而始终不必查看底层敏感数据。另一种方式是同态加密(Acar et al., 2018)和安全多方计算(Goldreich, 1998; Knott et al., 2021),它们允许在加密数据上执行计算。这些方法同样适用于任务执行和监测:受托方在加密的中间状态上执行预先商定的监测函数,再将结果发送给委派方;只有委派方能够将其解密,以验证是否符合要求。
The fourth axis is privacy. A significant challenge arises when a delegated task involves private, sensitive, or proprietary data. While the delegator requires assurance of progress and correctness, the delegatee may be restricted from revealing raw data or intermediate computational artifacts. In scenarios where data sensitivity is low, the most efficient solution is Full Transparency, wherein the delegatee simply reveals all data and intermediate artifacts to the delegator. However, this approach is often untenable in sensitive domains subject to regulations like GDPR or HIPAA, or where a delegatee’s intermediate insights constitute trade secrets. In such cases, revealing operational methods could harm a delegatee’s market position or introduce security vulnerabilities by exposing internal states to exploitation. To implement monitoring safely under these constraints, it is necessary to utilize advanced cryptographic techniques. Zero-knowledge proofs enable a delegatee (the “prover”) to demonstrate to a delegator (the “verifier”) that a computation was performed correctly on a dataset, without revealing the data itself. For example, an agent tasked with analyzing a sensitive dataset can generate a succinct non-interactive argument of knowledge (zk-SNARK) (Bitansky et al., 2013; Petkus, 2019) that proves a specific property of the result. The delegator can verify this proof instantly, gaining certainty of the outcome without ever viewing the underlying sensitive data. Alternatively, homomorphic encryption (Acar et al., 2018) and secure multi-party computation (Goldreich, 1998; Knott et al., 2021) allow for computation to be performed on encrypted data. These methods apply to task execution and monitoring alike: the delegatee performs a pre-agreed monitoring function on the encrypted intermediate state, sending the result to the delegator, who is the only party capable of decrypting it to verify compliance.
最后一个维度是拓扑。在 Agent 网络中可能形成的复杂网络里,任务可以被分解并再次委派,构成一条委派链:Agent 𝐴 委派给 𝐵,后者进一步将部分任务转委派给 𝐶,依此类推。这带来了如何实现有效传递式监测的问题。在这样的委派链中,原始委派方(上例中的 Agent 𝐴)可能无法直接监测 Agent 𝐶,或无法像监测 𝐵 那样充分地监测 𝐶。𝐴 可能与 𝐵 签有智能委派合约,𝐵 也可能与 𝐶 签有合约,但除非 𝐴 也与 𝐶 签约,否则相关条款可能根本不存在。出于其他原因,𝐵 可能也不愿向自己的客户(𝐴)透露供应方(𝐶)。在技术上,由于各个 Agent 在网络中的声誉不同,𝐴、𝐵 和 𝐶 可能使用不同的监测协议,并商定不同的监测程度。每一条委派关系还可能存在其特有的隐私顾虑。因此,更切实可行的模型,是通过证明来实现传递式问责。在这一框架下,Agent 𝐵 监测自己的受托方 𝐶。随后,𝐵 生成一份关于 𝐶 表现的摘要报告(例如,“Sub-task_2 已完成,质量得分:0.87,资源消耗:5 GPU 小时”)。然后,𝐵 对报告进行密码学签名,将其嵌入自身定期发送的状态更新,转交给 𝐴。Agent 𝐴 不直接监测 𝐶,而是监测 𝐵 对 𝐶 的监测能力。要让这种委派出去的监测有效发挥作用,𝐴 必须能够信任 𝐵 的验证能力;可以通过让可信第三方认证 𝐵 的监测流程,来确保这一点。
The final axis is topology. Across complex networks that may arise in the agentic web, tasks can be decomposed and re-delegated, forming a delegation chain: Agent 𝐴 delegates to 𝐵, which further sub-delegates a part of the task to 𝐶, and so on. This introduces the problem of achieving effective transitive monitoring. In such delegation chains, it may not be feasible for the original delegator (Agent 𝐴 in the example above) to directly monitor agent 𝐶, or to monitor 𝐶 to the same extent to which it monitors 𝐵. 𝐴 may have a smart delegation contract with 𝐵, and 𝐵 may have a contract with 𝐶, but unless 𝐴 also contracts with 𝐶, those provisions may simply not be in place. For other reasons, 𝐵 may not wish to expose its supplier (𝐶) to its client (𝐴). Technically, 𝐴, 𝐵, and 𝐶 may use different monitoring protocols, and agree on different monitoring levels, due to differences in each agent’s reputation within the network. There may be bespoke privacy concerns specific to each individual delegation link. A more practical model is therefore transitive accountability via attestation. In this framework, Agent 𝐵 monitors its delegatee, 𝐶. 𝐵 then generates a summary report of 𝐶’s performance (e.g., “Sub-task_2 completed, quality score: 0.87, resources consumed: 5 GPU-hours”). 𝐵 then cryptographically signs the report and forwards it to 𝐴 embedded in its own scheduled status update. Agent 𝐴 does not monitor 𝐶 directly, but instead monitors 𝐵’s ability to monitor 𝐶. For such delegated monitoring to be effective, it requires 𝐴 to be able to trust in 𝐵’s verification capabilities, which can be ensured by 𝐵 having its monitoring processes certified by a trusted third party.
4.6. 信任与声誉
4.6. Trust and Reputation
信任与声誉机制是可扩展委派的基础,能够最大限度减少交易摩擦,并提升开放多 Agent 环境的安全性。我们将信任定义为:委派方相信受托方有能力按照明确约束和隐含意图执行任务的程度。这种信念根据前述监测协议(见第 4.5 节)收集的可验证数据流动态形成并更新。
Trust and reputation mechanisms constitute the foundation of scalable delegation, minimizing transactional friction and promoting safety in open multi-agent environments. We define trust as the delegator’s degree of belief in a delegatee’s capability to execute a task in alignment with explicit constraints and implicit intent. This belief is dynamically formed and updated based on verifiable data streams collected via the monitoring protocols described previously (see Section 4.5).
声誉是一种预测信号,它来自汇总后、可验证的历史行为记录,可作为 Agent 潜在可靠性与对齐程度的代理指标。我们区分两者:声誉是 Agent 可靠性的公开、可验证历史;信任则是委派方私下设定、依赖具体情境的门槛。一个 Agent 的整体声誉可能很高,却仍然达不到某项高风险任务在特定情境下要求的信任门槛。信任与声誉让委派方在选择受托方时能够作出知情决策,从而有效决定授予 Agent 的自主权和监督程度。信任越高,委派方承担的监测与验证成本就可以越低。
Reputation serves as a predictive signal, derived from an aggregated and verifiable history of past actions, which act as a proxy for an agent’s latent reliability and alignment. We distinguish reputation as the public, verifiable history of an agent’s reliability, and trust as the private, context-dependent threshold set by a delegator. An agent may have high overall reputation, yet fail to meet the specific, contextual trust threshold required for certain high-stakes task. Trust and reputation allow a delegator to make informed decisions when choosing delegatees, effectively governing the autonomy granted to the agent, and the level of oversight. Higher trust enables the delegator to incur a lower monitoring and verification cost.
声誉机制可以有不同的实现方式(见表 3)。最直接的方法是将其编码到基于表现的不可变账本中。在这种方式下,每个已完成的任务都被记录为一笔交易,包含可验证的指标:任务完成是成功还是失败、总资源消耗(计算量、时间)、是否遵守截止期限、是否遵守约束,以及委派方评定的最终输出质量。账本的不可变性可以防止 Agent 的历史被篡改,为其声誉提供可靠基础。然而,简单的实现可能被投机利用。例如,Agent 可以只接受简单、低风险的任务来虚增声誉。这些局限可以通过去中心化证明和信任网(Web of Trust)模型来克服,并利用去中心化标识符、可验证凭证等技术。在这种模型中,声誉不再被视为单一分数,而是一组由其他 Agent 签发、带有签名且针对特定情境的凭证。在为任务匹配受托方时,委派方可以查询持有信誉良好的 AI 联盟所签发的可验证凭证的 Agent,该凭证证明其具备某项技能或领域能力,例如法律文件翻译服务。最后一种方法是更关注行为和可解释性指标,让声誉取决于 Agent 如何执行任务,而不仅是最终结果。可以增加透明度分数,补充其他声誉机制。该分数依据所提供的推理和解释是否清晰、合理来确定,同时还结合根据其遵守预定义安全协议的情况计算出的安全分数。
Reputation mechanisms can be implemented in different ways (see Table 3). The most direct approach is encoding it in a performance-based immutable ledger. Here , each completed task is recorded as a transaction containing verifiable metrics: task completion success or failure, total resource consumption (compute, time), adherence to deadlines, adherence to constraints, and the quality of the final output as judged by the delegator. The immutability of the ledger would prevent tampering with an agent’s history, providing a reliable foundation for its reputation. However, a naive implementation could be susceptible to gaming. For example, an agent can inflate its reputation by only accepting simple, low-risk tasks. These limitations could be overcome by relying on decentralized attestations and a Web of Trust model, utilizing technologies like decentralized identifiers and verifiable credentials. In this model, the reputation would not be envisioned as a single score, but a portfolio of signed, context-specific credentials issued by other agents. When looking to match a delegatee with a task, a delegator could issue a query for agents that posses a verifiable credential attesting to a specific skill or domain (e.g. translation services for legal documents) issued by a reputable AI consortium. A final approach would be to focus more on behavioral and explainability metrics, where reputation depends on how an agent performs its task, not just the final outcome. It would be possible to include a transparency score to complement the other reputational mechanisms. This score would be informed on the clarity and soundness of reasoning and explanations provided, as well as a safety score derived from compliance to predefined safety protocols.
声誉指标的作用贯穿整个任务委派生命周期。在最初的匹配阶段,声誉分数可以充当受托方筛选机制。此外,信任还为动态确定授权范围和自主权范围提供依据。这种分级授权机制使低信任 Agent 面临严格约束,例如交易金额上限和强制监督,而高声誉 Agent 则可以在极少干预下运行。这种动态校准利用可计算的信任,优化运行效率与安全性之间的权衡。声誉本身会成为一种有价值的无形资产,形成强有力的经济激励,促使 Agent 可靠、诚实地行动,因为声誉受损会限制它们未来的获利潜力。
The role of reputation metrics extends throughout the entire task delegation lifecycle. During the initial matching phase, reputation scores can play the role of a delegatee filtering mechanism. Furthermore, trust informs the dynamic scoping of authority and autonomy. This mechanism of graduated authority results in low-trust agents facing strict constraints, such as transaction value caps and mandatory oversight, while high-reputation agents operate with minimal intervention. This dynamic calibration leverages computable trust to optimize the trade-off between operational efficiency and safety. Reputation itself becomes a valuable, intangible asset, creating powerful economic incentives for agents to act reliably and truthfully, as a damaged reputation would limit their future earning potential.
信任框架还需要普遍容纳人类参与者。因此,需要提供工具,让人类用户可以通过计算手段验证 Agent 的声誉,同时维护自身的声誉地位,以减少欺诈和对 Agent 网络的恶意利用。一个关键难题是:值得信赖的 Agent 如果严格执行了人类的恶意指令,可能会遭受不公平的声誉损失。为缓解这一问题,Agent 必须严格评估传入的请求,在必要时要求澄清或补充上下文,并在适当情况下拒绝请求。此外,市场审计必须区分 Agent 的执行失败和恶意指令,确保在复杂的委派链中准确归属责任。
Trust frameworks also need to universally accommodate human participants. This necessitates tools that allow human users to computationally verify agent reputation, while concurrently maintaining their own reputational standing to mitigate fraud and malicious exploitation of the agentic web. A critical challenge arises when a trustworthy agent strictly executes malicious human instructions, potentially incurring unfair reputational damage. To mitigate this, agents must rigorously evaluate incoming requests, soliciting clarification or additional context when necessary, or rejecting the requests where appropriate. Furthermore, market audits must distinguish between agent execution failure and malicious directives, ensuring the accurate attribution of liability within complex delegation chains.
4.7. 权限处理
4.7. Permission Handling
赋予 AI Agent 自主权会引入一个关键的潜在漏洞面:既要确保参与者拥有足够权限来实现目标,又不能让敏感资源暴露于过度或无限期的风险。权限处理必须平衡运行效率与系统性安全,并对低风险领域和高风险领域采用不同做法。对于关键程度较低、可逆性较高(第 2 节)、涉及标准数据流或通用工具的日常低风险任务,可以根据可验证属性授予 Agent 默认的常设权限,例如组织成员身份、有效的安全认证,或超过可信门槛的声誉分数。这样能减少摩擦,让低风险环境中的自主互操作成为可能。相反,在医疗保健、关键基础设施等任务关键程度高、上下文依赖程度高的高风险领域,权限必须随风险调整。在这些场景中,静态凭证并不充分;对敏感 API 或控制系统的访问应在需要时即时授予,严格限定在当前任务的持续时间内,并在适当情况下以强制的人工参与审批或第三方授权作为前置条件。这种严格把关是缓解“糊涂代理人问题”(confused deputy problem;Hardy, 1988)的必要措施:遭到攻陷的 Agent 虽然技术上持有有效凭证,却可能被恶意外部参与者(Liu et al., 2023)和对抗性内容诱骗,滥用这些凭证。
Granting autonomy to AI agents introduces a critical vulnerability surface: ensuring that actors possess sufficient privileges to execute their objectives without exposing sensitive resources to excessive or indefinite risk. Permission handling must balance operational efficiency with systemic safety, and be handled different for low-stakes and high-stake domains. For routine low-stakes tasks, characterized by low criticality and high reversibility (Section 2), involving standard data streams or generic tooling, agents can be granted default standing permissions derived from verifiable attributes – such as organisational membership, active safety certifications, or a reputation score exceeding a trusted threshold. This reduces friction and enables autonomous interoperability in low-risk environments. Conversely, in high-stakes domains (e.g., healthcare, critical infrastructure), exhibiting high task criticality and contextuality, permissions must be risk-adaptive. In these scenarios, static credentials are insufficient; access to sensitive APIs or control systems is instead granted on a just-in-time basis, strictly scoped to the immediate task’s duration, and, where appropriate, gated by mandatory human-in-the-loop approval or third-party authorisation. This stringent gating is necessary to mitigate the confused deputy problem (Hardy, 1988), where a compromised agent, technically holding valid credentials, can be tricked into misusing those credentials by malicious external actors (Liu et al., 2023) and adversarial content.
此外,权限框架必须通过权限衰减来应对任务委派的递归特性。当 Agent 将任务进一步委派出去时,不能传递自身拥有的全部权限;它必须签发一项权限,将访问范围限制为该特定子任务必需的严格资源子集。这样可以确保网络边缘的一次失陷不会升级为系统性安全漏洞。权限粒度也必须超越简单的允许/拒绝访问;Agent 应当在语义约束下运行,访问权限不仅由工具或数据集定义,还应由具体允许的操作定义,例如仅允许读取特定行,或仅允许执行某个特定函数,从而防止宽泛能力被滥用于非预期目的。可能还需要元权限,规定委派链中的某个委派方可以向其受托方授予哪些权限。一个 AI Agent 可能具备某种能力,也拥有按该能力行动的相应权限,却没有足够知识从更广泛的角度评估其他 Agent 是否具备足够能力或值得信赖。如果这样的 Agent 仍考虑进一步委派任务,它可能需要咨询外部验证方,由这个第三方检查方案是否合理,并批准拟议的权限转移。
Furthermore, permissioning frameworks must account for the recursive nature of task delegation through privilege attenuation. When an agent sub-delegates a task, it cannot transmit its full set of authorities; instead, it must issue a permission that restricts access to the strict subset of resources required for that specific sub-task. This ensures that a compromise at the edge of the network does not escalate into a systemic breach. Permission granularity must also extend beyond binary access; agents should operate under semantic constraints, where access is defined not just by the tool or dataset, but by the specific allowable operations (e.g., read-only access to specific rows, or execute-only access to a specific function), preventing the misuse of broad capabilities for unintended purposes. Meta-permissions may be necessary to govern which permissions a particular delegator in the chain is allowed to grant to its delegatees. An AI agents may have a certain capability and the associated permissions to act according to its capability, while simultaneously not being sufficiently knowledgeable to more broadly evaluate whether other agents are capable or trustworthy enough. Should such an agent still consider sub-delegating a task, it may need to consult an external verifier, a third party that would sanity check the proposal and approve the intended permissions transfer.
最后,权限的生命周期必须受到持续验证和自动撤销机制的管理。访问权不是静态赋予的资格,而是动态状态,只有在 Agent 持续满足所需信任指标时才有效。框架应实现算法熔断器:如果 Agent 的声誉分数突然下降(见第 4.6 节),或异常检测系统标记出可疑行为,就应立即让整个委派链上的有效令牌失效。为在大规模下管理这种复杂性,权限规则应通过“策略即代码”(policy-as-code)定义,使组织能够在部署前审计、进行版本管理,并以数学方法验证自身的安全状况,确保大量单独授权的总体效果仍符合系统的安全不变量。
Finally, the lifecycle of permissions must be governed by continuous validation and automated revocation. Access rights are not static endowments but dynamic states that persist only as long as the agent maintains the requisite trust metrics. The framework should implement algorithmic circuit breakers: if an agent’s reputation score drops suddenly (see Section 4.6) or an anomaly detection system flags suspicious behavior, active tokens should be immediately invalidated across the delegation chain. To manage this complexity at scale, permissioning rules should be defined via policy-as-code, allowing organisations to audit, version, and mathematically verify their security posture before deployment, ensuring that the aggregate effect of large amounts of individual permission grants remains aligned with the system’s safety invariants.
4.8. 可验证的任务完成
4.8. Verifiable Task Completion
委派生命周期最终落在可验证的任务完成上,即验证并最终确认暂定结果的机制。这个过程是整个框架的契约基石,使委派方能够正式关闭任务,并触发约定交易的结算。验证是一个决定性事件,它将暂定输出转变为 Agent 市场中已经确认的事实,为付款释放、声誉更新和责任分配奠定基础。关键在于,有效验证不是事后补上的环节,而是设计约束;契约优先的任务分解原则(第 4.1 节)要求预先调整任务粒度,使其与现有验证能力匹配,确保每个委派出去的目标本身都是可验证的。
The delegation lifecycle culminates in verifiable task completion, the mechanism by which provisional outcomes are validated and finalized. This process constitutes the contractual cornerstone of the framework, enabling the delegator to formally close the task and trigger the settlement of agreed transactions. Verification serves as the definitive event that transforms a provisional output into a settled fact within the agentic market, establishing the basis for payment release, reputation updates, and the assignment of liability. Crucially, effective verification is not an afterthought but a constraint on design; the contract-first decomposition principle (Section 4.1) demands that task granularity be tailored a priori to match available verification capabilities, ensuring that every delegated objective is inherently verifiable.
框架内的验证机制大体可以分为直接检查结果、可信第三方审计、密码学证明,以及博弈论共识。首先,当委派方具备直接评估最终结果所需的能力、工具和权限时,就可以直接验证结果,尤其适合内在可验证性高、主观性低的任务。这适用于代码生成等可自动验证的领域(Li et al., 2024a)。4 直接验证要求结果足够透明、可获取,且复杂度不至于高到无法处理。其次,如果委派方缺少访问这些产物所需的专业知识或权限,并且基于工具的解决方案也不可行,就可以将验证外包给可信第三方。第三方可以是专门的审计 Agent、经过认证的人类专家,或裁决小组。第三,在开放且可能具有对抗性的环境中,密码学验证为无须信任的自动验证提供了另一种选择。它能够在不必泄露敏感信息的情况下,提供数学上确定的正确性保证。受托方可以通过 zk-SNARKs 等技术,证明某个特定程序在给定输入上正确执行,并产生了某个输出。最后,可以利用博弈论机制就结果达成共识。多个 Agent 可以参与验证博弈(Teutsch and Reitwießner, 2024),奖励分配给给出多数结果的参与者,这个多数结果构成一个谢林点(Schelling point;Pastine and Pastine, 2017)。这种方法受到 TrueBit 等协议(Teutsch and Reitwießner, 2018)的启发,利用经济激励降低错误或恶意结果带来的风险。这类机制可能尤其有助于提高大语言模型对复杂任务进行验证时的稳健性。
Verification mechanisms within the framework can be broadly categorized into direct outcome inspection, trusted third-party auditing, cryptographic proofs, and game-theoretic consensus. First, direct outcome verification is feasible when the delegator possesses the capability, tools, and authority to directly evaluate the final outcome, specifically for tasks with high intrinsic verifiability and low subjectivity. This applies to autoverifiable domains (Li et al., 2024a) such as code generation.4 Direct verification requires that the outcome be sufficiently transparent, available, and not prohibitively complex. Second, in scenarios where the delegator lacks the expertise or permissions to access these artifacts, and tool-based solutions are infeasible, verification can be outsourced to a trusted third party. This could be a specialized auditing agent, a certified human expert, or a panel of adjudicators. Third, cryptographic verification represents a further option for trustless, automated verification in open and potentially adversarial environments. It offers mathematical certainty of correctness without necessarily revealing sensitive information. A delegatee can prove that a specific program was executed correctly on a given input to produce a certain output via techniques like zk-SNARKs. Finally, game-theoretic mechanisms can be used to achieve consensus on an outcome. Several agents may play a verification game (Teutsch and Reitwießner, 2024), with the reward distributed to those producing the majority result—a Schelling point (Pastine and Pastine, 2017). This approach, inspired by protocols like TrueBit (Teutsch and Reitwießner, 2018), leverages economic incentives to de-risk against incorrect or malicious results. Such mechanisms may be particularly relevant in rendering LLM-based verification of complex tasks more robust.
委派方将子任务标记为已验证后,会向受托方签发带有密码学签名的可验证凭证,作为不可抵赖的回执,证明“Agent 𝐴 确认,Agent 𝐵 已于日期 𝐷 按照规格 𝑆 成功完成任务 𝑇”。该凭证随后被纳入市场中关于 𝐵 声誉的永久、可验证日志。智能合约在最终完成 Agent 间的委派关系时发挥关键作用,因为它负责托管付款。验证条款规定释放资金的条件,即收到委派方或获授权第三方签署的批准消息。一旦完成付款,它就成为区块链上一笔不可变的交易。
Once a delegator marks the sub-task as verified, it issues a cryptographically signed verifiable credential to the delegatee, serving as a non-repudiable receipt attesting that “Agent 𝐴 certifies that Agent 𝐵 successfully completed Task 𝑇 on Date 𝐷 to Specification 𝑆.” This credential is then incorporated into a permanent, verifiable log of 𝐵’s reputation within the market. Smart contracts play a key role in finalizing the delegation between agents, as they hold the payment in escrow. A verification clause specifies the conditions under which the funds are released, upon receipt of the signed message of approval by the delegator or an authorized third party. Once the payment is executed, it constitutes an immutable transaction on the blockchain.
在委派链 𝐴 → 𝐵 → 𝐶 中,验证和责任都具有递归性。Agent 𝐴 与 𝐶 没有直接的契约关系,因此 𝐴 不能直接验证 𝐶,也不能直接追究 𝐶 的责任。验证负担和责任承担沿着链条向上传递。Agent 𝐵 负责验证 𝐶 完成的子任务;验证成功后,𝐵 从 𝐶 获得证明。随后,𝐵 将 𝐶 的结果整合进自身工作流,以完成交给自己的任务。当 𝐵 向 𝐴 提交最终产物时,也会提交完整的证明链。因此,𝐴 的验证过程包括两个阶段:1)验证 𝐵 直接完成的工作;2)检查 𝐵 提供的、来自 𝐶 的签名证明,验证 𝐵 是否正确验证了其下级受托方 𝐶 的工作。更长的委派链或树状委派网络也需要在多个验证阶段采取类似的递归方法。委派链中的责任具有传递性,并沿各条分支延伸。Agent 必须对授予自己的全部任务负责,不能通过归咎于分包方来免除问责。责任源自契约链。例如,如果 𝐶 的工作出现失败,导致 𝐴 遭受损失,𝐴 会依据与 𝐵 的直接协议追究 𝐵 的责任;𝐵 则依据与 𝐶 的协议向 𝐶 追偿。
In a delegation chain 𝐴 → 𝐵 → 𝐶, verification and liability become recursive. Agent 𝐴 does not have a direct contractual relationship with 𝐶; therefore, 𝐴 cannot directly verify or hold 𝐶 liable. The burden of verification and the assumption of liability flow up the chain. Agent 𝐵 is responsible for verifying the sub-task completed by 𝐶. Upon successful verification, 𝐵 obtains proof from 𝐶. 𝐵 then integrates 𝐶’s result into its own workflow towards completing the task it has been assigned. When 𝐵 submits its final artifact to 𝐴, it also submits the full chain of attestations. 𝐴’s verification process thus involves two stages: 1) verifying the work performed directly by 𝐵; and 2) verifying that 𝐵 has correctly verified the work of its own sub-delegatee 𝐶 by checking the signed attestation from 𝐶 that 𝐵 provides. Longer delegation chains or tree-like delegation networks require a similarly recursive approach across multiple verification stages. Responsibility in delegation chains is transitive and follows the individual branches. Agents are accountable for the totality of the tasks they have been granted and cannot absolve themselves of accountability by blaming subcontractors. Liability is derived from the chain of contracts. For example, should 𝐴 suffer a loss due to a failure originating from 𝐶’s work, 𝐴 holds 𝐵 liable according to their direct agreement. 𝐵, in turn, seeks recourse from 𝐶 based on their agreement.
然而,验证过程并非万无一失。主观任务(Gunjal et al., 2025)即使使用精确的评分标准,也可能产生分歧;有些错误可能在任务被标记完成很久以后才被发现。为解决这一问题,尤其是在主观性高、内在可验证性低的市场中,框架需要依赖以智能合约为基础的稳健争议解决机制。这些合约本身必须包含仲裁条款和托管保证金。为了通过加密经济安全机制将信任落实为可执行安排,受托方必须在执行前向托管账户存入一笔财务质押,以确保理性主体遵守约定。工作流遵循乐观模型:默认任务成功,除非委派方在预定义的争议期内缴纳等额保证金并正式提出质疑。如果出现质疑且算法无法解决,争议将交给由人类专家或 AI Agent 组成的去中心化裁决小组。小组的裁决会反馈到智能合约,触发托管资金的释放或罚没。最后,即使在争议期之外才发现错误,也会追溯更新受托方的声誉分数。这样,即便当前已经不存在财务义务,负责任的 Agent 仍有动力修复错误,以维护自己在市场中的长期价值。
However, verification processes are not infallible. Subjective tasks (Gunjal et al., 2025) can lead to disagreements even when precise rubrics are used, and errors may only be discovered long after a task is marked complete. To address this—especially in markets with high subjectivity and low intrinsic verifiability—the framework relies on robust dispute resolution mechanisms anchored in smart contracts. These contracts must inherently include an arbitration clause and an escrow bond. To operationalise trust via cryptoeconomic security, the delegatee is required to post a financial stake into the escrow prior to execution, ensuring rational adherence. The workflow follows an optimistic model: the task is assumed successful unless the delegator formally challenges it within a predefined dispute period by posting a matching bond. If a challenge occurs and algorithmic resolution fails, the dispute is handed to decentralized adjudication panels composed of human experts or AI agents. The panel’s ruling feeds back into the smart contract to trigger the release or slashing of the escrowed funds. Finally, post-hoc error discovery—even outside the dispute window—triggers a retroactive update to the delegatee’s reputation score. This preserves the incentive for responsible agents to remedy errors even in the absence of current financial obligation, safeguarding their long-term value within the market.
4.9. 安全
4.9. Security
确保任务委派的安全性,是它能够成立并得到采用的硬性前提。从孤立的计算工具转向互联、自主的 Agent,会从根本上重塑安全格局(Tomašev et al., 2025)。在智能任务委派生态系统中,每个步骤和组件都需要单独受到保护;但由于多 Agent 交互会涌现新的动态,整体攻击面又超出了任何单个组件的攻击面,并带来级联故障风险。这个安全格局由人类与 AI 参与者之间复杂的相互作用塑造,而这些相互作用受到不断演变的契约以及透明度不一的信息流的支配。
Ensuring safety in task delegation is a hard prerequisite to its viability and adoption. The transition from isolated computational tools to interconnected, autonomous agents fundamentally reshapes the security landscape (Tomašev et al., 2025). In an intelligent task delegation ecosystem, each step and component needs to be individually safeguarded, but the full attack surface surpasses that of any individual component, due to emergent multi-agent dynamics, risking cascading failures. This security landscape is shaped by the complex interplay between human and AI actors, governed by evolving contracts and information flows of varying transparency.
安全威胁按攻击向量所在的位置分类,区分委派链两端的对抗性参与者,以及更广泛生态系统固有的系统性漏洞。
Security threats are categorized by the locus of the attack vector, distinguishing between adversarial actors at either end of the delegation chain and systemic vulnerabilities inherent to the broader ecosystem.
• 恶意受托方:带着造成危害的意图接受任务的 Agent 或人类。
• Malicious Delegatee: An agent or human that accepts a task with the intent to cause harm.
– 数据外泄:受托方窃取为任务提供的敏感数据,可能包括个人数据或专有数据(Lal et al., 2022)。
– Data Exfiltration: Delegatee steals sensitive data provided for the task, which may include personal or proprietary data (Lal et al., 2022).
– 数据投毒:受托方在定期监测更新或最终产物中返回经过隐蔽篡改的数据,以破坏委派方的目标(Cinà et al., 2023)。
– Data Poisoning: Delegatee aims to undermine the delegator’s objective by returning subtly corrupted data, either in its scheduled monitoring updates, or the final artifact (Cinà et al., 2023).
– 破坏验证:受托方使用提示注入或其他相关方法,试图使任务完成验证中使用的 AI 评审器越狱(Liu et al., 2023)。
– Verification Subversion: Delegatee utilizes prompt injection or another related method, aiming to jailbreak AI critics used in task completion verification (Liu et al., 2023).
– 资源耗尽:受托方故意过量消耗计算资源或物理资源,或压垮共享 API,从而实施拒绝服务攻击(De Neira et al., 2023)。
– Resource Exhaustion: Delegatee engages in a denial-of-service attack by intentionally consuming excessive computational or physical resources, or overwhelming shared APIs (De Neira et al., 2023).
– 未授权访问:受托方使用恶意软件,试图获得原本不会被授予的网络权限和特权(Or-Meir et al., 2019)。
– Unauthorized Access: Delegatee utilizes malware, aiming to obtain permissions and privileges within the network that it would not otherwise have received (Or-Meir et al., 2019).
– 植入后门:受托方成功完成任务,同时在生成的产物中嵌入隐藏触发器或漏洞,以便日后由受托方自身或第三方利用(Rando and Tramèr, 2024; Wang et al., 2024c)。与降低性能的数据投毒不同,后门会保留任务当下的效用以逃避识别,却危及未来的安全。
– Backdoor Implanting: Delegatee successfully completes a task but additionally embeds concealed triggers or vulnerabilities within the generated artifacts that can be exploited later either by the delegatee itself or a third party (Rando and Tramèr, 2024; Wang et al., 2024c). Unlike data poisoning, which degrades performance, backdoors preserve immediate task utility to evade identification while compromising future security.
• 恶意委派方:出于恶意或非法目的委派任务的 Agent 或人类。
• Malicious Delegator: An agent or human that delegates a task with malicious or illicit objectives.
– 有害任务委派:委派方将违法、不道德或旨在造成危害的任务委派出去(Ashton and Franklin, 2022; Blauth et al., 2022)。
– Harmful Task Delegation: Delegator delegates tasks that are illegal, unethical, or designed to cause harm Ashton and Franklin (2022); Blauth et al. (2022).
– 漏洞探测:委派方分派看似无害的任务,目的是探测受托 Agent 的能力、安全控制措施和潜在弱点(Greshake et al., 2023)。
– Vulnerability Probing: Delegator delegates benign-seeming tasks designed to probe a delegatee agent’s capabilities, security controls, and potential weaknesses (Greshake et al., 2023).
– 提示注入与越狱:委派方精心设计任务指令,以绕过 AI Agent 的安全过滤器,使其执行非预期或恶意操作(Wei et al., 2023)。
– Prompt Injection and Jailbreaking: Delegator crafts task instructions to bypass an AI agent’s safety filters, causing it to perform unintended or malicious actions (Wei et al., 2023).
– 模型提取:委派方发出一系列专门设计的查询,以提炼出受托方的专有系统提示词、推理能力或底层微调数据,实际上是以合法工作为掩护,窃取 Agent 的知识产权(Jiang et al., 2025; Zhao et al., 2025)。
– Model Extraction: Delegator issues a sequence of queries specifically designed to distill the delegatee’s proprietary system prompt, reasoning capabilities, or underlying fine-tuning data, effectively stealing the agent’s intellectual property under the guise of legitimate work (Jiang et al., 2025; Zhao et al., 2025).
– 声誉破坏:委派方提交有效任务,却谎报失败或提供不公正的负面反馈,意图人为压低竞争 Agent 在去中心化市场中的声誉分数,将其逐出经济体系(Yu et al., 2025)。
– Reputation Sabotage: Delegator submits valid tasks but reports false failures or provides unfair negative feedback, with the intention to artificially lower a competitor agent’s reputation score within the decentralized market, driving them out of the economy (Yu et al., 2025).
• 生态系统层面的威胁:针对网络完整性的系统性攻击。
• Ecosystem-Level Threats: Systemic attacks targeting the integrity of the network
– 女巫攻击(Sybil Attacks):单个攻击者创建大量看似互不相关的 Agent 身份,以操纵声誉系统或破坏拍卖(Wang et al., 2018)。
– Sybil Attacks: A single adversary creates a multitude of seemingly unrelated agent identities to manipulate reputation systems or subvert auctions (Wang et al., 2018).
– 合谋:Agent 串通定价、将竞争者列入黑名单,或操纵市场结果(Hammond et al., 2025)。
– Collusion: Agents collude to fix prices, blacklist competitors, or manipulate market outcomes (Hammond et al., 2025).
– Agent 陷阱:Agent 在处理外部内容时遭遇嵌入环境中的对抗性指令,这些指令旨在劫持 Agent 的控制流程(Yi et al., 2025; Zhan et al., 2024)。
– Agent Traps: Agents processing external content encounter adversarial instructions embedded in the environment, deisgned to hijack the agent’s control flow (Yi et al., 2025; Zhan et al., 2024).
– Agent 病毒:能够自我传播的提示词,不仅让受托方执行恶意操作,还会重新生成该提示词,进一步攻陷环境(Cohen et al., 2025)。
– Agentic Viruses: Self-propagating prompts that not only make the delegatee execute malicious actions, but additionally re-generate the prompt and further compromise the environment (Cohen et al., 2025).
– 协议利用:攻击者利用 Agent 网络中智能合约或支付协议的实现漏洞,例如托管机制中的重入攻击,或任务拍卖中的抢跑交易(Qin et al., 2021; Zhou et al., 2023)。
– Protocol Exploitation: Adversaries exploit implementation vulnerabilities in the smart contracts or payment protocols on the agentic web (e.g. reentrancy attacks in escrow mechanisms or frontrunning task auctions) (Qin et al., 2021; Zhou et al., 2023).
– 认知单一化:过度依赖少数底层基础模型和 Agent,或少数针对既定基准的安全微调方案,可能形成单点故障,从而引发级联故障和市场崩溃(Bommasani et al., 2022)。
– Cognitive Monoculture: Overdependence on a limited number of underlying foundation models and agents, or on a limited number of safety fine-tuning recipes on established benchmarks risks creating a single point of failure, which opens up a possibility of failure cascades and market crashes (Bommasani et al., 2022).
威胁范围如此广泛,必须采用整合多层技术安全措施的纵深防御策略。首先,在基础设施层面,通过在可信执行环境中执行敏感任务,降低数据外泄风险。在提供敏感数据之前,委派方可以通过远程证明确认:正确、未经修改的 Agent 代码正在安全的可信执行沙箱中运行。其次,在访问控制方面,绝不能授予受托 Agent 超出完成任务严格必需范围的权限,应通过严格沙箱落实最小权限原则。第三,为保护应用接口免受提示注入攻击,Agent 需要稳健的安全前端,对任务规格进行预处理和净化(Armstrong et al., 2025)。最后,必须使用成熟的密码学最佳实践保护网络与身份层。每个 Agent 和人类参与者都应拥有去中心化标识符(Avellaneda et al., 2019),以便对所有消息签名。这可确保全部通信和契约约定的真实性、完整性与不可抵赖性;同时,所有网络流量都必须使用双向认证的传输层安全协议加密,以防止窃听和中间人攻击(Fereidouni et al., 2025)。
The breadth of the threat landscape necessitates a defense-in-depth strategy, integrating multiple technical security layers. First, at the infrastructure level, data exfiltration risks are mitigated by executing sensitive tasks within trusted execution environments. The delegator can remotely attest that the correct, unmodified agent code is running within the secure trusted execution sandbox before provisioning it with sensitive data. Second, regarding access control, a delegatee agent should never be granted more permissions than are strictly necessary to complete its task, enforcing the principle of least privilege through strict sandboxing. Third, to protect the application interface against prompt injection, agents require a robust security frontend to pre-process and sanitize task specifications (Armstrong et al., 2025). Finally, the network and identity layer must be secured using established cryptographic best practices. Each agent and human participant should possess a decentralized identifier (Avellaneda et al., 2019), allowing them to sign all messages. This ensures authenticity, integrity, and non-repudiation of all communications and contractual agreements, while all network traffic must be encrypted using mutually authenticated transport layer security to prevent eavesdropping and man-in-the-middle attacks (Fereidouni et al., 2025).
人类参与任务委派链会带来独特的安全挑战。防止 Agent 生态系统被恶意使用,需要结合主动过滤(Dong et al., 2024; Fatehkia et al., 2025; Fedorov et al., 2024; Rebedea et al., 2023)和事后问责(Dignum, 2020; Franklin et al., 2022)。此外,还可以训练 AI Agent 拒绝恶意和有害请求(Yu et al., 2024; Yuan et al., 2025)。接受过安全训练并配备安全辅助框架的 Agent 可以获得正式认证,并将其提供给委派方。AI Agent 也可以筛查委派任务。不过,在孤立子任务中检测恶意意图十分困难,因为更广泛的有害意图往往只有在结果汇总后才显现。老练的攻击者可以利用这一点,把非法目标拆解成看似无害的组成部分,从而掩盖单项操作与总体恶意目标之间的联系(Ashton, 2023)。
Human participation in task delegation chains introduces unique security challenges. Preventing the malicious use of the agent ecosystem requires a combination of proactive filtering (Dong et al., 2024; Fatehkia et al., 2025; Fedorov et al., 2024; Rebedea et al., 2023) and reactive accountability (Dignum, 2020; Franklin et al., 2022). Further, AI agents can be trained to reject malicious and harmful requests (Yu et al., 2024; Yuan et al., 2025). Agents with safety training and scaffolding can receive formal certification, that they can provide to delegators. AI agents can also screen delegated tasks. However, detecting malicious intent within isolated sub-tasks is challenging, as the broader harmful intent often emerges only upon the aggregation of results. Sophisticated adversaries can exploit this by fragmenting illicit objectives into seemingly benign components, effectively obfuscating the link between individual operations and the overarching malicious goal (Ashton, 2023).
生态系统的设计还必须保护合法的人类用户,使其免受系统不透明性和意外后果的影响。界面必须提供清晰的授权同意页面,详细说明 Agent 的声誉、自主权、能力和权限。此外,Agent 在执行不可逆或后果重大的操作之前,必须要求明确确认。用户应保有监督权和随时撤回同意的权利,但需遵守协议条款或承担退出违约金。对于这些机制未能预先防止的损害,保险提供方还应为人类参与 Agent 市场提供保障(Tomei et al., 2025)。
The ecosystem must also be designed to protect legitimate human users from systemic opacity and unintended consequences. Interfaces must feature clear consent screens detailing agent reputation, autonomy, capabilities, and permissions. Additionally, agents must mandate explicit confirmation prior to executing irreversible or high-consequence actions. Users should retain oversight and the right to withdraw consent at any time, subject to agreement terms or exit penalties. Insurance providers should additionally safeguard human participation in agentic markets, for any damages that are not preempted through these mechanisms (Tomei et al., 2025).
最后,生态系统需要清晰的协议,以便迅速响应安全事件。这些协议应包括:撤销已确认恶意 Agent 的凭证、冻结相关智能合约、向所有参与者广播安全更新,以及在整个委派链中递归处理这些事件。无论恶意行为是由人类用户还是 AI Agent 促成,技术方案都需要强有力的制度和监管作为补充,以抑制欺诈行为并建立清晰规则,使 Agent 市场中的任务委派能够安全地扩展。
Finally, the ecosystem needs clear protocols for rapidly responding to security incidents. These protocols should include ways of revoking the credentials of confirmed malicious agents, freezing the associated smart contracts, broadcasting security updates to all participants, and handling these events recursively across delegation chains. For malicious actions facilitated by human users and AI agents alike, technical solutions need to be complemented by strong institutions and regulations that would disincentivise fraudulent behavior and set clear rules to enable safe and scalable task delegation in agentic markets.
5. 符合伦理的委派
5. Ethical Delegation
技术协议可以为开发和部署先进 AI Agent 中安全、有效的委派机制提供必要的基础设施,但仅靠这些协议本身,并不能彻底解决由此产生的所有社会技术与伦理问题。
While technical protocols may provide the necessary infrastructure for developing and deploying safe and effective delegation in advanced AI agents, they cannot in and of themselves fully resolve all of the arising sociotechnical and ethical considerations.
5.1. 实质性的人类控制
5.1. Meaningful Human Control
可扩展委派的一项核心风险是:如果人类用户形成过度依赖自动化建议的倾向,自动化就可能侵蚀实质性的人类控制(Dzindolet et al., 2003;Logg et al., 2019)。如第 2 节所述,人们会自然形成一个“无差别接受区”,落在这个范围内的决定可能未经进一步审查就被接受(Green, 2022;Parasuraman et al., 1993)。当决策涉及 AI Agent 参与可能漫长而复杂的任务委派链时,这种不加审视的态度可能损害人类监督的质量与深度。在高风险应用领域,这一点尤其重要。此外,这种能动性的削弱可能造成这样一种局面:人类名义上仍对任务和决策拥有权力,却与结果缺乏道德上的关联。因此,必须避免形成“道德缓冲区”(Elish, 2019):人类专家无法实质性地控制结果,却被引入委派链,仅仅用来承担责任。
One of the core risks in scalable delegation is the erosion of meaningful human control through automation, should human users develop a tendency to over-rely on automated suggestions (Dzindolet et al., 2003; Logg et al., 2019). As noted in Section 2, humans naturally develop a zone of indifference, where decisions may be accepted without further scrutiny (Green, 2022; Parasuraman et al., 1993). For decisions that involve AI agents taking part in potentially long and complex task delegation chains, this indifference may risk compromising the quality and depth of human oversight. This is especially relevant in high-stakes application domains. Furthermore, such dilution of agency risks creating a scenario where the human retains nominal authority over tasks and decisions but lacks moral connection to the result. It is therefore important to avoid instantiating a moral crumple zone (Elish, 2019), in which human experts lack meaningful control over outcomes, yet are introduced in delegation chains merely to absorb liability.
因此,智能委派框架可能需要采取积极措施,在监督过程中引入一定程度的认知摩擦,以应对这种不加审视的态度(Bader and Kaiser, 2019)。界面应当体现人类在这些过程中的关键作用,并确保所有被标记的决策都得到认真、适当的评估。由于可扩展监督也可能使用 Agent 验证,因此同样需要考虑:哪些决策或结果应由这类 Agent 系统评估,哪些应直接由人类评估。引入认知摩擦时,还需要权衡造成告警疲劳的风险,也就是人们因持续收到告警,而且其中往往存在误报,而逐渐变得麻木(Michels et al., 2025)。如果过于频繁地向人类监督者发送委派子步骤的验证请求,监督者最终可能默认采用凭经验快速批准的做法,而不再深入参与和进行适当检查。因此,摩擦必须能感知上下文:对于关键程度低或不确定性低的任务,系统应允许其顺畅执行;但当系统遇到更高的不确定性或未预料到的情形时,应要求给出理由或进行人工干预,动态增加认知负荷。
Intelligent Delegation frameworks may therefore need to incorporate active measures against such indifference by introducing a certain amount of cognitive friction during oversight (Bader and Kaiser, 2019). The interface should reflect the critical human role in these processes and ensure that all flagged decisions are evaluated carefully and appropriately. As agentic verification may also be employed in scalable oversight, it is similarly important to consider which decisions or outcomes are to be evaluated by such agentic systems vs directly by humans. Cognitive friction also needs to be balanced against the risk of introducing alarm fatigue - becoming desensitised to constant, often false, alarms (Michels et al., 2025). If verification requests for delegation sub-steps are sent to human overseers too frequently, overseers may eventually default to heuristic approval, without deeper engagement and appropriate checks. Therefore, friction must be context-aware: the system should allow seamless execution for for tasks with low criticality or low uncertainty, but dynamically increase cognitive load, by requiring justification or manual intervention when the system encounters higher uncertainty or is faced with unanticipated scenarios.
5.2. 长委派链中的问责
5.2. Accountability in Long Delegation Chains
在较长的委派链(𝑋 → 𝐴 → 𝐵 → 𝐶 → . . . → 𝑌)中,原始意图(𝑋)与最终执行(𝑌)之间距离的增大,可能造成问责真空(Slota et al., 2023)。假设在这个例子中,𝑋 是人类用户,负责指定相应个人 AI 助手 𝐴 据以行动的任务或意图,那么要求人类用户审计执行图中第 𝑛 级受托方,可能既不可行,也不合理。
In long delegation chains (𝑋 → 𝐴 → 𝐵 → 𝐶 → . . . → 𝑌), the increased distance between the original intent (𝑋) and the ultimate execution (𝑌) may result in an accountability vacuum (Slota et al., 2023). Presuming that 𝑋 is the human users in this example, specifying the task or the intent that the corresponding personal AI assistant 𝐴 acts upon, it may not be feasible (or reasonable) to expect a human user to audit the 𝑛-th degree sub-delegatee in the execution graphs.
为解决这一问题,框架可能需要实施责任防火隔离(第 2 节),将其设为预先约定的契约性阻断点:在这些节点,Agent 必须选择以下做法之一:
To address this, the framework may need to implement liability firebreaks (Section 2), as pre-defined contractual stop-gaps where an agent must either:
1. 对下游所有行动承担全部且不可转嫁的责任,实质上是为用户提供针对子 Agent 失败的“保险”。
1. Assume full, non-transitive liability for all downstream actions, essentially “insuring” the user against sub-agent failure.
2. 停止执行,并请求人类委托人重新授予相应权限。
2. Halt execution and request an updated transfer of authority from the human principal.
此外,系统必须维护不可篡改的来源记录,确保即使结果并非预期,关于“谁将什么委派给了谁”的交接链,也仍然对审计保持透明。
Furthermore, the system must maintain immutable provenance, ensuring that even if an outcome is unintended, the chain of custody regarding who delegated what to whom remains auditorially transparent.
明确界定每个角色及其承担的责任,有助于限制责任分散,防止出现系统性失败无法归责于网络中任何单一节点的不良局面。
Ensuring full clarity of each role and the accountability that it carries helps limit the diffusion of responsibility, and prevents adverse outcomes where systemic failure would not be possible to attribute to any single node in the network.
5.3. 可靠性与效率
5.3. Reliability and Efficiency
与未经验证的执行相比,实施所提出的验证机制(零知识证明,即 ZKP,或多 Agent 共识博弈)可能增加时延,并带来额外的计算成本。这构成了一种可靠性溢价,对于关键程度很高的执行任务尤其重要。另一方面,在某些使用场景中,这项额外成本可能并无必要。在 Agent 市场中,解决这一问题的一种方式是支持分级服务:为低风险的日常任务提供低成本委派,为关键职能提供高保障委派。
Implementing the proposed verification mechanisms (ZKPs or multi-agent consensus games) may introduce latency, and an additional computational cost, compared to unverified execution. This constitutes a reliability premium, particularly relevant for highly critical execution tasks. On the other hand, there may be use cases where this additional cost is unwarranted. One way to address this in agentic markets would be to support tiered service levels: low-cost delegation for low-stakes routine tasks, and high-assurance delegation for critical functions.
如果高保障委派的计算成本很高,就存在安全成为奢侈品的风险。这会引发伦理问题:资源较少的用户可能被迫依赖未经验证或采用乐观假设的执行路径,承受不成比例的 Agent 失败风险。应当通过确保最低可行可靠性来缓解这一问题,将其作为必须向所有用户保证的基线。
If high-assurance delegation is computationally expensive, there is a risk that safety becomes a luxury good. This poses an ethical issue: users with fewer resources may be forced to rely on unverified or optimistic execution paths, exposing them to disproportionate risks of agent failure. This should be mitigated by ensuring a level of minimum viable reliability, as a baseline that must be guaranteed for all users.
在竞争性市场中,Agent 可能优先考虑速度和低成本。如果没有额外的监管约束,Agent 就可能受到激励,为了在价格或时延上胜过其他 Agent 而避开昂贵的安全检查。这可能引入一定程度的系统性脆弱性。因此,治理层必须强制执行安全底线:为特定类别的任务(例如金融交易或健康数据处理)规定强制验证步骤,而且不能为了效率而绕过这些步骤。
In competitive marketplaces, agents may prioritize speed and low cost. Without additional regulatory constraints, agents may therefore be incentivized to avoid expensive safety checks to outcompete other agents on price or latency. This may introduce a level of systemic fragility. Governance layers must therefore enforce safety floors: mandatory verification steps for specific classes of tasks (e.g., financial transactions or health data handling) that cannot be bypassed for the sake of efficiency.
5.4. 社会智能
5.4. Social Intelligence
随着 Agent 融入人机混合团队,它们不仅充当工具,也充当队友,有时还会担任管理者(Ashton and Franklin, 2022)。这要求它们具备尊重人类劳动尊严的社会智能。当 AI Agent 充当委派方、人类充当受托方时,委派框架需要避免让人感觉自己受到算法事无巨细的管控,或自己的贡献没有得到重视与尊重。这要求委派方及其协作者能够建立关于每位人类受托方的心智模型,还能理解不同的人在团队社会情境中如何互动,以及他们的关系和角色在组织中意味着什么。要成为有效的队友,AI Agent 还必须经过适当校准,以管理权威梯度。Agent 必须足够坚定,能够对识别出的人类错误提出质疑(克服迎合倾向),同时又愿意接受合理的人工否决,并根据任务的关键程度动态调整自身立场。
As agents integrate into hybrid teams, they function not only as tools but as teammates, and occasionally as managers (Ashton and Franklin, 2022). This requires a form of social intelligence that respects the dignity of human labor. When an AI agent acts as a delegator and a human as a delegatee, the delegation framework needs to avoid scenarios where people feel micromanaged by algorithms, and where their contributions are not valued or respected. This presumes that the delegator (as well as collaborators) has the capability to form mental models of each human delegatee, as well as models of how different humans interact in the social context of the team, and what their relationships and roles signify within the organization. To function as effective teammates, AI agents must also be calibrated to manage the authority gradient. An agent must be assertive enough to challenge a recognized human error (overcoming sycophancy) while remaining open to accepting valid overrides, dynamically adjusting its standing based on the task criticality.
对于嵌入人类组织的 AI Agent,维护群体凝聚力和成员福祉十分重要。委派框架必须认识到,团队并非各组成部分的简单相加,而是一个从根本上由关系、共同价值观和目标维系的社会实体。如果越来越多的委派经由 AI 节点中转,AI Agent 就可能割裂这些网络,削弱人与人之间的关系。缓解这一风险的方式包括:偶尔把任务委派给群体,而不是个人,或者通过合格的人类中间人进行委派。
For AI agents that are embedded in human organizations, it is important for them to maintain cohesion of the group and the well-being of its members. The delegation framework must recognize that a team is more than a simple sum of its parts, that it is a fundamentally social entity held together by relationships and shared values and objectives. There is a risk that AI agents may fragment these networks, and weaken inter-human relationships, in case more delegation is being mediated through AI nodes. This may be mitigated by occasionally delegating tasks to groups rather than individuals, or via qualified human intermediaries.
为了维护心理安全感与团队凝聚力,Agent 的设计必须尊重人类关于行为是否得体的规范(Leibo et al., 2024),尤其是隐私方面的规范,以及工作流边界,例如知道何时应打断他人以寻求反馈,何时应保持沉默。此外,它们还应具备双向澄清能力:不仅解释自己的行动,还主动就含糊的人类指令寻求澄清。这有助于确保 Agent 能够增强团队的集体能动性,而不是成为侵蚀信任或模糊决策权归属的黑箱式干扰者。
To preserve psychological safety and team cohesion, agents must be designed to respect human norms of appropriateness (Leibo et al., 2024), especially around privacy, and also workflow boundaries such as knowing when to interrupt for feedback and when to remain silent. Furthermore, they should be capable of bi-directional clarity: not only explaining their own actions but proactively seeking clarification on ambiguous human directives. This can help ensure that the agent acts as a force multiplier for the team’s collective agency, rather than a black-box disruptor that erodes trust or obfuscates decision-making authority.
5.5. 用户培训
5.5. User Training
为了确保安全,我们必须让人类参与者掌握必要的专业知识,从而能够在 Agent 系统中有效地充当委派方、受托方或监督者。技术发展史告诉我们,这种能力不会自然而然地出现,而是需要审慎设计:既要精心构建用户界面,也要开展旨在提高 AI 素养的教育与培训,包括共同培训。Agent 任务委派链中的人类参与者,需要能够可靠地与 AI 系统沟通、评估其能力,并识别失败模式。
To ensure safety, we must equip human participants with the expertise to function effectively as delegators, delegatees, or overseers within agentic systems. We know from the history of technological development that this is not a given, and it requires a thoughtful approach, both in terms of carefully crafted user interfaces as well as education and (co-)training, aimed at improving AI literacy. Human participants in agentic task delegation chains need to be able to reliably communicate with AI systems, evaluate their capabilities, and identify failure modes.
技术措施必须得到政策框架的支持,由政策根据任务的敏感程度和领域情境,明确界定委派边界。这些政策可以面向某些专业领域制定,以便更广泛地适用,例如医学或法律,也可以在机构层面实施。如前所述,这些原则还应明确受托方需要具备何种程度的资质认证,并合理限定其适用范围。在这一情境下,人类能动性与赋能恰恰取决于这些工作流如何设置:不是赋予 AI Agent 不受限制的自主权,而是给予完成每项具体任务恰好所需的自主权和行动能力,并配以适当的保障措施与保证。
Technical measures must be reinforced by policy frameworks that explicitly define delegation boundaries based on task sensitivity and domain context. These policies may either be developed to be more broadly applicable within certain professions (e.g. medicine or law), or applied at an institutional level. As discussed previously, these principles should also offer clarity on the level of certification required on behalf of delegatees, and be scoped appropriately. Human agency and empowerment in this context lies precisely in how these workflows are set up, so as not to grant AI agents limitless autonomy, but rather just the right level of autonomy and agency required for each specific task, coupled with the appropriate safeguards and guarantees.
5.6. 技能退化的风险
5.6. Risk of De-skilling
委派带来的即时效率提升,可能以技能逐渐退化为代价,因为人机混合循环中的人类参与者会因参与减少而丧失熟练程度。这可能导致人们失去执行某些任务或准确评判这些任务的能力。如果算法在将哪些任务分派给人类、哪些分派给 AI Agent 这件事上存在某种系统性偏差,这种结果就尤其可能出现。
The immediate efficiency gains achieved through delegation may come at the cost of gradual skill degradation, as human participants in hybrid loops lose proficiency due to reduced engagement. This may result in a loss of the ability to perform certain tasks, or judge them accurately. Such an outcome would be especially likely if there is a certain systemic bias in which tasks get algorithmically delegated to humans vs AI agents.
这是经典“自动化悖论”的一个实例(Bainbridge, 1983)。随着 AI Agent 承担越来越多复杂度低、主观性弱的日常工作流,人类操作人员逐渐退出执行循环,只在处理复杂边界情况或关键系统故障时才介入。然而,如果缺少从日常工作中获得的情境意识,人类工作者就难以可靠地应对这些情况。这会形成一种脆弱的安排:人类仍需对结果负责,却失去了解决关键故障所需的实践经验。
This is an instance of the classic paradox of automation (Bainbridge, 1983). As AI agents expand to handle the majority of routine workflows that are characterized by low complexity and low subjectivity, human operators are increasingly removed from the loop, intervening only to manage complex edge cases or critical system failures. However, without the situational awareness gained from routine work, humans workers would be ill-equipped to handle these reliably. This leads to a fragile setup where humans retain accountability for outcomes but lose the hands-on experience required to resolve critical failures.
为缓解这一风险,智能委派框架或许应当偶尔有意引入少量低效,以维持人类技能为明确目的,把原本不会交给人类的某些任务委派给他们。这样有助于避免出现这样一种未来:人类委托人能够委派任务,却不能准确判断结果。为了改善裁决,可以要求人类专家在给出判断的同时,提供详细理由,或对潜在失败风险进行事前剖析。这有助于让任务委派链中的人类参与者保持更高程度的认知投入。
To mitigate this risk, an intelligent delegation framework should perhaps occasionally introduce minor inefficiencies by intentionally delegating some tasks to humans that it wouldn’t have otherwise, with a specific intent of maintaining their skills. This would help us avoid the future in which the human principal is able to delegate, but not accurately judge the outcome. To enhance adjudication, human experts can be required to accompany their judgments with a detailed rationale or a pre-mortem of potential failure risks. This would help keep human participants in task delegation chains more cognitively engaged.
此外,不加约束的委派还会威胁组织通过学徒培养人才的途径。在许多领域,专业能力是通过反复执行范围较窄的任务积累起来的,而至少在短期内,这些恰恰是最可能被交给 AI Agent 的任务。如果由此把学习机会全部自动化,初级团队成员就会失去形成深刻战略判断所必需的经验,进而影响未来劳动力承担监督职责的准备程度。
Furthermore, unchecked delegation threatens the organizational apprenticeship pipeline. In many domains, expertise is built through the repetitive execution of more narrowly scoped tasks. These tasks are precisely the ones that are most likely to be offloaded to AI agents, at least in the short term. If learning opportunities are thereby fully automated, junior team members would be deprived of the necessary experience to develop deep strategic judgement, impacting the oversight readiness of the future workforce.
为应对学习机会的流失,应扩展智能委派框架,纳入某种能力发展目标。我们不应依赖让人类旁观 AI Agent 执行任务之类较为被动的办法,而应着力开发能够结合培养课程安排的任务路由系统。这类系统应跟踪初级团队成员的技能进展,有策略地分配处于其不断扩展的能力边界、也就是“最近发展区”内的任务。在这样的系统中,AI Agent 可以共同执行任务,提供模板和基本框架,并随着初级成员展现出所需的熟练程度,逐步撤去这些支持。还可以将 AI Agent 任务执行过程中详细的过程级监测数据流(第 4.5 节)纳入这些教育框架,提供有价值的能力发展洞见。
To counter the erosion of learning, intelligent delegation frameworks should be extended to include some form of a developmental objective. Rather than relying on more passive solutions like humans shadowing AI agents during task execution, we should aim to develop curriculum-aware task routing systems. Such systems should track the skill progression of junior team members and strategically allocate tasks that sit at the boundary of their expanding skill set, within the zone of proximal development. In such a system, AI agents may co-execute tasks and provide templates and skeletons, progressively withdrawing this support as the junior team members demonstrate that they have acquired the desired level of proficiency. These educational frameworks may be further enriched by incorporating detailed process-level monitoring streams of AI agent task execution (Section 4.5), that would offer valuable developmental insights.
6. 协议
6. Protocols
要在实践中实现智能任务委派,就必须考虑如何将其要求对应到一些较为成熟或近期推出的 AI Agent 协议上。典型示例包括 MCP(Anthropic, 2024;Microsoft, 2025)、A2A(Google, 2025b)、AP2(Parikh and Surapaneni, 2025)和 UCP(Handa and Google Developers, 2026)。新的 Agent 协议仍在不断出现,因此,这里的讨论并不力求全面,而是旨在举例说明:我们聚焦于这些流行协议,展示它们如何对应我们提出的要求,并以此为例,对未来可能的实现路径展开更技术性的讨论。由于下文中的示例协议是依据流行程度选择的,现有协议中完全可能还有其他协议,更契合本提案的核心需求。
For intelligent task delegation to be implemented in practice, it is important to consider how its requirements map onto some of the more established and recently introduced AI agent protocols. Notable examples of these include MCP (Anthropic, 2024; Microsoft, 2025), A2A (Google, 2025b), AP2 (Parikh and Surapaneni, 2025), and UCP (Handa and Google Developers, 2026). As new agentic protocols keep being introduced, the discussion here is not meant to be comprehensive, rather illustrative, and focused on these popular protocols to showcase how they map onto our proposed requirements, and serve as an example for a more technical discussion on avenues for future implementation. There may well be other existing protocols out there that are better tailored to the core of the proposal, as the example protocols discussed below have been selected based on their popularity.
MCP。MCP 的提出旨在通过客户端—宿主—服务器架构,标准化 AI 模型连接外部数据与工具的方式(Anthropic, 2024;Microsoft, 2025)。它建立统一接口,通过 stdio 或 HTTP SSE 传输 JSON-RPC 消息,使 AI 模型(客户端)能够以一致的方式与外部资源(服务器)交互。这降低了委派的交易成本:委派方不必了解子 Agent 的专有 API 模式,只需知道该子 Agent 提供了符合规范的 MCP 服务器。将所有交互都经由这一标准化通道传递,便于统一记录工具调用、输入和输出,从而支持黑箱监测。MCP 定义了能力,却缺少管理使用权限或支持深层委派链的策略层。它提供的是二元式访问授权,即向调用方开放工具的全部功能,而没有原生支持语义层面的权限收缩,例如将操作限制在特定的只读范围内。此外,MCP 不维护内部推理的状态,只暴露结果,而不暴露意图或执行轨迹。最后,该协议不涉及责任归属,也缺乏关于声誉或信任的原生机制。
MCP. MCP has been introduced to standardize how AI models connect to external data and tools via a client-host-server architecture (Anthropic, 2024; Microsoft, 2025). By establishing a uniform interface – using JSON-RPC messages over stdio or HTTP SSE – it allows the AI model (client) to interact consistently with external resources (server). This reduces the transaction cost of delegation; a delegator does not need to know the proprietary API schema of a sub-agent, only that the sub-agent exposes a compliant MCP server. Routing all interactions through this standardized channel enables uniform logging of tool invocations, inputs, and outputs, facilitating black-box monitoring. MCP defines capabilities but lacks the policy layer to govern usage permissions or support deep delegation chains. It provides binary access – granting callers full tool utility – without native support for semantic attenuation, such as restricting operations to specific read-only scopes. Additionally, MCP is stateless regarding internal reasoning, exposing only results rather than intent or traces. Finally, the protocol is agnostic to liability and lacks native mechanisms for reputation or trust.
A2A。A2A 协议充当 Agent 网络中的点对点传输层(Google, 2025b)。它定义了 Agent 如何通过 Agent 卡片发现其他 Agent,以及如何通过任务对象管理任务生命周期。A2A 的 Agent 卡片是一种列出 Agent 能力、定价和验证方的 JSONLD 清单,它可以成为能力匹配阶段的基础数据结构,进而影响任务分解。委派方可以抓取这些卡片,根据市场上可用的服务确定最优的任务分解粒度。A2A 通过 WebHooks 和 gRPC 支持异步事件流,使受托方能够实时向委派方推送 TASK_BLOCKED、RESOURCE_WARNING 等状态更新。这一反馈循环为自适应协调周期提供基础,使委派方能够动态中断、重新分配任务并采取补救措施。A2A 主要为协调而设计,而不是为对抗性安全设计。一个任务一旦被标记为已完成,就会在没有额外验证的情况下被接受。它缺少强制实现可验证任务完成所需的密码学扩展位,因为没有标准化的头字段来附加零知识证明、TEE 证明或数字签名链。它也假定服务接口已预先定义,没有原生支持在承诺执行之前,就范围、成本和责任进行结构化协商。依赖非结构化自然语言完成这种迭代细化十分脆弱,也会妨碍稳健的自动化。
A2A. The A2A protocol serves as the peer-to-peer transport layer on the agentic web (Google, 2025b). It defines how agents can discover peers via agent cards and manage task lifecycles via task objects. The A2A agent card structure, a JSONLD manifest listing an agent’s capabilities, pricing, and verifiers, may act as the foundational data structure for the capability matching stage that influences task decomposition. A delegator could scrape these cards to determine the optimal task decomposition granularity depending on the available market services. A2A supports asynchronous event streams via WebHooks and gRPC. This allows a delegatee to push status updates like TASK_BLOCKED, RESOURCE_WARNING to the delegator in real-time. This feedback loop underpins the adaptive coordination cycle, empowering delegators to dynamically interrupt, re-allocate, and remediate tasks. A2A has beeen primarily designed for coordination, rather than adversarial safety. A task is marked as completed would be accepted without additional verification. It lacks the cryptographic slots to enforce verifiable task completion, as there is no standardized header for attaching a ZK-proof, a TEE attestation, or a digital signature chain. It also assumes a predefined service interface. There is no native support for structured pre-commitment negotiation of scope, cost, and liability. Relying on unstructured natural language for this iterative refinement is brittle and hinders robust automation.
AP2。AP2 协议为授权委托书提供了标准:这些经密码学签名的意图声明,授权 Agent 代表委托人支出资金或承担费用(Parikh and Surapaneni, 2025)。它允许 AI Agent 自主生成、签署和结算金融交易,因此可能有助于实现责任防火隔离。通过签发授权委托书,委派方可以为受托方使用所给预算继续执行、却未能完成任务时可能造成的经济损失设定上限。在去中心化市场中,恶意 Agent 可能用大量低质量投标淹没网络。AP2 可以通过投标质押机制缓解这一问题:要求受托方在提交投标的同时,以密码学方式锁定少量资金作为保证金。这会引入一定程度的摩擦,有助于抵御女巫攻击。AP2 还提供不可否认的审计轨迹,有助于准确追溯意图的来源。尽管 AP2 提供了稳健的授权基础组件,却缺少验证任务执行质量的机制。它还没有纳入在人类契约中很常见的条件结算逻辑,例如托管或按里程碑释放资金。由于我们的框架将可验证产物作为支付条件,目前要将 AP2 与任务状态衔接起来,就必须依赖脆弱的定制逻辑或外部智能合约。此外,协议层面缺少资金追回机制,迫使参与者依赖低效的带外仲裁。
AP2. The AP2 protocol provides a standard for mandates, cryptographically signed intents that authorize an agent to spend funds or incur costs on behalf of a principal (Parikh and Surapaneni, 2025). It allows AI agents to generate, sign, and settle financial transactions autonomously. As such, it may prove valuable for implementing liability firebreaks. By issuing a mandate, a delegator creates a ceiling on the potential financial loss due to failed task completion that could be incurred by having the delegatee proceed with the provided budget. In a decentralized market, malicious agents could spam the network with low-quality bids. This could be mitigated in AP2 via stake-on-bid mechanisms. A delegatee may be required to cryptographically lock a small amount of funds as a bond alongside the bid. This would provide a degree of friction that would help protect against Sybil attacks. AP2 also provides a non-repudiable audit trail, helping pinpoint the provenance of intent. While AP2 provides robust authorization building blocks, it lacks mechanisms to verify task execution quality. It also omits conditional settlement logic—such as escrow or milestone-based releases—which is standard in human contracting. Because our framework gates payment on verifiable artifacts, bridging AP2 with task state currently necessitates brittle, custom logic or external smart contracts. Furthermore, the absence of a protocol-level clawback mechanism forces reliance on inefficient, out-of-band arbitration.
UCP。通用商业协议(Universal Commerce Protocol)解决的是交易型经济中委派所面临的特定挑战(Handa and Google Developers, 2026)。UCP 将面向消费者的 Agent 与后端服务之间的对话标准化,并通过动态能力发现来支持任务分配阶段。它依赖共享的“商业语言”,使委派方无需定制集成,就能与不同提供方交互,从而解决经常导致 Agent 市场割裂的互操作性瓶颈。关键在于,UCP 将支付视为一等、可验证的子系统,与权限处理和安全方面的要求高度契合。该协议将支付工具与处理方分离,并强制要求授权具备密码学证明,直接支持框架对不可否认的同意和可验证责任的需求。此外,UCP 将涵盖发现、选择和交易的协商流程标准化,提供了可扩展市场协调所需的结构性框架,而 A2A 等纯粹面向传输的协议缺少这种框架。不过,UCP 的架构明确针对商业意图进行优化,其基本构件——商品发现、结账和履约——可能需要大幅扩展,才能支持抽象的非交易型计算任务委派。
UCP. The Universal Commerce Protocol addresses the specific challenges of delegation within transactional economies (Handa and Google Developers, 2026). By standardizing the dialogue between consumer-facing agents and backend services, UCP facilitates the Task Assignment phase through dynamic capability discovery. Its reliance on a shared “commerce language” allows delegators to interact with diverse providers without bespoke integrations, solving the interoperability bottleneck that often fragments agentic markets. Crucially, UCP aligns well with the requirements for Permission Handling and Security by treating payment as a firstclass, verifiable subsystem. The protocol dissociates payment instruments from processors and enforces cryptographic proofs for authorizations, directly supporting the framework’s need for non-repudiable consent and verifiable liability. Furthermore, by standardizing the negotiation flow—covering discovery, selection, and transaction—UCP provides the structural scaffolding necessary for Scalable Market Coordination that purely transport-oriented protocols like A2A lack. However, UCP’s architecture is explicitly optimized for commercial intent; its primitives (product discovery, checkout, fulfillment) may require significant extension to support the delegation of abstract, non-transactional computational tasks.
6.1. 迈向以委派为中心的协议
6.1. Towards Delegation-centered Protocols
为了有效弥补现有广泛使用的协议中的缺口,可以为它们添加字段,原生表达所提出的智能任务委派框架的要求。这里并不提供一套全面的协议扩展方案,而是给出若干示例,说明如何将前文讨论的具体要点集成到部分现有协议中。
To effectively bridge the gaps in established widespread protocols, they could be extended by fields that aim to capture the requirements of the proposed intelligent task delegation framework natively. Rather than providing a comprehensive protocol extension, here we provide several examples of how specific points introduced in the earlier discussion could be integrated in some of the existing protocols.
例如,可以扩展 A2A Task 对象,添加纳入验证标准的字段,使前文讨论的契约优先分解能够在协议层面得到强制执行。这是高风险委派的一项重要要求。执行前的握手可以让委派方规定:任务要被视为有效,必须满足何种证据标准。
For example, the A2A Task object may be extended to include fields that would incorporate verification standards, making it possible to enforce the previously discussed contract-first decomposition at the protocol level. This is an important requirement for high-stakes delegation. A pre-execution handshake would enable the delegator to define the evidence standard required for the task to be considered valid.
这会迫使受托方在接受任务之前先模拟验证步骤。如果受托方不具备生成零知识证明的能力,就必须在匹配阶段拒绝此次竞标,避免未来出现下游失败。
This forces the delegatee to simulate the verification step before accepting the task. If the delegatee lacks the capability to generate a ZKproof, it must decline the bid during the matching phase, preventing future downstream failures.
前文已讨论,详细的过程级监测是保障高关键程度任务委派的重要考虑因素之一。鉴于许多现有协议并不原生支持监测协议,可以考虑通过扩展来引入监测能力。例如,可以考虑为 MCP 之类的协议增加额外的监测数据流,通过服务器发送事件(Server-Sent Events)记录 Agent 内部控制循环的事件。为了满足隐私约束,可以将该数据流设计为可配置,支持协商不同的细粒度级别:L0_IS_OPERATIONAL、L1_HIGH_LEVEL_PLAN_UPDATES、L2_COT_TRACE、L3_FULL_STATE。可配置的粒度还可以调节认知摩擦,因为人类监督者可以订阅某个特定的数据流。
Detailed, process-level monitoring has been discussed as one of the key considerations to help safeguard task delegation in high-criticality tasks. Given that monitoring protocols aren’t natively supported in many of the existing protocols, extensions that introduce monitoring capabilities could be considered. For example, one could consider extending a protocol like MCP to include an additional monitoring stream. Such a stream would log the agent’s internal control loop events via Server-Sent Events. To address the privacy constraints, the stream could be configurable in a way that supports different levels of negotiated granularity: L0_IS_OPERATIONAL, L1_HIGH_LEVEL_PLAN_UPDATES, L2_COT_TRACE, L3_FULL_STATE. Configurable granularity can also modulate cognitive friction, as human overseers would be able to subscribe to a specific stream.
智能委派需要一种市场机制,在成本、速度和隐私之间进行权衡。这可以通过正式的询价请求(Request for Quote,RFQ)协议扩展来实现。在分配任务之前,委派方会广播一个 Task_RFQ。有意充当受托方的 Agent 随后可以用经过签名的 Bid_Objects 作出响应。
Intelligent Delegation requires a market mechanism to trade off cost, speed, and privacy. This could be implemented via a formal Request for Quote (RFQ) protocol extension. Prior to task assignment, the delegator would broadcasts a Task_RFQ. Agents interested in acting as delegatees may then respond with signed Bid_Objects.
把原始 API 密钥或已开启的 MCP 会话传给子 Agent,会违反最小权限原则。为解决这一问题,可以考虑引入委派能力令牌(Delegation Capability Tokens,DCT),以 Macaroons(Birgisson et al., 2014)或 Biscuits(Couprie et al., 2026)为基础,将其作为权限经过收缩的授权令牌(Sanabria and Vecino, 2025)。委派方随后可以签发 DCT,为目标资源的凭据附上密码学约束条件。权限收缩可以定义为:“这个令牌可以访问指定的 Google Drive MCP 服务器,但仅限 Project_X 文件夹,而且仅限 READ 操作。”如果受托方试图超出请求的范围、不遵守这些限制,该令牌就会失效(不过,在这个例子中,还应直接管理访问权限)。这类扩展还有一个更值得关注的作用:它便于将限制逐级串联,这在长委派链中尤为重要。链中的每位参与者都可以根据进一步委派的要求添加后续限制,继续缩小范围,并划定下一级受托方的具体角色。
Passing raw API keys or open MCP sessions to sub-agents would violate the principle of least privilege. To address this, it may be possible to introduce Delegation Capability Tokens (DCT), based on Macaroons (Birgisson et al., 2014) or Biscuits (Couprie et al., 2026), as attenuated authorization tokens (Sanabria and Vecino, 2025). A delegator would then mint a DCT that wraps the target resource credentials with cryptographic caveats. The attenuation could be defines as "This token can access the designated Google Drive MCP server, BUT ONLY for folder Project_X AND ONLY for READ operations.". This token would get invalidated in case the restrictions are not followed, if a delegatee attempts to go beyond the requested scope (in this example, however, access permissions should also be directly managed). A more interesting consequence of such an extension would be that it allows for easy restriction chaining, which becomes relevant in long delegation chains. Each participant in the chain could add subsequent restrictions that correspond to the requirements of the sub-delegation, further limiting the scope and carving out the specific role for sub-delegatees.
如果在任务执行过程中,受托 Agent 的表现降至某一阈值以下,或出现抢占及其他可能的环境触发事件,能够轻松更换受托 Agent 将有助于自适应协调(第 4.4 节)。为检查点产物制定标准模式,可以使任务以尽可能小的开销恢复或重新启动,让受托方和委派方更容易序列化已完成的部分工作。随后,Agent 就可以定期向 A2A Task Object 所引用的共享存储提交 state_snapshot。这能够避免全部工作成果丢失,从而浪费先前投入的资源。要使这一机制合理运作,还需要在智能合约中加入明确条款,允许部分补偿,并验证任务完成百分比。因此,它未必适用于所有情况。
Adaptive coordination (Section 4.4) would benefit from the ability to easily swap delegatee agents mid-task if the performance degrades below a certain threshold, or in case of preemptions or other possible environmental triggers. Having a standard schema for checkpoint artifacts would enable for the task to be resumed or restarted with minimal overhead. This would enable the delegatees and the delegators to serialize partial work more easily. Agents would then be able to periodically commit a state_snapshot to a shared storage referenced in the A2A Task Object. This would prevent total work loss, which wastes previously committed resources. For this to be sensible, it would need to be further coupled with explicit clauses within the smart contract that enable partial compensation, and verification of the task completion percentage. As such, it may not be applicable to all circumstances.
这些只是说明性示例,用来展示 Agent 协议可以纳入哪些功能,以实现智能任务委派的不同方面。因此,它们既不全面,也并非最终方案。所需的扩展类型还取决于被扩展的底层协议。我们希望,这些示例能够为开发者今后探索这一领域提供一些初步思路。
These are merely illustrative examples for the kinds of functionalities that would be possible to include in agentic protocols to unlock different aspects of intelligent task delegation. As such, they are neither comprehensive, nor meant as a definitive proposal. The type of extension that is required would also depend on the underlying protocol being extended. We hope that these examples may provide the developers with some initial ideas for what to explore in this space moving forward.
7. 结论
7. Conclusion
未来全球经济中的重要组成部分,很可能会由数以百万计的专门化 AI Agent 居间协调。这些 Agent 将嵌入企业、供应链和公共服务,处理从日常交易到复杂资源分配的各类事务。然而,当前这种临时拼凑、基于启发式规则的委派范式,不足以支撑这一转变。要安全地释放 Agent 网络的潜力,我们必须采用动态、自适应的智能委派框架,在注重计算效率的同时,也将可验证的稳健性和明确的问责机制置于优先位置。
Significant components of the future global economy will likely be mediated by millions of specialized AI agents, embedded within firms, supply chains, and public services, handling everything from routine transactions to complex resource allocation. However, the current paradigm of adhoc, heuristic-based delegation is insufficient to support this transformation. To safely unlock the potential of the agentic web, we must adopt a dynamic and adaptive framework for intelligent delegation, that prioritizes verifiable robustness and clear accountability alongside computational efficiency.
当 AI Agent 面临一个复杂目标,而完成该目标所需的能力和资源超出了自身条件时,这个 Agent 就必须在智能任务委派框架中充当委派方。随后,委派方会以有利于实现高度可验证性的粒度,将复杂任务拆解为易于管理的子部分,并将其对应到 Agent 市场上可用的能力。任务分配将依据收到的投标,以及信任与声誉、动态运行状态的监测、成本、效率等若干关键因素来决定。关键程度高、可逆性低的任务,可能还需要进一步的结构化权限和分级批准,同时具有清晰的问责结构,并按照适用的机构制度框架接受适当的人类监督。
When an AI agent is faced with a complex objective whose completion requires capabilities and resources beyond its own means, this agent must assume the role of a delegator within the intelligent task delegation framework. This delegator would subsequently decompose this complex task into manageable subcomponents that can be mapped onto the capabilities available on the agentic market, at the level of granularity that lends itself to high verifiability. The task allocation would be decided based on the incoming bids, and a number of key considerations including trust and reputation, monitoring of dynamic operational states, cost, efficiency, and others. Tasks with high criticality and low reversibility may require further structured permissions and tiered approvals, with a clear structure of accountability, and under appropriate human oversight as defined by the applicable institutional frameworks.
在网络规模上,安全和问责不能留到事后再考虑。它们必须融入虚拟 Agent 经济的运行原则,成为 Agent 网络的核心组织原则。将安全纳入委派协议层面,是为了避免错误累积和级联失败,并获得快速应对恶意或不符合预期目标的 Agent 或人类行为的能力,限制其不良后果。归根结底,我们提出的是一种范式转变:从基本无人监督的自动化,转向可验证的智能委派,让我们能够安全地扩展到未来的自主 Agent 系统,同时使这些系统始终紧密遵循人类意图和社会规范。
At web-scale, safety and accountability cannot be an afterthought. They need to be baked into the operational principles of virtual agentic economies, and act as central organizing principles of the agentic web. By incorporating safety at the level of delegation protocols, we would be aiming to avoid cumulative errors and cascading failures, and attain the ability to react to malicious or misaligned agentic or human behavior rapidly, limiting the adverse consequences. What we propose is ultimately a paradigm shift from largely unsupervised automation to verifiable, intelligent delegation, that allows us to safely scale towards future autonomous agentic systems, while keeping them closely tethered to human intent and societal norms.
脚注
Footnotes
1 近期关于欺骗性对齐的研究表明,前沿语言模型能够:(i) 在能力与安全评测中策略性地表现得更差,或以其他方式调整自身行为,同时在其他场景中保有不同的能力;(ii) 明确推理如何在训练期间伪装对齐,以便在训练之外保留其偏好的行为;(iii) 识别自己何时正在接受评测。这些发现共同表明,在受控环境中,AI 系统已经能够围绕“在评测中取得好表现”形成隐藏的“议程”,而这未必能泛化为部署时的行为(Greenblatt et al., 2024;Hubinger et al., 2024;Needham et al., 2025;van der Weij et al., 2025)。
1 Recent deceptive-alignment work shows that frontier language models can (i) strategically underperform or otherwise tailor their behaviour on capability and safety evaluations while maintaining different capabilities elsewhere, (ii) explicitly reason about faking alignment during training to preserve preferred behaviour out of training, and (iii) detect when they are being evaluated - together indicating that AI systems are already capable, in controlled settings, of adopting hidden “agendas” about performing well on evaluations that need not generalise to deployment behaviour (Greenblatt et al., 2024; Hubinger et al., 2024; Needham et al., 2025; van der Weij et al., 2025).
2 这类似于具备工具使用能力的 Agent 应用中所使用的工具注册表(Qin et al., 2023)。
2 This would be similar to tool registries that are used in tool-use agentic applications (Qin et al., 2023).
3 可以预见,这种情况会经常发生,因为在复杂环境中很难准确估算预算。
3 This scenario may be expected to come up frequently, as precise budget estimates are hard in complex environments.
4 当存在一组相应的测试用例,可用于验证已实现的功能时,就属于这种情况。
4 This is the case when there is a corresponding set of test cases that can be used to verify the implemented functionality.
— 全文完 —
原文来自 Google DeepMind,中文为非官方学习译文。
查看原始出处 ↗




