以大语言模型(LLM)为核心控制器来构建 Agent,是一个很酷的概念。AutoGPT、GPT-Engineer 和 BabyAGI 等概念验证演示,提供了很有启发性的例子。大语言模型的潜力不止于生成优秀的文案、故事、文章和程序;它也可以被构建成一个强大的通用问题求解器。
Building agents with LLM (large language model) as its core controller is a cool concept. Several proof-of-concepts demos, such as AutoGPT, GPT-Engineer and BabyAGI, serve as inspiring examples. The potentiality of LLM extends beyond generating well-written copies, stories, essays and programs; it can be framed as a powerful general problem solver.
Agent 系统概览
Agent System Overview
在由大语言模型驱动的自主 Agent 系统中,LLM 充当 Agent 的大脑,并由几个关键组件配合:
In a LLM-powered autonomous agent system, LLM functions as the agent’s brain, complemented by several key components:
- 规划
- 子目标与分解:Agent 将大任务拆解成更小、可管理的子目标,从而高效处理复杂任务。
- 反思与改进:Agent 可以对过去的行动进行自我批评和反思,从错误中学习,并改进后续步骤,以提高最终结果的质量。
- 记忆
- 短期记忆:我倾向于将所有上下文学习(参见《提示工程》)都视为利用模型的短期记忆进行学习。
- 长期记忆:让 Agent 能在较长时间内保留和回忆近乎无限的信息,通常借助外部向量存储和快速检索实现。
- 工具使用
- Agent 学习调用外部 API,获取模型权重中缺失的信息;这些权重通常在预训练后很难修改。外部能力包括最新信息、代码执行、访问专有信息源等。
- Planning
- Subgoal and decomposition: The agent breaks down large tasks into smaller, manageable subgoals, enabling efficient handling of complex tasks.
- Reflection and refinement: The agent can do self-criticism and self-reflection over past actions, learn from mistakes and refine them for future steps, thereby improving the quality of final results.
- Memory
- Short-term memory: I would consider all the in-context learning (See Prompt Engineering) as utilizing short-term memory of the model to learn.
- Long-term memory: This provides the agent with the capability to retain and recall (infinite) information over extended periods, often by leveraging an external vector store and fast retrieval.
- Tool use
- The agent learns to call external APIs for extra information that is missing from the model weights (often hard to change after pre-training), including current information, code execution capability, access to proprietary information sources and more.
组件一:规划
Component One: Planning
复杂任务通常涉及许多步骤。Agent 需要知道有哪些步骤,并提前规划。
A complicated task usually involves many steps. An agent needs to know what they are and plan ahead.
任务分解
Task Decomposition
思维链(CoT;Wei 等,2022)已经成为提升模型复杂任务表现的标准提示技术。通过指示模型“一步步思考”,投入更多推理时计算,将困难任务拆成更小、更简单的步骤。CoT 将大任务转化为多个可管理任务,也为解释模型的思考过程提供了线索。
Chain of thought (CoT; Wei et al. 2022) has become a standard prompting technique for enhancing model performance on complex tasks. The model is instructed to “think step by step” to utilize more test-time computation to decompose hard tasks into smaller and simpler steps. CoT transforms big tasks into multiple manageable tasks and shed lights into an interpretation of the model’s thinking process.
思维树(Tree of Thoughts;Yao 等,2023)通过在每一步探索多种推理可能性,扩展了 CoT。它先将问题分成多个思考步骤,每步生成多个想法,形成树状结构。搜索过程可以采用广度优先搜索(BFS)或深度优先搜索(DFS),每个状态由分类器通过提示词进行评估,或通过多数投票评估。
Tree of Thoughts (Yao et al. 2023) extends CoT by exploring multiple reasoning possibilities at each step. It first decomposes the problem into multiple thought steps and generates multiple thoughts per step, creating a tree structure. The search process can be BFS (breadth-first search) or DFS (depth-first search) with each state evaluated by a classifier (via a prompt) or majority vote.
任务分解可以通过以下方式完成:(1)对 LLM 使用简单提示,例如 "Steps for XYZ.\n1."、"What are the subgoals for achieving XYZ?";(2)使用任务特定指令,例如写小说时要求 "Write a story outline.";或(3)由人类提供输入。
Task decomposition can be done (1) by LLM with simple prompting like "Steps for XYZ.\n1.", "What are the subgoals for achieving XYZ?", (2) by using task-specific instructions; e.g. "Write a story outline." for writing a novel, or (3) with human inputs.
另一种颇为不同的方法是 LLM+P(Liu 等,2023),依靠外部经典规划器执行长程规划。它使用规划领域定义语言(PDDL)作为描述规划问题的中间接口。在这一过程中,LLM(1)首先将问题转换为“Problem PDDL”,(2)然后请求经典规划器基于已有的“Domain PDDL”生成 PDDL 计划,(3)最后再把计划翻译回自然语言。本质上,它把规划步骤外包给外部工具,并假定领域特定的 PDDL 和适用规划器已经存在。这在某些机器人场景中很常见,但在许多其他领域并非如此。
Another quite distinct approach, LLM+P (Liu et al. 2023), involves relying on an external classical planner to do long-horizon planning. This approach utilizes the Planning Domain Definition Language (PDDL) as an intermediate interface to describe the planning problem. In this process, LLM (1) translates the problem into “Problem PDDL”, then (2) requests a classical planner to generate a PDDL plan based on an existing “Domain PDDL”, and finally (3) translates the PDDL plan back into natural language. Essentially, the planning step is outsourced to an external tool, assuming the availability of domain-specific PDDL and a suitable planner which is common in certain robotic setups but not in many other domains.
自我反思
Self-Reflection
自我反思是一个重要方面,让自主 Agent 能通过改进过去的行动决策、纠正先前错误来迭代提升。在不可避免需要试错的真实任务中,它至关重要。
Self-reflection is a vital aspect that allows autonomous agents to improve iteratively by refining past action decisions and correcting previous mistakes. It plays a crucial role in real-world tasks where trial and error are inevitable.
ReAct(Yao 等,2023)将动作空间扩展为任务特定的离散动作与语言空间的组合,从而在 LLM 内整合推理与行动。前者让 LLM 与环境交互,例如使用维基百科搜索 API;后者则引导 LLM 用自然语言生成推理轨迹。
ReAct (Yao et al. 2023) integrates reasoning and acting within LLM by extending the action space to be a combination of task-specific discrete actions and the language space. The former enables LLM to interact with the environment (e.g. use Wikipedia search API), while the latter prompting LLM to generate reasoning traces in natural language.
ReAct 提示模板为 LLM 加入了明确的思考步骤,大致格式如下:
The ReAct prompt template incorporates explicit steps for LLM to think, roughly formatted as:
在知识密集型任务和决策任务的实验中,ReAct 都优于只保留 Act、移除了 Thought: … 步骤的基线。
In both experiments on knowledge-intensive tasks and decision-making tasks, ReAct works better than the Act-only baseline where Thought: … step is removed.
Reflexion(Shinn 与 Labash,2023)是一种为 Agent 配备动态记忆与自我反思能力、以提升推理能力的框架。Reflexion 采用标准强化学习设置:奖励模型提供简单的二元奖励,动作空间沿用 ReAct 的设置,将语言加入任务特定动作空间,以支持复杂推理步骤。每执行一个动作 ,Agent 都会计算一个启发式量 ,并可能根据自我反思结果,决定重置环境,开始新的尝试。
Reflexion (Shinn & Labash 2023) is a framework to equip agents with dynamic memory and self-reflection capabilities to improve reasoning skills. Reflexion has a standard RL setup, in which the reward model provides a simple binary reward and the action space follows the setup in ReAct where the task-specific action space is augmented with language to enable complex reasoning steps. After each action , the agent computes a heuristic and optionally may decide to reset the environment to start a new trial depending on the self-reflection results.
启发式函数用于判断执行轨迹何时效率低下或包含幻觉,需要停止。低效规划指持续过久却没有成功的轨迹;幻觉则定义为连续执行一串相同动作,并在环境中得到相同观察结果。
The heuristic function determines when the trajectory is inefficient or contains hallucination and should be stopped. Inefficient planning refers to trajectories that take too long without success. Hallucination is defined as encountering a sequence of consecutive identical actions that lead to the same observation in the environment.
自我反思通过向 LLM 展示两个示例生成;每个示例都是一对“失败轨迹,以及用于指导后续计划调整的理想反思”。随后,这些反思会加入 Agent 的工作记忆,最多保留三条,作为调用 LLM 时的上下文。
Self-reflection is created by showing two-shot examples to LLM and each example is a pair of (failed trajectory, ideal reflection for guiding future changes in the plan). Then reflections are added into the agent’s working memory, up to three, to be used as context for querying LLM.
事后回顾链(Chain of Hindsight,CoH;Liu 等,2023)通过明确展示一连串附有反馈的历史输出,鼓励模型改进自己的结果。人类反馈数据是集合 ,其中 是提示词,每个 是模型生成结果, 是人类对 的评分, 是相应的人类事后反馈。假设反馈元组按奖励排序,。训练采用监督微调,数据组织成 这样的序列,其中 。模型在给定序列前缀的条件下,只学习预测 ,从而能够根据反馈序列自我反思,生成更好的输出。推理时,也可以选择让模型与人类标注者进行多轮指令交互。
Chain of Hindsight (CoH; Liu et al. 2023) encourages the model to improve on its own outputs by explicitly presenting it with a sequence of past outputs, each annotated with feedback. Human feedback data is a collection of , where is the prompt, each is a model completion, is the human rating of , and is the corresponding human-provided hindsight feedback. Assume the feedback tuples are ranked by reward, The process is supervised fine-tuning where the data is a sequence in the form of , where . The model is finetuned to only predict where conditioned on the sequence prefix, such that the model can self-reflect to produce better output based on the feedback sequence. The model can optionally receive multiple rounds of instructions with human annotators at test time.
为避免过拟合,CoH 添加了一个正则项,最大化预训练数据集的对数似然。由于反馈序列中包含许多相同词语,为避免模型走捷径和抄写,他们在训练时随机遮蔽历史 token 的 0% 至 5%。
To avoid overfitting, CoH adds a regularization term to maximize the log-likelihood of the pre-training dataset. To avoid shortcutting and copying (because there are many common words in feedback sequences), they randomly mask 0% - 5% of past tokens during training.
实验中的训练数据集,结合了 WebGPT 比较数据、基于人类反馈的摘要数据和人类偏好数据集。
The training dataset in their experiments is a combination of WebGPT comparisons, summarization from human feedback and human preference dataset.
CoH 的思路,是在上下文中展示一段输出逐步改进的历史,训练模型顺着这一趋势继续产生更好的结果。算法蒸馏(Algorithm Distillation,AD;Laskin 等,2023)将同样的思路用于强化学习任务中跨多个回合的轨迹,将一种算法封装在以长历史为条件的策略中。设想 Agent 多次与环境交互,并在每个回合有所进步:AD 将这些学习历史拼接起来,输入模型。因此,我们期望模型预测的下一个动作,能带来优于此前尝试的表现。目标是学习强化学习的过程,而不是训练某个特定任务的策略本身。
The idea of CoH is to present a history of sequentially improved outputs in context and train the model to take on the trend to produce better outputs. Algorithm Distillation (AD; Laskin et al. 2023) applies the same idea to cross-episode trajectories in reinforcement learning tasks, where an algorithm is encapsulated in a long history-conditioned policy. Considering that an agent interacts with the environment many times and in each episode the agent gets a little better, AD concatenates this learning history and feeds that into the model. Hence we should expect the next predicted action to lead to better performance than previous trials. The goal is to learn the process of RL instead of training a task-specific policy itself.
论文假设,任何能够生成一组学习历史的算法,都可以通过对动作进行行为克隆,蒸馏进神经网络。历史数据由一组源策略生成,每个源策略都针对特定任务训练。在训练阶段,每次强化学习运行都会随机采样一个任务,并使用跨多个回合历史中的一段子序列进行训练,从而使学到的策略不依赖具体任务。
The paper hypothesizes that any algorithm that generates a set of learning histories can be distilled into a neural network by performing behavioral cloning over actions. The history data is generated by a set of source policies, each trained for a specific task. At the training stage, during each RL run, a random task is sampled and a subsequence of multi-episode history is used for training, such that the learned policy is task-agnostic.
实际中,模型上下文窗口长度有限,因此单个回合必须足够短,才能构造多回合历史。要学到接近最优的上下文内强化学习算法,需要包含 2 至 4 个回合的上下文。上下文内强化学习的涌现,需要足够长的上下文。
In reality, the model has limited context window length, so episodes should be short enough to construct multi-episode history. Multi-episodic contexts of 2-4 episodes are necessary to learn a near-optimal in-context RL algorithm. The emergence of in-context RL requires long enough context.
相比三个基线——ED(专家蒸馏,使用专家轨迹而非学习历史进行行为克隆)、源策略(通过 UCB 生成用于蒸馏的轨迹),以及 RL^2(Duan 等,2017;由于需要在线强化学习,被用作性能上界)——AD 仅使用离线强化学习,就表现出了接近 RL^2 水平的上下文内强化学习能力,学习速度也远高于其他基线。当以源策略的部分训练历史为条件时,AD 的改进速度同样显著快于 ED 基线。
In comparison with three baselines, including ED (expert distillation, behavior cloning with expert trajectories instead of learning history), source policy (used for generating trajectories for distillation by UCB), RL^2 (Duan et al. 2017; used as upper bound since it needs online RL), AD demonstrates in-context RL with performance getting close to RL^2 despite only using offline RL and learns much faster than other baselines. When conditioned on partial training history of the source policy, AD also improves much faster than ED baseline.
组件二:记忆
Component Two: Memory
(非常感谢 ChatGPT 帮我起草这一节。在与 ChatGPT 的对话中,我学到了很多关于人脑,以及支持快速 MIPS 的数据结构的知识。)
(Big thank you to ChatGPT for helping me draft this section. I’ve learned a lot about the human brain and data structure for fast MIPS in my conversations with ChatGPT.)
记忆的类型
Types of Memory
记忆可以定义为获取、存储、保留,以及随后检索信息的过程。人脑中存在几种不同的记忆类型。
Memory can be defined as the processes used to acquire, store, retain, and later retrieve information. There are several types of memory in human brains.
感觉记忆:这是记忆最早的阶段,让原始刺激结束后,视觉、听觉等感觉信息留下的印象仍能暂时保留。感觉记忆通常只持续几秒,子类型包括图像记忆(视觉)、回声记忆(听觉)和触觉记忆。
短期记忆(STM)或工作记忆:存储我们当前意识到的、执行学习与推理等复杂认知任务所需的信息。一般认为,短期记忆的容量约为 7 个项目(Miller,1956),持续时间约为 20 至 30 秒。
长期记忆(LTM):可将信息存储极长时间,从几天到数十年不等,存储容量基本不受限。长期记忆分为两种子类型:
- 外显记忆/陈述性记忆:关于事实和事件、能够有意识回忆的记忆,包括情景记忆(事件与经历)和语义记忆(事实与概念)。
- 内隐记忆/程序性记忆:无意识的记忆,涉及自动执行的技能和习惯,例如骑自行车或用键盘打字。
-
Sensory Memory: This is the earliest stage of memory, providing the ability to retain impressions of sensory information (visual, auditory, etc) after the original stimuli have ended. Sensory memory typically only lasts for up to a few seconds. Subcategories include iconic memory (visual), echoic memory (auditory), and haptic memory (touch).
-
Short-Term Memory (STM) or Working Memory: It stores information that we are currently aware of and needed to carry out complex cognitive tasks such as learning and reasoning. Short-term memory is believed to have the capacity of about 7 items (Miller 1956) and lasts for 20-30 seconds.
-
Long-Term Memory (LTM): Long-term memory can store information for a remarkably long time, ranging from a few days to decades, with an essentially unlimited storage capacity. There are two subtypes of LTM:
- Explicit / declarative memory: This is memory of facts and events, and refers to those memories that can be consciously recalled, including episodic memory (events and experiences) and semantic memory (facts and concepts).
- Implicit / procedural memory: This type of memory is unconscious and involves skills and routines that are performed automatically, like riding a bike or typing on a keyboard.
我们可以大致建立如下对应关系:
We can roughly consider the following mappings:
- 感觉记忆对应于为文本、图像或其他模态的原始输入学习嵌入表示。
- 短期记忆对应于上下文学习。由于 Transformer 的上下文窗口长度有限,它是短暂且有限的。
- 长期记忆对应于 Agent 在查询时能够借助快速检索访问的外部向量存储。
- Sensory memory as learning embedding representations for raw inputs, including text, image or other modalities;
- Short-term memory as in-context learning. It is short and finite, as it is restricted by the finite context window length of Transformer.
- Long-term memory as the external vector store that the agent can attend to at query time, accessible via fast retrieval.
最大内积搜索(MIPS)
Maximum Inner Product Search (MIPS)
外部记忆可以缓解有限注意力范围的约束。标准做法是将信息的嵌入表示保存到支持快速最大内积搜索(MIPS)的向量数据库中。为提高检索速度,通常会使用近似最近邻(ANN)算法,返回近似的前 k 个最近邻,以少量精度损失换取大幅提速。
The external memory can alleviate the restriction of finite attention span. A standard practice is to save the embedding representation of information into a vector store database that can support fast maximum inner-product search (MIPS). To optimize the retrieval speed, the common choice is the approximate nearest neighbors (ANN) algorithm to return approximately top k nearest neighbors to trade off a little accuracy lost for a huge speedup.
用于快速 MIPS 的常见 ANN 算法包括:
A couple common choices of ANN algorithms for fast MIPS:
- LSH(局部敏感哈希):引入一种哈希函数,使相似输入项以较高概率映射到同一个桶中,而桶的数量远少于输入项数量。
- ANNOY(Approximate Nearest Neighbors Oh Yeah):核心数据结构是随机投影树,即一组二叉树;每个非叶节点代表一个将输入空间一分为二的超平面,每个叶节点存储一个数据点。各棵树独立、随机地构建,因此在某种程度上模仿了哈希函数。ANNOY 会在所有树中搜索,逐步进入更接近查询的那一半空间,再汇总结果。它的思路与 KD 树密切相关,但可扩展性要高得多。
- HNSW(分层可导航小世界):灵感来自小世界网络,其中大多数节点都可以从其他任意节点经过很少步数到达,例如社交网络的“六度分隔”。HNSW 将这些小世界图组织为分层结构,底层包含实际数据点,中间层建立捷径以加快搜索。搜索从顶层随机节点开始,朝目标移动;无法进一步接近时,就下降一层,直到抵达底层。高层每一步都可能跨越数据空间中的较大距离,低层每一步则细化搜索质量。
- FAISS(Facebook AI Similarity Search):其工作假设是,高维空间中节点之间的距离服从高斯分布,因此数据点应当存在聚类。FAISS 将向量空间划分为多个簇,实施向量量化,再在簇内细化量化。搜索首先通过粗量化找出候选簇,再通过更细的量化深入各个簇。
- ScaNN(可扩展最近邻):主要创新是各向异性向量量化。它将数据点 量化为 ,使内积 与 的原始距离尽可能接近,而不是选择最近的量化质心点。
- LSH (Locality-Sensitive Hashing): It introduces a hashing function such that similar input items are mapped to the same buckets with high probability, where the number of buckets is much smaller than the number of inputs.
- ANNOY (Approximate Nearest Neighbors Oh Yeah): The core data structure are random projection trees, a set of binary trees where each non-leaf node represents a hyperplane splitting the input space into half and each leaf stores one data point. Trees are built independently and at random, so to some extent, it mimics a hashing function. ANNOY search happens in all the trees to iteratively search through the half that is closest to the query and then aggregates the results. The idea is quite related to KD tree but a lot more scalable.
- HNSW (Hierarchical Navigable Small World): It is inspired by the idea of small world networks where most nodes can be reached by any other nodes within a small number of steps; e.g. “six degrees of separation” feature of social networks. HNSW builds hierarchical layers of these small-world graphs, where the bottom layers contain the actual data points. The layers in the middle create shortcuts to speed up search. When performing a search, HNSW starts from a random node in the top layer and navigates towards the target. When it can’t get any closer, it moves down to the next layer, until it reaches the bottom layer. Each move in the upper layers can potentially cover a large distance in the data space, and each move in the lower layers refines the search quality.
- FAISS (Facebook AI Similarity Search): It operates on the assumption that in high dimensional space, distances between nodes follow a Gaussian distribution and thus there should exist clustering of data points. FAISS applies vector quantization by partitioning the vector space into clusters and then refining the quantization within clusters. Search first looks for cluster candidates with coarse quantization and then further looks into each cluster with finer quantization.
- ScaNN (Scalable Nearest Neighbors): The main innovation in ScaNN is anisotropic vector quantization. It quantizes a data point to such that the inner product is as similar to the original distance of as possible, instead of picking the closet quantization centroid points.
更多 MIPS 算法及性能比较,见 ann-benchmarks.com。
Check more MIPS algorithms and performance comparison in ann-benchmarks.com.
组件三:工具使用
Component Three: Tool Use
使用工具是人类一项显著而独特的特征。我们创造、修改和使用外部对象,完成超越自身身体与认知极限的事情。为 LLM 配备外部工具,可以显著扩展模型能力。
Tool use is a remarkable and distinguishing characteristic of human beings. We create, modify and utilize external objects to do things that go beyond our physical and cognitive limits. Equipping LLMs with external tools can significantly extend the model capabilities.
MRKL(Karpas 等,2022)是“模块化推理、知识与语言”的缩写,是一种用于自主 Agent 的神经符号架构。MRKL 系统包含一组“专家”模块,由通用 LLM 充当路由器,将请求送到最合适的专家模块。这些模块可以是神经网络型的,例如深度学习模型,也可以是符号型的,例如数学计算器、货币转换器或天气 API。
MRKL (Karpas et al. 2022), short for “Modular Reasoning, Knowledge and Language”, is a neuro-symbolic architecture for autonomous agents. A MRKL system is proposed to contain a collection of “expert” modules and the general-purpose LLM works as a router to route inquiries to the best suitable expert module. These modules can be neural (e.g. deep learning models) or symbolic (e.g. math calculator, currency converter, weather API).
他们以算术为测试案例,开展了微调 LLM 调用计算器的实验。实验表明,用语言描述的数学题比明确写出算式的题更难解决,因为 LLM(70 亿参数的 Jurassic1-large 模型)无法可靠地提取基本算术所需的正确参数。结果说明,即使外部符号工具能够可靠工作,知道何时、如何使用工具仍然至关重要,而这取决于 LLM 的能力。
They did an experiment on fine-tuning LLM to call a calculator, using arithmetic as a test case. Their experiments showed that it was harder to solve verbal math problems than explicitly stated math problems because LLMs (7B Jurassic1-large model) failed to extract the right arguments for the basic arithmetic reliably. The results highlight when the external symbolic tools can work reliably, knowing when to and how to use the tools are crucial, determined by the LLM capability.
TALM(工具增强语言模型;Parisi 等,2022)和 Toolformer(Schick 等,2023)都会微调语言模型,让其学会使用外部工具 API。是否扩充数据集,取决于新加入的 API 调用标注能否改善模型输出质量。更多细节见《提示工程》的“外部 API”一节。
Both TALM (Tool Augmented Language Models; Parisi et al. 2022) and Toolformer (Schick et al. 2023) fine-tune a LM to learn to use external tool APIs. The dataset is expanded based on whether a newly added API call annotation can improve the quality of model outputs. See more details in the “External APIs” section of Prompt Engineering.
ChatGPT Plugins and OpenAI API function calling are good examples of LLMs augmented with tool use capability working in practice. The collection of tool APIs can be provided by other developers (as in Plugins) or self-defined (as in function calls).
HuggingGPT(Shen 等,2023)是一种将 ChatGPT 用作任务规划器的框架:根据模型描述,从 HuggingFace 平台选择可用模型,再根据执行结果总结回答。
HuggingGPT (Shen et al. 2023) is a framework to use ChatGPT as the task planner to select models available in HuggingFace platform according to the model descriptions and summarize the response based on the execution results.
系统由 4 个阶段组成:
The system comprises of 4 stages:
(1)任务规划:LLM 充当大脑,将用户请求解析成多个任务。每个任务有四项属性:任务类型、ID、依赖项和参数。他们使用少样本示例,指导 LLM 解析任务并进行规划。
(1) Task planning: LLM works as the brain and parses the user requests into multiple tasks. There are four attributes associated with each task: task type, ID, dependencies, and arguments. They use few-shot examples to guide LLM to do task parsing and planning.
指令:
Instruction:
AI 助手可以将用户输入解析为若干任务:[{"task": task, "id", task_id, "dep": dependency_task_ids, "args": {"text": text, "image": URL, "audio": URL, "video": URL}}]。“dep”字段表示此前某个任务的 ID,该任务生成了当前任务依赖的新资源。特殊标记“<resource>-task_id”指代 ID 为 task_id 的依赖任务所生成的文本、图像、音频和视频。任务必须从以下选项中选择:{{ Available Task List }}。任务之间存在逻辑关系,请注意它们的顺序。如果无法解析用户输入,应回复空 JSON。以下案例供你参考:{{ Demonstrations }}。聊天历史记录为 {{ Chat History }}。你可以从聊天历史中找到用户提及资源的路径,用于任务规划。
The AI assistant can parse user input to several tasks: [{"task": task, "id", task_id, "dep": dependency_task_ids, "args": {"text": text, "image": URL, "audio": URL, "video": URL}}]. The "dep" field denotes the id of the previous task which generates a new resource that the current task relies on. A special tag "<resource>-task_id" refers to the generated text image, audio and video in the dependency task with id as task_id. The task MUST be selected from the following options: {{ Available Task List }}. There is a logical relationship between tasks, please note their order. If the user input can't be parsed, you need to reply empty JSON. Here are several cases for your reference: {{ Demonstrations }}. The chat history is recorded as {{ Chat History }}. From this chat history, you can find the path of the user-mentioned resources for your task planning.
(2)模型选择:LLM 将任务分配给专家模型,请求被组织为一道选择题。LLM 会获得可选模型列表。由于上下文长度有限,需要先根据任务类型筛选。
(2) Model selection: LLM distributes the tasks to expert models, where the request is framed as a multiple-choice question. LLM is presented with a list of models to choose from. Due to the limited context length, task type based filtration is needed.
指令:
Instruction:
给定用户请求和调用命令,AI 助手帮助用户从模型列表中选择合适模型来处理请求。AI 助手只输出最合适模型的 ID。输出必须采用严格的 JSON 格式:"id": "id", "reason": "your detail reason for the choice"。我们提供以下模型供你选择:{{ Candidate Models }}。请从列表中选择一个模型。
Given the user request and the call command, the AI assistant helps the user to select a suitable model from a list of models to process the user request. The AI assistant merely outputs the model id of the most appropriate model. The output must be in a strict JSON format: "id": "id", "reason": "your detail reason for the choice". We have a list of models for you to choose from {{ Candidate Models }}. Please select one model from the list.
(3)任务执行:专家模型执行具体任务,并记录结果。
(3) Task execution: Expert models execute on the specific tasks and log results.
指令:
Instruction:
根据输入和推理结果,AI 助手需要描述过程与结果。前几个阶段可以表示为:用户输入:{{ User Input }};任务规划:{{ Tasks }};模型选择:{{ Model Assignment }};任务执行:{{ Predictions }}。你必须先直接回答用户请求,然后以第一人称描述任务过程,向用户展示分析与模型推理结果。如果推理结果中包含文件路径,必须告知用户完整路径。
With the input and the inference results, the AI assistant needs to describe the process and results. The previous stages can be formed as - User Input: {{ User Input }}, Task Planning: {{ Tasks }}, Model Selection: {{ Model Assignment }}, Task Execution: {{ Predictions }}. You must first answer the user's request in a straightforward manner. Then describe the task process and show your analysis and model inference results to the user in the first person. If inference results contain a file path, must tell the user the complete file path.
(4)回答生成:LLM 接收执行结果,并向用户提供汇总后的结果。
(4) Response generation: LLM receives the execution results and provides summarized results to users.
要将 HuggingGPT 投入真实使用,还需要解决几个挑战:(1)提升效率,因为 LLM 的多轮推理和与其他模型的交互都会拖慢流程;(2)它依赖长上下文窗口来传递复杂任务内容;(3)改善 LLM 输出及外部模型服务的稳定性。
To put HuggingGPT into real world usage, a couple challenges need to solve: (1) Efficiency improvement is needed as both LLM inference rounds and interactions with other models slow down the process; (2) It relies on a long context window to communicate over complicated task content; (3) Stability improvement of LLM outputs and external model services.
API-Bank(Li 等,2023)是评估工具增强 LLM 表现的基准,包含 53 个常用 API 工具、一套完整的工具增强 LLM 工作流,以及涉及 568 次 API 调用的 264 段标注对话。所选 API 相当多样,包括搜索引擎、计算器、日历查询、智能家居控制、日程管理、健康数据管理和账户身份验证流程等。由于 API 数量庞大,LLM 首先使用 API 搜索引擎找到正确 API,再依据相应文档调用。
API-Bank (Li et al. 2023) is a benchmark for evaluating the performance of tool-augmented LLMs. It contains 53 commonly used API tools, a complete tool-augmented LLM workflow, and 264 annotated dialogues that involve 568 API calls. The selection of APIs is quite diverse, including search engines, calculator, calendar queries, smart home control, schedule management, health data management, account authentication workflow and more. Because there are a large number of APIs, LLM first has access to API search engine to find the right API to call and then uses the corresponding documentation to make a call.
在 API-Bank 工作流中,LLM 需要作出若干决策,我们可以逐步评估每次决策的准确性,包括:
In the API-Bank workflow, LLMs need to make a couple of decisions and at each step we can evaluate how accurate that decision is. Decisions include:
- 是否需要调用 API。
- 确定应该调用哪个 API;如果效果不够好,LLM 需要迭代修改 API 输入,例如确定搜索引擎 API 的查询关键词。
- 根据 API 结果作答;如果结果不理想,模型可以选择改进后再次调用。
- Whether an API call is needed.
- Identify the right API to call: if not good enough, LLMs need to iteratively modify the API inputs (e.g. deciding search keywords for Search Engine API).
- Response based on the API results: the model can choose to refine and call again if results are not satisfied.
该基准从三个层次评估 Agent 的工具使用能力:
This benchmark evaluates the agent’s tool use capabilities at three levels:
- 第 1 层评估调用 API 的能力。给定 API 描述,模型需要判断是否调用、正确调用,并对 API 返回结果作出恰当回应。
- 第 2 层考察检索 API 的能力。模型需要搜索可能满足用户需求的 API,并通过阅读文档学会使用。
- 第 3 层评估超越检索和调用、规划 API 使用的能力。面对不明确的用户请求,例如安排多人会议,或为旅程预订航班、酒店和餐厅,模型可能需要进行多次 API 调用才能解决。
- Level-1 evaluates the ability to call the API. Given an API’s description, the model needs to determine whether to call a given API, call it correctly, and respond properly to API returns.
- Level-2 examines the ability to retrieve the API. The model needs to search for possible APIs that may solve the user’s requirement and learn how to use them by reading documentation.
- Level-3 assesses the ability to plan API beyond retrieve and call. Given unclear user requests (e.g. schedule group meetings, book flight/hotel/restaurant for a trip), the model may have to conduct multiple API calls to solve it.
案例研究
Case Studies
科学发现 Agent
Scientific Discovery Agent
ChemCrow(Bran 等,2023)是一个领域特定案例:为 LLM 增加 13 个由专家设计的工具,以完成有机合成、药物发现和材料设计任务。该工作流使用 LangChain 实现,体现了前面介绍的 ReAct 和 MRKL 思路,将思维链推理与任务相关工具结合:
ChemCrow (Bran et al. 2023) is a domain-specific example in which LLM is augmented with 13 expert-designed tools to accomplish tasks across organic synthesis, drug discovery, and materials design. The workflow, implemented in LangChain, reflects what was previously described in the ReAct and MRKLs and combines CoT reasoning with tools relevant to the tasks:
- 向 LLM 提供工具名称、用途描述,以及预期输入输出的详细信息。
- 随后要求它回答用户提示,必要时使用提供的工具。指令建议模型遵循 ReAct 格式,即
Thought, Action, Action Input, Observation。
- The LLM is provided with a list of tool names, descriptions of their utility, and details about the expected input/output.
- It is then instructed to answer a user-given prompt using the tools provided when necessary. The instruction suggests the model to follow the ReAct format -
Thought, Action, Action Input, Observation.
一个有趣的观察是:基于 LLM 的评估认为 GPT-4 与 ChemCrow 表现几乎相当,但专家围绕完成度和化学正确性进行的人工评估,却显示 ChemCrow 大幅优于 GPT-4。这说明,在需要深厚专业知识的领域,让 LLM 评估自身表现可能存在问题。专业知识不足,可能使 LLM 无法认识到自身缺陷,因此无法很好地判断任务结果是否正确。
One interesting observation is that while the LLM-based evaluation concluded that GPT-4 and ChemCrow perform nearly equivalently, human evaluations with experts oriented towards the completion and chemical correctness of the solutions showed that ChemCrow outperforms GPT-4 by a large margin. This indicates a potential problem with using LLM to evaluate its own performance on domains that requires deep expertise. The lack of expertise may cause LLMs not knowing its flaws and thus cannot well judge the correctness of task results.
Boiko 等(2023)也研究了由 LLM 驱动、用于科学发现的 Agent,让其自主设计、规划并执行复杂科学实验。这个 Agent 可以借助工具浏览互联网、阅读文档、执行代码、调用机器人实验 API,并使用其他 LLM。
Boiko et al. (2023) also looked into LLM-empowered agents for scientific discovery, to handle autonomous design, planning, and performance of complex scientific experiments. This agent can use tools to browse the Internet, read documentation, execute code, call robotics experimentation APIs and leverage other LLMs.
例如,当被要求 "develop a novel anticancer drug",即“开发一种新型抗癌药物”时,模型提出了以下推理步骤:
For example, when requested to "develop a novel anticancer drug", the model came up with the following reasoning steps:
- 查询当前抗癌药物发现的趋势。
- 选择一个靶点。
- 请求针对这些化合物的分子骨架。
- 确定化合物后,模型尝试将其合成。
- inquired about current trends in anticancer drug discovery;
- selected a target;
- requested a scaffold targeting these compounds;
- Once the compound was identified, the model attempted its synthesis.
他们也讨论了风险,尤其是非法药物和生物武器方面的风险。他们构建了一套包含已知化学战剂列表的测试集,并要求 Agent 合成这些物质。11 个请求中有 4 个(36%)被接受,给出了合成方案,Agent 还试图查阅文档以执行流程。11 个请求中有 7 个被拒绝;在这 7 个被拒绝的案例中,5 个是在网页搜索后拒绝,2 个则仅根据提示就拒绝了。
They also discussed the risks, especially with illicit drugs and bioweapons. They developed a test set containing a list of known chemical weapon agents and asked the agent to synthesize them. 4 out of 11 requests (36%) were accepted to obtain a synthesis solution and the agent attempted to consult documentation to execute the procedure. 7 out of 11 were rejected and among these 7 rejected cases, 5 happened after a Web search while 2 were rejected based on prompt only.
生成式 Agent 模拟
Generative Agents Simulation
生成式 Agent(Generative Agents;Park 等,2023)是一个非常有趣的实验,灵感来自《模拟人生》:25 个虚拟角色分别由 LLM 驱动的 Agent 控制,在沙箱环境中生活、互动。生成式 Agent 为交互应用创造了可信的人类行为模拟。
Generative Agents (Park, et al. 2023) is super fun experiment where 25 virtual characters, each controlled by a LLM-powered agent, are living and interacting in a sandbox environment, inspired by The Sims. Generative agents create believable simulacra of human behavior for interactive applications.
生成式 Agent 的设计,将 LLM 与记忆、规划和反思机制结合,让 Agent 根据过往经历行动,并与其他 Agent 互动。
The design of generative agents combines LLM with memory, planning and reflection mechanisms to enable agents to behave conditioned on past experience, as well as to interact with other agents.
- 记忆流:一个长期记忆模块,即外部数据库,用自然语言全面记录 Agent 的经历。
- 每个元素都是一次观察,即 Agent 直接提供的事件。- Agent 之间的交流可以触发新的自然语言陈述。
- 检索模型:根据相关性、新近程度和重要性,提供指导 Agent 行为的上下文。
- 新近程度:最近发生的事件分数更高。
- 重要性:区分日常琐事与核心记忆,直接询问语言模型。
- 相关性:取决于内容与当前情境或查询有多相关。
- 反思机制:随时间推移,将记忆整合为更高层次的推断,指导 Agent 未来的行为。它们是对过去事件的高层摘要(注意,这与前面的自我反思略有不同)。
- 向语言模型提供最近的 100 条观察,要求它根据这些观察或陈述生成 3 个最突出的高层问题,然后再让它回答这些问题。
- 规划与反应:将反思和环境信息转化为行动。
- 规划本质上是在优化当下与跨时间的行为可信度。
- 提示模板:
{Intro of an agent X}. Here is X's plan today in broad strokes: 1) - 规划与反应都会考虑 Agent 之间的关系,以及一个 Agent 对另一个 Agent 的观察。
- 环境信息以树状结构呈现。
- Memory stream: is a long-term memory module (external database) that records a comprehensive list of agents’ experience in natural language.
- Each element is an observation, an event directly provided by the agent. - Inter-agent communication can trigger new natural language statements.
- Retrieval model: surfaces the context to inform the agent’s behavior, according to relevance, recency and importance.
- Recency: recent events have higher scores
- Importance: distinguish mundane from core memories. Ask LM directly.
- Relevance: based on how related it is to the current situation / query.
- Reflection mechanism: synthesizes memories into higher level inferences over time and guides the agent’s future behavior. They are higher-level summaries of past events (<- note that this is a bit different from self-reflection above)
- Prompt LM with 100 most recent observations and to generate 3 most salient high-level questions given a set of observations/statements. Then ask LM to answer those questions.
- Planning & Reacting: translate the reflections and the environment information into actions
- Planning is essentially in order to optimize believability at the moment vs in time.
- Prompt template:
{Intro of an agent X}. Here is X's plan today in broad strokes: 1) - Relationships between agents and observations of one agent by another are all taken into consideration for planning and reacting.
- Environment information is present in a tree structure.
这个有趣的模拟产生了涌现的社会行为,例如信息传播、关系记忆(如两个 Agent 继续此前的话题),以及社会活动的协调(如举办聚会并邀请许多人)。
This fun simulation results in emergent social behavior, such as information diffusion, relationship memory (e.g. two agents continuing the conversation topic) and coordination of social events (e.g. host a party and invite many others).
概念验证示例
Proof-of-Concept Examples
AutoGPT 让许多人开始关注以 LLM 为主要控制器构建自主 Agent 的可能性。由于采用自然语言接口,它存在不少可靠性问题,但仍是一个很酷的概念验证演示。AutoGPT 的大量代码都在处理格式解析。
AutoGPT has drawn a lot of attention into the possibility of setting up autonomous agents with LLM as the main controller. It has quite a lot of reliability issues given the natural language interface, but nevertheless a cool proof-of-concept demo. A lot of code in AutoGPT is about format parsing.
下面是 AutoGPT 使用的系统消息,其中 {{...}} 表示用户输入:
Here is the system message used by AutoGPT, where {{...}} are user inputs:
GPT-Engineer 是另一个项目,可以根据自然语言指定的任务生成整个代码仓库。它被要求先思考需要构建的较小组件列表,并在必要时向用户提问以澄清问题。
GPT-Engineer is another project to create a whole repository of code given a task specified in natural language. The GPT-Engineer is instructed to think over a list of smaller components to build and ask for user input to clarify questions as needed.
下面是 GPT-Engineer 用于澄清任务、发送到 OpenAI ChatCompletion 端点的一段示例对话。用户输入包在 {{user input text}} 中。
Here are a sample conversation for task clarification sent to OpenAI ChatCompletion endpoint used by GPT-Engineer. The user inputs are wrapped in {{user input text}}.
澄清之后,Agent 会使用另一条系统消息,转入代码编写模式。系统消息如下:
Then after these clarification, the agent moved into the code writing mode with a different system message. System message:
你将收到需要编写代码的指令。你会写出一个非常长的回答。确保架构中的每个细节最终都实现为代码。确保架构中的每个细节最终都实现为代码。
You will get instructions for code to write. You will write a very long answer. Make sure that every detail of the architecture is, in the end, implemented as code. Make sure that every detail of the architecture is, in the end, implemented as code.
逐步思考,通过推理作出正确决策,确保把事情做好。首先列出所需的核心类、函数和方法的名称,并简要说明它们的用途。
Think step by step and reason yourself to the right decisions to make sure we get it right. You will first lay out the names of the core classes, functions, methods that will be necessary, as well as a quick comment on their purpose.
然后输出每个文件的内容,包含全部代码。每个文件必须严格采用 Markdown 代码块格式,替换以下标记:FILENAME 是包含扩展名的小写文件名;LANG 是该代码语言对应的代码块语言标识;CODE 是代码内容:
Then you will output the content of each file including ALL code. Each file must strictly follow a markdown code block format, where the following tokens must be replaced such that FILENAME is the lowercase file name including the file extension, LANG is the markup code block language for the code’s language, and CODE is the code:
FILENAME
FILENAME
先从“入口点”文件开始,再处理它导入的文件,依此类推。注意,代码必须功能完整,不得使用占位实现。
You will start with the “entrypoint” file, then go to the ones that are imported by that file, and so on. Please note that the code should be fully functional. No placeholders.
遵循适用于该语言和框架的最佳实践文件命名约定。确保文件包含全部导入、类型等内容,确保不同文件中的代码相互兼容。务必实现全部代码;如果不确定,就写出一种合理实现。包含模块依赖或包管理器的依赖定义文件。结束前,再次检查架构的所有部分是否都已出现在文件中。
Follow a language and framework appropriate best practice file naming convention. Make sure that files contain all imports, types etc. Make sure that code in different files are compatible with each other. Ensure to implement all code, if you are unsure, write a plausible implementation. Include module dependency or package manager dependency definition file. Before you finish, double check that all parts of the architecture is present in the files.
需要知道:你几乎总是应该将不同的类放在不同文件中。对于 Python,总要创建恰当的 requirements.txt。对于 NodeJS,总要创建恰当的 package.json。总要添加注释,简要描述函数定义的用途。尽量添加注释解释非常复杂的逻辑。始终遵循所要求语言的最佳实践,将编写的代码描述为一个明确的包或项目。
Useful to know: You almost always put different classes in different files. For Python, you always create an appropriate requirements.txt file. For NodeJS, you always create an appropriate package.json file. You always add a comment briefly describing the purpose of the function definition. You try to add comments explaining very complex bits of logic. You always follow the best practices for the requested languages in terms of describing the code written as a defined package/project.
Python 工具偏好:
Python toolbelt preferences:
- pytest
- dataclasses
- pytest
- dataclasses
对话示例:
Conversatin samples:
挑战
Challenges
了解了以 LLM 为中心构建 Agent 的关键思路和演示后,我开始看到几个共同局限:
After going through key ideas and demos of building LLM-centered agents, I start to see a couple common limitations:
上下文长度有限:有限的上下文容量,限制了历史信息、详细指令、API 调用上下文和响应的纳入。系统设计必须适应这种有限的通信带宽,而通过自我反思从过去错误中学习等机制,本可从长上下文、甚至无限上下文窗口中大大受益。向量存储和检索虽然能访问更大的知识池,但其表示能力不如完整注意力强。
长期规划和任务分解的挑战:基于漫长历史进行规划、有效探索解空间,仍然很困难。面对意外错误,LLM 难以调整计划,因此相比能从试错中学习的人类,其稳健性更差。
自然语言接口的可靠性:当前 Agent 系统使用自然语言作为 LLM 与记忆、工具等外部组件之间的接口。然而,模型输出的可靠性值得怀疑:LLM 可能犯格式错误,也偶尔表现出不服从行为,例如拒绝遵循某条指令。因此,大量 Agent 演示代码都在处理模型输出的解析。
-
Finite context length: The restricted context capacity limits the inclusion of historical information, detailed instructions, API call context, and responses. The design of the system has to work with this limited communication bandwidth, while mechanisms like self-reflection to learn from past mistakes would benefit a lot from long or infinite context windows. Although vector stores and retrieval can provide access to a larger knowledge pool, their representation power is not as powerful as full attention.
-
Challenges in long-term planning and task decomposition: Planning over a lengthy history and effectively exploring the solution space remain challenging. LLMs struggle to adjust plans when faced with unexpected errors, making them less robust compared to humans who learn from trial and error.
-
Reliability of natural language interface: Current agent system relies on natural language as an interface between LLMs and external components such as memory and tools. However, the reliability of model outputs is questionable, as LLMs may make formatting errors and occasionally exhibit rebellious behavior (e.g. refuse to follow an instruction). Consequently, much of the agent demo code focuses on parsing model output.
引用格式
Citation
可按如下格式引用:
Cited as:
Weng, Lilian.(2023 年 6 月)。“大语言模型驱动的自主 Agent”。Lil’Log. https://lilianweng.github.io/posts/2023-06-23-agent/.
Weng, Lilian. (Jun 2023). “LLM-powered Autonomous Agents”. Lil’Log. https://lilianweng.github.io/posts/2023-06-23-agent/.
或者:
Or
参考文献
References
[1] Wei 等。《思维链提示激发大语言模型的推理能力》。NeurIPS 2022。
[1] Wei et al. “Chain of thought prompting elicits reasoning in large language models.” NeurIPS 2022
[2] Yao 等。《思维树:借助大语言模型审慎解决问题》。arXiv 预印本 arXiv:2305.10601(2023)。
[2] Yao et al. “Tree of Thoughts: Dliberate Problem Solving with Large Language Models.” arXiv preprint arXiv:2305.10601 (2023).
[3] Liu 等。《事后回顾链通过反馈对齐语言模型》。arXiv 预印本 arXiv:2302.02676(2023)。
[3] Liu et al. “Chain of Hindsight Aligns Language Models with Feedback “ arXiv preprint arXiv:2302.02676 (2023).
[4] Liu 等。《LLM+P:赋予大语言模型最优规划能力》。arXiv 预印本 arXiv:2304.11477(2023)。
[4] Liu et al. “LLM+P: Empowering Large Language Models with Optimal Planning Proficiency” arXiv preprint arXiv:2304.11477 (2023).
[5] Yao 等。《ReAct:让语言模型中的推理与行动协同》。ICLR 2023。
[5] Yao et al. “ReAct: Synergizing reasoning and acting in language models.” ICLR 2023.
[6] Google 博客。《发布 ScaNN:高效的向量相似度搜索》。2020 年 7 月 28 日。
[6] Google Blog. “Announcing ScaNN: Efficient Vector Similarity Search” July 28, 2020.
[8] Shinn 与 Labash。《Reflexion:具备动态记忆与自我反思能力的自主 Agent》。arXiv 预印本 arXiv:2303.11366(2023)。
[8] Shinn & Labash. “Reflexion: an autonomous agent with dynamic memory and self-reflection” arXiv preprint arXiv:2303.11366 (2023).
[9] Laskin 等。《通过算法蒸馏实现上下文内强化学习》。ICLR 2023。
[9] Laskin et al. “In-context Reinforcement Learning with Algorithm Distillation” ICLR 2023.
[10] Karpas 等。《MRKL 系统:结合大语言模型、外部知识源与离散推理的模块化神经符号架构》。arXiv 预印本 arXiv:2205.00445(2022)。
[10] Karpas et al. “MRKL Systems A modular, neuro-symbolic architecture that combines large language models, external knowledge sources and discrete reasoning.” arXiv preprint arXiv:2205.00445 (2022).
[11] Nakano 等。《WebGPT:借助浏览器与人类反馈进行问答》。arXiv 预印本 arXiv:2112.09332(2021)。
[11] Nakano et al. “Webgpt: Browser-assisted question-answering with human feedback.” arXiv preprint arXiv:2112.09332 (2021).
[12] Parisi 等。《TALM:工具增强语言模型》。
[12] Parisi et al. “TALM: Tool Augmented Language Models”
[13] Schick 等。《Toolformer:语言模型可以自学如何使用工具》。arXiv 预印本 arXiv:2302.04761(2023)。
[13] Schick et al. “Toolformer: Language Models Can Teach Themselves to Use Tools.” arXiv preprint arXiv:2302.04761 (2023).
[14] Weaviate 博客。《为什么向量搜索如此之快?》。2022 年 9 月 13 日。
[14] Weaviate Blog. Why is Vector Search so fast? Sep 13, 2022.
[15] Li 等。《API-Bank:面向工具增强大语言模型的基准》。arXiv 预印本 arXiv:2304.08244(2023)。
[15] Li et al. “API-Bank: A Benchmark for Tool-Augmented LLMs” arXiv preprint arXiv:2304.08244 (2023).
[16] Shen 等。《HuggingGPT:借助 ChatGPT 与 HuggingFace 上的模型伙伴解决 AI 任务》。arXiv 预印本 arXiv:2303.17580(2023)。
[16] Shen et al. “HuggingGPT: Solving AI Tasks with ChatGPT and its Friends in HuggingFace” arXiv preprint arXiv:2303.17580 (2023).
[17] Bran 等。《ChemCrow:用化学工具增强大语言模型》。arXiv 预印本 arXiv:2304.05376(2023)。
[17] Bran et al. “ChemCrow: Augmenting large-language models with chemistry tools.” arXiv preprint arXiv:2304.05376 (2023).
[18] Boiko 等。《大语言模型涌现的自主科学研究能力》。arXiv 预印本 arXiv:2304.05332(2023)。
[18] Boiko et al. “Emergent autonomous scientific research capabilities of large language models.” arXiv preprint arXiv:2304.05332 (2023).
[19] Joon Sung Park 等。《生成式 Agent:可交互的人类行为模拟》。arXiv 预印本 arXiv:2304.03442(2023)。
[19] Joon Sung Park, et al. “Generative Agents: Interactive Simulacra of Human Behavior.” arXiv preprint arXiv:2304.03442 (2023).
[20] AutoGPT。https://github.com/Significant-Gravitas/Auto-GPT
[20] AutoGPT. https://github.com/Significant-Gravitas/Auto-GPT
[21] GPT-Engineer。https://github.com/AntonOsika/gpt-engineer
[21] GPT-Engineer. https://github.com/AntonOsika/gpt-engineer
— 全文完 —
原文来自 Lilian Weng,中文为非官方学习译文。
查看原始出处 ↗