资料馆/评测与改进
Live-SWE-agent阅读档案 · 非官方中文译文

Live-SWE-agent:首个可在运行时自我演化的软件工程 AgentLive-SWE-agent | The First Live AI Software Agent

本站收录 本站更新
完整译文与原文逐段对应。图片、图注、表格和代码保留原文。A complete reading edition. Figures, captions, tables and code are preserved from the source.
中文译文ENGLISH ORIGINAL
live-swe-agent banner

Live-SWE-agent 是首个能够在运行时自我演化的动态软件工程 Agent,在处理真实世界问题的同时,即时扩展和调整自身能力。我们的关键洞察是:软件 Agent 本身也是软件系统,而现代基于大语言模型的 Agent 已经具备在运行时扩展或修改自身行为的内在能力。

Live-SWE-agent is the first live, runtime self-evolving software engineering agent that expands and revises its own capabilities on the fly while working on a real-world issue. Our key insight is that software agents are themselves software systems, and modern LLM-based agents already possess the intrinsic capability to extend or modify their own behavior at runtime.

📣 动态

📣 News

  • [2025 年 11 月 24 日]:Claude Opus 4.5 + Live-SWE-agent 在 SWE-bench Verified 上取得了 79.2% 的成绩,领先于目前所有开源运行框架,并且已经非常接近 Anthropic 为 Opus 4.5 手工设计的内部运行框架!!
  • [2025 年 11 月 20 日]:Gemini 3 Pro + Live-SWE-agent 在 SWE-bench Verified 上取得了 77.4% 的成绩,超越所有已可用的模型(包括 Claude 4.5)!
  • [2025 年 11 月 17 日]:Live-SWE-agent 在 SWE-Bench Pro 上实现了 45.8% 的问题解决率,刷新了最佳水平!
  • [2025 年 11 月 17 日]:我们发布了 Live-SWE-agent 1.0.0!
  • [Nov 24th, 2025]: Claude Opus 4.5 + Live-SWE-agent scores 79.2% on SWE-bench Verified, leading all current open-source scaffolds and coming very close to Anthropic’s internal, manually engineered scaffold for Opus 4.5!!
  • [Nov 20th, 2025]: Gemini 3 Pro + Live-SWE-agent scores 77.4% on SWE-bench Verified, outperforming all available models (including Claude 4.5)!
  • [Nov 17th, 2025]: Live-SWE-agent achieves the new state-of-the-art solve rate of 45.8% on SWE-Bench Pro!
  • [Nov 17th, 2025]: We've released Live-SWE-agent 1.0.0!

🏆 排行榜

🏆 Leaderboard

在软件任务中,近期的大语言模型通常使用手工设计的专有 Agent 运行框架进行基准评测,这使得公平比较不同模型的真实能力变得困难。

For software tasks, recent LLMs are often benchmarked using manually engineered, proprietary agent scaffolds, which makes it difficult to compare the true capabilities of different models fairly.

Live-SWE-agent 不仅表明,一个极简、开放且能在运行时演化的运行框架已经有能力超越专有框架,还提供了统一而强大的平台,让未来发布的模型可以在相同条件下进行真正公平的比较。

Live-SWE-agent not only demonstrates that a minimal, open, and live scaffold already has the ability to outperform proprietary scaffolds, but also offers a unified and powerful platform that enables genuinely fair, apples-to-apples comparisons for future model releases.

如下所示,在我们针对近期模型建立的排行榜中(所有模型均使用 Live-SWE-agent 进行评测),Claude Opus 4.5 以 SWE-bench Verified 上 79.2% 的得分大幅领先,继续位居第 1。

As shown below, on our leaderboard of recent models (all evaluated with Live-SWE-agent), Claude Opus 4.5 retains the #1 spot with a score of 79.2% on SWE-bench Verified by a large margin.

更多模型的得分即将公布!详情请访问我们的排行榜。欢迎提交你的模型评测结果,帮助我们建设一个更加全面、公平的基准评测平台!

More model scores are coming soon! For more details, please visit our leaderboard. Feel free to submit your model's evaluation results to help build a more comprehensive and fair benchmarking platform!

📊 对比

📊 Comparison

下图展示了 Live-SWE-agent 与最先进的开源方案及专有商业 Agent 运行框架在 SWE-bench Verified 和 SWE-Bench Pro 上的对比。

Below shows the comparison graph between Live-SWE-agent and state-of-the-art open-source solutions and proprietary commercial agent scaffolds on SWE-bench Verified and SWE-Bench Pro.

🚀 安装与配置

🚀 Setup

我们基于广受欢迎的 mini-swe-agent 框架构建了 Live-SWE-agent,只做了极少量修改。

We built Live-SWE-agent on top of the popular mini-swe-agent framework with very minimal modifications.

要使用 Live-SWE-agent,只需先按照这份指南安装 mini-swe-agent,再使用自定义的 Live-SWE-agent 配置:

To use Live-SWE-agent, simply install mini-swe-agent first using this guide and use the custom Live-SWE-agent config:

mini --config config/livesweagent.yaml # using custom Live-SWE-agent config

更多详情请查看 config 文件夹。

See the config folder for more details.

⚙️ 实验产物

⚙️ Artifacts

你可以在我们的 v1.0.0 发布页面下载 Live-SWE-agent 的完整执行轨迹、补丁和结果:

You can download the complete trajectories, patches, and results of Live-SWE-agent in our v1.0.0 release:

  • swebench_verified:在 SWE-bench Verified 上的完整运行记录
  • swebench_pro:在 SWE-Bench Pro 上的完整运行记录
  • swebench_verified: complete runs on SWE-bench Verified
  • swebench_pro: complete runs on SWE-Bench Pro

你也可以从我们的 🤗 Hugging Face 数据集获取这些内容。

You also obtain them in our 🤗 huggingface datasets

📜 引用

📜 Attribution

@article{livesweagent,
  author    = {Xia, Chunqiu Steven and Wang, Zhe and Yang, Yan and Wei, Yuxiang and Zhang, Lingming},
  title     = {Live-SWE-agent: Can Software Engineering Agents Self-Evolve on the Fly?},
  year      = {2025},
  journal   = {arXiv preprint},
}

🙏 致谢

🙏 Acknowledgements

— 全文完 —

原文来自 Live-SWE-agent,中文为非官方学习译文。
查看原始出处 ↗

点击空白处或按 Esc 关闭