Chengshu@skadai · 2026.09.30
11,415 字 · 9,121 词 · 约 70 分钟

《当 AI 开始建造自己》中英对照全译:逐句翻译、动效复刻与批注

把 Anthropic Institute 的《When AI builds itself》整篇译成中文并转载:每个段落先英文后中文,复刻了原文的滚动驱动时间线与点阵 hero 动效,并在关键观点处加了批注。

这是 The Anthropic Institute 于 2026 年发布的文章 《When AI builds itself》(当 AI 开始建造自己)的全文翻译。原文的主线只有一句话:Anthropic 正在把越来越多的 AI 研发工作交给 AI 自己做,而这条曲线如果继续延伸,终点是一个能完全自主地设计并开发自己后继者的系统——也就是所谓的递归自我改进(recursive self-improvement)。

为了让它既能当译文读、又能当原文读,这份转载做了四件事:

  1. 逐句逐段全译,没有跳段、没有概述式省略。每个段落都先给英文原文,紧跟中文翻译。
  2. 复刻原文的 HTML 动效:顶部 hero 是一块随页面滚动/加载运行的点阵(canvas)动画,正文中段是一条滚动驱动的时间线——页面滚过多长的距离,时间线上的阶段就亮到第几个。
  3. 对重要观点加了批注,用带 批注 标记的方框标出;批注是译者的解读,不是原文内容。
  4. 图表、脚注、引语、作者名单都保留并翻译;原文里出现的链接原样保留。

原文标题When AI builds itself

副标题Our progress toward recursive self-improvement, and its implications.

发布方The Anthropic Institute

作者Marina Favaro、Jack Clark(Santi Ruiz 提供编辑支持)

原文出处anthropic.com/institute/recursive-self-improvement

本站说明本文版权归 Anthropic 所有,此处为全文中文翻译转载,仅供学习交流;译文为译者所加,如与原文有出入,以原文为准。

EN 英文原文 中 中文翻译 批注 译者解读

The Anthropic Institute · 2026

When AI builds itself

当 AI 开始建造自己

Our progress toward recursive self-improvement, and its implications.

我们在递归自我改进上的进展,以及它的含义

For most of AI’s history, humans drove every step in its development cycle. But at Anthropic, we are delegating a growing share of AI development to AI systems themselves, which is speeding up our work.

在 AI 发展史的大部分时间里,人类驱动着研发循环中的每一步。但在 Anthropic,我们正把越来越大比例的 AI 研发工作交给 AI 系统自己做,而这正在加快我们的工作。

Taken far enough, and given enough compute, that trend points to an AI system capable of fully autonomously designing and developing its own successor. This is called recursive self-improvement. We are not there yet, and recursive self-improvement is not inevitable. But it could come sooner than most institutions are prepared for.

如果这条路走得足够远、并且有足够的算力,这个趋势指向的将是一个能够完全自主地设计并开发自己后继者的 AI 系统。这被称为递归自我改进(recursive self-improvement)。我们还没有到那一步,递归自我改进也不是必然会发生。但它到来的时间,可能比大多数机构准备好的时间更早。

Using public benchmarks and previously unreported data from within Anthropic, The Anthropic Institute is showing that AI is already accelerating the development of AI systems. To take just one example: today, Anthropic engineers on average ship 8x as much code per quarter as they did from 2021-2025.

借助公开基准,以及此前未公开过的 Anthropic 内部数据,The Anthropic Institute 正在说明:AI 已经在加速 AI 系统的研发。仅举一例:今天,Anthropic 工程师平均每季度交付的代码量,是 2021–2025 年间的 8 倍。

The technical trends discussed in this piece suggest that AI systems are going to become much more capable in coming years. These trends have huge implications. AI that can build itself would be a major development in the history of technology—one that could bring enormous good for the world in science, healthcare, and beyond. But full recursive self-improvement also might increase the risks of humans losing control over AI systems. If systems are capable of fully building their own successors, the ways we secure them, monitor them, and shape their behavior all grow much more important.

本文讨论的技术趋势表明,未来几年 AI 系统会变得强大得多。这些趋势影响巨大。能够建造自己的 AI,将是技术史上的一个重大进展——它可能在科学、医疗以及更多领域为世界带来巨大的好处。但完全的递归自我改进,也可能加大人类失去对 AI 系统控制的风险。如果系统能够完全自主地构建自己的后继者,那么我们如何保护它们、监控它们、塑造它们的行为,都会变得重要得多。

时间线:从写代码的人,到改进自己的 Claude

原文在开头放了一条滚动驱动的时间线:它把 AI 研发的演进压成五个阶段,页面往下滚,阶段依次点亮,最后的节点合上「递归」这个环。这里按同样的机制复刻了一版。

人 码 验 模 Claude 设定目标 写代码 跑实验 训练下一代 改进 Claude

2021–2023

Building the first Claude

构建第一个 Claude

In the early days, work at Anthropic looked like work at any other tech company: people writing code and docs on laptops.

早期,Anthropic 的工作和任何一家科技公司没什么两样:人们在笔记本电脑上写代码、写文档。

2023–2025

Chatbots

聊天机器人

People used early chatbots to help with parts of the process, like generating short code snippets and copying the output into text editors.

人们用早期的聊天机器人协助流程中的某些环节,比如生成简短的代码片段,再把结果复制进文本编辑器。

2025–2026

Coding agents

编码智能体

As the agents became more capable, they were able to write and edit code on their own, sometimes entire files.

随着智能体能力增强,它们能自己编写和修改代码,有时是整个文件。

Today

Autonomous agents

自主智能体

Agents can now run code themselves and delegate hours of work to other agents.

如今,智能体可以自己运行代码,并把数小时的工作委派给其他智能体。

20XX?

Closing the loop

合上这个环

In the future, agents could become capable enough to build and train models themselves. If this happens, future versions of Claude could be continuously improved by Claude itself.

未来,智能体可能强大到足以自行构建和训练模型。若如此,未来的 Claude 版本就可能由 Claude 自己持续改进。

Evidence from the outside world|来自外部的证据

The rate at which AI models improve is accelerating. The length of tasks that they can reliably complete on their own has been doubling roughly every four months, up from an earlier trend of doubling every seven months. In March 2024, Claude Opus 3 could complete software tasks that take humans about four minutes to complete. A year later, Claude Sonnet 3.7 managed tasks that took about an hour and a half. A year after that, Claude Opus 4.6 managed 12-hour tasks.1 If this trend holds, tasks that take a skilled person days could come into range this year. In 2027, AI systems could be capable of tasks that take a person weeks.

AI 模型变强的速度正在加快。它们能够独立可靠完成的任务时长,大约每四个月翻一倍——而此前的趋势是每七个月翻一倍。2024 年 3 月,Claude Opus 3 能完成人类约四分钟就能搞定的软件任务;一年后,Claude Sonnet 3.7 能处理大约一个半小时的任务;又过了一年,Claude Opus 4.6 已经能完成 12 小时的任务。1 如果这一趋势延续下去,需要熟练的人花上数天的任务,今年就可能落入它的能力范围;到 2027 年,AI 系统或许就能胜任需要一个人花上数周的任务。

The same pattern appears on coding and research benchmarks. Benchmarks measure the performance of models in a given domain, and they’re “saturated” when models achieve close to 100% performance.2 SWE-bench is a standard test of real-world software engineering: it hands a model an actual open-source codebase and a real bug report, and asks it to write a code change that fixes the issue and passes the project’s own tests. Models have gone from scoring in the low single digits to saturating the benchmark in two years.

同样的模式也出现在编码与研究的基准测试上。基准衡量的是模型在某一领域的表现;当模型的表现接近 100% 时,我们说这个基准被「饱和」(saturated)了。2 SWE-bench 是真实世界软件工程的一项标准测试:它把一个真实的开源代码库和一份真实的缺陷报告交给模型,要求它写出一个代码改动,既能修好这个问题,又能通过项目自己的测试。在两年之内,模型从个位数低段的得分,一路涨到把这个基准打满。

CORE-Bench tests whether a model can reproduce existing research, a prerequisite for them to conduct original research. It gives an AI model the code and data behind a published paper, and asks it to rerun everything and confirm it can replicate the paper’s results. AI systems went from succeeding at reproducing the results roughly 20% of the time in 2024 to saturating the benchmark fifteen months later. METR, which runs the benchmark measuring how well models can complete long-duration tasks, found that Claude Mythos Preview could work for “at least” 16 hours and was “at the upper end of what [METR] can measure without new tasks.”

CORE-Bench 检验的是模型能否复现已有的研究——这是它们开展原创研究的前提。它把一篇已发表论文背后的代码和数据交给 AI 模型,要求它把所有流程重跑一遍,确认能复现论文的结果。AI 系统在 2024 年时能复现成功的比例约为 20%,十五个月后就把这个基准打满了。运营「长时程任务完成度」基准的 METR 发现,Claude Mythos Preview 能连续工作「至少」16 小时,并且已经「达到了 [METR] 在没有新任务的情况下所能测量的上限」。

Public benchmarks say a lot about the capabilities of these systems. But they can’t reveal the impact AI systems are having on speeding up AI development itself. For that, we need direct evidence from within AI companies like Anthropic.

公开基准能说明这些系统能力的很多方面,但它们无法揭示 AI 系统对「加速 AI 研发本身」产生了什么影响。要了解这一点,我们需要来自 Anthropic 这类 AI 公司内部的直接证据。

Evidence from within Anthropic|来自 Anthropic 内部的证据

Building a frontier model takes two broad categories of work. There is engineering: writing the code, standing up the infrastructure, and overseeing the model training. And there is research: deciding what experiments to run, interpreting what comes back, and figuring out which ideas to try next.

打造一个前沿模型需要两大类工作。一类是工程:写代码、搭建基础设施、监督模型训练;另一类是研究:决定跑哪些实验、解读跑出来的结果、判断接下来该试哪个想法。

Across both engineering and research, the picture is consistent. In engineering, Claude can be handed an underspecified problem and figure out how to solve it; humans supply the goal, but they no longer need to supply the method. In research, Claude can already match or outperform skilled humans at executing a well-specified experiment. However, large performance gaps persist when it comes to Claude exercising judgement in choosing goals in both engineering and research. That’s the gap between AI today and a future system that could autonomously design its own successor.

在工程和研究这两类工作里,图景是一致的。在工程上,你可以把一个定义得很粗糙的问题丢给 Claude,它能自己想办法解决;人类给出目标,但不再需要给出方法。在研究上,对于一个定义清晰的实验,Claude 已经能追平、甚至超过熟练的人类。然而,只要涉及「在工程和研究中判断该选什么目标」,Claude 与人类之间仍存在很大的性能差距。而这,正是今天的 AI 与那个「能自主设计自己后继者」的未来系统之间的差距。

It’s common for employees at Anthropic to receive more open-ended and important tasks as they gain more experience. Early on, they execute a task someone else specified, like, “The export button isn’t working, please fix it.” With experience, they’re handed a goal and design the approach themselves, such as, “Investigate why the network slows down under heavy load.” At the most senior levels, they are deciding which problems are worth working on at all: “What should the team build next quarter?” We can use internal Anthropic data to see how far Claude has come in being able to handle these different kinds of tasks.

在 Anthropic,员工的资历越深,接到的任务就越开放、越重要,这是常态。刚入行时,他们执行的是别人已经指定好的任务,比如「导出按钮坏了,请修一下」。有了经验之后,别人只给一个目标,方法由他们自己设计,比如「查一下为什么网络在高负载下会变慢」。到了最资深的一层,他们要决定的是「哪些问题值得做」,比如「下个季度团队该做什么?」。我们可以用 Anthropic 的内部数据,看看 Claude 在这几类任务上已经走到了哪一步。

Claude writes a significant proportion of Anthropic’s code. As of May 2026, more than 80% of the code we merge into Anthropic’s codebase was authored by Claude.3 Before Claude Code launched in research preview in February 2025, this number was in the low single digits. That shift also shows up in the amount of output per engineer. Lines of code merged per engineer per day stayed constant through Anthropic’s first four years (2021-2024), then began to climb upward in 2025 when Claude began to run code rather than just suggesting it for an engineer to copy and paste. The slope steepened again in 2026 when models began to work autonomously over longer time horizons. These two inflection points are shown in the chart below. In the second quarter of 2026, the typical engineer was merging 8× as much code per day as they were in 2024.4 This is because much of the code is written by Claude, with the engineer directing and reviewing, rather than typing it themselves.

Claude 写了 Anthropic 相当大一部分代码。截至 2026 年 5 月,合并进 Anthropic 代码库的代码里,超过 80% 由 Claude 编写。3 在 2025 年 2 月 Claude Code 以研究预览版发布之前,这个数字还停留在个位数低段。这一变化也体现在每位工程师的产出上。在 Anthropic 的前四年(2021–2024),每位工程师每天合并的代码行数基本恒定;到 2025 年,Claude 开始自己运行代码、而不只是给出建议让人复制粘贴,这条线开始往上爬;2026 年,模型开始在更长的时间尺度上自主工作,斜率再次变陡。下图标出了这两个拐点。2026 年第二季度,普通工程师每天合并的代码量是 2024 年的 8 倍。4 原因正是:大部分代码由 Claude 写,工程师负责指挥和审阅,而不是自己敲键盘。

Bar graph showing code contributed per person, per quarter, starting in Q2 2021 and ending in Q2 2026, with model release markers.
Code contributed per person, by quarter. Each bar is the average, over the days in that quarter, of lines of code merged per active contributor, shown as a multiple of the pre-2025 average. 图:人均代码贡献量,按季度统计。每根柱子表示该季度内「每位活跃贡献者每天合并的代码行数」的平均值,以 2025 年前的平均值为 1 倍来折算。

A caveat: Lines of code is an imperfect measure, as it measures quantity over quality. So 8× lines of code/engineer/day in the second quarter of 2026 is almost certainly an overstatement of the true productivity gain. Nonetheless, it indicates an acceleration. At Anthropic, we don’t reward people for how many lines of code they write; rather, team members are producing more code simply because they’re using AI systems to write more code.

需要加一条说明:代码行数是个不完美的指标,它衡量的是数量而非质量。所以,把 2026 年第二季度「每位工程师每天 8 倍代码行数」直接等同于 8 倍真实生产力提升,几乎肯定是高估了。尽管如此,它确实说明加速正在发生。在 Anthropic,我们并不按写了多少行代码来奖励人;团队成员产出更多代码,只是因为他们用 AI 系统写出了更多代码。

The increase in lines of code written lines up with subjective impressions of large productivity increases. In a March 2026 poll of 130 employees from across Anthropic research teams, the median respondent estimated that they produced around 4x as much output with Mythos Preview as they would have without access to any AI models, on the kinds of projects they would have been working on regardless.5 We expect that the true degree of uplift in March was somewhat lower.6 Nevertheless, we find the overall claim plausible, and in line with our other observations: a significant fraction of Anthropic technical staff is accomplishing their core work multiple times faster than they could without AI assistance.

代码行数的增长,与「生产力大幅提升」的主观感受相互印证。2026 年 3 月对 Anthropic 各研究团队 130 名员工做的一项调查显示,受访者的中位数估计是:在那些他们无论如何都会推进的项目上,借助 Mythos Preview 产出的成果,大约是不使用任何 AI 模型时的 4 倍。5 我们认为,3 月真实的提升幅度应该比这略低一些。6 但总体判断是可信的,也与其他观察一致:Anthropic 技术团队中有相当一部分人,正以数倍于无 AI 辅助的速度完成自己的核心工作。

We also see evidence that people at Anthropic are using Claude to do work that simply wouldn’t have happened otherwise, like building exploratory tooling and addressing long-deferred cleanup. For example, in April 2026, Claude shipped over 800 fixes that reduced a class of API errors by a factor of one thousand. The engineer overseeing Claude estimated that a human would have taken four years to complete this work; solving other people’s bugs is slow and painstaking, and humans struggle to hold that much unfamiliar context in their head at once.

我们还看到证据表明,Anthropic 的人在让 Claude 去做那些「否则根本不会有人做」的工作,比如搭建探索性工具、清理长期积压的欠账。举个例子:2026 年 4 月,Claude 提交了 800 多个修复,把某一类 API 错误降低了三个数量级(1000 倍)。负责监督 Claude 的工程师估计,这活儿换成人来做要花四年;修别人的 bug 又慢又磨人,而人类很难把那么多陌生的上下文同时装在脑子里。

“I started leaning hard into Claudifying about a year ago. That’s been a crazy adventure and it’s now been ~5 months since I last wrote any code myself.”

「大约一年前,我开始大力『Claude 化』。那是一段疯狂的旅程,到现在我已经差不多五个月没亲手写过一行代码了。」

—— Anthropic employee*|Anthropic 员工*

The code that Claude writes is “good” and improving. “Good code” means two things: it works, and it is written in a manner that allows another engineer to understand it and build upon it. On the first criterion, the evidence is clear. The rate at which Anthropic staff correct, redirect, or take over mid-task from Claude has been falling steadily for a year, including on the most complex and open-ended tasks. This means problems with no clear specification, where the engineer isn’t sure what the answer looks like. This is evident in Claude’s success rate over time on tasks of different difficulties, as shown in the graph below. Claude writes code that works.

Claude 写的代码是「好」的,而且在变好。「好代码」有两层含义:一是它能跑通;二是它的写法能让另一位工程师看懂、并在其上继续构建。第一条标准上,证据是清楚的。Anthropic 员工纠正、改向、或者中途从 Claude 手里接管任务的比例,已经持续下降了一年——在最复杂、最开放的任务上也是如此。所谓「最开放」,指的是那些没有明确规格、连工程师都不确定答案长什么样的问题。Claude 在不同难度任务上的成功率随时间的变化(见下图)说明了这一点。Claude 写的代码能跑通。

Line graph showing the Claude Code session success rate on trivial, routine, substantial, and open-ended tasks across six models.
Claude Code session success rate over time, by task difficulty. How to read this: session success is determined by a Claude judge; a session is deemed successful if the Claude Code agent clearly succeeded at the user’s tasks without requiring corrections. 图:Claude Code 会话成功率随时间的变化,按任务难度分组。读法说明:会话是否成功由一个 Claude 裁判判定;只有当 Claude Code 智能体明确完成了用户的任务、且无需人工纠正时,该会话才算成功。

On the most open-ended tasks, Claude’s success rate reached 76% in May 2026, up 50 percentage points in six months. To give an example of tasks in this difficulty tier, a routine upgrade began crashing tens of thousands of training jobs. An engineer pointed Claude at the live incident with little more than some text content and cluster access. Working through the running jobs and testing one environment setting at a time, Claude isolated the single obscure debugging flag that was triggering the crash, reproduced it reliably, and confirmed a fix. In about two hours, Claude delivered what would normally be two to three days of work.

在最开放的那一类任务上,Claude 的成功率在 2026 年 5 月达到 76%,六个月内上升了 50 个百分点。举一个这个难度档位的例子:一次例行升级开始让成千上万个训练任务崩溃。一位工程师把 Claude 指向这个线上故障,给它的东西不过是几段文字内容和集群访问权限。Claude 逐个排查正在运行的任务,一次只测一个环境设置,最终定位到那个触发崩溃、极其冷僻的调试开关,稳定复现了问题,并确认了修复方案。大约两个小时,它完成了正常情况下需要两到三天的工作。

The second criterion is writing code that another engineer can understand and build on. Here the gap between humans and AI persists, but is closing fast. There isn’t full consensus among staff at Anthropic, but many believe that the Claude-written code was still worse in quality than human-written code at Anthropic in late 2025, and is roughly at parity today. We expect it to be better within the year.

第二条标准是写出另一位工程师能看懂、能接着往下做的代码。在这条上,人类与 AI 之间的差距仍然存在,但正在迅速收窄。Anthropic 内部并未完全达成共识,但不少人认为:2025 年底时,Claude 写的代码质量还比不上 Anthropic 人类写的代码;而今天两者大致已经持平。我们预计,一年之内 Claude 会反超。

“Claude-written code was somewhat worse than human-written code at Anthropic in late 2025, is roughly at parity today, and we expect it to be strictly better within the year.”

「在 Anthropic,2025 年底时 Claude 写的代码还略逊于人类写的代码;今天两者大致持平;我们预计一年之内它会明确地更好。」

This has changed the way that Anthropic now reviews its own code. Proposed changes to our codebase are now read by an automated Claude reviewer that looks for bugs, security flaws, and other defects before it can merge. Using this tool, we ran a retrospective analysis, and found that an automated Claude review of every change to our codebase would have caught roughly a third of the bugs behind past incidents on claude.ai before they ever reached production. The engineers who wrote that code are among the best in the world at building these systems. Claude is now catching the mistakes that they missed.

这也改变了 Anthropic 现在审查自己代码的方式。提交到代码库的改动,在合并之前会先由一个自动化的 Claude 审查员阅读,查找 bug、安全缺陷和其他瑕疵。借助这个工具,我们做了一次回溯分析,发现如果用自动化的 Claude 审查每一处代码改动,过去 claude.ai 上那些事故背后大约三分之一的 bug,本可以在进入生产环境之前就被拦下。写出那些代码的工程师,是这个世界上最擅长构建这类系统的人之一。而 Claude 现在能抓住他们漏掉的错误。

Claude is good at running experiments to hit a goal that someone else has set. Every time Anthropic releases a model, we run the same test: we give Claude some code that trains a small AI model, and ask it to make that code run as fast as possible while still passing the same correctness checks. The goal and the success metrics are fixed in advance, so Claude’s job is to find speedups by rewriting the code, running it, timing it, and repeating. It’s a miniature version of an experimental research loop. In May 2025, Claude Opus 4 averaged a ~3x speedup over the starting code. By April 2026, Claude Mythos Preview was achieving ~52x. For calibration, a skilled human researcher would need four to eight hours to reach 4x.7 In this part of the research workflow—optimizing steps within a clearly defined experiment—Claude has gone from super helpful to superhuman in under a year.

Claude 很擅长为别人设定好的目标做实验。每次 Anthropic 发布新模型,我们都会跑同一个测试:给 Claude 一段训练小型 AI 模型的代码,要求它在通过同样的正确性检查的前提下,让这段代码跑得尽可能快。目标和成功指标都是预先固定的,所以 Claude 要做的就是改写代码、运行、计时、再重复,从中找出加速的办法。这是一个微缩版的实验研究闭环。2025 年 5 月,Claude Opus 4 相对初始代码平均加速约 3 倍;到 2026 年 4 月,Claude Mythos Preview 已经能跑到约 52 倍。作为参照,一位熟练的人类研究员要花四到八个小时才能达到 4 倍。7 在研究流程的这个环节——在定义清晰的实验里优化步骤——Claude 在不到一年里从「超级好用的助手」变成了「超人」。

“The shape of stuff today is roughly ‘humans have ideas, and the models are able to implement, test and evaluate them an [order of magnitude] faster than before.’”

「今天的大致格局是:『人类出想法,而模型能以(快一个数量级的)速度把它们实现、测试和评估。』」

Claude is getting better at proposing its own experiments. In April 2026, Anthropic published the first demonstration of Claude running an open-ended research project end to end. Claude-powered agents were given an open problem in AI safety—roughly, can a weaker model reliably supervise a stronger one?—and were left to solve it. This involved proposing hypotheses, testing them, sharing findings with parallel agents, and iterating. The task has a clear performance “floor” and “ceiling”: the floor is how well the weak supervisor would do on its own; the ceiling is how the strong model does when trained on correct answers. Two human researchers, over about a week, recovered roughly 23% of that gap; the agents recovered 97% over 800 cumulative hours and used roughly $18,000 in compute. There are some caveats to this work; the result didn’t transfer cleanly to production-scale models, and humans still chose the problem and created the scoring rubric. But within those bounds, the agents designed every experiment themselves. Direction-setting was the only meaningful role a human played.

Claude 越来越会自己提出实验了。2026 年 4 月,Anthropic 发表了首个「由 Claude 端到端跑完一个开放式研究项目」的演示。他们给一组由 Claude 驱动的智能体抛出一个 AI 安全领域的开放问题——大致是「一个较弱的模型能否可靠地监督一个更强的模型?」——然后放手让它去解决。这个过程包括提出假设、验证假设、与并行的智能体共享发现、反复迭代。这个任务有明确的性能「下限」和「上限」:下限是弱监督者独自能做到的水平;上限是强模型在正确答案上训练后的表现。两名人类研究员花了大约一周,填补了这个差距的大约 23%;而这些智能体累计跑了 800 小时,用掉大约 18,000 美元算力,填补了 97%。这项工作有几处保留:结果没能干净地迁移到生产规模的模型上,而且问题和评分标准仍然由人类选定。但在这些边界之内,每一个实验都是智能体自己设计的。人类唯一真正有意义的角色,是设定方向。

“Claude did all of this with pretty minimal help from me over the course of 1-2 days. I think if [a junior colleague] came back to me with results like this in the same span of time, I would be mildly impressed. The future is now.”

「这一切都是 Claude 在一两天之内、几乎没怎么需要我帮忙就做完的。我想,如果(一位初级同事)在同样的时间里给我带回这样的结果,我会略微刮目相看。未来已经来了。」

Claude is getting better at steering research sessions towards research findings. We examined real Claude Code sessions (between January and March 2026) where Anthropic researchers were working with Claude on an open-ended investigative problem, like figuring out why a training run kept crashing, or why a model scored poorly on a benchmark. In each case, we found a moment where the researcher took a detour: they pursued a direction that sent the session sideways before it eventually got back on track. We then showed various Claude models only the work from before the session went off-course and asked what it would do next. A separate Claude that was able to see how the session eventually turned out then judged whether the AI or the human suggested the better next step.8

Claude 越来越擅长把研究过程引向真正的发现。我们考察了一批真实的 Claude Code 会话(2026 年 1 月至 3 月之间)。在这些会话里,Anthropic 的研究员和 Claude 一起攻关一个开放式调查问题,比如查清一次训练为什么总是崩溃,或者某个模型为什么在基准上得分很低。在每一个案例里,我们都能找到一个研究员走弯路的瞬间:他们追着一个方向,把整段会话带偏,之后才绕回正轨。然后,我们只把「走偏之前」的工作展示给不同版本的 Claude 模型,问它下一步会怎么做;再由另一个能看到会话最终走向的 Claude 来评判:AI 和人类,谁提出的下一步更好。8

Because we deliberately picked moments (n=129) where we know the human’s choice had room for improvement, this isn’t a like-for-like comparison between model and human judgement. What these moments give us is a set of realistic, challenging situations where the right next step is not obvious, and where the human’s choice serves as a useful yardstick to compare model performance over time. On this measure, our best model in November 2025 (Opus 4.5) beat the human choice 51% of the time; in April 2026 (Mythos Preview), this grew to 64%. The day-to-day work of research is largely a chain of these next-step decisions, making this a relevant measure of the model’s ability to eventually run an investigation of its own. We view this result as an early signal that AI systems are getting better at making the kinds of judgement calls that AI research depends on.

由于我们刻意挑选的是那些「已知人类的选择还有改进空间」的瞬间(n=129),这并非一场模型与人类判断力的对等比较。这些瞬间提供给我们的,是一组真实、有挑战、正确答案并不明显的情境;而人类的选择可以充当一把有用的标尺,用来纵向比较模型的表现。在这个指标上,我们 2025 年 11 月最好的模型(Opus 4.5)有 51% 的时候比人类的选择更好;到 2026 年 4 月(Mythos Preview),这个比例涨到 64%。研究的日常工作,很大程度上就是一串这样的「下一步」决策,所以这个指标能反映模型最终能否独立开展一项研究。我们认为,这是一个早期信号:AI 系统正在变得善于做出 AI 研究赖以进行的那类判断。

Bar graph titled ‘Where a researcher went wrong, could Claude have done better?’ showing the performance of nine models.
Where a researcher went wrong, could Claude have done better? How to read this: the practical ceiling line measures an “ideal” answer written by a model that could see the whole session (including how it ended). 图:研究员走错的地方,Claude 能做得更好吗?读法说明:「实际天花板」这条线,代表一个能看到整段会话(包括结局)的模型所写出的「理想答案」。

“The comparative advantage of humans as of right now is still in seeing the bigger picture and thinking beyond the confines of the immediate task.”

「就目前而言,人类仍然占优的地方,是看清全局、并跳出眼前任务的局限去思考。」

What might the future of work at Anthropic look like?|Anthropic 未来的工作会是什么样?

The evidence suggests that the human role is narrowing at each step in the AI development process. Once human- and AI-authored code quality reach parity, humans will stop writing code entirely, and shift to only reviewing it. But if they can’t review code as quickly as Claude can generate it, human review will become the bottleneck to AI development. Similarly, once Claude can run experiments, the question shifts towards “Which of these experiments is worth running?” Put simply: the doing (i.e., writing the code, running the experiment, producing the result) now costs almost nothing in human time, even if it still has costs in compute.

证据表明,在 AI 研发流程的每一步上,人类扮演的角色都在收窄。一旦人类和 AI 写的代码质量持平,人类就会完全停止写代码,转为只做审查。但如果人类审查代码的速度赶不上 Claude 生成代码的速度,人类审查就会变成 AI 研发的瓶颈。同理,一旦 Claude 能跑实验,问题就变成「这些实验里,哪一个值得跑?」。简单说:「做」(写代码、跑实验、产出结果)现在在人类时间上几乎不花成本,尽管在算力上仍然要花成本。

An area of human comparative advantage, for now, is research taste and judgment, including choosing which problems matter, which results to trust, and when an approach is a dead end.

就目前而言,人类的比较优势在于研究的品味和判断力——包括判断哪些问题重要、哪些结果可信,以及一条路什么时候已经走到了死胡同。

“Work (and life) ran on a gift economy of small favors between humans. ‘Can you help me get this script running?’ [...] each one created a little debt, a little mutual awareness. [Claude is] faster, it creates zero debt, but each of these is a lost bid for human collaboration.”

「工作(以及生活)过去靠的是人与人之间小恩小惠式的礼物经济。『你能帮我让这个脚本跑起来吗?』……每一次都制造了一点点亏欠、一点点相互的体认。(Claude)更快、不制造任何亏欠,但每一次这样的时刻,也都是一次人类协作机会的落空。」

“On days where everything works well, I can’t help but think nothing I do matters, everything is automated and better and faster than I ever will be. But then there are days where everything breaks and I don’t understand why and I realize I have no idea what I’ve been up to anymore.”

「在一切顺利的日子里,我忍不住会想:我做的任何事都不重要,一切都是自动的,比我更好、更快。但也会有这样的日子:所有东西都坏了,我不明白为什么,然后就意识到,我已经完全不知道自己一直在忙些什么了。」

What if we’re wrong?|如果我们错了呢?

A natural objection to the evidence presented above is that the work that is still in human hands—choosing which problems to work on—is what matters most. Without that judgment, Claude is a capable assistant, but not a system that could drive AI progress on its own.

对上面这些证据,一个自然的反驳是:还留在人类手里的那部分工作——选择做什么问题——恰恰是最重要的。没有这种判断力,Claude 充其量是个能干的助手,而不是一个能独自驱动 AI 进步的系统。

It is genuinely unclear whether today’s training methods and architectures could unlock that capacity. But AI is rarely advanced by “eureka!” moments. There have been a few of these in AI’s recent history, like the Transformer architecture, or mixture-of-experts models, but paradigm-shifting ideas arrive years apart. In between, most progress is incremental: we scale something up, see what breaks, fix it, and try again. That is exactly the kind of workflow Claude now excels at. Edison said that genius is 1% inspiration and 99% perspiration. But we see perspiration becoming increasingly automated. It’s becoming clear that much of what advances the frontier is automatable; large-scale research progress is mostly a function of tools and resources, which dictate how fast you can run experiments, how many you can run at once, and how quickly you can get results.

今天的训练方法和架构能否解锁这种能力,确实还说不清。但 AI 的进步很少靠「尤里卡!」式的灵光一现。AI 的近期历史里有过少数几次这样的时刻,比如 Transformer 架构,或者混合专家(MoE)模型,但范式级的想法往往相隔数年才出现一次。在这之间,大部分进步是渐进的:把某个东西做大,看哪里崩了,修好,再试一次。而这恰恰是 Claude 现在极其擅长的流程。爱迪生说,天才是 1% 的灵感加 99% 的汗水。但我们看到,汗水正在变得越来越自动化。越来越清楚的是,推动前沿的很多东西是可以自动化的;大规模的研究进展,很大程度上是工具和资源的函数——它们决定了你能多快跑实验、能同时跑多少个、能多快拿到结果。

Even if we suppose that Claude never achieves good research taste, a conservative reading of our evidence still implies compounding acceleration. If humans spend most of their time on the single-digit fraction of work that is direction-setting, while Claude handles the rest, that means each engineer or researcher is steering far more work than before. The evidence we see suggests that people at Anthropic are both moving faster and covering a broader surface. In practice, this means that AI already makes Anthropic move much faster than it did before the advent of effective AI tools.

即便我们假设 Claude 永远培养不出好的研究品味,对我们这些证据做一种保守的解读,仍然指向复利式的加速。如果人类把大部分时间花在「定方向」这不到百分之十的工作上,其余全交给 Claude,那就意味着每一位工程师或研究员所驾驭的工作量都远超以往。我们看到的证据表明,Anthropic 的人既跑得更快,覆盖的面也更广。实际上这意味着:AI 已经让 Anthropic 跑得比「有效的 AI 工具出现之前」快得多。

The less conservative reading is that the early evidence on Claude’s improving research judgment—narrow as it is today—is an indicator that this capability is improving as well. “Research taste” might be just another AI capability that AI systems fail at for a time, then get good at. We’ve seen a similar pattern with other qualitative skills, like AI systems being able to explain why a joke is funny, demonstrate theory of mind, and solve linguistic riddles.

不那么保守的解读是:关于 Claude 研究判断力正在提升的那些早期证据——尽管今天还很单薄——本身就是这种能力正在改善的信号。「研究品味」也许只是又一项 AI 能力:AI 系统先是在它上面屡屡失败,然后忽然就擅长了。我们在其他定性能力上见过类似的模式,比如 AI 系统能解释一个笑话为什么好笑、能展现心智理论(theory of mind)、能解语言谜题。

Possible futures|可能的未来

What happens next depends on two things: whether the trend continues, and what we choose to do if it does. We can imagine at least three future scenarios:

接下来会发生什么,取决于两件事:这个趋势会不会延续,以及如果延续,我们选择怎么做。我们至少可以设想三种未来情景:

  1. The trend stalls, but today’s AI capabilities are widely diffused. This article features many exponential trajectories. But these trajectories may actually turn out to be S-curves. We may be approaching the bend in the curve, where returns to scale diminish and the line straightens, then flattens. The judgment that separates a competent researcher from a great one might be a capability that cannot come from scaling up training inputs like compute and data. If so, getting past this bottleneck would require a new idea, like an architectural approach that supplants the Transformer architecture that all current frontier models use. Alternately, the binding constraint to AI progress could be in the supply chain, not the model: advancing and diffusing the frontier may require more energy and compute than presently exists. The pace of chip fabrication, grid expansion, or interconnect bandwidth may be the constraint, rather than intelligence itself. We also cannot rule out an exogenous shock to the AI ecosystem that dramatically slows things, like a sudden diminishment in the supply of compute or electricity, either of which would slow progress and make forward investment by labs more expensive. Or we may not be anticipating some other barrier to progress. Even if model capabilities were frozen at today’s level, we would expect major changes to occur in the world. Project Glasswing is one early sign: in its first weeks, Mythos Preview found more than ten thousand high- and critical-severity software vulnerabilities across the world’s most important systems—enough that the bottleneck in cyber defense has already shifted from finding vulnerabilities to patching them fast enough. And we are still early in the diffusion of today’s models into the wider economy, where a 100-person company can increasingly do the work of a 1,000-person one, because each employee will sit atop a pyramid of agents. We include this scenario for completeness, but we don’t believe it’s likely. Every capability we can measure, including those that feel “squishier,” like quality of code and success on open-ended tasks, has so far followed the same curve. We have not yet seen that curve bend. Of the three futures we consider, this one would give governments and societies the most time to adapt. We are more worried about the next two, which would move faster and leave far less room for preparation.

    趋势停滞,但今天的 AI 能力已经广泛扩散。这篇文章里有很多指数曲线。但这些曲线,实际上也可能只是 S 形曲线。我们也许正接近曲线的弯折处:规模收益递减,线先是走平,然后变平。把「能胜任的研究员」和「伟大的研究员」区分开的那种判断力,也许根本无法从算力、数据这类训练投入的规模化中获得。若果真如此,要越过这个瓶颈就需要一个新想法,比如一种取代 Transformer 架构的架构方案——目前所有前沿模型都建立在 Transformer 之上。另一种可能是:制约 AI 进步的瓶颈不在模型,而在供应链。把前沿推进并扩散出去,需要的能源和算力可能超过当下已有的水平。真正的约束也许是芯片制造、电网扩容或互联带宽的速度,而不是智能本身。我们也无法排除 AI 生态遭遇外生冲击、让一切明显放缓的可能,比如算力或电力供给突然萎缩——这二者中任何一个都会拖慢进展,也会让实验室的前期投资变得更贵。又或者,还有某个我们尚未预料到的障碍。即便模型能力冻结在今天的水平,我们依然预期世界会发生重大变化。Project Glasswing 就是一个早期的信号:在它上线的头几周里,Mythos Preview 在全球最重要的系统里发现了超过一万个高危和严重级别的软件漏洞——多到网络防御的瓶颈已经从「找漏洞」变成了「打补丁的速度」。而且,今天的模型向更广泛经济体的扩散才刚刚开始:在那里,一家 100 人的公司越来越能做 1000 人公司的活,因为每个员工都能坐拥一整个智能体金字塔。我们把这个情景列出来是为了完整,但并不认为它很可能发生。迄今,我们能测量的每一种能力——包括那些感觉更「软」的,比如代码质量和开放式任务的成功率——都遵循同一条曲线,我们还没有看到这条曲线弯折。在我们考虑的三种未来中,这一种会给政府和社会留出最多的适应时间。我们更担心的是后面两种:它们来得更快,留给准备的空间也小得多。

  2. AI labs continue to see compounding efficiency gains. In this scenario, AI development becomes substantially automated, but humans continue to set research directions and judge results. Organizations that use AI systems would become much more efficient as time goes on, so we could expect to see significant productivity multipliers on each person in this organization. 100-person companies could do the work of 10,000- or 100,000-person organizations. This would revolutionize knowledge work and government services, but could also be turned to harmful ends, from authoritarian surveillance of whole populations to influence operations that tailor manipulation to each individual and run at a scale no human team could match. The role of humans at companies like Anthropic would shift. People would partner with AI systems to scale up research and generate new insights, and together they would build the systems needed to verify that AI outputs can be trusted. The evidence we’ve laid out here suggests that we’re likely heading into this scenario. But speeding up one part of a process often just shifts the bottleneck elsewhere: overall pace is capped by the parts that haven’t sped up. In computing, this is known as Amdahl’s law, and the same logic can apply to organizations. Anthropic has already encountered one signature of Amdahl’s law: as we’ve begun to push more code around the organization, human code review has become a new bottleneck. We’ve also encountered this friction outside engineering. There has been an explosion of new ideas, initiatives, tools, and simulations, as a result of Anthropic employees working with highly capable models—far more than we have the capacity to pursue. The rate at which organizations can spot and fix these bottlenecks may be a skill that improves over time, and it may become the most important skill for any organization.

    AI 实验室持续获得复利式的效率提升。在这个情景里,AI 研发在很大程度上实现自动化,但人类继续设定研究方向、评判结果。使用 AI 系统的组织会随时间变得越来越高效,因此可以预期,组织中的每个人都会带来显著的生产力倍数。100 人的公司可以完成 10000 人乃至 100000 人组织的工作。这将彻底改变知识工作和政府服务,但也可能被用于有害的目的:从对全体民众的威权式监控,到为每个人量身定制操纵手段、并以任何人类团队都无法企及的规模运行的舆论影响行动。像 Anthropic 这样的公司里,人类的角色会发生转变:人们与 AI 系统协作,扩大研究规模、产生新的洞见,并共同构建「验证 AI 输出是否可信」所需要的系统。我们在这里摆出的证据表明,我们很可能正在走入这个情景。但加速流程中的某一环,往往只是把瓶颈挪到别处:整体速度被那些没有加速的环节卡住。在计算领域,这被称为阿姆达尔定律(Amdahl’s law),同样的逻辑也适用于组织。Anthropic 已经遇到了阿姆达尔定律的一个典型症状:随着我们在组织内部推动越来越多代码流动,人类的代码审查已经成了新的瓶颈。我们在工程之外也遇到了这种摩擦。由于 Anthropic 的员工与能力极强的模型协作,新的想法、计划、工具和模拟呈爆炸式涌现——远远超出我们有余力去推进的数量。一个组织发现并修复这些瓶颈的速度,也许会成为一种随时间提升的能力,而且可能成为任何组织最重要的能力。

  3. AI systems themselves become capable of full recursive self-improvement, and begin building their successors. If technical trends in advancing capabilities continue, and AI systems are able to develop the capabilities inherent to transformative human ingenuity, then it is plausible that AI systems could design and refine themselves. In this world, the pace of progress in AI development becomes determined entirely by the availability of compute (or the speed of discovering various efficiencies in algorithmic training or inference) for AI systems. Humans play a substantially diminished role in their development, likely moving most of our effort towards oversight, validation, and verification of an expanding “virtual lab” run by AI systems. We expect that systems capable of automated AI research and development would have skills that would transfer to the rest of science, allowing them to begin to revolutionize other fields. How the alignment problem gets solved—or not—in this future is something we are least certain about. Models could prove to be sufficiently aligned and capable enough of research taste that they discover and implement novel solutions that we have not yet reached. They could also be sufficiently wise to halt development if not. Alternatively, the rare occurrences of misalignment present in today’s models could compound as the models build their successors, growing more frequent but less understood until we lose control of them. It’s possible that we can’t build, integrate, and verify the tools that we’d need to understand which trendline we are actually on. We do not have good intuitions for what this world would look like, because our economy is currently driven by humans and human-built tools. By its nature, a world driven by fast recursive self-improvement could become dominated by the self-improving model as its capabilities fully eclipse those of humans and the model proliferates across the broader economy. It is difficult to predict what the economy looks like if human labor stops being competitive. Even if model development became fully automated and recursive, we can’t predict what that would mean for most humans’ daily lives. Amdahl’s law applies here as well. Recursive intelligence could lead to achieving many of the benefits outlined in Machines of Loving Grace, quickly in some domains. We expect that embodied intelligence (i.e., robotics) might quickly follow recursive intelligence, and follow a similar path of increasing returns at decreasing cost. More powerful intelligence might help us build things in the physical world more quickly, run more productive clinical trials of lifesaving drugs, and develop novel forms of coordination. But achieving recursive improvement alone does not suggest an immediate change in how industrial production occurs, societies organize, or markets function. More intelligence can’t learn what a drug does over decades of use, can’t hold elections sooner than a constitution dictates, and can’t turn a stranger into an old friend in a weekend. For most people, the felt pace of this future will still be set by the bottlenecks, even if the laboratory upstream runs at the speed of compute. That collision, where recursive intelligence building itself ever faster meets the world of humans, relationships, and governance, is another part of this future we can’t predict.

    AI 系统本身具备了完全的递归自我改进能力,并开始构建自己的后继者。如果能力提升的技术趋势延续下去,并且 AI 系统能够发展出「变革性的人类创造力」所固有的那些能力,那么 AI 系统设计并打磨自己就是可能的。在这样的世界里,AI 研发的推进速度将完全取决于 AI 系统可获得的算力(或者说,发现算法训练或推理中各种效率提升的速度)。人类在其发展中的作用会大幅缩小,我们的精力很可能大部分转向对一个不断扩张的、由 AI 系统运行的「虚拟实验室」做监督、验证与核验。我们预期,具备自动化 AI 研发能力的系统,也会拥有可以迁移到其他科学领域的技能,从而开始革新其他领域。在这个未来里,对齐问题会被解决、还是不会被解决,是我们最没有把握的部分。模型可能被证明足够对齐、也足够有研究品味,从而发现并实现我们尚未达到的新解法;它们也可能足够明智,在做不到时停下开发。反过来说,今天的模型里那些罕见的失准(misalignment)事件,也可能随着模型构建后继者而累积——变得更频繁,却更不被理解,直到我们失去对它们的控制。也有可能,我们无法构建、整合并验证那些「用来判断我们究竟在哪条趋势线上」所必需的工具。我们对这个世界会是什么样,缺乏好的直觉,因为我们的经济目前是由人和人造工具驱动的。就其本性而言,一个由快速递归自我改进驱动的世界,可能会被那个自我改进的模型所主宰——当它的能力全面超越人类、并在更广泛的经济中扩散开来时。如果人类劳动不再有竞争力,经济会是什么样,很难预测。即便模型开发变得完全自动、完全递归,我们也无法预测这对大多数人的日常生活意味着什么。阿姆达尔定律在这里同样适用。递归智能可能会让《Machines of Loving Grace》里描绘的许多好处在某些领域迅速实现。我们预计,具身智能(也就是机器人)可能会紧随递归智能之后,走上一条「回报递增、成本递减」的相似道路。更强大的智能也许能帮我们更快地在物理世界里建造东西、开展更有产出的救命药物临床试验、发展出新的协作形式。但仅仅实现递归改进,并不意味着工业生产方式、社会组织方式或市场运行方式会立刻改变。更多的智能无法提前知道一种药物在几十年使用中的效果,无法比宪法规定的更早举行选举,也无法在一个周末把陌生人变成老朋友。对大多数人来说,这个未来被感受到的速度,仍将由那些瓶颈决定——即便上游的实验室正以算力的速度运转。递归智能以越来越快的速度建造自己,与人类、关系与治理的世界相撞——这场碰撞,是这个未来里我们同样无法预测的另一部分。

What should we do?|我们该做什么?

If it were possible to effectively slow the development of this technology to give ourselves more time to deal with its immense implications, we think that would likely be a good thing. But if a slowdown simply lets the least cautious actors catch up technologically, it could leave everyone less safe. Without a global coordination mechanism, companies and governments will have to make difficult decisions about safety while under competitive and geopolitical pressures.

如果有可能有效地放慢这项技术的发展,好让我们有更多时间去应对它巨大的影响,我们认为这很可能是一件好事。但如果「放慢」只是让最不谨慎的行动者从技术上追赶上来,那反而会让所有人更不安全。没有一套全球协调机制,公司和政府就不得不在竞争压力和地缘政治压力之下,做出艰难的安全决策。

We believe it would be good for the world to have the option to slow or temporarily pause frontier AI development to enable societal structures and alignment research to keep up with the advance of the technology. The Anthropic Institute will conduct research—in collaboration with many others—and take actions to help build the systems that a credible slowdown or pause would require. These systems would enable frontier AI developers to verify that others globally have actually stopped or slowed, and that a bad actor could not use the auspices of a coordinated slowdown to jump ahead in secret. If such systems existed, we expect that we would slow down or temporarily pause, if other developers at or near the frontier also did so in a verifiable manner.

我们认为,让世界拥有一个「可以放慢、或临时暂停前沿 AI 开发」的选项,是有益的——这样社会结构和对齐研究才能跟上技术的推进。The Anthropic Institute 将(与许多其他方合作)开展研究、并采取行动,帮助构建「可信的放缓或暂停」所需的那些系统。这些系统要能让前沿 AI 开发者核验:全球其他开发者确实已经停下或放慢了;并且坏人无法假借「协调放缓」之名偷偷抢跑。如果这样的系统存在,我们预期,只要处于或接近前沿的其他开发者也都以可核验的方式这么做,我们就会放慢、或临时暂停。

A meaningful slowdown or pause would require multiple well-resourced labs at or near the frontier, in multiple countries, agreeing to stop under the same conditions. It would also require that each can verify that the others have actually stopped. Due to the unique characteristics of AI systems, the detectability (a lower standard than verifiability) element of this arms control problem is much more challenging than with other technologies. Training runs are far easier to conceal than missile silos, their inputs are general-purpose, and the incentive to defect quietly is enormous, because whoever continues while others pause could inherit the lead. A credible pause also has to specify what triggers it, what lifts it, and who adjudicates.

一场有意义的放缓或暂停,需要多个国家中多个资源充足、处于或接近前沿的实验室,同意在相同条件下停下来;还需要每一方都能核验其他方确实已经停下。由于 AI 系统的独特性质,这个军控问题中「可探测性」(比「可核验」更低的一条标准)这一环,比其他技术要困难得多。训练任务比导弹发射井容易隐藏得多,它们的投入是通用的,而且「悄悄违约」的动机极大——因为谁在别人暂停时继续推进,谁就可能继承领先地位。一场可信的暂停还必须说明:什么触发它、什么解除它、由谁来裁定。

None of this is necessarily impossible in principle—the world has built verification regimes for other complex technologies (e.g., the Intermediate-Range Nuclear Forces Treaty)—but those regimes took decades to build both the infrastructure and the trust. We don’t have that long. A unilateral pause by one lab, by contrast, is achievable immediately, but accomplishes much less: it would change who the front-runner is, but it would not create the wider deliberative process that is currently missing.

原则上,这一切并非必然不可能——世界已经为其他复杂技术建立过核验机制(例如《中程核力量条约》),但这些机制花了数十年才建起基础设施和信任。我们没有那么长的时间。相比之下,由一家实验室单方面暂停是可以立刻做到的,但作用小得多:它会改变谁是领跑者,却无法创造出目前缺失的那种更广泛的商议过程。

In the coming months, we will organize conversations where policymakers, researchers, civil society, and other AI companies can help answer some of the questions this piece raises, especially around full recursive self-improvement and how to create better options for coordination and deliberation. We’ll publish what comes out of it. The window to investigate the questions together is here, and people outside AI companies should be involved in this deliberation.

在接下来的几个月里,我们将组织一些对话,让政策制定者、研究者、公民社会和其他 AI 公司一起来帮助回答这篇文章提出的部分问题,尤其是围绕「完全的递归自我改进」以及「如何为协调与商议创造更好的选项」。我们会把结果公开发表。共同审视这些问题的窗口就在当下,而 AI 公司之外的人也应当参与到这场商议中来。

Marina Favaro and Jack Clark co-authored this piece, with editorial support from Santi Ruiz. Shan Carter, Romello Goodman, and Nikki Makagiansar created the visuals from data collected by Brian Calvert and Jun Shern Chan. Daniel Freeman, Jim Baker, Max Young, Sarah Pollack, Francesco Mosconi, Holden Karnofsky, Andy Jones, Kevin Troy, Chloe Lubinski, Anton Korinek, Meg Tong, Andrew Ho, Dan Altman, Drake Thomas, Jack Shen, Sasha de Marigny, and Avital Balwit provided feedback.

本文由 Marina Favaro 和 Jack Clark 共同撰写,Santi Ruiz 提供编辑支持。Shan Carter、Romello Goodman 和 Nikki Makagiansar 依据 Brian Calvert 与 Jun Shern Chan 收集的数据制作了图表。Daniel Freeman、Jim Baker、Max Young、Sarah Pollack、Francesco Mosconi、Holden Karnofsky、Andy Jones、Kevin Troy、Chloe Lubinski、Anton Korinek、Meg Tong、Andrew Ho、Dan Altman、Drake Thomas、Jack Shen、Sasha de Marigny 和 Avital Balwit 提供了反馈。

Update 9/18/2026|更新:2026 年 9 月 18 日

原文在这一天追加了一次更新,给出截止到 2026 年 9 月的最新一张成功率曲线(下方图表)。原文在此处未附正文,只有图表与读法说明;此段对图表的描述为译者依据图注补充。

This update adds the latest Claude Code session-success chart, extending the series through September 2026: success rates rise across all four task types and converge around 88–92%, with open-ended problems improving the most (from about 26% to 91%).

这次更新补上了最新的 Claude Code 会话成功率曲线,把数据延伸到 2026 年 9 月:四种任务类型的成功率全面上升,并在 88%–92% 附近收敛,其中开放式问题的进步最大(从约 26% 升到 91%)。

Line graph of Claude Code session success rate from August 2025 to September 2026 across four task types, converging around 88 to 92 percent.
Claude Code session success rate, August 2025 – September 2026, by task type (trivial, routine, substantial, open-ended). How to read this: session success is determined by a Claude judge; a session is deemed successful if the Claude Code agent clearly succeeded at the user’s tasks without requiring corrections. Changes in workloads can lead to short-term fluctuations in success rates. 图:Claude Code 会话成功率,2025 年 8 月至 2026 年 9 月,按任务类型分组(琐碎、常规、重要、开放式)。读法说明:会话是否成功由一个 Claude 裁判判定;只有当 Claude Code 智能体明确完成了用户的任务、且无需人工纠正时,该会话才算成功。工作负载的变化可能造成成功率短期波动。

Footnotes|脚注

  1. METR’s key measure tells you the time horizon over which AI systems can be 50% reliable at a basket of tasks, though the trendline looks the same at 80% reliability.

    METR 的核心指标衡量的是:在一组任务上,AI 系统能以 50% 的可靠度完成的时间跨度;不过在 80% 可靠度上,趋势线的形状是一样的。

  2. Especially as they shift toward more open-ended formats and more difficult tasks (e.g., Olympiad-level mathematics), benchmarks often saturate below 100% due to errors in the question and answer sets like ambiguous problem statements and unsolvable questions.

    尤其是当基准转向更开放的题型和更难的题目(例如奥赛级别的数学)时,基准常常会在 100% 以下就饱和,因为题库和答案本身含有错误,比如题意含糊、题目无解。

  3. Anthropic leadership have publicly estimated that 90% or more of our code is written by Claude, including scripts and experimental code. Our >80% figure measures the share of lines merged to production that can be attributed to Claude. This is a more conservative measurement in two ways: our attribution pipeline has gaps, and the lines not attributed to Claude include auto-generated code and other artifacts that were not hand-written by humans either.

    Anthropic 领导层曾公开估计,我们有 90% 以上的代码由 Claude 编写,其中包括脚本和实验代码。我们采用的「>80%」这个数字,衡量的是被合并进生产环境、且可归因于 Claude 的代码行占比。这是一个更保守的口径,原因有二:我们的归因流程本身有缺口;而那些未被归因于 Claude 的行里,还包含自动生成的代码和其他同样不是人手写的产物。

  4. This surge in code production is straining the infrastructure everyone shares. GitHub—the platform most of the world’s software is built on—saw roughly one billion code commits in all of 2025; by mid-2026 it saw 275 million a week, on pace for roughly 14 billion over the year. The company’s COO has said that it is “pushing incredibly hard” on capacity just to keep up.

    这股代码产出的激增,正在让所有人共享的基础设施吃紧。GitHub——世界上大多数软件都构建于其上的平台——在 2025 年全年大约有 10 亿次代码提交;到 2026 年年中,它一周就有 2.75 亿次提交,按这个速度全年大约 140 亿次。该公司 COO 表示,为了跟上,他们在容量上「拼得非常狠」。

  5. Additional details on the methodology of this survey are discussed in section 2.3.5 of the Claude Opus 4.7 System Card.

    关于这项调查方法的更多细节,见 Claude Opus 4.7 系统卡第 2.3.5 节。

  6. Many respondents may not have thought carefully about how to account for various biases or subtleties in the question definition, and recent research by METR shows that developer estimates of AI productivity uplift can be overestimated.

    许多受访者可能并没有仔细想过如何排除问题定义中的各种偏差和微妙之处;而 METR 最近的研究表明,开发者对 AI 生产力提升的估计可能会偏高。

  7. How large the speedup gets depends heavily on how much room for improvement the starting code leaves, and it should not be read as a real-world training speedup. So the absolute multiple is not the figure to anchor on here. What is more informative is the like-for-like comparison that this experimental setup makes possible, both across models (~3x to ~52x over the past year) and against a skilled human (~4x in four to eight hours on the same task).

    加速的幅度在很大程度上取决于初始代码留出了多少优化空间,因此它不应被理解为真实世界里的训练加速。所以在这里,绝对值倍数不是应该锚定的数字。更有信息量的是这套实验设置所能提供的「同条件对比」:既包括模型之间的纵向对比(过去一年从约 3 倍到约 52 倍),也包括与熟练人类的对比(同一个任务上,人类用四到八小时达到约 4 倍)。

  8. As a check on judge bias, we ran the same test on a separate set of 127 moments where the human’s next move was already strong (as opposed to the original set, where the human’s direction had room for improvement). There, the models’ suggestions were judged better only about 20% of the time.

    为检验裁判偏差,我们在另一组 127 个时刻上跑了同样的测试——这些时刻里,人类下一步的选择本来就很强(与原先那组「人类方向还有改进空间」的情形相反)。在这组里,模型的建议只有大约 20% 的时候被判为更好。

* Quotes from Anthropic employees throughout this article are drawn from internal discussions and used with permission. They reflect individual views as of May 2026, not official company positions.

* 本文中所有来自 Anthropic 员工的引语,均取自内部讨论并获授权使用。它们反映的是个人在 2026 年 5 月的看法,不代表公司的官方立场。

原文链接When AI builds itself · The Anthropic Institute

译文说明全文中文翻译为译者所加,按「先英文、后中文」逐段对照;页首点阵动画与中段滚动时间线为本站按原文机制复刻;带「批注」标记的方框为译者解读。版权归 Anthropic 所有。

讨论

这里是静态站点,没有内嵌评论区。如果这篇文章对你有用,欢迎通过 RSS 订阅后续更新。