Menu Close

What Anthropic’s R&D Automation Index measures

Educational vertical Epoch AL0–AL5 ladder diagram for Anthropic R&D Automation Index explainer, with side plate on why the metric matters for recursive self-improvement and pacing visibility; not the parent KPI metric-panel hero.

Anthropic’s new index rates how much of its own AI research Claude performs on an Epoch AL0–AL5 scale. Here is what “leads,” “collaborates,” and “fully autonomous” mean — and why the metric sits in the pacing debate.

Frontier labs increasingly use their models to help build the next models. That feedback loop is useful for speed and for running more tests before release. It is also the heart of the recursive self-improvement worry: if AI does more of AI R&D end to end, outsiders have a harder time seeing how fast capability is compounding. Anthropic’s September institute post and Reuters coverage of the same release introduced public measurements meant to illuminate that pace. Our news brief is here: https://www.aitechdaily.com/anthropic-claude-leads-rd-automation/. This piece explains what the R&D Automation Index measures, how the automation levels work, and why the company says the numbers matter.

What the index is

The Anthropic R&D Automation Index is a prototype score of how much of Anthropic’s AI research and development Claude performs. The company says it catalogs kinds of AI R&D work inside the lab, rates how automated each task currently is, and aggregates those ratings.

Methodologically, Anthropic describes a bottom-up map built from work records. In its appendix, it says Claude research agents sampled staff weeks in July 2026, producing on the order of 15,000 granular model R&D tasks, then organized them into a hierarchical tree with 542 nodes. Those task categories are frozen as a basket so later measurements compare like with like. Separate Claude agents gather evidence on how each category is done; an independent Claude judge assigns an automation level. Person-time weights decide which categories count more in the aggregate. Anthropic also reports checks against staff ratings and plans for third-party evaluators with access comparable to internal risk teams.

The index is therefore an internal production metric made public — not a consumer product score, and not a claim about what Claude can do for outside customers.

The Epoch AL0–AL5 ladder

Anthropic uses an automation rating scale developed by Epoch AI. It runs from AL0 (no AI involvement) to AL5 (AI operates fully autonomously, with no human in the loop). The middle rungs are the ones that matter for reading the headlines:

  • **AL3 — collaborates:** AI can do large chunks of work under close human direction. Humans stay actively involved when new problems appear.
  • **AL4 — leads:** AI can complete most of a task end-to-end from a high-level prompt while a human supervises. The human may review and decide whether something ships, but does not have to stay tuned in for every surprise.
  • **AL5 — fully autonomous:** AI would monitor, investigate, fix, test, and deploy without needing a human to bring the issue to its attention. Anthropic says it has not measured this level for any subset of the work in the index.

Anthropic’s own footnotes make the difference concrete with a broken data-pipeline example: at AL3 the engineer stays in the loop when new issues appear; at AL4 Claude works through the failure and tags a human for a ship decision; at AL5 Claude would not need the human to notice the failure at all.

What Anthropic reported for August 2026

As of August 2026, Anthropic and Reuters reported:

  • Claude “leads” **26%** of Anthropic’s AI R&D work (AL4 share), up from about **1%** in March.
  • The share of work at or above “AI collaborates” (AL3+) is **at least 90%**.
  • Claude is **not** operating fully autonomously (AL5) for any measured subset of AI R&D work.

Those three sentences are the core of the index release. They say a lot of research work already involves Claude as a collaborator, a rising minority is work Claude leads under supervision, and none of the measured basket is fully hands-off.

Anthropic also published two companion measurement families that sit next to the index. On agent oversight, it said about 30,000 agents were doing research and engineering work on its main internal platform at a time in August; every action those agents take is screened before execution; of more than a billion decisions that month, about 1 in 47,000 was blocked. On compute, in a sample July week, about 6% of compute used for AI R&D went to safety work, and about 12% of compute used for AI-driven AI R&D went to safety — figures Anthropic called conservative because dual-purpose work was counted as capability.

Why the metric matters

Capability evals ask what a model can do. The Automation Index asks a different question: how much of the lab’s own model-building work is already being done by the model. That is closer to the recursive self-improvement and pacing debate than a single coding benchmark.

If AL4 share rises quickly while AL5 remains zero, the story is still “AI leads more of the work, humans still supervise.” If AL5 ever appears in a measured subset, the story changes: some slice of R&D would no longer require a human to initiate or stay in the loop. Anthropic framed the measurements as public visibility tools for a moment when executives and policymakers are arguing over whether to slow frontier development. The company says it expects the numbers to shift if there were coordination on pacing, and that third-party evaluators will help verify safety practices and monitor metrics like these.

What the index does not do is invent a countdown to runaway self-improvement. It does not say Claude is fully autonomous today. It does not, by itself, prove that other labs’ internal automation looks the same. Cross-lab comparison still needs a common methodology and external checks — points Anthropic itself lists as obstacles.

For the company methodology and the Reuters wire figures, start with Anthropic’s institute post and our Story so far brief above. Read the AL labels as Anthropic’s application of Epoch’s scale, not as an independent audit of every task in the tree.

Sources

Story so far

0 0 votes
Article Rating
Subscribe
Notify of
0 Comments
Inline Feedbacks
View all comments
0
Would love your thoughts, please comment.x
()
x