Anthropic Says Claude Now Leads 26% of Its Own AI R&D Work

Anthropic Says Claude Now Leads 26% of Its Own AI R&D Work

In This Article

  1. What Anthropic disclosed
  2. The AL0-AL5 scale behind the number
  3. How the number was actually produced
  4. The caveats Anthropic put on its own number
  5. The scale of the operation behind the number
  6. What this does and doesn't mean

Key Takeaways

What Anthropic disclosed

On September 17, 2026, Anthropic published the first results of what it calls the R&D Automation Index: a self-measured accounting of how much of Anthropic's own AI research and development work Claude currently performs. The headline number is that Claude "leads" 26% of that work as of August 2026, up from under 1% in February. More than 90% of the work now involves Claude at or above a "collaborates" level, though Anthropic is explicit that no measured category has reached full autonomy.

The AL0-AL5 scale behind the number

Anthropic didn't invent the measurement scale itself. It adopted an Automation Level (AL) framework developed by the independent research nonprofit Epoch AI, running from AL0 (no AI involvement) to AL5 (fully autonomous, no human in the loop). AL3 means Claude completes large chunks of work under close human direction; AL4 ("leads") means Claude completes most of a task end-to-end from a high-level prompt while a human supervises; AL5 hasn't been reached in any category Anthropic measured. Using an external scale rather than an in-house rubric is the point — it's meant to let other labs report comparable numbers on the same axis later.

How the number was actually produced

The methodology, per Anthropic's report, sampled 20% of staff across relevant departments every week through July 2026. A Claude research agent gathered evidence on roughly 15,000 granular tasks, organized into 542 hierarchical nodes and weighted by how much person-time each task actually takes. A separate Claude "judge" then assigned automation levels to each one. Those AI-assigned ratings matched staff's own self-ratings exactly 59% of the time, and fell within one level 97% of the time.

The caveats Anthropic put on its own number

Anthropic is unusually direct about the limits here. The report states plainly that using its own models to evaluate its own systems "could mean that the 'judge' model could make the same kinds of errors as the model it is checking." No outside party has independently verified any of the figures, and Anthropic says it's now working to embed third-party evaluators for future rounds. The company also reported that around 6% of AI R&D compute went to safety work in the week of July 13-20, rising to about 12% for AI-driven R&D specifically — but it cautions that one week of data isn't enough to call that a trend.

The scale of the operation behind the number

Separately from the automation percentage, Anthropic disclosed that roughly 30,000 agents were running simultaneously on its primary internal platform at any given point in August. Of more than 1 billion agent decisions, online monitoring blocked about 1 in 47,000, with 100% real-time coverage and offline review flagging roughly 1 to 2 transcripts per 1,000 for closer inspection.

What this does and doesn't mean

Read narrowly, this is not a claim that Claude is building its own successor unsupervised. Anthropic states outright that Claude is not operating autonomously in any measured category of AI R&D — the 26% figure describes work where a human still reviews the output, not work happening with no one watching. What it does show, on Anthropic's own numbers, is a real and fast-moving shift: from a company where AI barely touched its research pipeline in February to one where a quarter of that pipeline is now driven by Claude with a human in the loop, seven months later.

Sources: Anthropic — Measurements for understanding the pace of AI development inside frontier labs; implicator.ai — Anthropic Says Claude Leads 26% of Its AI R&D Work. Analysis and framing by Precision AI Academy.

Common questions

Does this mean Claude is building itself with no human involved? No. Anthropic states Claude is not operating fully autonomously (AL5) in any measured category. The 26% "leads" figure is AL4, meaning a human still supervises the work end-to-end.

Who verified Anthropic's numbers? No outside party has independently verified the figures yet. Anthropic disclosed the numbers itself and says it is now working to embed third-party evaluators for future measurement rounds.

What is the AL0-AL5 scale? An Automation Level framework built by the independent nonprofit Epoch AI, running from AL0 (no AI involvement) to AL5 (fully autonomous). Anthropic adopted it rather than building its own rubric so the numbers are comparable if other labs report on the same scale.

About Precision AI Academy

Precision AI Academy publishes practical AI news, plain-language analysis, and free courses for builders and working professionals. It is a sister site of Precision Federal, a federal software and AI firm. We verify the numbers, cite the primary sources, and skip the hype.