In This Article
Key Takeaways
- Self-reported AI automation percentages are becoming a standard PR move — the number alone tells you almost nothing without the definition behind it.
- Anthropic's September 2026 claim that Claude "leads" 26% of its R&D is a useful worked example because Anthropic published its methodology, not just the headline.
- The four questions that matter: who measured it, against what definition, was it verified externally, and what's explicitly excluded from the claim.
- The presence of caveats in a company's own disclosure is a stronger trust signal than the size of the number itself.
Why this matters right now
"AI now does X% of our work" is becoming a standard sentence in AI company communications, and the number is almost always self-reported, unaudited, and defined by the company making the claim. That doesn't make it worthless — it makes it a claim you need to know how to read, the same way you'd read a vendor's own benchmark result or a startup's own retention numbers. Anthropic's R&D Automation Index, published September 17, 2026, is a genuinely useful worked example because Anthropic did something most companies making this kind of claim don't: it published the actual methodology alongside the headline.
The worked example: Claude "leads" 26% of R&D
Anthropic's claim is that Claude "leads" 26% of the company's AI research and development work as of August 2026, up from under 1% in February. It used an external Automation Level (AL) scale built by the nonprofit Epoch AI, running AL0 (no AI) through AL5 (fully autonomous). "Leads" specifically means AL4: Claude completes most of a task end-to-end from a high-level prompt, with a human still supervising. According to reporting on the methodology, Anthropic sampled 20% of staff weekly, had a Claude agent catalog roughly 15,000 tasks, and had a separate Claude judge rate the automation level of each — matching staff's own ratings exactly 59% of the time and within one level 97% of the time.
Four questions to ask before you trust the number
Who measured it, and did anyone outside the company check it? Anthropic states directly that no outside party has verified its figures and that it is now working to embed third-party evaluators. A claim with that admission attached is more trustworthy than a clean number with no methodology at all — the absence of a caveat is usually the red flag, not the presence of one.
What exactly does the top-line word mean? "Leads" sounds like autonomy. It isn't — it's AL4, work a human still supervises end to end. Anthropic is explicit that no category reached AL5, fully autonomous. Vendors who use a word like "automated" or "AI-driven" without defining the tier are asking you to fill in the strongest possible interpretation yourself.
Who built the measurement scale? Anthropic used a third-party framework (Epoch AI's) instead of inventing its own rubric. That's a meaningfully different claim than a company grading itself on a scale it also designed, because an external scale is at least theoretically comparable across companies later.
What's the sample and how was it judged? Anthropic's own number came from an AI judge (a Claude agent) rating tasks against staff self-reports, agreeing exactly 59% of the time. That's a real, disclosed limitation — a model checking a model's own work can share the same blind spots. A claim that doesn't tell you who or what did the judging isn't giving you enough to evaluate it.
A quick red-flag / green-flag checklist
- Green flag: the company names the measurement scale, methodology, and sample size.
- Green flag: the disclosure includes explicit exclusions ("not autonomous in any category," "one week isn't a trend").
- Red flag: a percentage with no definition of what counts as "automated," "AI-led," or "AI-driven."
- Red flag: no mention of who verified the number, or whether it was verified at all.
- Red flag: the claim conflates "collaborates with AI" (a human still does the work, with AI assistance) and "AI leads" (AI does the work, a human supervises) as if they're the same thing.
Applying this to vendor claims you actually evaluate
The same four questions work on any vendor pitch that leans on an automation percentage — a coding-assistant company claiming its tool "writes 40% of production code," a customer-service AI vendor claiming it "resolves 70% of tickets," or a federal AI offeror citing an internal efficiency metric in a proposal. Ask what tier the number describes, who checked it, and what's explicitly excluded. If a vendor can't answer those in one sentence each, treat the number as marketing until proven otherwise. Anthropic's disclosure is a good template precisely because it survives that scrutiny better than most — not because the 26% is impressive on its own.
Sources: Anthropic — Measurements for understanding the pace of AI development inside frontier labs; implicator.ai — Anthropic Says Claude Leads 26% of Its AI R&D Work. Analysis and framing by Precision AI Academy.
Common questions
Is a self-reported AI automation number ever trustworthy? It can be a useful signal if the company discloses its methodology, definitions, and limitations — as Anthropic did. Treat an unexplained percentage with no methodology as marketing, not evidence.
What's the difference between AI "collaborating" and AI "leading" on a task? In Anthropic's framework (built on Epoch AI's AL scale), "collaborates" (AL3) means AI does large chunks of work under close human direction. "Leads" (AL4) means AI completes most of a task end-to-end from a high-level prompt, with a human supervising the result — not the same as full autonomy (AL5).
Why does it matter who built the measurement scale? A scale built by an independent third party (like Epoch AI) is at least theoretically comparable if other companies report on the same axis later. A company grading itself on its own bespoke rubric has no such check.