Technology

Claude Leads a Quarter of Anthropic's AI Research Work

Martin HollowayPublished 21h ago4 min readBased on 5 sources
Reading level
Claude Leads a Quarter of Anthropic's AI Research Work
Image by tstokes from Pixabay

Anthropic says its chatbot Claude now "leads" 26 percent of its AI research and development work. The figure was shared in a blog post proposing standard ways to report the pace of AI progress Engadget.

Anthropic defines "leads" as Claude completing most of a task end to end from a high-level prompt while a human supervises. That stops short of full autonomy. It is supervised work by an AI agent, a system that can carry out multi-step tasks, rather than independent research.

Anthropic placed the 26 percent inside a wider picture. On more than 90 percent of its research, AI does at least large chunks of work under close human direction, including the 26 percent Claude leads. At the same time, Anthropic said Claude is "not operating fully autonomously for any measured subset of AI R&D work".

The disclosure came in a post titled "Measurements for understanding the pace of AI development inside ..." published at Anthropic. Anthropic framed it as a proposal for how AI companies should report automation, oversight and resource inputs in a comparable way.

The first of three proposed measurements covers "AI-led AI R&D". To quantify it, Anthropic said it used an index of how much AI research is done by Claude and an automation rating scale from Epoch AI to plot Claude's automation level since August 2025. The intent is tracking over time, not a one-time snapshot.

The second proposal covers oversight of AI agents. Anthropic proposed measuring how much agent activity is monitored, how long review takes, and how often behavior is flagged. For teams running coding and research agents, those translate to coverage, review latency, the time to review, and flag rate.

The third proposal covers inputs. Anthropic proposed tracking how much compute, the processing power used to train and run models, is devoted to AI research as a way to assess if frontier development is appropriately paced. Compute can be measured and checked from outside in a way internal workflow data often cannot.

Earlier company data provides context without replacing the 26 percent figure. In research published in December 2025, Anthropic reported that 27% of Claude-assisted work at the company consists of tasks that wouldn't have been done otherwise Anthropic. In its March 2026 economic index report, the company reported that in February, the top 10 tasks made up 19% of all Claude.ai traffic, down from 24% in November, indicating broadening use. Separately, the system card for Claude Opus 5 stated that Claude Opus 5 does not cross the automated AI R&D capability threshold set out in Anthropic's RSP.

The broader context here is that self-measurement is becoming expected of leading labs. External benchmarks measure model capability. They do not measure how work actually happens inside the lab. What share is agent-led, how tightly it is supervised, and how much compute supports it can only be reported by the developer.

In my view, the most useful part is pairing automation level with oversight burden. A 26 percent lead rate means little without review coverage and flag rates. Supervised end-to-end work can still cost much human attention if outputs need line-by-line checks, revised prompts, or rework. Teams using code generation with automated testing will recognize the pattern. Output rises, review becomes the bottleneck, and quality depends on catching small errors quickly.

Looking at what this means for practitioners, the Epoch AI scale and the compute measure point to discussion that is easier to audit. If labs consistently report automation share, oversight delay and research compute, customers, auditors and regulators can compare claims instead of parsing blog posts. That does not settle questions about test validity or selective disclosure. It does create a baseline to track month to month.

Worth flagging is what Anthropic did not claim. No measured work is fully autonomous. More than 90 percent involves large AI contributions under close direction, but direction stays human. That keeps the present in the area of high-leverage tools and supervised agents, not self-improving systems without human review. The long-term optimism still holds. Systems that reliably handle routine research subtasks free researchers for design, interpretation and judgment, where breakthroughs tend to start.