NEWS / SEP.2026
Anthropic internally estimates that Claude leads 26 % of measured, weighted AI R&D work
Anthropic estimates that Claude led, under human supervision, 26 % of the weighted AI R&D work measured at the company in August 2026. Its new internal index distinguishes this delegation from full autonomy, which none of the assessed categories reached.

According to Anthropic, Claude leads 26 % of measured AI R&D work, under human supervision
Anthropic announced on September 17, 2026 that, according to its internal estimate, Claude had led 26 % of the weighted artificial intelligence research and development work measured at the company in August, under human supervision. The announcement, reported the same day by Reuters, is based on an internal prototype called the Anthropic R&D Automation Index.
The AL4 level assigned to this share of the work corresponds to delegation under supervision. A person provides a general objective, Claude performs most of the task, and a human supervises. The intended benefit for researchers and engineers is less intervention as the work proceeds.
On the Epoch AI scale used by Anthropic, AL3 denotes large portions of work performed by AI under close human direction. More than 90 % of the measured work reaches AL3 or above, a proportion that includes the 26 % classified as AL4. No measured category reaches AL5, full autonomy with no human required. The index therefore describes delegated work involving human intervention.
The engineer decides whether to deploy
Anthropic illustrates the difference between collaboration and leading work under supervision with an illustrative example of an overnight data pipeline failure. At AL3, the engineer directs the investigation and decides how to handle additional difficulties.
In the AL4 version of this example, the engineer passes the alert and an objective to Claude. The agent examines the execution logs, finds the cause, and writes a fix, which it tests. It then checks a copy of the data and documents the changes. The engineer reads the report and then decides whether to deploy.
If this delegation expands, teams could run more experiments and have more results to validate. However, the published figures concern the distribution of work and do not quantify hours saved.
A July basket for measuring August
The basket of categories is based on work from July 2026. Each week, a random sample of 20 % of employees in each relevant department was used to compile the inventory. Claude consulted Slack and internal documentation to identify approximately 15,000 tasks and group them into categories.
This basket is then fixed to allow comparisons. The weighting assigns one unit to each person per week, distributes it equally among their tasks, and then adds these fractions together by category. The August results therefore use a basket and weights constructed in July. The index approximates a share of human work, rather than simply counting tasks or measuring jobs eliminated.
One Claude agent investigates how the work in the categories is performed. Another Claude agent acts as the judge and assigns them an automation level. To assess a month, the judges use only information from that month or earlier months.
Agents’ actions pass through an automated filter
In a separate measurement, Anthropic reports approximately 30,000 agents simultaneously active in research and engineering on its most widely used internal platform in August. The lab says that all their actions undergo an automated check before execution, generally within a few seconds. Monitoring after execution also analyzes their traces. Each action is checked automatically, while human supervision focuses on the task assigned to Claude.
Of more than a billion decisions analyzed in August, the pre-execution monitor reportedly blocked approximately 0.002 %, or nearly one decision in 47,000, according to Anthropic.
Internal measurements to be replicated
Using Claude as a judge exposes the index to errors shared across models. Anthropic also acknowledges debatable boundaries between AL3 and AL4, as well as approximate categories and weighting. The fixed basket tracks the automation of work already inventoried and provides an imperfect account of new activities.
For monitoring, the blocking rate measures how frequently the monitor intervenes. Its effectiveness also depends on the dangerous actions it lets through. Tests must measure these errors. Anthropic mentions earlier monitoring trials by METR and plans to embed external evaluators.
These measurements complement Anthropic’s promise to host external evaluators, whose implementation remains to be documented.
Independent replication would allow the R&D figures and the sensitivity of the checks to be tested. To make a subsequent measurement comparable, the basket would need to be retained or its version change documented.