Skip to content
context Other

Anthropic published metrics on autonomy in the development of Claude models; the methodology shows inconsistencies

only one source so far

Anthropic says that Claude leads 26 % of the work on developing its own models (level AL4), compared with less than 1 % in February 2026. However, agreement of only 33% between human evaluators and 59% between the model and humans calls the methodology into question.

Anthropic published metrics describing how autonomously the development of its own models proceeds. It uses the AL0 to AL5 scale from Epoch AI for this purpose. According to the company, 26 % of work was at level AL4 in August 2026, an increase from less than 1 % in February 2026. More than 90 % of work reaches at least level AL3. According to the company, no work reaches level AL5.

Epoch AI calls AL4 “AI leads”, but according to the article, this does not mean full autonomy – that corresponds only to level AL5. Anthropic gives an example: an engineer passes a bug description to Claude, the model analyzes, fixes and tests it without further questions, but it is not allowed to deploy the fix on its own – a human makes that decision. A human thus still defines the task and the direction of the work; the main difference from level AL3 (“collaborates”) is that the model does not wait for help when a complication arises.

The evaluation was conducted by Claude itself: agents collected supporting material from Slack and internal documents, and another Claude model then assigned the individual levels. Anthropic itself points out that this “evaluator” model may make the same mistakes as the system it assesses. A review showed that when two employees evaluated the same area of work, they agreed on the level in only about 33 % of cases. The Claude model that assigned the official scores agreed with the human evaluation in 59 % of cases; in 97 % of cases, they differed by at most one level – but that one level corresponds exactly to the difference between “collaborates” (AL3) and “leads” (AL4). Moreover, the metric is based on how many hours of human work a given activity takes, rather than how many decisions the model actually makes or how much influence it has on the direction of research.

The publication is linked to a call by Anthropic CEO Dario Amodei for a coordinated slowdown of development at the frontier of AI capabilities; the company says the public needs more information for this and that these metrics complement tests of model capabilities.

What changed

Why it matters

It appears that AI companies can evaluate themselves using metrics whose methodology has significant weaknesses – low agreement between human evaluators (33 %) and between the model and humans (59 %) means that the presented figure of 26 % is largely a matter of interpretation. This is relevant to anyone assessing claims by AI companies about the capabilities or autonomy of their systems, as well as to the debate about regulation and a voluntary slowdown in AI development that Anthropic itself is initiating.

Two audiences, two different impacts

What this means

01

For individuals

When assessing similar claims about the performance or autonomy of AI tools, the measurement methodology needs to be verified, not just the resulting percentage – here, the reported 59% and 33% agreement between evaluators is evidence that the figure of 26 % can be interpreted in different ways.

What to do When evaluating claims by AI companies about the degree of autonomy of their models, verify what methodology was used and who conducted the evaluation.
More practical updates →
02

For a business

Companies considering how much to trust their own and third-party metrics about AI systems “leading” work should bear in mind that a model evaluating itself (as Anthropic did) may have low reliability, which is also relevant to debates about AI safety regulation.

Risks and compliance
What to decide When deciding how much to trust safety or performance claims from AI vendors, consider low agreement between evaluators as a sign of uncertainty in the methodology.
More business impacts →
AI autonomie Anthropic Claude Epoch AI metriky vývoje

Check the original

Event sources

only one source so far · 1 publisher, 1 independent. We count feeds from the same owner only once.

1
The Decoder (daily AI news) independent context · first detected Anthropic wants you to know Claude leads a quarter of its research, but "lead" doesn't mean what you think