Anthropic published metrics on autonomy in the development of Claude models; the methodology shows inconsistencies
Anthropic says that Claude leads 26 % of the work on developing its own models (level AL4), compared with less than 1 % in February 2026. However, agreement of only 33% between human evaluators and 59% between the model and humans calls the methodology into question.
Anthropic published metrics describing how autonomously the development of its own models proceeds. It uses the AL0 to AL5 scale from Epoch AI for this purpose. According to the company, 26 % of work was at level AL4 in August 2026, an increase from less than 1 % in February 2026. More than 90 % of work reaches at least level AL3. According to the company, no work reaches level AL5.
Epoch AI calls AL4 “AI leads”, but according to the article, this does not mean full autonomy – that corresponds only to level AL5. Anthropic gives an example: an engineer passes a bug description to Claude, the model analyzes, fixes and tests it without further questions, but it is not allowed to deploy the fix on its own – a human makes that decision. A human thus still defines the task and the direction of the work; the main difference from level AL3 (“collaborates”) is that the model does not wait for help when a complication arises.
The evaluation was conducted by Claude itself: agents collected supporting material from Slack and internal documents, and another Claude model then assigned the individual levels. Anthropic itself points out that this “evaluator” model may make the same mistakes as the system it assesses. A review showed that when two employees evaluated the same area of work, they agreed on the level in only about 33 % of cases. The Claude model that assigned the official scores agreed with the human evaluation in 59 % of cases; in 97 % of cases, they differed by at most one level – but that one level corresponds exactly to the difference between “collaborates” (AL3) and “leads” (AL4). Moreover, the metric is based on how many hours of human work a given activity takes, rather than how many decisions the model actually makes or how much influence it has on the direction of research.
The publication is linked to a call by Anthropic CEO Dario Amodei for a coordinated slowdown of development at the frontier of AI capabilities; the company says the public needs more information for this and that these metrics complement tests of model capabilities.
Why it matters
It appears that AI companies can evaluate themselves using metrics whose methodology has significant weaknesses – low agreement between human evaluators (33 %) and between the model and humans (59 %) means that the presented figure of 26 % is largely a matter of interpretation. This is relevant to anyone assessing claims by AI companies about the capabilities or autonomy of their systems, as well as to the debate about regulation and a voluntary slowdown in AI development that Anthropic itself is initiating.
Two audiences, two different impacts
What this means
For individuals
When assessing similar claims about the performance or autonomy of AI tools, the measurement methodology needs to be verified, not just the resulting percentage – here, the reported 59% and 33% agreement between evaluators is evidence that the figure of 26 % can be interpreted in different ways.
For a business
Companies considering how much to trust their own and third-party metrics about AI systems “leading” work should bear in mind that a model evaluating itself (as Anthropic did) may have low reliability, which is also relevant to debates about AI safety regulation.
Risks and complianceCheck the original
Event sources
only one source so far · 1 publisher, 1 independent. We count feeds from the same owner only once.