Skip to content

Anthropic · The most demanding information tasks

Claude Opus 5

A choice for complex analysis, long tasks involving multiple steps and outputs where quality matters more than speed and cost.

Compare with another model

Costs
High · $5 per million input tokens and $25 per million output tokens; a token is a small piece of text
Speed
Rather slow at the most thorough setting
Availability
Online service and apps from Anthropic
Input length
Very long input · up to 1 million tokens (small pieces of text)
Inputs
Text and image inputs

Data checked . Release date unverified.

Indicative capability profile

Where the model is strong

The five levels are our clear summary of the results below. They are not a ranking that applies to every task.

Analysis Among the best at solving specialist problems.
Coding Among the best at coding with additional tools.
Autonomous work across multiple steps Strong at long, multi-step tasks.
Writing and content High-quality structure and analysis, though not always the fastest way to produce text.
Speed The most thorough setting can add tens of minutes of waiting.

Ideal use

When to choose it

  • Complex research and supporting information for decisions
  • Long tasks the model solves in several steps
  • Advanced coding and independent code review

Usage boundaries

When to choose another model

  • Routine sorting and simple transcriptions
  • Interactive tasks requiring an immediate response

Traceable supporting sources

Data and measurement sources

Distinguish between the manufacturer's documentation and the results of a specific test. Measurements also depend on the settings and the task set used.

Artificial Analysis

Intelligence Index v4.1

61 points

At the top of the aggregate measurement as of 24 July 2026.

Open original source ↗
Artificial Analysis

AA-Briefcase

1720 Elo points

A higher number is better. The model led in working with information; its tools were also tested.

Open original source ↗
Artificial Analysis

Terminal-Bench v2.1

89 %

Very good result in coding and terminal work.

Open original source ↗

Trends over time

Related events from AI Radar

Search more →
Anthropic worth noting

Astra and Claude Opus 5 models help decipher two long-unsolved Enigma messages

The difference in approach is notable: the Astra model worked with minimal input and sourced the context itself, while the Claude Opus 5 model needed more concrete guidance from a human – so when assigning similar AI research tasks, it's worth considering how much initiative to expect from the model.

The case shows that current state-of-the-art models can autonomously carry out multi-stage archival research, build simulations, and solve tasks without a clearly defined procedure – a relevant signal for companies considering deploying AI agents on complex analytical and research tasks where no ready-made methodology exists.

1 source
Anthropic worth noting update

Anthropic challenges US approach in the Claude Fable 5 and Mythos 5 export control case, responds with jailbreak severity framework

New: Export control issued due to concerns about a jailbreak and the model's cyber capabilities; Anthropic challenged the legality of the export control and the lack of a clear process; Anthropic responded by publishing a jailbreak severity framework

Security researchers and developers now have access to a formal Anthropic program and framework for reporting and assessing the severity of jailbreaks affecting Claude models.

Companies dependent on the Claude Fable 5 or Mythos 5 models face the risk that US export controls could temporarily restrict their access to the model without a clearly defined process, which is relevant for supply chain and geopolitical risk planning.

✓ 2
Anthropic worth noting

Anthropic pushes for mandatory AI system safety audits in the US

The proposed legislation (e.g., the Massachusetts bond bill) would introduce an obligation to disclose safety frameworks and undergo independent third-party audits for companies developing or deploying AI systems, which would increase compliance costs and risk-documentation requirements.

1 source

Next step

Compare the model using your actual task.

A benchmark narrows the selection. A short trial on your data determines the choice.

Open comparison →