Skip to content

Model radar

Choose a model based on the work, not the logo.

We supplement verified vendor data with editorial recommendations: which tasks to entrust to the model, when to choose a cheaper option, and when to combine multiple models.

There is no single best model. A more capable model may not be faster or more cost-effective. Compare use cases, pricing and availability. Check speed on your own task.

8 current profiles 4 combinations 4 methodologies

Choose by task ↓ Compare two models →

Start with your task

What do you need to do?

Choose the type of work. We will show models suited to it based on documented capabilities.

Clear filter ×

Practical selection

Models for: Write or fix code

Starting with the latest documented release. Entries with the same date are sorted by name; unknown dates come last. Newer does not mean better.

Open comparison →
Data for the current selection checked The selection is neither a popularity ranking nor a complete catalog.

Online model prices are standard API rates per million tokens (pieces of text), excluding discounts for repeated input and batch processing. App subscriptions and tool usage are billed separately.

Anthropic More complex analysis and development

Claude Opus 5.5

Published

The current version of the Opus series for working with long prompts, text and code.

Usage and availability
Best used for
  • Programming and multi-step tasks
  • Analysis of long documents and images

Claude API, Amazon Bedrock, Google Cloud and Microsoft Foundry

Costs
$4 input / $20 output per million tokens in standard API mode.
Strengths and supporting evidence →
OpenAI Everyday work and programming

GPT-6 Sol

Published

A versatile choice for working with text, code and images.

Usage and availability
Best used for
  • Writing and fixing code
  • Processing documents and multi-step tasks

OpenAI API

Costs
$2 input / $10 output per million tokens. Above 272 thousand input tokens: 2× input and 1.5× output.
Strengths and supporting evidence →
xAI Code and working with information

Grok 4.7

Published

A model for coding and multistep tasks involving text and images.

Usage and availability
Best used for
  • Code development and fixes
  • Analysis of text and image materials

xAI API, Cursor and Grok Build

Costs
$2 input / $6 output per million tokens below 200 thousand input tokens; above this threshold, $4 / $12.
Strengths and supporting evidence →
OpenAI Complex tasks from brief to result

GPT-6 Astra

Published

A model designed for complex analysis, development and tool use.

Usage and availability
Best used for
  • Analysis of complex problems
  • Research, document creation and larger programming tasks

OpenAI API

Costs
$10 input / $50 output per million tokens. Above 272 thousand input tokens: 2× input and 1.5× output.
Strengths and supporting evidence →
Google Documents, media and automation

Gemini 3.8 Flash

Published

Processes text, images, audio and video; the output is text.

Usage and availability
Best used for
  • Analysis of documents, audio and video
  • Programming and multi-step business tasks

Gemini API and Google AI Studio

Costs
$0.75 input / $3.75 output per million tokens through December 31, 2026. From January 1, 2027: $1.50 / $7.50.
Strengths and supporting evidence →
Anthropic Long and demanding tasks

Claude Fable 5.1

Published

A variant for demanding work where longer processing times and higher costs are acceptable.

Usage and availability
Best used for
  • Complex tasks with multiple sequential steps
  • Working with extensive context and code

Claude API, Amazon Bedrock, Google Cloud and Microsoft Foundry

Costs
$10 input / $50 output per million tokens in standard API mode.
Strengths and supporting evidence →
BottleCap AI Downloadable model subject to license terms

ThinkingCap-Qwen3.8-27B

Release date unverified

A modification of Qwen3.8-27B focused on shorter internal task processing.

Usage and availability
Best used for
  • Testing the model on your own hardware
  • Token usage comparison with the default model

Hugging Face after accepting the access terms. PolyForm Small Business 1.0.0 license and additional permission for personal use.

Costs
Self-hosting costs. Commercial licensing under the terms of BottleCap AI; no uniform public price is listed.
Strengths and supporting evidence →
Older profiles · 10

Historical overview from 31. 7. 2026. Prices, availability and recommendations here have not been verified as current. This is not a list of discontinued models.

How to combine multiple models

One model does not have to do everything.

The following workflows are editorial suggestions to try. A low-cost model can sort simple tasks, a stronger one can solve difficult cases, and fixed rules can check the result.

01

Many tasks without unnecessary costs

Sorting and processing hundreds to thousands of similar inputs.

A low-cost model handles routine work, and the more expensive one gets only the truly difficult cases.

  1. 1. Sorting

    It identifies the task type, finds the necessary information and estimates its confidence.

  2. 2. Difficult cases

    When the first model is not confident enough, a stronger model takes over.

  3. 3. Checking

    Fixed rules in the program and a human check both the format and a sample of the results.

Caution: Use your own data to verify when a stronger model should take over the work. A model cannot reliably assess its own answer.

02

Quick code edits with a second check

Modifying a program in use where an error could cause harm.

One system makes the change, another looks for blind spots, and tests decide.

  1. 1. Editing

    Reviews the code, proposes the smallest possible change and runs the necessary tests.

  2. 2. Independent review

    It receives the task, the changes made and the test results, but not the original rationale.

  3. 3. Decision

    Automated checks, tests and human approval take precedence over model opinions.

Caution: A second model will not help if both receive the same incorrect assumption without evidence.

03

Documents → analysis → verification

Research, strategy or decision support based on a large number of documents.

A fast model prepares the material, and a stronger model focuses on analyzing it.

  1. 1. Preparing materials

    Finds important claims in documents and media and adds links to their sources.

  2. 2. Analysis

    It builds arguments, identifies contradictions and formulates recommendations from traceable source material.

  3. 3. Verification

    It looks for contradictions, but facts are ultimately checked against the original documents.

Caution: Agreement between two models is not proof; the cited inputs are what matters.

04

Sensitive data under your own control

Sensitive materials cannot be sent in full to an external online service.

The sensitive part stays on your own server, and the online service receives only the necessary data.

  1. 1. Processing on my own server

    A local model can help find facts. Verify the removal of sensitive data with a separate check before sending.

  2. 2. Optional analysis

    The online model receives only text without sensitive data and only the information it needs.

  3. 3. Return and review

    The result is combined with internal data only on your own server.

Caution: First, check the license for your intended use. Being able to run a model yourself does not in itself guarantee privacy. Server security and what data is stored also matter.

How to read the supporting material

Manufacturer data and firsthand experience carry different weight.

The current profiles are based on the linked documentation from the providers. We do not yet have our own comparative measurements of these versions. The following websites are additional sources for verification, not evidence for the ratings of all the models above. Some tests also evaluate the model together with the tools it uses.

How AI Radar works →
What we are not measuring yet

Public tests will not tell you how a model will handle your own texts, data and workflows. Before deciding, try it on a few real tasks. Prices are often quoted per token, meaning small pieces of text.