Skip to content
AI Radar analysis price change

We calculated the impact of the new GPT-5.6 prices: how Luna and Terra compare with competitors

Author AI Radar

A price cut of 80 or 20 percent sounds straightforward, but the actual bill depends on the ratio of input, cached and output tokens. We therefore prepared a comparison of three identical hypothetical usage scenarios — from a smaller assistant to a larger business deployment — and placed the new prices alongside public rates for comparable competing models.

What we calculated

The same workload instead of just a percentage

For each model, we calculate exactly the same volume of input, cached and output tokens. A toggle lets you change the hypothetical scenario and the usage multiplier. The result is an estimated cost based on public API pricing, rather than a measurement of response quality, energy consumption or the total cost of implementing AI.

Same workload, same units

GPT-5.6 Luna: impact on costs

1 million input and 250 thousand output tokens per period.

Editorially verified July 31, 2026
Before the change 2.5 USD
After the change 0.5 USD

Modelled cost decrease of 2 USD (80 %).

OpenAI · Before the change GPT-5.6 Luna
2.5 USD 100.0 %
OpenAI: Advancing the price-performance frontier with GPT-5.6 ↗ Rate effective from to
OpenAI · After the change GPT-5.6 Luna
0.5 USD 20.0 %
OpenAI API: GPT-5.6 Luna model pricing ↗ Rate effective from
Google · Comparison Gemini 3.1 Flash-Lite
0.625 USD 25.0 %
Google AI for Developers: Gemini API pricing ↗ Rate effective from
Anthropic · Comparison Claude Haiku 4.5
2.25 USD 90.0 %
Claude Platform Docs: pricing ↗ Rate effective from
xAI · Comparison Grok 4.3
1.88 USD 75.0 %
xAI Docs: pricing ↗ Rate effective from

What changed in pricing

Cost componentBeforeAfterChange
Input per 1 million tokens 1 USD 0.2 USD -80 %
Cached input per 1 million tokens 0.1 USD 0.02 USD -80 %
Output per 1 million tokens 6 USD 1.2 USD -80 %
How we calculate comparisons

We calculate the same volume of new input, cached input and output tokens using standard API prices in USD. We exclude taxes, batch and priority modes, long context, tools and hourly cache storage. The comparison shows the price of the same token workload, not the same output quality; providers also use different tokenization. The old Luna price is derived exactly from the new official price and the announced 80% price reduction.

An incomplete price list, a missing mandatory component or a different currency is not filled in with estimates for an exact comparison.

How to use the result

Impact on users and businesses

01

For individuals and smaller teams

For individuals, developers and smaller teams, the most noticeable change is with GPT-5.6 Luna. In scenarios with repeated context, the cached input rate also has a significant effect alongside the base price. The calculation therefore shows a specific bill rather than just a marketing percentage.

02

For businesses and larger operations

Companies should not apply the announced discount percentage directly to their entire AI budget. Actual savings depend on the usage mix, cache utilization, context length and the selected service mode. The same calculation can inform budgeting and a fresh build-versus-buy comparison.

AI Radar verdict

What the change means for competition

The new prices increase pressure on competitors, especially for high-volume API tasks, but the cheapest model may not automatically be the best choice. The decision should combine the price of the same workload with output quality, latency, limits and the cost of any change of provider.