We calculated the impact of the new GPT-5.6 prices: how Luna and Terra compare with competitors
A price cut of 80 or 20 percent sounds straightforward, but the actual bill depends on the ratio of input, cached and output tokens. We therefore prepared a comparison of three identical hypothetical usage scenarios — from a smaller assistant to a larger business deployment — and placed the new prices alongside public rates for comparable competing models.
What we calculated
The same workload instead of just a percentage
For each model, we calculate exactly the same volume of input, cached and output tokens. A toggle lets you change the hypothetical scenario and the usage multiplier. The result is an estimated cost based on public API pricing, rather than a measurement of response quality, energy consumption or the total cost of implementing AI.
Same workload, same units
GPT-5.6 Luna: impact on costs
1 million input and 250 thousand output tokens per period.
Modelled cost decrease of 2 USD (80 %).
What changed in pricing
| Cost component | Before | After | Change |
|---|---|---|---|
| Input per 1 million tokens | 1 USD | 0.2 USD | -80 % |
| Cached input per 1 million tokens | 0.1 USD | 0.02 USD | -80 % |
| Output per 1 million tokens | 6 USD | 1.2 USD | -80 % |
How we calculate comparisons
We calculate the same volume of new input, cached input and output tokens using standard API prices in USD. We exclude taxes, batch and priority modes, long context, tools and hourly cache storage. The comparison shows the price of the same token workload, not the same output quality; providers also use different tokenization. The old Luna price is derived exactly from the new official price and the announced 80% price reduction.
An incomplete price list, a missing mandatory component or a different currency is not filled in with estimates for an exact comparison.
How to use the result
Impact on users and businesses
For individuals and smaller teams
For individuals, developers and smaller teams, the most noticeable change is with GPT-5.6 Luna. In scenarios with repeated context, the cached input rate also has a significant effect alongside the base price. The calculation therefore shows a specific bill rather than just a marketing percentage.
For businesses and larger operations
Companies should not apply the announced discount percentage directly to their entire AI budget. Actual savings depend on the usage mix, cache utilization, context length and the selected service mode. The same calculation can inform budgeting and a fresh build-versus-buy comparison.
AI Radar verdict
What the change means for competition
The new prices increase pressure on competitors, especially for high-volume API tasks, but the cheapest model may not automatically be the best choice. The decision should combine the price of the same workload with output quality, latency, limits and the cost of any change of provider.