OpenAI introduces Ultrafast mode for the GPT-6 Astra model in the Responses API
OpenAI has added a new "ultrafast" service tier to the Responses API for the GPT-6 Astra model, which reduces latency between generated output tokens. The feature does not support EU or other regional data residency.
On 29 September 2026, OpenAI added a so-called Ultrafast mode to the Responses API for the GPT-6 Astra model. It is activated via the service_tier parameter set to the value "ultrafast" and, according to the company, reduces the time between individual generated output tokens.
The feature is available to API customers, is subject to rate limits, and uses global processing with data storage in the USA. Neither EU nor other regional residency for inference is supported, which is relevant for customers with data processing localization requirements.
According to the company, Ultrafast mode has its own pricing structure different from the standard tier, but the source does not state specific prices. For details, see the source article.
Why it matters
Lower latency between tokens is relevant for applications dependent on fast response, such as voice assistants or real-time chat interfaces. However, the lack of EU data residency support means that companies with regulatory requirements for data storage in Europe cannot yet deploy this option.
Two audiences, two different impacts
What this means
For individuals
Developers building their own applications on top of the GPT-6 Astra model API can speed up response generation in latency-sensitive features by enabling service_tier ultrafast.
For a business
Companies running products with the GPT-6 Astra model via API may consider switching to the ultrafast tier for faster response times; however, this option is currently unavailable for customers with an EU data residency requirement.
DevelopmentCheck the original
Event sources
clearly official source · 1 publisher, 0 independent. We count feeds from the same owner only once.