Skip to content
context New models

OpenAI has released new speech-to-text models gpt-transcribe and gpt-transcribe-live

only one source so far

OpenAI has released two new speech-to-text models – gpt-transcribe for recordings and gpt-transcribe-live for low-latency live transcription. According to the company, they are more accurate and support 57 languages. Currently only via the API.

OpenAI has released two new models for transcribing spoken words into text, called gpt-transcribe and gpt-transcribe-live. The model gpt-transcribe is designed to transcribe recorded audio files, while gpt-transcribe-live is optimized for low latency and is therefore better suited to transcribing live speech.

According to OpenAI, both models are more accurate than previous models and support 57 languages. They are currently available only via the API, and it is unclear whether and when they will also be made available in other ways, for example directly in an application.

You can find details in the source article.

What changed

Why it matters

Developers and companies building products based on speech transcription (meeting notes, captions, call transcription) gain access to a model that, according to OpenAI, improves output accuracy and expands language coverage to 57 languages. Because it is currently available only via the API, ordinary users cannot use it directly until it appears in a specific application or service.

Relevant practical impact

What this means

01

For a business

According to OpenAI, companies developing products with speech transcription gain access to a more accurate API model with support for 57 languages, but there is no ready-made end-user application yet – integration requires custom development via the API.

Development
What to decide Evaluate deploying the models gpt-transcribe and gpt-transcribe-live via the API for products that use speech transcription (e.g. call transcription, captions, meeting notes).
More business impacts →
API ASR gpt-transcribe OpenAI speech-to-text voice transcription

Check the original

Event sources

only one source so far · 1 publisher, 1 independent. We count feeds from the same owner only once.

1
Živě.cz independent context · first detected OpenAI má nový model pro přepis hlasu na text. Je přesnější než dřív a podporuje 57 jazyků