OpenAI has released new speech-to-text models gpt-transcribe and gpt-transcribe-live
OpenAI has released two new speech-to-text models – gpt-transcribe for recordings and gpt-transcribe-live for low-latency live transcription. According to the company, they are more accurate and support 57 languages. Currently only via the API.
OpenAI has released two new models for transcribing spoken words into text, called gpt-transcribe and gpt-transcribe-live. The model gpt-transcribe is designed to transcribe recorded audio files, while gpt-transcribe-live is optimized for low latency and is therefore better suited to transcribing live speech.
According to OpenAI, both models are more accurate than previous models and support 57 languages. They are currently available only via the API, and it is unclear whether and when they will also be made available in other ways, for example directly in an application.
You can find details in the source article.
Why it matters
Developers and companies building products based on speech transcription (meeting notes, captions, call transcription) gain access to a model that, according to OpenAI, improves output accuracy and expands language coverage to 57 languages. Because it is currently available only via the API, ordinary users cannot use it directly until it appears in a specific application or service.
Relevant practical impact
What this means
For a business
According to OpenAI, companies developing products with speech transcription gain access to a more accurate API model with support for 57 languages, but there is no ready-made end-user application yet – integration requires custom development via the API.
DevelopmentCheck the original
Event sources
only one source so far · 1 publisher, 1 independent. We count feeds from the same owner only once.