Switch language한국어
Back to the list

OpenAI releases recorded and real-time transcription voice model "GPT-Live-Transcribe"

TL;DR AI

Key summary

2 min read
  1. OpenAI launched two new API transcription models: GPT-Live-Transcribe for low-latency live speech and GPT-Transcribe for recorded audio and batch jobs.

  2. Both models can use context like topic, names, jargon, and language hints to improve accuracy on accents, mixed-language speech, numbers, and technical terms.

  3. GPT-Live-Transcribe streams partial results as audio arrives, while GPT-Transcribe supports files up to 25 MB, multiple formats, streaming output, and language detection.

  4. OpenAI also highlighted related tools for speaker diarization, subtitles, and translation, and said the new models performed better in its evaluations.

Read the original