GPT Transcribe improves on its predecessor but can't catch ElevenLabs, Google, or Mistral on error rates
OpenAI introduces GPT Transcribe and GPT Live Transcribe, offering faster pre-recorded audio processing (34x real-time) and low-latency real-time streaming. GPT Transcribe achieves a word error rate (WER) of 3.31% on the AA-WER benchmark, improving by 0.7 percentage points over its predecessor, with a 25% price reduction to $0.0045 per minute. Despite improvements, OpenAI lags behind competitors like ElevenLabs Scribe v2 (2.3% WER), Google Gemini 3 Pro (2.9% WER), and Mistral Voxtral Small (3% W
Analysis
TL;DR
- OpenAI introduces GPT Transcribe and GPT Live Transcribe, offering faster pre-recorded audio processing (34x real-time) and low-latency real-time streaming.
- GPT Transcribe achieves a word error rate (WER) of 3.31% on the AA-WER benchmark, improving by 0.7 percentage points over its predecessor, with a 25% price reduction to $0.0045 per minute.
- Despite improvements, OpenAI lags behind competitors like ElevenLabs Scribe v2 (2.3% WER), Google Gemini 3 Pro (2.9% WER), and Mistral Voxtral Small (3% WER).
- The models support text context, keywords, and multiple languages, aligning with OpenAI’s broader Realtime model generation initiative.
Why It Matters
This update reflects the rapid evolution in speech recognition technology, where accuracy, speed, and cost efficiency are critical for enterprise and developer adoption. For AI practitioners, understanding these benchmarks helps inform model selection based on use case—whether prioritizing low latency, high throughput, or competitive pricing. The continued pressure from rivals underscores the need for continuous innovation in multimodal AI systems.
Technical Details
- GPT Transcribe: Processes pre-recorded audio at ~34x real-time speed; supports text context, keyword injection, and multilingual input; achieves 3.31% WER on AA-WER benchmark.
- GPT Live Transcribe: Designed for real-time streaming with minimal latency, suitable for live applications such as meetings or customer service interactions.
- Pricing Model: Reduced to $0.0045 per minute of audio processed—a 25% drop from prior pricing—making it more accessible for large-scale deployments.
- Benchmark Performance: Evaluated via Artificial Analysis’ AA-WER metric, which measures transcription accuracy across diverse audio conditions and speaker demographics.
- Integration: Part of OpenAI’s Realtime ecosystem, complementing GPT-Realtime-Whisper and enabling seamless API-based deployment within existing workflows.
Industry Insight
The competitive landscape in speech-to-text is intensifying, with Mistral undercutting pricing at $0.003/minute while maintaining strong performance—suggesting that cost leadership may become a key differentiator alongside accuracy. Developers should evaluate trade-offs between model fidelity, latency, and budget when selecting transcription services. As real-time capabilities grow in demand for interactive AI applications, providers must balance innovation with scalability and affordability to maintain market relevance.
Disclaimer: The above content is generated by AI and is for reference only.