AI News AI资讯 7h ago Updated 2h ago 更新于 2小时前 46

GPT Transcribe improves on its predecessor but can't catch ElevenLabs, Google, or Mistral on error rates GPT Transcribe 改进了其前身,但在错误率上仍无法与 ElevenLabs、Google 或 Mistral 相媲美

OpenAI introduces GPT Transcribe and GPT Live Transcribe, offering faster pre-recorded audio processing (34x real-time) and low-latency real-time streaming. GPT Transcribe achieves a word error rate (WER) of 3.31% on the AA-WER benchmark, improving by 0.7 percentage points over its predecessor, with a 25% price reduction to $0.0045 per minute. Despite improvements, OpenAI lags behind competitors like ElevenLabs Scribe v2 (2.3% WER), Google Gemini 3 Pro (2.9% WER), and Mistral Voxtral Small (3% W OpenAI发布GPT Transcribe和GPT Live Transcribe两款语音识别模型,分别处理预录音频与实时流媒体。 GPT Transcribe错误率3.31%,较前代提升0.7个百分点,价格下调25%至每分钟0.0045美元。 在AA-WER基准测试中,OpenAI落后于ElevenLabs(2.3%)、Google(2.9%)和Mistral(3%)。 Mistral以更低定价($0.003/分钟)推出Voxtral Transcribe V2,加剧市场竞争。 新模型支持多语言输入、文本上下文及关键词辅助转录,并集成至OpenAI实时生成生态。

65
Hot 热度
70
Quality 质量
60
Impact 影响力

Analysis 深度分析

TL;DR

  • OpenAI introduces GPT Transcribe and GPT Live Transcribe, offering faster pre-recorded audio processing (34x real-time) and low-latency real-time streaming.
  • GPT Transcribe achieves a word error rate (WER) of 3.31% on the AA-WER benchmark, improving by 0.7 percentage points over its predecessor, with a 25% price reduction to $0.0045 per minute.
  • Despite improvements, OpenAI lags behind competitors like ElevenLabs Scribe v2 (2.3% WER), Google Gemini 3 Pro (2.9% WER), and Mistral Voxtral Small (3% WER).
  • The models support text context, keywords, and multiple languages, aligning with OpenAI’s broader Realtime model generation initiative.

Why It Matters

This update reflects the rapid evolution in speech recognition technology, where accuracy, speed, and cost efficiency are critical for enterprise and developer adoption. For AI practitioners, understanding these benchmarks helps inform model selection based on use case—whether prioritizing low latency, high throughput, or competitive pricing. The continued pressure from rivals underscores the need for continuous innovation in multimodal AI systems.

Technical Details

  • GPT Transcribe: Processes pre-recorded audio at ~34x real-time speed; supports text context, keyword injection, and multilingual input; achieves 3.31% WER on AA-WER benchmark.
  • GPT Live Transcribe: Designed for real-time streaming with minimal latency, suitable for live applications such as meetings or customer service interactions.
  • Pricing Model: Reduced to $0.0045 per minute of audio processed—a 25% drop from prior pricing—making it more accessible for large-scale deployments.
  • Benchmark Performance: Evaluated via Artificial Analysis’ AA-WER metric, which measures transcription accuracy across diverse audio conditions and speaker demographics.
  • Integration: Part of OpenAI’s Realtime ecosystem, complementing GPT-Realtime-Whisper and enabling seamless API-based deployment within existing workflows.

Industry Insight

The competitive landscape in speech-to-text is intensifying, with Mistral undercutting pricing at $0.003/minute while maintaining strong performance—suggesting that cost leadership may become a key differentiator alongside accuracy. Developers should evaluate trade-offs between model fidelity, latency, and budget when selecting transcription services. As real-time capabilities grow in demand for interactive AI applications, providers must balance innovation with scalability and affordability to maintain market relevance.

TL;DR

  • OpenAI发布GPT Transcribe和GPT Live Transcribe两款语音识别模型,分别处理预录音频与实时流媒体。
  • GPT Transcribe错误率3.31%,较前代提升0.7个百分点,价格下调25%至每分钟0.0045美元。
  • 在AA-WER基准测试中,OpenAI落后于ElevenLabs(2.3%)、Google(2.9%)和Mistral(3%)。
  • Mistral以更低定价($0.003/分钟)推出Voxtral Transcribe V2,加剧市场竞争。
  • 新模型支持多语言输入、文本上下文及关键词辅助转录,并集成至OpenAI实时生成生态。

为什么值得看

该资讯揭示了当前语音识别领域的竞争格局与技术演进趋势,对关注API服务定价策略、模型性能对比及实时转录应用场景的从业者具有参考价值。同时反映了大模型厂商在语音任务上的差异化布局与市场压力。

技术解析

  • GPT Transcribe专为离线音频文件设计,处理速度达实时速度的34倍;GPT Live Transcribe面向低延迟实时流式转录,适配动态交互场景。
  • 基于AA-WER benchmark评估,GPT Transcribe词错误率(WER)为3.31%,优于其前身GPT-4o Transcribe,但仍逊于头部竞品。
  • 支持多语言输入、外部文本作为上下文提示、关键词高亮等增强功能,提升特定领域转录准确性。
  • 定价体系优化:从旧版降至$0.0045/min,响应市场对性价比需求,但未达到Mistral的$0.003/min水平。
  • 与OpenAI Realtime生态协同,包括GPT-Realtime-Whisper模型,形成“批处理+实时流”双轨语音解决方案。

行业启示

  • 语音识别赛道进入精细化竞争阶段,错误率与成本成为核心指标,厂商需在精度、速度与价格间寻找平衡点。
  • 新兴玩家如Mistral通过激进定价切入市场,迫使传统巨头加速迭代或调整商业模式,推动整体服务下沉。
  • 企业用户在选择语音API时,应结合具体场景(如客服录音 vs 会议直播)权衡不同模型的延迟、精度与费用特性。

Disclaimer: The above content is generated by AI and is for reference only. 免责声明:以上内容由 AI 生成,仅供参考。

Speech 语音 Product Launch 产品发布 Evaluation 评测