Gradium AI Releases New Default TTS Model: 81.0% Hard-Case Pass Rate at 216 ms Time-to-First-Audio
Gradium AI released a new default TTS model (live as of August 31, 2026) achieving an 81.0% human-rated pass rate on a 500-sentence hard-case benchmark across five languages, outperforming Cartesia Sonic 3.6 (75.1%), ElevenLabs v3 Conversational (65.4%), Fish Audio S2.1 Pro (49.5%), and Inworld TTS 1.5 Max (46.5%) The model delivers a 216 ms P50 time to first audio on Coval's TTS benchmark with an exceptionally tight 30 ms interquartile range (p75-p25) across 480 runs, representing the lowest va
Analysis
TL;DR
- Gradium AI released a new default TTS model (live as of August 31, 2026) achieving an 81.0% human-rated pass rate on a 500-sentence hard-case benchmark across five languages, outperforming Cartesia Sonic 3.6 (75.1%), ElevenLabs v3 Conversational (65.4%), Fish Audio S2.1 Pro (49.5%), and Inworld TTS 1.5 Max (46.5%)
- The model delivers a 216 ms P50 time to first audio on Coval's TTS benchmark with an exceptionally tight 30 ms interquartile range (p75-p25) across 480 runs, representing the lowest variance among tested models
- The 500-sentence evaluation set is open-sourced on Hugging Face under CC BY 4.0, covering 10 criteria (7 atomic: spelling, acronyms, alphanumeric tokens, dates, regular numbers, large/floating numbers, emails; 3 composite: Orders, IT Ticket, Claims) across English, German, French, Spanish, and Portuguese
- No migration is required for existing users; the model was switched on as default across Gradium's API and Studio, with existing voices and custom clones continuing to work unchanged
- The model reads phone numbers, emails, IBANs, and reference codes with no text normalization required, directly addressing the failure points that matter most in voice agent calls
Why It Matters
Voice agents consistently fail on the most critical parts of a call—order numbers, callback digits, email addresses—making this benchmark directly relevant to anyone building production voice AI systems. The combination of high accuracy (81.0% pass rate) and low latency (216 ms P50) with minimal variance positions Gradium's model as a strong candidate for real-world deployment where both correctness and responsiveness are non-negotiable. The open-sourcing of the evaluation dataset also sets a potential new standard for transparent, human-rated TTS benchmarking in the industry.
Technical Details
- Benchmark methodology: 500-sentence hard-case set with strict human rating—each sentence must pass all 10 criteria (7 atomic + 3 composite) to be counted as a pass; a single dropped digit fails the entire sentence. Audio was loudness-normalized, order randomized, and raters capped at 40 comparisons with enforced breaks to prevent fatigue.
- Latency performance: 216 ms P50 time to first audio on Coval's benchmark, 170 ms faster than the previous default model. The interquartile range of 30 ms (p75-p25 across 480 runs) is the tightest among the five compared models, indicating highly consistent inference behavior.
- Multilingual coverage: Evaluated across five languages (EN, DE, FR, ES, PT) with equal weighting in the pooled pass rate. The model handles alphanumeric tokens, IBANs, emails, and large/floating-point numbers without requiring text normalization preprocessing.
- Deployment: Switched on as the default model across Gradium's API and Studio on August 31, 2026. Existing voice IDs and custom clones remain compatible with zero migration. New users can integrate via Python SDK and WebSocket TTS endpoint.
- Evaluation transparency: The full 500-sentence dataset is open-sourced on Hugging Face under CC BY 4.0, and Gradium is offering 1M credits for complete hard-case failure reports submitted via Discord, encouraging community-driven validation.
Industry Insight
- The emphasis on "hard-case" accuracy (numbers, emails, alphanumeric tokens) signals a shift in the voice AI market from naturalness-focused benchmarks to utility-focused ones—practitioners should prioritize models that handle structured data correctly over those that merely sound pleasant.
- The tight latency variance (30 ms IQR) is as important as the median latency itself; callers experience tail turns, not medians, so models with low spread provide a more predictable and professional user experience in production voice agents.
- Open-sourcing evaluation datasets (as Gradium did) may become a competitive differentiator, pushing the industry toward more transparent and comparable TTS benchmarks rather than proprietary, opaque scoring systems.
Disclaimer: The above content is generated by AI and is for reference only.