Cartesia Ships Sonic-3.6: A Streaming TTS Model That Now Leads Both Artificial Analysis Speech Arenas
Cartesia released Sonic-3.6, a real-time streaming text-to-speech model that now ranks #1 on both Artificial Analysis speech leaderboards (1,283 Elo Provider Voice, 1,123 Elo Controlled Voice) The model uses state space models instead of transformers, achieving sub-90ms time-to-first-audio latency Sonic-3.6's top ranking on the Controlled Voice board indicates genuine synthesis engine improvement rather than just a larger voice catalog Available in beta as a hosted API only; no self-hosted weigh
Analysis
TL;DR
- Cartesia released Sonic-3.6, a real-time streaming text-to-speech model that now ranks #1 on both Artificial Analysis speech leaderboards (1,283 Elo Provider Voice, 1,123 Elo Controlled Voice)
- The model uses state space models instead of transformers, achieving sub-90ms time-to-first-audio latency
- Sonic-3.6's top ranking on the Controlled Voice board indicates genuine synthesis engine improvement rather than just a larger voice catalog
- Available in beta as a hosted API only; no self-hosted weights or open-source release
- Pricing at $49 per 1M characters (per Artificial Analysis), half the cost of ElevenLabs Eleven v3, with production-focused features like inline expression tags, instant voice cloning from ~10 seconds, and native alphanumerics handling
Why It Matters
Sonic-3.6 represents a significant milestone in real-time TTS by demonstrating that state space models can outperform transformer-based architectures on both naturalness and latency benchmarks, challenging the industry assumption that these tradeoffs are inevitable. For AI practitioners building voice agents, the model's production-ready features—inline expression tags, instant voice cloning, and sub-90ms TTFA—make it immediately deployable for conversational AI applications without extensive preprocessing. The Controlled Voice leaderboard victory is particularly noteworthy because it isolates engine quality from voice catalog breadth, providing a more rigorous measure of genuine model improvement.
Technical Details
- Architecture: Sonic-3.6 runs on state space models (SSMs) rather than transformers, which Cartesia frames as resolving the traditional speed-versus-naturalness tradeoff through architectural innovation rather than incremental scaling
- Latency: Vendor-stated sub-90ms time-to-first-audio (TTFA) for TTS and 100ms transcript latency for the companion Ink-2 speech-to-text model; these are model-level latencies, not end-to-end round-trip measurements
- Benchmarks: #1 on both Artificial Analysis speech leaderboards—1,283 Elo on Provider Voice (where each model uses its own voices) and 1,123 Elo on Controlled Voice (where all models are cloned onto the same eight reference voices, isolating synthesis engine quality)
- Production Features: Inline expression tags (e.g., [laughter]), instant voice cloning from ~10 seconds of audio, custom pronunciation dictionaries with IPA overrides, exposed speed/volume/emotion parameters via API, LiveKit Agents plugin integration, and native handling of alphanumerics without preprocessing
- Deployment: Hosted API only in beta; no self-hosted weights or Hugging Face repository; Sonic 3.5 remains the stable version in partner ecosystems
Industry Insight
- The rise of state space models in production TTS signals a potential architectural shift away from transformers for latency-sensitive audio generation tasks; practitioners should evaluate SSM-based models for real-time voice agent deployments where sub-100ms TTFA is critical
- Sonic-3.6's Controlled Voice leaderboard dominance suggests the next competitive frontier in TTS is synthesis engine quality rather than voice catalog size, pushing vendors to invest in fundamental model architecture improvements over data collection
- The pricing landscape ($49/M characters for Sonic-3.6 vs. $100/M for ElevenLabs vs. $10/M for Speechify Simba 3.2 at lower Elo) indicates a mid-tier value opportunity; teams should benchmark actual round-trip latency and quality in their specific deployment context rather than relying on vendor-stated model latency figures
Disclaimer: The above content is generated by AI and is for reference only.