Cartesia Ships Sonic-3.6: A Streaming TTS Model That Now Leads Both Artificial Analysis Speech Arenas
Published · Aug 18 · Tue Source · MarkTechPost

Cartesia Ships Sonic-3.6: A Streaming TTS Model That Now Leads Both Artificial Analysis Speech Arenas

Cartesia launched Sonic-3.6, a streaming audio generation model utilizing state space architectures. The system achieved top rankings on Artificial Analysis leaderboards, securing 1,283 and 1,123 Elo scores respectively.

KeywordsCartesiaShipsSonic-3.6StreamingTTSModelThatNow

Cartesia has unveiled Sonic-3.6, a new streaming text-to-speech model designed for efficient audio generation. The release marks a significant update in the company's lineup of generative voice technologies.

Distinct from prevalent transformer-based designs, this iteration employs state space models. This architectural choice suggests a focus on different computational efficiencies or latency characteristics inherent to streaming applications.

Benchmarks show the model leading the Provider Voice category with 1,283 Elo and the Controlled Voice category with 1,123 Elo. These metrics indicate the system currently holds the top position on both Artificial Analysis speech leaderboards.

Such performance highlights the evolving landscape of synthetic voice technology. As developers seek alternatives to standard transformer-based systems, models like Sonic-3.6 demonstrate viable pathways for high-quality synthetic speech.

This page provides an editorial summary based on publicly available information. It is not a republished article. Use the source link below for the original report.