AIHOT 于 2026-08-18 收录了“Cartesia 发布 Sonic-3.6 实时文本转语音模型,登顶两大语音榜单”这一公开动态。以下先呈现从来源页面抓取的正文,再给出 AIHOT 摘要与 TopoReduce 编辑解读。
PUBLIC SOURCE CONTENT
已抓取公开正文公开原文内容
Cartesia Ships Sonic-3.6: A Streaming TTS Model That Now Leads Both Artificial Analysis Speech Arenas - MarkTechPost
- Editors Pick
- Agentic AI
- Technology
- AI Shorts
- Artificial Intelligence
- Applications
- Language Model
- Audio Language Model
- Large Language Model
- New Releases
- Staff
- Tech News
- TTS
- Uncategorized
- Voice AI
Cartesia has released Sonic-3.6, the newest version of its real-time text-to-speech model. It arrives roughly three months after Sonic-3.5. The new change is naturalness, and this one is independently checkable. Sonic 3.6 now holds #1 on both Artificial Analysis speech leaderboards — 1,283 Elo on the Provider Voice board and 1,123 on the Controlled Voice board. The second result matters more. That board clones every model onto the same eight reference voices, which isolates the synthesis engine from the voice catalog. Sonic-3.6 leads it, with Sonic-3.5 second and ElevenLabs Eleven v3 third. The model runs on state space models rather than transformers, and Cartesia states sub-90ms time-to-first-audio. It is available in beta.
Is it deployable?
YES, it is available in beta and as a hosted API. Not as self-hosted weights.
Sonic is a closed, commercial model. There are no open weights and no Hugging Face repo. You rent it.
- Company level: Solo developers and startups (Free/Pro $5 tiers), scaleups running contact centers (Startup $49 / Scale $299), and regulated enterprises needing DPAs, BAAs, and SSO.
- Industries: Financial services, healthcare, retail and e-commerce, logistics, recruiting, SaaS support, consumer companion apps, media localization
- Applications: Inbound support agents, outbound qualification calls, IVR replacement, appointment reminders, sales-training simulators, audio localization, in-product voice UI
The Architecture
Sonic runs on state space models rather than transformers. Cartesia’s launch page frames the usual tradeoffs — speed versus naturalness, accuracy versus cost — as architectural, not inevitable.
The practical output is time-to-first-audio. Cartesia states sub-90ms TTS latency, and 100ms transcript latency for its Ink-2 speech-to-text model. Both are vendor-stated model latency, not measured end-to-end round trips.
Interactive explainer
Features that matter in production
Sonic exposes controls built for agent transcripts rather than narration:
- Inline expression tags. Non-verbal expressions like [laughter] go directly in the transcript.
- Instant voice cloning from about 10 seconds of audio.
- Custom pronunciation dictionaries, including IPA overrides such as <<s|ə|ˈ|p|i|n|ə>> for subpoena.
- Speed, volume, and emotion parameters exposed through the API and integrations like the LiveKit Agents plugin.
- Native alphanumerics. Order numbers, phone numbers, and confirmation codes read correctly without preprocessing.
Cartesia’s launch demos show English with natural pauses and filler words, plus Hinglish code-switching between Hindi and English.
Pricing reality
Artificial Analysis normalizes Sonic 3.6 at $49.00 per 1M characters. That is half of ElevenLabs Eleven v3 at $100.00, and well above Speechify Simba 3.2 at $10.00 for a 1,240 Elo.
Cartesia sells credits, not characters. Scale at $299 per month includes roughly 10,667 TTS minutes and 15 concurrent requests. Line voice agents bill separately at $0.06 per minute.
Key Takeaways
- Sonic-3.6 is #1 on both Artificial Analysis speech arenas — 1,283 Elo Provider Voice, 1,123 Controlled Voice.
- Winning the Controlled board means the engine improved, not just the voice catalog.
- It is beta on Cartesia’s API only; docs still list Sonic 3.5 as stable, and partners carry 3.5.
- Deployable as a hosted API, not self-hosted weights; commercial use starts at the $5 Pro tier.
- Latency claims (sub-90ms TTFA) are vendor-stated model latency, so benchmark your own round trip.
Check out the Project Page-Cartesia Sonic, Cartesia launch page, Cartesia pricing, Cartesia docs, Artificial Analysis Speech Arena and @cartesia on X. All figures verified August 18, 2026.. Also, feel free to follow us on Twitter and don’t forget to join our 150k+ML SubReddit and Subscribe to our Newsletter. Wait! are you on telegram? now you can join us on telegram as well.
Need to partner with us for promoting your GitHub Repo OR Hugging Face Page OR Product Release OR Webinar etc.? Connect with us
Asif Razzaq
Website | + postsBio
Asif Razzaq is the CEO of Marktechpost Media Inc.. As a visionary entrepreneur and engineer, Asif is committed to harnessing the potential of Artificial Intelligence for social good. His most recent endeavor is the launch of an Artificial Intelligence Media Platform, Marktechpost, which stands out for its in-depth coverage of machine learning and deep learning news that is both technically sound and easily understandable by a wide audience. The platform boasts of over 2 million monthly views, illustrating its popularity among audiences.
- Asif Razzaq
NVIDIA Releases TensorRT Model Connect in Public Preview: Hugging Face Checkpoint to Native C++ Inference in Two Commands
- Asif Razzaq
ByteDance Seed and Tsinghua AIR Introduces CUDA Agent: A Large-Scale Agentic RL System for CUDA Kernel Generation
- Asif Razzaq
MiniMax Releases MiniMax-Music3: An Open-Weights Music Model Generating Complete Five-Minute Songs From Lyrics and a Structured Caption
- Asif Razzaq
DeepSeek AI Releases DeepSeek Harness in Developer Preview: An MIT-Licensed Agent Harness Where Everything is a Plugin
- Agentic AI
- Technology
- AI Shorts
- Artificial Intelligence
- Applications
- Language Model
- Audio Language Model
- Large Language Model
- New Releases
- Staff
- Tech News
- TTS
- Uncategorized
- Voice AI
Cartesia has released Sonic-3.6, the newest version of its real-time text-to-speech model. It arrives roughly three months after Sonic-3.5. The new change is naturalness, and this one is independently checkable. Sonic 3.6 now holds #1 on both Artificial Analysis speech leaderboards — 1,283 Elo on the Provider Voice board and 1,123 on the Controlled Voice board. The second result matters more. That board clones every model onto the same eight reference voices, which isolates the synthesis engine from the voice catalog. Sonic-3.6 leads it, with Sonic-3.5 second and ElevenLabs Eleven v3 third. The model runs on state space models rather than transformers, and Cartesia states sub-90ms time-to-first-audio. It is available in beta.
Is it deployable?
YES, it is available in beta and as a hosted API. Not as self-hosted weights.
Sonic is a closed, commercial model. There are no open weights and no Hugging Face repo. You rent it.
- Company level: Solo developers and startups (Free/Pro $5 tiers), scaleups running contact centers (Startup $49 / Scale $299), and regulated enterprises needing DPAs, BAAs, and SSO.
- Industries: Financial services, healthcare, retail and e-commerce, logistics, recruiting, SaaS support, consumer companion apps, media localization
- Applications: Inbound support agents, outbound qualification calls, IVR replacement, appointment reminders, sales-training simulators, audio localization, in-product voice UI
The Architecture
Sonic runs on state space models rather than transformers. Cartesia’s launch page frames the usual tradeoffs — speed versus naturalness, accuracy versus cost — as architectural, not inevitable.
The practical output is time-to-first-audio. Cartesia states sub-90ms TTS latency, and 100ms transcript latency for its Ink-2 speech-to-text model. Both are vendor-stated model latency, not measured end-to-end round trips.
Interactive explainer
Features that matter in production
Sonic exposes controls built for agent transcripts rather than narration:
- Inline expression tags. Non-verbal expressions like [laughter] go directly in the transcript.
- Instant voice cloning from about 10 seconds of audio.
- Custom pronunciation dictionaries, including IPA overrides such as <<s|ə|ˈ|p|i|n|ə>> for subpoena.
- Speed, volume, and emotion parameters exposed through the API and integrations like the LiveKit Agents plugin.
- Native alphanumerics. Order numbers, phone numbers, and confirmation codes read correctly without preprocessing.
Cartesia’s launch demos show English with natural pauses and filler words, plus Hinglish code-switching between Hindi and English.
Pricing reality
Artificial Analysis normalizes Sonic 3.6 at $49.00 per 1M characters. That is half of ElevenLabs Eleven v3 at $100.00, and well above Speechify Simba 3.2 at $10.00 for a 1,240 Elo.
Cartesia sells credits, not characters. Scale at $299 per month includes roughly 10,667 TTS minutes and 15 concurrent requests. Line voice agents bill separately at $0.06 per minute.
Key Takeaways
- Sonic-3.6 is #1 on both Artificial Analysis speech arenas — 1,283 Elo Provider Voice, 1,123 Controlled Voice.
- Winning the Controlled board means the engine improved, not just the voice catalog.
- It is beta on Cartesia’s API only; docs still list Sonic 3.5 as stable, and partners carry 3.5.
- Deployable as a hosted API, not self-hosted weights; commercial use starts at the $5 Pro tier.
- Latency claims (sub-90ms TTFA) are vendor-stated model latency, so benchmark your own round trip.
Check out the Project Page-Cartesia Sonic, Cartesia launch page, Cartesia pricing, Cartesia docs, Artificial Analysis Speech Arena and @cartesia on X. All figures verified August 18, 2026.. Also, feel free to follow us on Twitter and don’t forget to join our 150k+ML SubReddit and Subscribe to our Newsletter. Wait! are you on telegram? now you can join us on telegram as well.
Need to partner with us for promoting your GitHub Repo OR Hugging Face Page OR Product Release OR Webinar etc.? Connect with us
Asif Razzaq
Website | + postsBio
Asif Razzaq is the CEO of Marktechpost Media Inc.. As a visionary entrepreneur and engineer, Asif is committed to harnessing the potential of Artificial Intelligence for social good. His most recent endeavor is the launch of an Artificial Intelligence Media Platform, Marktechpost, which stands out for its in-depth coverage of machine learning and deep learning news that is both technically sound and easily understandable by a wide audience. The platform boasts of over 2 million monthly views, illustrating its popularity among audiences.
- Asif Razzaq
NVIDIA Releases TensorRT Model Connect in Public Preview: Hugging Face Checkpoint to Native C++ Inference in Two Commands
- Asif Razzaq
ByteDance Seed and Tsinghua AIR Introduces CUDA Agent: A Large-Scale Agentic RL System for CUDA Kernel Generation
- Asif Razzaq
MiniMax Releases MiniMax-Music3: An Open-Weights Music Model Generating Complete Five-Minute Songs From Lyrics and a Structured Caption
- Asif Razzaq
DeepSeek AI Releases DeepSeek Harness in Developer Preview: An MIT-Licensed Agent Harness Where Everything is a Plugin
AIHOT 摘要
Cartesia 发布实时文本转语音模型 Sonic-3.6,在 Artificial Analysis 两大语音榜单均位列第一,其中 Controlled Voice 榜 1,123 Elo,超越 ElevenLabs Eleven v3。该模型基于状态空间模型,首音频延迟低于 90ms,现以 beta 版和托管 API 形式提供,定价为每 100 万字符 49 美元。
为什么值得关注
在受控语音榜单上夺冠比供应商榜单更有参考价值,它隔离了音色库差异,提示合成引擎本身自然度提升,而半价于 ElevenLabs 的标准化定价会影响实时语音服务的选型。
工程化解读
从 TopoReduce 的工程视角看,这条信息属于“多模态与端侧”主题。它的价值不只在于一个新产品或新观点本身,还在于说明 AI 系统正在如何影响模型接入、智能体协作、研发流程、基础设施和团队决策。实际采用前,应结合原文确认版本、适用范围、价格和运行条件。
- 发布时间:2026-08-18;AIHOT 分类:多模态与端侧。
- AIHOT 标签:
- AIHOT 判断:在受控语音榜单上夺冠比供应商榜单更有参考价值,它隔离了音色库差异,提示合成引擎本身自然度提升,而半价于 ElevenLabs 的标准化定价会影响实时语音服务的选型。
- AIHOT 评分:62;评分用于站内排序,不等同于独立评测结论。
TopoReduce 编辑观察
当 AI 动态进入真实生产环境,团队需要同时关注能力边界、数据来源、调用成本、权限控制和可回滚性。把单条新闻放回完整工程链路中阅读,比只看标题更有助于判断它是否适合自己的产品和工作流。