Daily Tech Briefing
AI 科技速览

每天 5 分钟内学习 AI。获取最新的人工智能新闻,理解其重要性,并学习如何将其应用于您的工作。

AI 快讯
The Decoder · 2026/7/21 11:31:15
Alibaba's Qwen Audio 3.0 TTS Plus tops the competition in the text-to-speech rankings

Alibaba's Qwen Audio 3.0 TTS Plus tops the competition in the text-to-speech rankings

AI 中文解读
阿里云最新发布的Qwen Audio 3.0 TTS Plus文本转语音模型,在权威评测排行榜上拿下第一,超越了SpeechifyAI的Simba 3.2等对手。它有两个版本:Flash版延迟仅300毫秒,适合实时对话;Plus版则追求更高质量的语音输出。这个模型支持16种语言,包括他加禄语、马来语、泰语和越南语等较少覆盖的语言,还能用自然语言调整说话风格,或者加入“[生气]”“[笑声]”这类非语言提示。在声音克隆方面,它处理嘈杂或带回声的录音也比之前版本更稳。不过它的速度是短板,每秒只能生成16个字符,远低于对手的120或30个字符,价格是每百万字符27.6美元。对普通用户来说,这意味着未来用AI朗读文章、制作有声内容或语音助手时,会听到更自然、更懂情绪的声音,甚至能模仿特定人的语气,只是等待时间可能稍长一点。
Alibaba's Qwen Audio 3.0 TTS Plus tops the competition in the text-to-speech rankings Matthias Bastian View the LinkedIn Profile of Matthias Bastian Jul 21, 2026 Alibaba's new text-to-speech model, Qwen-Audio-3.0-TTS-Plus, leads Artificial Analysis' Speech Arena leaderboard for provider voices. With an Elo score of 1,236, it sits just ahead of Simba 3.2 (1,234). Gemini 3.1 Flash TTS (1,214) and Sonic 3.5 (1,207) follow behind. Alibaba's Qwen-Audio-3.0-TTS-Plus takes the top spot on Artificial Analysis' Text to Speech Leaderboard for provider voices, edging out SpeechifyAI's Simba 3.2 by just two Elo points. | Image: Artificial Analysis The model comes in two versions. Flash is built for real-time interaction with about 300 milliseconds of latency, while Plus targets high-quality speech output. It supports 16 languages, including less commonly covered ones like Tagalog, Malay, Thai, and Vietnamese, along with several Chinese dialects. Users can steer the speaking style with natural language or add nonverbal cues using tags like "[angry]" or "[giggles]." Alibaba also says the model handles noisy or echo-heavy reference recordings better than previous versions when cloning voices. Speed is a weak spot: At 16 characters per second, it trails Sonic 3.5 (120) and Simba 3.2 (30.2) by a wide margin. Pricing lands at $27.60 per million characters through Alibaba Cloud Model Studio. A collection of audio samples is available here.AdDEC_D_Incontent-1Ad AI News Without the Hype – Curated by Humans Subscribe to THE DECODER for ad-free reading, a weekly AI newsletter, our exclusive "AI Radar" frontier report six times a year, full archive access, and access to our comment section. Subscribe now Source: Artificial Analysis Ask about this article… Search
分享
阅读原文