Daily Tech Briefing
AI 科技速览

每天 5 分钟内学习 AI。获取最新的人工智能新闻,理解其重要性,并学习如何将其应用于您的工作。

AI 快讯
Apple ML Research · 2026/7/28 00:00:00

Memory Efficient Audio Synthesis with Decoupled Temporal Depth Diffusion Transformers

AI 中文解读
苹果最新发布了一项语音合成技术,让Siri的声音终于“活”了起来。核心亮点是:在不联网、不上传数据的情况下,手机就能实时生成自然流畅、富有情感的语音,而且不占太多内存和算力。以前AI语音总带着机械感,现在苹果用一个名为“AFM 3 Core Advanced”的端侧模型,把语义信息转换成高质量音频。简单说,就像把一句话浓缩成“语义密码”,再用一个轻巧的解码器,在手机芯片上瞬间还原成声情并茂的声音,整个过程都在设备内部完成,既保护隐私又节省资源。这项技术最直接的影响是:以后你用Siri导航、听新闻、发语音消息,听到的不再是冷冰冰的电子音,而是更像真人在跟你说话。同时,因为是用AI直接合成,以后甚至可以让Siri用你喜欢的声音读文章、讲故事,甚至根据上下文调整语气。对于普通用户来说,这意味着更自然的人机对话体验;对开发者而言,则打开了低功耗、高保真语音应用的大门,比如离线语音助手、有声书生成等,都能在手机上流畅跑起来了。
Siri Expressive Voices synthesize rich, configurable speech in real time and entirely on device, powered by AFM 3 Core Advanced, Apple’s most powerful on-device foundation model. This work presents the memory-efficient audio synthesis architecture behind that capability: a detokenizer that converts the semantic audio tokens emitted by the foundation model into high-fidelity audio within the tight compute and memory budget of the Apple Matrix Coprocessor (AMX). We convert semantic audio tokens to a residual vector quantization (RVQ) representation with a three-component design—a streaming…
分享
阅读原文
Memory Efficient Audio Synthesis with Decoupled Temporal Depth Diffusion Transformers | BriefSum AI 情报