Daily Tech Briefing
AI 科技速览

每天 5 分钟内学习 AI。获取最新的人工智能新闻,理解其重要性,并学习如何将其应用于您的工作。

AI 快讯
TechCrunch AI · 2026/7/28 14:00:00
Fish Audio raises $52M seed to build AI voice models for creators and enterprises

Fish Audio raises $52M seed to build AI voice models for creators and enterprises

AI 中文解读
Fish Audio获得5200万美元种子轮融资,打造更逼真、可定制的AI语音模型,让机器说话也能带感情。这家由前英伟达研究员创立的公司,开源了多个语音模型,吸引了800万用户,年收入已达2100万美元。他们拥有超过1.5万种自然语言控制参数,能灵活调整语调、情感和语速——游戏公司可以用来给角色配出喜怒哀乐,企业客服可以生成标准又自然的应答,视频创作者则能一秒克隆自己的声音。不过,此前因用户声音被未经同意上传引发争议,公司已自动完善了删除流程。对普通人来说,这意味着以后听到的AI语音不再生硬单调:虚拟主播会更亲切,语音助手会更像真人,游戏里的NPC也会更有灵魂,甚至你用少量音频就能生成专属声音模型,让数字分身替你说话。
The market for AI-generated voice models is massive. Creative use cases require AI voice models to be more expressive, while enterprises looking to automate customer support and sales ops need them to be more steerable. Palo Alto-based Fish Audio wants to cater to all of those use cases with its library of more than 15,000 natural language controls. Since launching last year, the startup now has more than 8 million people using the open source or hosted versions of its models, and generates annual recurring revenue of $21 million. To continue building on that traction, the startup on Tuesday said it has raised $52 million in a seed round that was led by Coreline Ventures and Capital Today. The funding also saw participation from 359 Capital, Parable, Play Time, Alphalist Partners, Bayhouse Ventures, Carya Venture Partners, and HF0. Fish Audio started as a small project by former Nvidia researcher Shijia Liao, who, frustrated by non-expressive synthetic voices available on the market, trained a voice-generation model on a single GPU, which he then open-sourced. The Fish Speech repository on GitHub now has more than 31,000 stars and is used by indie developers, video game designers, and creators. The company has launched five models in the last year: four speech-generation models and one speech-to-text model. It has open-sourced three of its speech-generation models, but its latest S2.1 Pro model is available only through its paid API. Fish Audio offers paid monthly plans suited for creators and teams that unlock a set number of minutes of generation plus voice-cloning features. The company also offers an enterprise version of its APIs and platform, and says organizations like HeyGen and Sanas are already using it. “Every enterprise has different use cases and different preferences. For example, companies like HeyGen, which use our voices to power AI avatars, want realism in voices; a gaming studio would want expressive voices for their characters; and voice agent companies like LiveKit want more natural-sounding and low-latency voices that are expressive enough for calls,” Fish Audio’s CEO and co-founder Rissa Cao said. var playerInstance_jwplayer_6a695b3903eb9 = jwplayer( "jwplayer_6a695b3903eb9" ); playerInstance_jwplayer_6a695b3903eb9.setup({ playlist: "https://cdn.jwplayer.com/v2/media/nsQAyeWN", }); One way the startup has built its library of voices is by asking users to submit their own voices for training its models, and compensating them if their voices are used. That resulted in some trouble a few months ago, however, as some creators alleged that their voices were uploaded to Fish Audio without their consent. The startup had a DMCA takedown process in place to address such concerns, but the takedowns themselves took a long time. Cao told TechCrunch that the company has now automated the takedown process. Creators can submit a short voice sample or a contract to prove that an uploaded voice belongs to them, and their voice will be taken off the startup’s platform in less than three minutes, she said. Still, that doesn’t prevent anyone from uploading an artist’s voice without their knowledge. And until the artist finds out, their voice will continue to be used on the platform unless they file for its removal. Osuke Honda, a partner at Coreline Ventures, said a community-driven model only works when creators trust the platform. “A community-centric approach can only become a durable advantage if creators trust the platform. That means consent, transparency, and attribution must be built into the product rather than treated as afterthoughts. I believe the industry needs to move toward verified voice ownership, clear licensing terms, easy reporting and takedown processes, and eventually revenue-sharing models where creators benefit financially when their voices are licensed or used commercially,” he said. Cao said when the startup was only offering its p
分享
阅读原文