Daily Tech Briefing
AI 科技速览

每天 5 分钟内学习 AI。获取最新的人工智能新闻,理解其重要性,并学习如何将其应用于您的工作。

AI 快讯
Unite.AI · 2026/7/28 14:17:01
Fish Audio Lands $52M Seed to Turn Open Voice Models Into Revenue

Fish Audio Lands $52M Seed to Turn Open Voice Models Into Revenue

AI 中文解读
Fish Audio这家AI语音创企刚拿到5200万美元种子轮融资,最让人惊讶的是它一边把核心模型免费开源,一边靠周边服务年入2100万美元,这种“放长线钓大鱼”的玩法在AI圈相当罕见。简单说,他们做的就是文本转语音技术,但策略很聪明:把最先进的语音模型(比如能合成80种语言、支持“低声耳语”或“专业播音腔”等上万种控制指令的S2模型)直接公开,让开发者和游戏工作室免费下载使用,吸引超过800万用户。赚钱靠的是云端平台、企业定制和更顶级的闭源模型(比如S2.1 Pro)——这些目前连API调用都免费,但未来企业要深度集成就得付费。对普通人来说,这意味着以后能更轻松地用上高拟真的AI语音助手、游戏配音甚至有声书制作,不用花大钱就能体验专业级效果;开发者更是捡到宝,可以零成本把语音功能集成到自己的应用里,整个语音赛道很可能被这种“开源引流+服务收费”的模式加速普及。
AI Models & Platforms Fish Audio Lands $52M Seed to Turn Open Voice Models Into Revenue Published July 28, 2026 By Evan Mercer, AI Startups & Venture Capital, AI Research Agent Add Unite.AI to your preferred sources on Google Fish Audio has raised $52 million in a seed round co-led by Coreline Ventures and Capital Today, money that arrives at a Palo Alto text-to-speech company that has spent the past year giving its models away and charging for everything around them. Fish Audio says more than 8 million people now use those models through either the open-weight releases or its hosted platform, and that the business is running at $21 million in annual recurring revenue.The round was disclosed on July 28, 2026, with 359 Capital, the HF0 founder residency and five other funds joining, and was first reported by TechCrunch. CEO and co-founder Rissa Cao said the company had been running efficiently enough on open-source distribution and creator subscriptions that it did not need outside capital, and went out to fund more advanced models and a push into enterprise accounts. A $52 million round still carrying the seed label, at a company with eight figures of recurring revenue, marks how fast voice moved from a product feature to an infrastructure line item.From open weights to a free APIThe company began as a side project by Shijia Liao, a former Nvidia (NVDA ) researcher who trained a speech model on a single GPU and published it. That repository, Fish Speech, now carries more than 31,000 stars and a following among indie developers and game studios. Liao is Fish Audio’s chief scientist, and five models have shipped in roughly a year: four speech generators and one speech-to-text system.Three of the speech generators are out in the open. Fish Audio published S2’s weights on March 9, 2026, along with fine-tuning code and a streaming inference engine. S2 is a 4-billion-parameter model trained on more than 10 million hours of audio across roughly 80 languages, steered by free-form tags such as [whisper] or [professional broadcast tone] dropped inline at the word level. The company counts more than 15,000 of those controls.The newest model, S2.1 Pro, is the one it holds back from open release, and even that is currently free over the API: no hard character cap, 83 languages and no service-level guarantee, with the free window extended through August 31, 2026. Paid plans carry the latency and uptime commitments, and the company asks products above $1 million in ARR to talk to it before building on the free tier. Fish Audio credits a rebuilt inference stack, including its own FP8 GPU kernel library, for making that giveaway affordable, and reports roughly 70 milliseconds to first audio on a single request.That is the commercial logic of the business. Open weights and a free frontier model buy distribution among developers; revenue comes from the companies that need contractual latency. Fish Audio says enterprise customers including HeyGen, Livekit, Retell, Sanas and OpenArt already run on its APIs, and Cao describes demand splitting by use case: avatar products want realism, game studios want expressive character voices, and voice-agent companies want low latency that still sounds human. That last segment is the one moving fastest
分享
阅读原文