Daily Tech Briefing
AI 科技速览
每天 5 分钟内学习 AI。获取最新的人工智能新闻,理解其重要性,并学习如何将其应用于您的工作。
SiliconANGLE AI · 2026/7/28 13:30:55

Fish Audio makes a splash after raising $52M seed funding for AI voices
AI 中文解读
Fish Audio这家AI语音创业公司最近拿到5200万美元的种子轮融资,引来不少关注。它的诞生故事挺有意思:创始人曾是英伟达的研究员,因为受不了早期AI那种单调生硬的机器人声音,自己在卧室用一台笔记本电脑的显卡就开始训练模型,后来做成了开源项目Fish Speech,在GitHub上收获了3万多星标。现在他们的平台能根据1.5万种自然语言提示,让AI语音带上面部细微的情感变化——比如温柔、急切或调皮,不再像以前那样死板。更厉害的是,只需5秒的录音样本,15秒内就能完美克隆出一个人的声音,还支持83种语言。在盲测中,67%的听众觉得他们家的AI声音比竞品更自然。这项技术对普通人来说,意味着未来听到的AI语音助手、游戏角色旁白、甚至有声书和视频配音,都会像真人一样有血有肉,再也不用忍受那种让人出戏的机器人腔调了。
UPDATED 09:30 EDT / JULY 28 2026
AI
Fish Audio makes a splash after raising $52M seed funding for AI voices
by
Mike Wheatley
Fish Audio, an ambitious artificial intelligence startup that began life as a weekend project in its founder’s bedroom, said today it has raised an impressive $52 million in seed funding to try and establish voice as the default interface for every AI model.
Coreline Ventures and Capital Today led the round, which also saw participation from a host of other backers, including 359 Capital, Play Time, HF0, 645 Ventures, Parable, Carya Venture Partners, Alphalist Partners and a number of unnamed angels.
The startup, which is officially known as Hanabi AI Inc., emerged from a familiar developer origin story. Its co-founder and Chief Scientist Shijia Liao, who formerly worked as a video researcher at Nvidia Corp. and is a lifelong fan of Japanese anime, explained that he grew tired of having to listen to the flat and monotonously robotic synthetic voices of early AI models, and decided that he needed to do something about it.
Liao decided to start training his own voice AI models, and set out to do so with nothing more than a single graphics processing unit housed in the laptop in his bedroom. Despite the limited compute available to him, Liao managed to transform the initial text-to-speech and voice cloning models he developed into Fish Speech, an open-source project that quickly gained rapid traction on GitHub. It amassed more than 31,000 stars as it caught the attention of indie developers, content creators and video games designers who were desperate for livelier and more expressive voice generation tools.
Since its launch in 2023, Fish Audio has evolved to become one of the most comprehensive voice AI platforms in the business, used by developers to quickly build powerful text-to-speech and voice cloning systems as well as voice agents. The platform was designed specifically to eliminate the lifeless synthetic audio that characterized early AI models by providing granular, word-level emotion controls that are driven by over 15,000 natural language prompts. It allows developers to fine-tune the exact tone, inflection and pacing of their AI-generated voices.
Fish Audio’s platform is powerful. It claims to be able to clone a voice from a mere five-second audio clip in less than 15 seconds. It boasts native support for 83 languages too, and its current flagship model S2.1 Pro was able to outperform its top competitors in a series of blind listening tests. According to those tests, 67% of listeners preferred Fish Audio’s voice outputs ahead of those from other models. Those numbers help to explain why its user base has grown to over eight million and its annual recurring revenue now exceeds $21 million.
Create realistic voices on Fish Audio today!#fishaudio #TTS #aivoice pic.twitter.com/LOwqFtBH5d
— Fish Audio (@FishAudio) July 20, 2026
While initially targeted at video games developers and content creators, Fish Audio now caters to organizations in regulated industries such as healthcare and financial services, offering secure on-premises deployments with zero-data retention and HIPAA compliance to avoid compromising customer privacy. Today’s round positions Fish Audio as a formidable challenger to better known voice AI startups such as ElevenLabs Inc., which recently raised $500 million in a round that valued it at a staggering $11 billion.
Fish Audio Chief Executive Rissa Cao said he started the company along with Liao because he wanted to make AI voices that sound human to be accessible to everyone. “We make high-quality, human-sounding voices available to every user, from beginner creatives to million-d
分享
阅读原文 ↗