Technology

Fish Audio Raises $50M Seed to Scale AI Voice Models for Creators and Enterprises

Martin HollowayPublished 3d ago5 min readBased on 1 source
Reading level
Fish Audio Raises $50M Seed to Scale AI Voice Models for Creators and Enterprises

Fish Audio, a Palo Alto-based AI voice startup founded by former NVIDIA researcher Shijia Liao, raised $50 million in a seed funding round announced on July 28, 2026. The round was led by Coreline Ventures and Capital Today, with participation from 359 Capital, Parable, Play Time, Alphalist Partners, Bayhouse Ventures, Carya Venture Partners, and HF0. CEO and co-founder Rissa Cao leads the company. TechCrunch

The company has built a user base of more than 8 million across its open-source and hosted models, and generates $21 million in annual recurring revenue. Its Fish Speech repository on GitHub has accumulated over 31,000 stars, a measure of developer interest that signals real traction for an open-core voice generation platform. (Open-core means the underlying software is freely available, while advanced features or higher-capability models are sold as paid services.)

Fish Audio has shipped five models in the past year: four speech generation models and one speech-to-text model. Three of the speech generation models have been open-sourced, while the latest, S2.1 Pro, is available exclusively through the company's paid API. This split between open weights (freely downloadable model parameters) and proprietary API access follows a pattern that has become common across the AI landscape: open-source releases build developer adoption and credibility, while the most capable model sits behind a paid endpoint. The company also offers paid monthly plans for creators and teams that bundle generation minutes with voice cloning features.

On the enterprise side, Fish Audio counts HeyGen, Sanas, and Plaud among its customers. The company's model library contains more than 15,000 natural language controls, which let developers and creators direct speech output using plain-language descriptions rather than SSML, a verbose markup language traditionally used to control how synthesized speech sounds. Oskue Honda, a partner at Coreline Ventures, backed the round alongside Capital Today.

The funding lands against a backdrop of unresolved tension around voice cloning and consent. Fish Audio faced a controversy in which creators alleged their voices were uploaded to the platform without permission. In response, the company automated its DMCA voice takedown process to complete removals in under three minutes, a turnaround that addresses the mechanical bottleneck of executing takedowns but does not settle the harder question of how voice data enters the platform in the first place.

The broader context here is the structural choice Fish Audio has made by open-sourcing three of four speech generation models while gating S2.1 Pro behind the API. Community adoption and revenue capture are aligned by design: developers can self-host capable models, but the frontier model requires a paid relationship. Whether that gap between open and proprietary remains wide enough to sustain API revenue is an open question, as model capabilities in this space have been converging quickly across competing projects.

The consent controversy also raises a point about the voice synthesis category that goes beyond Fish Audio specifically. Voice cloning platforms operate at the intersection of utility and identity, and the technical barrier to cloning a voice from a short audio sample is now negligible. Automated takedown reduces harm after the fact, but proactive consent verification, if it becomes a technical and regulatory expectation, could materially reshape how these platforms onboard voice data. Companies that build robust consent infrastructure early may find themselves better positioned if regulation follows.

What the $50 million gives Fish Audio is runway to push model quality on the proprietary side while maintaining its open-source footprint. The enterprise customer roster, the ARR figure, and the GitHub star count together suggest the company has crossed the threshold from research project to viable business. The remaining challenge is whether it can scale its enterprise pipeline while managing the consent and provenance issues inherent to voice cloning. The capital and the customer base are in place; the governance question is not yet fully answered.