Fish Audio Raises $50M Seed for AI Voice Models Targeting Creators and Enterprises

Fish Audio, a Palo Alto-based AI voice startup founded by former NVIDIA researcher Shijia Liao, raised $50 million in a seed funding round announced on July 28, 2026. The round was led by Coreline Ventures and Capital Today, with participation from 359 Capital, Parable, Play Time, Alphalist Partners, Bayhouse Ventures, Carya Venture Partners, and HF0. CEO and co-founder Rissa Cao leads the company. TechCrunch
The company has built a substantial user base of more than 8 million users across its open-source and hosted models, and generates $21 million in annual recurring revenue. Its Fish Speech repository on GitHub has accumulated over 31,000 stars, reflecting meaningful traction within the developer community for an open-core voice generation stack.
Fish Audio has shipped five models in the past year: four speech generation models and one speech-to-text model. Three of the speech generation models have been open-sourced, while the latest, S2.1 Pro, is available exclusively through the company's paid API. This split between open weights and proprietary API access follows a pattern we have watched play out across the AI layer since the current cycle began: open-source releases build developer adoption and repository credibility, while the highest-capability model sits behind a monetized endpoint. The company also offers paid monthly plans for creators and teams that bundle generation minutes with voice cloning features.
On the enterprise side, Fish Audio counts HeyGen, Sanas, and Plaud among its customers. The company's model library contains more than 15,000 natural language controls, which let developers and creators direct speech output through descriptive prompts rather than SSML-style markup. Oskue Honda, a partner at Coreline Ventures, backed the round alongside Capital Today.
The funding lands against a backdrop of unresolved tension around voice cloning and consent. Fish Audio faced a controversy in which creators alleged their voices were uploaded to the platform without permission. In response, the company automated its DMCA voice takedown process to complete removals in under three minutes, a turnaround that addresses the mechanical bottleneck of takedown execution but does not settle the harder question of how voice data enters the platform in the first place.
The open-source-versus-proprietary split is worth examining. By open-sourcing three of four speech generation models while gating S2.1 Pro behind the API, Fish Audio is structurally aligning community adoption with revenue capture. Developers can self-host capable models, but the frontier model requires a paid relationship. Whether that gap between open and proprietary remains wide enough to sustain the API revenue is an open question; model capabilities in this space have been converging quickly across competing projects.
The consent controversy raises a broader point about the voice synthesis category that goes beyond Fish Audio specifically. Voice cloning platforms operate at the intersection of utility and identity, and the technical barrier to cloning a voice from a short sample is now negligible. Automated takedown reduces harm after the fact, but proactive consent verification, if it becomes a technical and regulatory expectation, could materially reshape how these platforms onboard voice data. The companies that build robust consent infrastructure early may find themselves better positioned if regulation follows.
What the $50 million gives Fish Audio is runway to push model quality on the proprietary side while maintaining its open-source footprint. The enterprise customer roster, the ARR figure, and the GitHub star count together suggest the company has crossed the threshold from research project to viable business, with the remaining challenge being whether it can scale its enterprise pipeline while managing the consent and provenance issues inherent to voice cloning. The capital and the customer base are in place; the governance question is not yet fully answered.


