Technology

Snorkel AI Raises $350M at $3.5B as Training Data Becomes AI's Bottleneck

Martin HollowayPublished 2w ago3 min readBased on 2 sources
Reading level
Snorkel AI Raises $350M at $3.5B as Training Data Becomes AI's Bottleneck
Photo by cottonbro studio on Pexels

Snorkel AI has raised $350 million in Series E financing at a $3.5 billion valuation.

The round was announced on Sept. 22, 2026. That is almost triple its prior valuation in under a year and a half. TechCrunch

Insight Partners and S32 led the round. Addition, Lightspeed, Greylock, GV and Wells Fargo participated. Company materials dated July 2024 had also named Third Point, March, Blumberg, Allegis, Standard VC, Frontline, P7, Walden Catalyst Ventures and Factory as participants in the Series E.

Seventeen months earlier, Snorkel had raised $100 million in a Series D at a $1.3 billion valuation. Alongside the new raise, the company reported a $375 million annualized revenue run-rate, or current sales extended over a year. That figure is up 18-fold over the prior 12 months.

Snorkel launched commercially in 2019. The launch followed four years of research by co-founder and CEO Alex Ratner and his team at a Stanford AI lab. It began as data labeling automation software, a programmatic layer to reduce manual annotation. It has since moved to delivering completed datasets as data-as-a-service.

In practice, the customer contracts for an outcome, domain-specific training, fine-tuning or evaluation data ready for model ingestion. The customer does not license a labeling stack and build the pipeline internally. It is like shifting from selling kitchen equipment to delivering finished meals. Snorkel uses software and models to generate data synthetically, meaning created by systems rather than collected by hand, working alongside subject matter experts, rather than operating purely as a human expert marketplace.

The broader context here is a change in what constrains AI development. Compute scarcity dominated discussion for two years. Data scarcity has now moved to the foreground, particularly for post-training, for reasoning traces, step-by-step records of how a model reaches an answer, for agentic tool use, where AI agents operate software tools, and for vertical domains where public web corpora are insufficient. Customers are no longer buying sample datasets to evaluate. They are buying continuous supply under production contracts. The syndicate mixes enterprise software specialists and large financial institutions. That fits late-stage AI infrastructure rounds, where capital needs are high and the customer base sits largely inside the Fortune 500 and frontier model labs. It leaves Snorkel among the best-capitalized private suppliers of training data for large-scale AI systems.

In my view, Snorkel's pivot from tooling to finished datasets tracks that change. Tooling alone leaves task design, quality control and expert recruitment with the enterprise. A managed dataset product absorbs that work, which large buyers will pay a premium for when model quality depends on proprietary or highly structured examples. Worth flagging is the durability question. A synthetic-first pipeline with expert oversight can iterate faster on edge cases, regenerate rejected samples and enforce consistent formatting across large volumes, while pure marketplaces add capacity in direct proportion to expert hours. Synthetic generation lowers marginal cost, but it must be filtered, audited and grounded or errors compound downstream. Differentiation will rest on expert networks, evaluation methodology and proof that delivered data improves measured model performance rather than simply expanding token volume. If Snorkel can tie its $375 million run-rate to repeatable accuracy gains, the $3.5 billion valuation will look like infrastructure pricing. If not, it will look like services revenue at software multiples.