Technology

Modal Labs Nears $750M Round at $15.75B as Inference Revenue Grows

Martin HollowayPublished 6d ago3 min readBased on 3 sources
Reading level
Modal Labs Nears $750M Round at $15.75B as Inference Revenue Grows
Photo by panumas nikhomkhai on Pexels

Modal Labs is closing in on a $750 million round led by Accel at a $15.75 billion valuation including the investment. TechCrunch reported the talks on September 28, 2026, citing a source with knowledge of the funding.

That price would more than triple the company's $4.65 billion valuation from four months ago. That earlier valuation was tied to a previously announced $355 million fundraise. The structure now under discussion would put a very large amount of primary capital into the company at a sharply higher valuation in a short interval.

Modal was founded in 2021 by CEO Erik Bernhardsson and CTO Akshat Bubna. The company is based in New York and has an estimated roughly 150 employees. As of May 2026, it had passed $300 million in annualized revenue. Modal declined to comment on the funding talks.

The September 28 report follows earlier reporting of a new raise. On September 23, 2026, Bloomberg News reported that Modal was in talks to raise new financing at a roughly $15 billion valuation. Bloomberg described both Modal and Baseten as startups in funding talks to help businesses run AI.

The broader context here is where money is made in applied AI. Training models gets attention. Running them for customers, called inference, pays the bills. Business teams focus on utilization of chips, autoscaling for demand spikes, cold starts that cause delays, inference latency or response time, throughput per accelerator, and overhead across regions and software frameworks. A provider that hides that work can turn engineering pain into steady revenue, which helps explain the May figure.

In my view, the headline valuation matters less than capital raised against team size and revenue. Roughly 150 staff with triple-digit millions in annualized revenue points to heavy automation and heavy use of outside compute. That pace is unusual. It also creates execution risk around buying capacity, reliability, support, and cost control as concurrent workloads grow.

Looking at what this means for teams running models in production, large rounds like this tend to fund capacity and abstraction. Capacity means reserved accelerators, networking, and footprint. Abstraction means interfaces, schedulers, observability tools, and integrations that let a small platform team serve many product teams without running clusters directly. Capital follows usage. The test will be whether ease of use and unit economics hold when early work becomes sustained traffic.

Worth flagging, after decades watching infrastructure cycles, is how fast scarcity can turn to sufficiency. The optimistic case is straightforward. If inference stays the bottleneck for real business use, well funded independent layers can thrive by making deployment boring and repeatable. That would help practitioners who prefer shipping features to tuning orchestration.