Technology

Infinity Raises $15M to Automate Inference Stack Development for Non-Nvidia Chips

Martin HollowayPublished 2d ago4 min readBased on 3 sources
Reading level
Infinity Raises $15M to Automate Inference Stack Development for Non-Nvidia Chips

AI infrastructure startup Infinity has raised $15 million at a $100 million valuation to build software that automates the creation of low-level inference stacks for AI chips. The round includes Touring Capital, Principal VC, and individual researchers from OpenAI and Anthropic. TechCrunch

Founded in 2025 by Jeremy Nixon, a former Google Brain researcher and creator of the AGI House hacker network community, Infinity operates under the domain infinity.inc. Nixon told TechCrunch the company emerged from his obsession with the concept of "automated invention," the belief that AI systems can function as a meta-technology. TechCrunch

Nixon previously developed a machine learning algorithm called Omega, which generated new ML algorithms and evaluated them automatically in a feedback loop. Infinity is now applying that philosophy to hardware. The company is building a universal inference library designed to run on all chips, enabling automated replication of state-of-the-art research results. TechCrunch

Central to Infinity's approach is Ignition, an AI research agent that writes the low-level code required for inference on Nvidia-alternative chips. Ignition tests, debugs, measures performance, and automatically rewrites code to improve it. Infinity claims Ignition constitutes a CUDA-level software stack. TechCrunch

AI chip maker D-Matrix, a would-be Nvidia challenger, is an Infinity customer. In a case study involving D-Matrix's Corsair platform, Ignition reduced a process that could have taken months or years down to hours or days. TechCrunch

Infinity's commercial model skips upfront licensing fees. Instead, the company takes a percentage of performance gains and cost savings, measured in tokens per second. TechCrunch

Nixon said Infinity is in talks with other major chip and cloud companies, though he did not name them. As of July 2026, the company employs 26 people across design, operations, and engineering. The TechCrunch report is based on an interview with Nixon and does not cite an external press release or a posting on Infinity's own domain. TechCrunch

The economics here are notable. By tying revenue to measured inference performance gains rather than licensing access, Infinity aligns its compensation directly with the hardware vendor's competitive positioning. If Ignition's automated code generation can genuinely match the quality of hand-tuned CUDA kernels, the cost structure of bringing a new AI accelerator to market shifts downward. That matters because the CUDA software ecosystem, not raw silicon capability, has been the primary competitive moat protecting Nvidia's data center dominance. TechCrunch

However, the scope of the claims warrants scrutiny. The D-Matrix case study is a single data point, and the details originate from Infinity's own research documentation rather than an independent benchmark. Automated code generation for novel architectures is a substantially different challenge than optimizing for established, well-documented hardware. The claim of a CUDA-level stack is ambitious and remains unverified by third parties. TechCrunch

If Ignition functions as described, it lowers the barrier to entry for silicon alternatives. Chips from smaller vendors or new entrants could reach production-ready inference performance faster, relying on automated optimization rather than large internal compiler and kernel engineering teams. That would not displace Nvidia overnight, but it could compress the timeline for viable alternatives to reach the market. The long-term implication is that the software ecosystem advantage narrows, and hardware performance and pricing become the deciding factors for AI accelerator selection. TechCrunch

For now, Infinity is a 26-person company with a working case study and a funding round backed by investors who understand the inference landscape. Whether Ignition scales across diverse architectures remains the open question, but the approach addresses a real and expensive bottleneck in AI hardware development.