Technology

OpenAI Built Its Own AI Chip — and It's Faster Than Nvidia's

Martin HollowayPublished 2month ago4 min readBased on 8 sources
Reading level
OpenAI Built Its Own AI Chip — and It's Faster Than Nvidia's
Photo by Pok Rie on Pexels

At the Hot Chips conference on August 25, 2026, OpenAI presented detailed technical specifications and the first benchmark results for Jalapeño, its custom AI chip codeveloped with Broadcom. The chip outperformed an Nvidia Blackwell system on SemiAnalysis's InferenceX benchmark, delivering more tokens per user and higher throughput per kilowatt than currently available state-of-the-art inference processors, according to TechCrunch.

Inference is the stage where a trained AI model actually generates responses — as opposed to training, where the model is first built. It is the most common and ongoing use of AI, and it consumes enormous amounts of computing power.

Richard Ho, OpenAI's head of hardware, said the benchmark results show "a very significant performance advance over state of the art," allowing Jalapeño to serve more AI work per unit of power with lower latency. The comparison was made against an Nvidia Blackwell system on InferenceX, an independent benchmark from SemiAnalysis that measures how well hardware serves AI requests. SemiAnalysis states that InferenceX has been widely reproduced, validated, and supported by nearly every major compute buyer, including Google Cloud and Microsoft Azure. The benchmark's v2 release, published in February 2026, evaluated NVIDIA Blackwell against AMD and Hopper GPUs.

Jalapeño's architecture is built around minimizing delays in the parts of inference that OpenAI identifies as the most common bottlenecks: prefill and communication. Prefill is the moment when the model first processes your input before it starts generating a response — a step that has gotten slower as AI models handle longer and longer conversations. In a blog post, OpenAI said the chip minimizes data movement and communication delays by keeping the model's working data close to where the computation happens, rather than shuffling it around. The chip is built to be as large as a single piece of silicon can be, packing in as much logic as physically possible. It was completed in an ultra-fast nine-month development cycle, as Tom's Hardware reported.

The chip was first announced in October 2025 and developed in close collaboration with Broadcom, with OpenAI's own AI models assisting in the development process. OpenAI's newsroom page, published in June 2026, introduced Jalapeño as a custom AI chip built to improve performance, efficiency, and scale across AI systems. At that time, OpenAI had received first samples of the chip and was testing how it handled AI workloads. The company said early testing showed performance per watt substantially better than current alternatives, though final performance measurements were still underway.

Those measurements have now arrived, at least in preliminary form. The Hot Chips presentation marks the transition from OpenAI's June claims of promising early results to independently structured benchmark figures on a recognized industry framework.

Ho estimated that Jalapeño will deploy in very small volumes at the end of 2026, with more significant deployment coming in 2027. OpenAI also said it plans to make Jalapeño a multigenerational platform, allowing AI products, models, chips, and memory to be developed together rather than one after the other.

The broader context here is that building your own chip is a major undertaking, and the nine-month timeline is striking. A chip of this complexity typically takes 18 to 24 months to design and produce. OpenAI's use of its own AI models to help with the design process, though described in general terms, suggests that AI-assisted tools are beginning to compress one of the slowest parts of chip development.

The industry question is whether a chip purpose-built for AI inference, owned and operated by a single company, changes the economics of running AI at scale enough to shift the competitive landscape. Nvidia's Blackwell platform remains the default for most operators, with mature software and broad support. OpenAI's decision to benchmark against Blackwell rather than an easier target signals confidence in the results, though the company is also the sole operator of Jalapeño, which means real-world deployment data will come only from OpenAI's own infrastructure.

For now, the benchmark numbers stand on their own. Whether Jalapeño's advantages hold at production scale, across diverse workloads and model architectures, will become clearer as the chip moves from limited deployment into broader use in 2027.