Stock OpenAI models saved Jalapeño over 13% of die area

One attention kernel on OpenAI's Jalapeño chip started out using under one percent of what the chip's memory and compute could deliver. After roughly forty hours of work with OpenAI's own models, it was running at nearly ninety percent, a result Richard Ho, one of the first engineers on OpenAI's hardware project, discussed at length in an interview published by More Than Moore on Substack.
The same models, used with no fine-tuning, also fixed a more basic problem. Logic the team had promised on the strength of simulation did not fit on the die, and in one example the models found a layout that saved over 13% of die area.
At a glance
- OpenAI's first in-house chip, Jalapeño, is an inference accelerator built with Broadcom as design partner and Celestica on board and rack integration, and HotChips 2026 showed the first measured results from working silicon.
- Each chip pairs 216 GiB of HBM4 at 15.4 TB/s with a compute die and an IO chiplet, and a full 2,048-accelerator system reaches 27 EFLOP/s at four-bit precision.
- The catch: OpenAI's benchmark sprint used single-token prediction on three open models, while production runs multi-token prediction, and Ho admits very long contexts may cause trouble he has not yet analyzed.
If you haven't followed Ho's career, it reads like a tour of custom silicon. He started in verification on John Hennessy's multiprocessor project at Stanford, built the Anton supercomputers at D. E. Shaw Research and was among the first engineers on Google's TPU, where he stayed for seven or eight generations. After a stint at the photonics startup Lightmatter he joined OpenAI in 2023. In June 2026, OpenAI confirmed Jalapeño.
Stock OpenAI models saved over 13% of Jalapeño's die area
The trouble came from ordinary schedule pressure. The team had promised performance targets based on simulation and on estimates of how much logic would fit in the available area, and those estimates were slightly off. Instead of taking the performance hit, the team turned to OpenAI's internal models. Ho says they were not fine-tuned at all and were only a little ahead of the public ones.
Ho thinks the models succeed because they can hold the whole chip in their context window and spot the soft places where logic can move, which is hard for a human to keep in their head. Much of the optimized code was written in XLS, a high-level synthesis tool with Rust-like syntax that team member Chris Leary open-sourced at Google. That helped, because the models understood software better than Verilog.
Ho still doesn't fully trust them. Sign-off still runs through standard EDA flows: on a 300 million gate design, a model that is 99.99% correct is not accurate enough to tape out. The team also didn't always understand why the models rewrote code the way they did, so every change still goes through the full validation flow.
One chip covers prefill and decode across a 2,048-accelerator system
Jalapeño is rated at 700 W peak, and its measured sustained draw is closer to 550 W. A local domain holds 128 accelerators, and a full system has 2,048. For a practitioner, the design choice matters more than the specs. Jalapeño is a single balanced part for the entire inference workload, while other vendors split prefill, speculative decode and full decode across hardware built for each stage.
Ho frames the choice at the level of the whole data center fleet. If you split the hardware, you have to commit capex and power to a fixed ratio between workload types, and those ratios shift a lot. A single fungible part, in his words, gives minimum regret.
The team expected this to be the most controversial claim of the talk. Asked whether a do-everything design pays a tax, Ho said there may be one, but it isn't identifiable yet and may only show up in real deployment.
OpenAI claims up to 104x over GB200 and GB300 at hard operating points
At HotChips 2026, Ho, Ravi Narayanaswami and Chris Leary compared Jalapeño with NVIDIA GB200 and GB300. They claimed 1.5x to 1.9x better performance per watt at peak throughput, 1.7x to 3.6x better latency, and up to 104x better results at operating points NVIDIA struggles to reach.
Silicon came back around mid-May, so the talk was a sprint. OpenAI used InferenceX, a third-party end-to-end benchmark, on three open-source models. It also invited SemiAnalysis to evaluate the chip, so that the methodology would be, as Ho put it, unassailable. There was no time to build draft models for speculative decoding, so the kernel team set itself a target: beat the published multi-token prediction numbers using single-token prediction. Ho says they did.
Why Jalapeño ties each HBM bank to its own core
In NVIDIA's architecture, as Ho describes it, many SMs share one L2, so any core can reach any memory in a single uniform space, which keeps programming simple. Jalapeño bets the other way. Each HBM bank is linked to specific cores, and as long as a core stays within its own bank, data doesn't have to move. There is a ring between them, but the chip mostly relies on its high-bandwidth local paths.
Think of an office with one shared file room compared with desks that each have their own cabinet. The shared room is easier to organize, but every trip there costs time. Private cabinets are faster, as long as each person's work fits in their own drawer.
When the project started, nobody knew whether real workloads would map onto this layout. Ho says the team has yet to find an open-source model it can't map, with the compiler, the kernels and the models doing much of the work. He expects even GPU makers to move toward this NUMA approach.
The benchmark picture is narrower than the headline multiples suggest. The comparison used three open models in single-token mode, production runs multi-token prediction on internal models nobody outside can test, and Ho says the team has moved on to production and may not publish much more. In our view, the weakest untested point for a one-chip-fits-all design is very long contexts, which Ho concedes may cause trouble and which the team has not yet analyzed.
B0, Gen 2 and Scotch Bonnet
The B0 stepping was planned from the start and is not a fix for bugs in A0. It is now in the lab going through qualification, and OpenAI is moving into volume ramp. No deployment date or fleet size has been given. The roadmap shown at HotChips runs through Gen 2 and Gen 3. Ho calls Jalapeño the low-hanging fruit and says later chips may add new architectural ideas. Asked whether the next chip might be called Scotch Bonnet, he said it could be.
Related stories
- Fewer than 100 people built OpenAI's Jalapeño chip
- OpenAI built idle silicon into Jalapeño on purpose
- Anthropic and OpenAI start shopping in 20–30 MW slices
- OpenAI pays $300M for the team behind Apple's Portrait Mode
- OpenAI's Habitat grew from a Python library to 22M req/sec
- AI-written GPU kernels ship without a line-by-line read
Comments
No comments yet. Be the first.
Join the conversation
Sign in with Google to leave a comment. Your name and avatar come from your Google profile, and the comment appears after moderation.
