openai
First results land for OpenAI's Jalapeño inference chip
Promtime
openaiOpenAI has published first results for Jalapeño, its custom inference chip, with reported throughput above 700 tokens per second per user on DeepSeek R1 at concurrency one. In its OpenAI announcement, the company describes higher throughput, lower latency and improved power efficiency for modern AI models.
The chip was designed specifically for inference, the stage where trained models generate answers, and uses a general architecture rather than targeting only OpenAI models. Initial testing covered several open models and included measurements from the earlier A0 silicon.
At a glance
- Jalapeño is a general-purpose inference accelerator designed to handle different models and workload mixes on the same chip pool.
- OpenAI reports more than 700 tokens per second per user on DeepSeek R1, and approximately 1,400 on Kimi-K2.5 and GPT-OSS.
- The results put power efficiency at the center of competition as AI operators face constrained data-center capacity and rising inference demand.
Inference hardware is increasingly judged by tokens produced per unit of power, not only by peak computational output. Jalapeño therefore represents a test of whether a new accelerator can compete with established platforms through hardware and software coordination, while avoiding a fixed split between prefill and decode capacity. That approach may be especially relevant as model workloads shift from knowledge tasks toward reasoning and agentic systems.
Jalapeño reached more than 700 tokens per second per user on DeepSeek R1
OpenAI’s benchmark results show Jalapeño exceeding 700 tokens per second per user on DeepSeek R1 at concurrency one. The measurements used single-token prediction, with no speculative decoding and no prefill-decode disaggregation. On Kimi-K2.5 and GPT-OSS, OpenAI showed throughput of approximately 1,400 tokens per second per user.
According to SemiAnalysis, Jalapeño beat the Nvidia, AMD and Google accelerators tested with its InferenceX suite across several open models. The comparison included Blackwell and used configurations in which competing chips relied on multi-token prediction, while Jalapeño did not. SemiAnalysis verified InferenceX runs in OpenAI’s lab, but did not run the full suite or evaluate the AgentX benchmark.
The first Jalapeño results use A0 silicon, while B0 is expected to improve efficiency
OpenAI began the chip program with Broadcom in the middle of 2024 and reached manufacturing tape-out in roughly 16 months. The A0 stepping used for the reported results was nine months into the program. A B0 stepping is already in fabrication and is expected to deliver about 25% better performance per watt than A0.
The B0 compute die is specified at 13.4 petaflops of MXFP4 performance on TSMC’s N3P process. Jalapeño will use HBM4 and an N3E I/O chiplet with 32 lanes of 800G SerDes. Twenty-four lanes provide 600 GB/s for local scale-up within a rack, while eight lanes provide 200 GB/s for a multi-rack domain of 2,048 XPUs.
OpenAI chose a homogeneous pool instead of separating prefill and decode chips
Jalapeño keeps the draft model and main model on the same chips and fabric rather than assigning separate pools to prefill and decode. The architecture also uses PCIe Gen 5 for system I/O connecting the accelerator to an x86 host CPU.
SemiAnalysis reported that Jalapeño’s output-token throughput per megawatt exceeded Vera Rubin results that used multi-token prediction, as well as earlier GB200 results. It also reported nearly 700 tokens per second per user for Kimi-K2.5, more than nine times the next-best result in that comparison. The benchmark workload was 8k input and 1k output tokens, without AgentX’s longer-context, multi-turn testing.
Custom inference silicon is becoming a broader industry direction. Anthropic told Business Insider that it is staffing an internal custom silicon team for Claude while continuing to use AWS, Google, Nvidia and AMD hardware. The Information reported that Alphabet is developing a chip called Frozen v2, with parts of Gemini etched into the silicon.
Jalapeño’s production timeline remains open OpenAI has not announced a release date or commercial availability for the chip. The results come from early silicon, and both OpenAI’s accelerator and the comparison platforms remain subject to software and hardware improvements. The next meaningful test will be whether Jalapeño can retain its efficiency on larger models and production-style agentic workloads beyond the initial benchmark configuration.
Comments
No comments yet. Be the first.
Join the conversation
Sign in with Google to leave a comment. Your name and avatar come from your Google profile, and the comment appears after moderation.
We only use your name and avatar from Google. We never store your email address.
