· via Hacker News – Front Page (native)
OpenAI's first custom chip Jalapeño beats Blackwell on efficiency in SemiAnalysis tests
OpenAI detailed Jalapeño, its first custom inference chip, at Hot Chips. SemiAnalysis benchmarks show it beating Nvidia's Blackwell on performance per watt, though Rubin may be the fairer rival.

OpenAI shows off Jalapeño
OpenAI has publicly detailed Jalapeño, its first custom chip for large language model inference, following its announcement at the Hot Chips conference. According to SemiAnalysis, which says it was invited to OpenAI's labs to run its InferenceX benchmark suite on the hardware, the chip was developed in partnership with Broadcom from a blank slate, with the program first made public in June and design work starting in mid-2024. The part went from initial team hiring to manufacturing tape-out in roughly 16 months, a cycle SemiAnalysis describes as exceptionally fast for an ASIC. On paper it lands in flagship territory thanks to HBM4 memory, the same memory class used by current top-end Nvidia and AMD accelerators.
A general-purpose inference chip, not an OpenAI special
Much of the media coverage has framed Jalapeño as silicon tuned narrowly for OpenAI's own models. SemiAnalysis disputes that reading, describing it as a general-purpose inference design that performs well across many models and workload types rather than being optimized for a single point on the latency-throughput curve. The publication says it ran InferenceX alongside OpenAI engineers in the lab, and that OpenAI, half-jokingly, demonstrated Doom running on the chip after porting it using only Codex prompts.
What the benchmarks show
SemiAnalysis's headline metric is token throughput per megawatt, and on it Jalapeño comes out ahead of every Nvidia, AMD and Google chip the publication has tested, across multiple leading open-source models. Notably, the Jalapeño figures were produced without multi-token prediction, while the competing numbers represent each rival's best configuration, all using MTP.
On interactivity, SemiAnalysis reports more than 700 tokens per second per user at a concurrency of one on DeepSeek R1, achieved with single-token prediction, no speculative decoding and no prefill-decode disaggregation. Kimi-K2.5 and GPT-OSS reportedly ran at roughly 1,400 tokens per second per user, and GSM8k evaluation scores matched what the same models produce on Nvidia hardware.
Important caveats
SemiAnalysis is explicit about the limits of these results. Every number was supplied by OpenAI; the publication verified some InferenceX runs in person but did not execute the full suite, nor its AgentX benchmark, which it prefers for chip comparisons because its long-context, multi-turn datasets better reflect the cache behavior of real production workloads.
The publication also argues that measuring Jalapeño against Blackwell is an incomplete and somewhat unfair comparison. The real rival is Nvidia's Vera Rubin, which likewise uses HBM4 and is already shipping to customers — SemiAnalysis notes Vera Rubin NVL72 delivers 5.4 times the performance per megawatt of GB200 NVL72 — while OpenAI currently has little beyond engineering samples of Jalapeño. Even so, the publication says Jalapeño's single-token-prediction throughput per megawatt exceeds the MTP-based Rubin figures Nvidia and CoreWeave published in July, and on performance per dollar the two are nearly identical, with Jalapeño not using speculative decoding. The tested models are also not the largest available; Nvidia and AMD have published AgentX results on DeepSeek V4 Pro and Kimi K3, and bigger, newer models are harder to bring up on new silicon.
Designed around the power wall
SemiAnalysis explains the design philosophy: OpenAI is constrained by datacenter power rather than budget or floor space, so tokens per megawatt is the number that matters most. Nvidia has been making the same argument — SemiAnalysis cites Jensen Huang's Computex 2026 keynote, where he said that with a fixed gigawatt of power, throughput per watt translates directly into revenue — and notes that adding grid capacity moves on a far slower timeline than adding hardware, which is pushing operators toward behind-the-meter generation.
Why it matters
First-generation chips almost never threaten incumbents; Jalapeño, at least on these numbers, does. If the vendor-supplied results hold up under independent testing, one of Nvidia's largest customers has produced silicon that leads on the metric that now governs AI infrastructure economics. The 16-month design cycle also lends weight to claims that AI tools are accelerating chip design itself. The honest framing is that this is early days: engineering samples, OpenAI-provided numbers and no AgentX results yet, set against a Rubin platform that is ramping now. But the direction is clear — the biggest AI operators are willing to build their own silicon, and that reshapes both Nvidia's pricing power and the hardware stack behind frontier models.
- #openai
- #ai-chips
- #custom-silicon
- #nvidia
- #inference