deniz.in

Markets

Weather

Loading weather

· via TechCrunch

OpenAI's first Jalapeño benchmarks show throughput-per-kilowatt lead over Nvidia Blackwell

OpenAI's first benchmarks for its Jalapeño inference chip, presented at Hot Chips, show more tokens per user and higher throughput per kilowatt than an Nvidia Blackwell system, with deployment due from late 2026.

OpenAI's first Jalapeño benchmarks show throughput-per-kilowatt lead over Nvidia Blackwell

OpenAI publishes first Jalapeño benchmarks

At the Hot Chips conference on Tuesday, OpenAI gave its most detailed public look yet at Jalapeño, the inference processor it has been building with Broadcom, and released the first benchmark results for the system. According to TechCrunch, the chip was tested on SemiAnalysis' InferenceX benchmark and came out ahead of the strongest inference processors currently on the market on two measures: tokens delivered per user and throughput per kilowatt.

Richard Ho, OpenAI's head of hardware, told journalists on a press call that the results represent a major step beyond today's state of the art. In his telling, the system does more AI work for every unit of power it consumes while also returning responses faster — efficient enough to serve a large number of customers at once, yet still capable of low latency.

One caveat sits just beneath the headline figures: the comparison is against an Nvidia Blackwell system, meaning hardware that exists today. Competitors are unlikely to stand still, and the baseline may look considerably different by the time Jalapeño is running at scale.

Design targets the slow phases of inference

Jalapeño was first announced last October and developed in close collaboration with Broadcom, with OpenAI's own models used during the design process. The company intends it as the first step in a longer-term platform in which AI products, models, chips and memory evolve together rather than as separate layers.

That full-stack control, OpenAI argues, is what allowed its engineers to go after stages of the inference pipeline that typically cause friction — especially the prefill phase, where an incoming prompt is first processed, and the communication phase between components. In a blog post presenting the results, the company said Jalapeño was engineered to reduce how much data moves around and how long communication takes. Model state, including the KV cache that accumulates while a response is generated, can be kept local, with compute, memory and networking resources matched to whichever phase of inference is running.

Limited volumes first, scale in 2027

Ho estimated that Jalapeño will be deployed in limited quantities at the end of 2026, with a much larger rollout following in 2027. That schedule frames how the benchmark results should be read: the advantage shown at Hot Chips is measured against hardware available now, and rival accelerator roadmaps will keep advancing before OpenAI's chip reaches meaningful deployment.

Why it matters

Inference has become the dominant running cost of large-scale AI, and power efficiency increasingly sets how much inference capacity a provider can economically serve. A processor that returns more tokens per kilowatt changes that arithmetic directly, which is why the per-energy figures may matter more than raw speed alone.

The results also signal that OpenAI's custom-silicon effort is producing measurable engineering outcomes. Designing inference hardware against a stack of its own models gives the company optimization levers that general-purpose accelerator vendors cannot easily replicate. The open question is durability: Jalapeño reaches volume only in 2027, against competitors who will not be standing still, and the Hot Chips numbers capture the lead only as it stands today.

  • #openai
  • #chips
  • #inference
  • #benchmarks
  • #broadcom

Related posts