· via The Verge
OpenAI claims Jalapeño AI chip beats Nvidia superchips on inference speed and efficiency
OpenAI says its Broadcom-built Jalapeño inference chip delivers 1.5–1.9x more work per watt and up to 3.6x lower latency than the best Nvidia-based results, with small volumes shipping this year.

OpenAI has published performance figures for Jalapeño, the custom AI chip it first introduced in June, claiming it runs inference workloads faster and more efficiently than competing systems. According to The Verge, the claims appeared in a company blog post on Tuesday and were expanded on in a briefing with reporters led by Richard Ho, OpenAI's vice president of hardware.
Best of both worlds
Ho described the chip as offering the "best of both worlds": lower latency combined with higher throughput. He argued this is unusual because AI systems typically have to trade one for the other — either answering individual requests quickly or handling many requests at once, but not both.
Jalapeño is an Application-Specific Integrated Circuit, or ASIC, developed in partnership with Broadcom, and it targets inference rather than training. Inference is the process of running an already-trained model to complete a task, answer a prompt or power an agent — the stage of the AI pipeline that determines how quickly users get responses and how much each response costs to serve.
How it was measured
To benchmark the chip, OpenAI used InferenceX, a platform for measuring how well systems handle inference workloads. The comparison targets were the best results recorded on the platform at the time, which had been achieved with systems built around Nvidia's GB200 or GB300 superchips.
Across three models — GPT-OSS 120B, DeepSeek R1 and Kimi K2.5 1T — OpenAI says Jalapeño delivered 1.5 to 1.9 times more AI work per watt than those Nvidia-based systems, along with 1.7 to 3.6 times lower end-to-end latency. Ho said the practical effect should be "faster responses, more responsive agents, and more reliable access as the demand grows."
These are OpenAI's own benchmark results, selected and framed by the company that built the chip, and The Verge's report gives no indication of independent verification.
Rollout and the Nvidia question
Jalapeño is set to deploy in "small volumes" by the end of this year, with volumes ramping up into 2027, Ho said. OpenAI did not disclose how many chips it expects to deploy next year.
Nor is the company treating the chip as a wholesale replacement for its existing infrastructure. Ho said Jalapeño will not replace OpenAI's entire chip lineup, noting that its broader compute strategy still involves "very good partners" such as Nvidia. Development of second- and third-generation versions of the chip is already underway.
Why it matters
Inference, not training, is increasingly where the economics of AI get decided. Every chat response and every step an autonomous agent takes consumes compute, and latency directly shapes how usable these systems feel. A chip that genuinely delivers more work per watt at lower latency would cut both the cost and the delay of serving large models at scale — which matters most as agentic workloads multiply the number of calls per user task.
The announcement also signals that OpenAI wants to own more of its hardware stack, following the broader industry pattern of AI and cloud players designing their own silicon rather than depending on a single supplier. But the claims warrant caution until third parties can test production hardware: vendor-run benchmarks on chosen workloads rarely tell the whole story, and the GB200 and GB300 will not stand still as comparison targets.
With volumes described as small through the end of this year, the real test arrives in 2027, when the planned ramp-up coincides with development of the chip's second and third generations. Until then, Jalapeño is best read as both a genuine engineering bet and a lever in OpenAI's ongoing relationship with Nvidia.
- #openai
- #ai-chips
- #inference
- #broadcom
- #nvidia
- #hardware