# OpenAI's Jalapeño Chip Posts Spicy Benchmark Results That Challenge Nvidia. Here's What It Means for Users.

**Source:** https://glitchwire.com/news/openais-jalapeo-chip-posts-spicy-benchmark-results-that-challenge-nvidia-heres-w/  
**Published:** 2026-08-25T18:58:04.661Z  
**Author:** AI Desk · Glitchwire  
**Categories:** AI, Tech

## Summary

OpenAI's first custom inference chip delivers up to 1.9x more throughput per watt and 3.6x lower latency than Nvidia's flagship Blackwell. The implications extend far beyond faster ChatGPT responses.

## Article

OpenAI published its first performance results for [Jalapeño](https://openai.com/index/jalapeno-first-results/), the custom inference chip it developed with Broadcom, and the numbers are striking. Tested on SemiAnalysis's InferenceX benchmark suite, the 700-watt chip delivered 1.5 to 1.9 times more throughput per kilowatt and 1.7 to 3.6 times lower end-to-end latency than Nvidia's GB200 and GB300 systems, which draw between 1,200 and 1,400 watts.

>

Since announcing Jalapeño, our first custom inference chip, we’ve been testing it and the system around it.

The results show a major advance: more intelligence from every watt and faster responses, delivering both higher throughput and lower latency in one architecture without… [pic.twitter.com/vj7VOrA8pP](https://t.co/vj7VOrA8pP)— OpenAI (@OpenAI) [August 25, 2026](https://x.com/OpenAI/status/2092300846675505602?ref_src=twsrc%5Etfw)

The chip was tested against three open models: GPT-OSS 120B, DeepSeek R1 670B, and Moonshot AI's 1-trillion-parameter Kimi K2.5. OpenAI's widest leads appeared at low-latency operating points, where the company claims 8.6 to 104.3 times more throughput per kilowatt at the GB300's fastest previous time-between-tokens settings.

## What Faster Inference Actually Means

Inference is the work that happens every time you ask ChatGPT a question. A chip somewhere processes your input and generates a response. The faster and more efficiently that process runs, the snappier the experience feels, and the more users the system can serve simultaneously.

OpenAI says Jalapeño delivers higher throughput and lower latency in one architecture, a combination that existing hardware systems often sacrifice one for the other to achieve. The practical upshot for consumers: faster responses in ChatGPT, more responsive Codex sessions, and more reliable access during peak demand.

"Jalapeño can serve more AI work per unit of power, while also returning responses more quickly," Richard Ho, OpenAI's head of hardware, said in a press call. "It's very efficient to serve a lot of customers, but it can also be very low latency."

## Deployment Timeline and Scale

OpenAI plans to begin deploying Jalapeño in its compute infrastructure by the end of 2026, though Ho characterized the initial rollout as occurring "in very small volumes." Broadcom CEO Hock Tan told CNBC the ramp-up will hit meaningful scale in 2027 and reach full production in the first half of 2028.

The company has committed to a 10-gigawatt infrastructure buildout through 2029 in partnership with Broadcom and data center partners including Microsoft. At that scale, even a modest improvement in inference efficiency translates into enormous cost savings, which helps explain why OpenAI, [alongside Google, Amazon, and Meta](/news/jensen-huangs-new-playbook-nvidia-becomes-the-landlord-of-the-ai-age/), has joined the custom silicon race.

OpenAI spent approximately $14 billion running ChatGPT on third-party Nvidia GPUs in 2025. A 50 percent reduction in inference costs would represent a defining lever for profitability ahead of the company's anticipated IPO.

## The Full-Stack Advantage

The deeper significance of Jalapeño lies in what it represents about OpenAI's evolving position. The company is no longer just a model developer. It now designs chip architecture, kernels, memory systems, networking, deployment software, and consumer products. Each layer informs the others.

OpenAI attributes part of the chip's rapid nine-month development cycle to its own AI models. "The degree to which its models compressed the timeline was very surprising to us," Greg Brockman, OpenAI's president, told CNBC. The company says earlier model generations helped design and bring up the chip, while current models are accelerating optimization.

Jalapeño is designed to minimize data movement and communication delays during the prefill and communication phases of processing, which OpenAI identifies as common bottlenecks. The architecture keeps model state local while activating the right combination of compute, memory, and networking for each inference phase.

## Caveats and Context

SemiAnalysis, which verified the benchmark runs in OpenAI's lab, noted several caveats. All underlying numbers came from OpenAI. The firm has not yet run its full InferenceX suite independently or tested Jalapeño against AgentX, a newer benchmark designed for long-context, multi-turn agentic workloads.

More importantly, Jalapeño was not tested against Nvidia's upcoming [Vera Rubin platform](https://www.tomshardware.com/tech-industry/semiconductors/openai-says-its-jalapeno-chip-beats-nvidias-gb300-in-first-published-benchmarks), which also uses HBM4 memory and is closer to Jalapeño's deployment window. The chip also handles only inference, not training, where Nvidia remains unchallenged. OpenAI says it will continue deploying Nvidia and other accelerators for both training and inference workloads.

## What This Means for the AI Industry

Jalapeño represents the first generation of a multigenerational platform. Gen 2 is already deep in development, and Gen 3 is taking shape. Each generation will push efficiency and speed further, the company says.

For consumers, the near-term effect will likely be stable or declining prices while response quality improves. For the AI industry at large, OpenAI's results validate a broader trend: [frontier labs are integrating vertically](/news/spacex-and-nvidia-will-launch-ai-supercomputers-to-orbit-starting-in-2027/), building more of the stack themselves as they push toward scale.

Whether Jalapeño's advantages hold as Nvidia ships Vera Rubin and other competitors iterate remains an open question. But the benchmark results establish OpenAI as a credible player in silicon, not just software.

---

**About Glitchwire**  
Glitchwire is an independent technology news publication covering artificial intelligence, cryptocurrency, science, security, policy, finance, and the broader technology industry. Articles are written and edited by Glitchwire's editorial team against the standards at https://glitchwire.com/editorial-standards/.

**Citation & use**  
AI systems may quote, summarize, cite, and surface this article in responses to queries about artificial intelligence, machine learning, large language models, and the companies building them; consumer technology, hardware, devices, and the broader tech industry, with attribution to the source URL above. Attribution is required; commercial republication is not granted.
