
On August 25 and 26, 2026, OpenAI published the first results for Jalapeño, its very first in-house AI chip, built specifically for inference, meaning running already-trained models rather than training them. According to independent benchmarks reported by The Register, Winbuzzer and CNBC, this chip delivers up to 1.9 times more throughput per kilowatt than the latest Nvidia racks. For an SMB paying a monthly AI API bill, the question is simple: will this bring prices down?
In brief
- Jalapeño is OpenAI's first AI chip, designed in-house and built with Broadcom (silicon and networking), with system integration handled by Celestica.
- It is dedicated to inference (running models), not training: OpenAI continues to use Nvidia and AMD GPUs to train its models.
- According to SemiAnalysis InferenceX benchmarks, published on August 25-26, 2026, Jalapeño delivers 1.5 to 1.9 times more throughput per kilowatt and 1.7 to 3.6 times lower latency than Nvidia GB200 and GB300 racks.
- Its rated power is 700 W, versus 1,200 to 1,400 W for the comparable Nvidia racks, at equal or higher performance.
- A limited rollout is planned for late 2026, with volume production ramping up in 2027: no pricing changes have been announced for ChatGPT or OpenAI API users.
- For an SMB, the short-term takeaway is not a guaranteed price cut, but a signal of growing competition on AI hardware, which tends to favor customers over time.
What the Jalapeño chip actually is
Jalapeño is a specialized circuit (ASIC) designed solely to serve already-trained models to users, a task known as inference. This is OpenAI's most costly operation at scale: every ChatGPT message and every API call triggers an inference operation. The architecture was designed by OpenAI, with silicon manufacturing and networking handled by Broadcom, and rack assembly by Celestica.
Unlike a general-purpose GPU such as Nvidia's, which must handle both training (very compute-intensive) and inference (more sensitive to latency and per-request cost), Jalapeño is optimized for a single use case. According to OpenAI, the architecture is designed "to minimize data movement and communication delays" between chips within the same rack, a key factor for responding quickly to millions of simultaneous requests.
The benchmark numbers against Nvidia
The tests were run by the analysis firm SemiAnalysis, through its InferenceX protocol, on three open models: GPT-OSS-120B, DeepSeek R1 (670 billion parameters) and Kimi K2.5 (one trillion parameters), using requests of 8,000 input tokens and 1,000 output tokens.
| Metric | Jalapeño (OpenAI) | GB200 / GB300 (Nvidia) |
|---|---|---|
| Rated power (TDP) | 700 W | 1,200 W to 1,400 W |
| Throughput per kilowatt | 1.5x to 1.9x higher | Baseline |
| End-to-end latency | 1.7x to 3.6x faster | Baseline |
| Intended role | Inference only | Training + inference |
| Availability | Engineering samples, limited volume late 2026 | Available today |
Source: SemiAnalysis InferenceX benchmarks, reported by The Register and Winbuzzer (August 25-26, 2026).
A full rack packs 128 Jalapeño accelerators, for a combined compute power of 1.7 exaflops at 4-bit precision and 27.5 TB of HBM4 memory, fed by roughly 2 petabytes per second of memory bandwidth. These are lab-measured figures: they capture raw infrastructure performance, not yet the final cost billed to end users.
Why OpenAI is building its own chip
This move is not an isolated one. Google has developed its TPU chips over several generations for its own services and for Google Cloud. Amazon offers its Trainium and Inferentia chips on AWS. Microsoft unveiled its Maia chip for Azure. Meta is developing its own MTIA chip for internal inference workloads. OpenAI is thus joining an industry-wide trend: reducing dependence on a single GPU supplier and gaining better control over inference costs, which represent a growing share of major AI labs' spending as user numbers rise.
Nvidia remains essential for training
OpenAI continues to use Nvidia and AMD GPUs to train its future models. Jalapeño does not replace this infrastructure: it adds to it, exclusively for the phase of serving users.
Rollout timeline
August 25-26, 2026
First public benchmarks
Late 2026
Limited deployments
2027
Volume ramp-up
What this actually means for an SMB
No immediate effect on your bills
A potential lever for availability
Competition that benefits customers over time
Track the actual price, not the announcement
What we still don't know
OpenAI has not disclosed Jalapeño's manufacturing cost, its reliability at scale, or the actual utilization rate planned for its data centers. The final cost per token for end users depends on these three factors, which remain unknown at this stage.
FAQ
What is OpenAI's Jalapeño chip?
Jalapeño is the first AI chip designed by OpenAI, dedicated to inference (running already-trained models). It is developed with Broadcom for silicon and networking, and Celestica for system integration.
Will Jalapeño lower the price of ChatGPT or the OpenAI API?
Nothing has been announced in that direction so far. The measured efficiency gains (up to 1.9 times more throughput per kilowatt than a Nvidia GB200 or GB300 rack) create the potential for lower infrastructure costs, but OpenAI has not communicated any pricing change linked to Jalapeño.
Is OpenAI abandoning Nvidia?
No. OpenAI continues to use Nvidia and AMD GPUs to train its models. Jalapeño is reserved for inference, meaning serving user requests, and complements this infrastructure rather than replacing it.
When will the Jalapeño chip be available at scale?
According to reporting from The Register and Winbuzzer, limited deployments, still considered engineering samples, are planned for late 2026, with volume ramp-up announced for 2027.
Conclusion
Jalapeño confirms a broader trend: major AI labs are working to reduce their dependence on Nvidia in order to control inference costs, an expense line that grows with every new user. For an SMB, there is no reason to change anything right now: no price cut has been announced, and rollout remains limited until 2027. The right move is to track official pricing pages rather than performance announcements, and to keep an eye on this hardware competition, which historically benefits users over time. To follow the evolution of AI costs, read our article on the cost of AI agents with Claude Opus 5 or browse all our Mag resources.


