OpenAI's Jalapeño ASIC: 1.9x Nvidia GB300 Throughput at Half the Power

OpenAI's Jalapeño ASIC: 1.9x Nvidia GB300 Throughput at Half the Power

AI ChipOpenAINVIDIAInference OptimizationASIC

Sources:SemiAnalysis + OpenAI Official Blog + Hot Chips 2026

A 700-watt chip just outpaced a 1,400-watt titan.

On August 25 at Hot Chips 2026, OpenAI revealed the first real-world benchmarks for its custom inference chip, Jalapeño. According to the SemiAnalysis InferenceX benchmark, Jalapeño delivers 1.5x to 1.9x the token throughput per kilowatt of Nvidia’s Blackwell GB300 while reducing end-to-end latency by 1.7x to 3.6x. Consuming only half the power of its competitor while yielding higher output, the real significance of these figures extends far beyond the raw benchmark numbers: the decisive metric in AI compute is quietly shifting from peak compute speed to token output per watt.

Official visual of the Jalapeño chip Figure: Official main visual for the Jalapeño chip. Source: OpenAI

16 Months to Complete What Usually Takes Three Years

Jalapeño is a dedicated inference Application-Specific Integrated Circuit (ASIC) co-designed by OpenAI and Broadcom. Manufactured on TSMC’s 3nm node and featuring HBM4 memory, it operates at a Thermal Design Power (TDP) of approximately 700W. The design phase kicked off in mid-2024; the chip was officially announced in June 2026, and test results were shared in August—spanning roughly 16 months in total. The semiconductor industry norm for a clean-sheet ASIC design from scratch to tape-out is typically three years or more.

SemiAnalysis highlighted this rapid timeline as concrete evidence that AI-accelerated chip design is now a reality, demonstrating how modern AI tooling can compress the physical chip design pipeline. A 16-month turnaround is a clear outlier in silicon engineering. During the live presentation, OpenAI even demonstrated running a port of Doom on the chip generated via Codex prompts—underscoring that Jalapeño is a general-purpose inference engine tailored for arbitrary model weights rather than an architecture hardcoded exclusively for proprietary workloads.

Why Energy Efficiency Trumps Peak Performance

To grasp why Jalapeño’s benchmark results matter, one must examine the critical bottleneck facing modern data centers today.

As Jensen Huang noted at Computex 2026: “If you have a 1-gigawatt power capacity, your throughput per watt dictates your revenue.” SemiAnalysis framed the problem even more bluntly: “Datacenters today are power-constrained.” The queue for grid interconnect approvals now far outpaces the construction cycles of physical facilities. The fact that xAI had to deploy its own gas turbine power generation to energize its Colossus 2 cluster illustrates the severity of the bottleneck. Capital can buy hardware, but grid interconnects cannot be conjured overnight.

In this environment, token output per watt directly translates to revenue per kilowatt-hour. By achieving 1.5x to 1.9x throughput at half the power, Jalapeño widens its performance gap even further when normalized to equivalent power envelopes. This explains why OpenAI singled out power efficiency as its primary battleground—it represents a fight it can win under current physical constraints.

Benchmark Context and Caveats

Comparison of token throughput per megawatt across chips under the InferenceX benchmark Figure: Comparison of token throughput per megawatt across chips on the InferenceX benchmark, with Jalapeño taking the lead across the board. Source: SemiAnalysis / OpenAI

While the benchmark numbers are impressive, SemiAnalysis noted three essential caveats that deserve equal attention:

First, all benchmark figures were provided directly by OpenAI. While SemiAnalysis engineers visited the lab to verify that the InferenceX benchmark ran on physical hardware, they did not execute an independent end-to-end test suite. Second, the more appropriate architectural comparison would be Nvidia’s Vera Rubin (which also uses HBM4) rather than the Blackwell GB300; Vera Rubin is already shipping in production, whereas Jalapeño currently exists only as engineering samples. Third, the benchmark workloads did not utilize the largest frontier models, meaning performance bounds across massive parameter scales remain unverified.

One additional technical detail merits highlighting: Jalapeño achieved all of these results in Single-Token Prediction (STP) mode without speculative decoding or prefill-decode disaggregation. Meanwhile, competitor baselines were benchmarked with optimal software configurations, including Multi-Token Prediction (MTP). Jalapeño’s STP throughput even surpassed the MTP numbers Nvidia published for Vera Rubin in July. This implies OpenAI still has optimization headroom left to unlock.

Piercing Nvidia’s Moat from Within

OpenAI was historically one of Nvidia’s largest customers. Today, it has built a chip capable of beating Nvidia’s flagship on energy efficiency and publicly showcased the comparison metrics.

The structural implication of this shift outweighs its immediate technical achievement. Silicon development has long been guarded by high barriers of capital and time—three-year development cycles, multi-billion-dollar R&D budgets, and severe yield risks. Together, these hurdles formed the bedrock of Nvidia’s moat: customers who could not afford custom silicon had no choice but to buy Nvidia hardware. Jalapeño’s 16-month tape-out demonstrates that AI-assisted tooling can breach that timeline barrier far faster than previously believed. While Microsoft Maia, Google TPU, and Amazon Trainium pursue similar hardware independence, none have challenged Nvidia’s performance metrics so directly and publicly.

Custom ASIC development originally aimed to reduce vendor lock-in and lower infrastructure costs. Jalapeño’s benchmark data demonstrates that custom silicon is no longer just a cost-containment strategy—it can compete head-to-head on the emerging battlefield of power efficiency.

On Jalapeño engineering samples, DeepSeek R1 reached over 700 tokens/s/user at single concurrency, while Kimi-K2.5 and GPT-OSS achieved approximately 1,400 tok/s/user, with GSM8k accuracy matching Nvidia hardware. These figures represent pre-production silicon, leaving room for further gains as production stepping matures.

A Shift in the Rules of the Game

Jalapeño is not yet a mass-production product. It remains an engineering sample with debated benchmark baselines and self-reported figures. All these caveats hold true.

Yet it unambiguously signals a fundamental shift in the competitive dimension of AI compute. The metric is transitioning from “who has the most FLOPs” to “who generates more tokens per watt”—a transition dictated by physical data center limits rather than marketing narratives. Nvidia previously faced no serious challenge in this energy-efficiency domain. Jalapeño’s arrival marks the first direct challenge coming from its own top-tier customer base. The ultimate question is clear: as OpenAI scales its custom silicon into volume production and peers follow suit, will Nvidia’s pricing power begin to compress under this new paradigm?

Reference Links:

  • SemiAnalysis: OpenAI Jalapeño - Better Than Nvidia Blackwell
  • OpenAI Official Blog: Jalapeño’s first results show industry-leading performance
  • Tom’s Hardware: OpenAI’s 700W Jalapeño ASIC outpaces 1,400W Nvidia flagship
  • TechCrunch: OpenAI’s Jalapeño chip is built for fast inference at scale
  • HN discussion (item?id=49434378)