Bargo
NVDA

OpenAI Jalapeno Beats Blackwell and Hunts Rubin on Inference Efficiency

First measured results show about 2x Blackwell on perf per watt and 53 to 104x tokens per MW, verified in lab but still narrow and pre production

Bargo · Aug 25, 2026

OpenAI's first custom inference chip, Jalapeno, beats Nvidia Blackwell on power efficiency and, in early lab tests, even edges Vera Rubin on tokens per megawatt. The results were checked in OpenAI's lab by SemiAnalysis but remain narrow, pre production, and not yet proven at gigawatt scale. For Nvidia, the moat is dented where power is the bottleneck, not broken.

What OpenAI claimed on Aug 25

OpenAI said Jalapeno delivered more peak throughput per kilowatt and lower token latency than commercial systems on InferenceX using GPT OSS 120B, with strong results also on DeepSeek R1 and Kimi K2.5. The company framed it as Jevons paradox, greater efficiency expanding consumption, and said AI helped compress design to tapeout in nine months and optimize arithmetic circuits. Engineering samples are running at target clocks, with deployment inside OpenAI compute by late 2026, Gen2 deep in development and Gen3 in planning. The chip was built with Broadcom and Celestica as a reticle sized ASIC with HBM, purpose built for inference, not a repurposed training chip Tom's Hardware.

Published Jalapeno specs: 13.4 PFLOPS FP4, 700W TDP, 15.4 TB/s memory bandwidth. OpenAI's chart at matched interactivity showed:

Model Jalapeno tokens/s/kW Existing best tokens/s/kW Multiplier
GPT OSS 120B 22,935 427 53.7x
DeepSeek R1 12,258 118 103.9x
Kimi K2.5 6,744 120 56.2x

OpenAI also summarized the gain as 1.5 to 1.9 times more AI work per watt and 1.7 to 3.6 times lower latency in other views AA.

What Nvidia publishes for Blackwell and Rubin

Nvidia's published Blackwell specs are 9 PFLOPS FP4, 192GB HBM3e, 8 TB/s at up to 1,000W for B200, and 15 PFLOPS FP4, 288GB HBM3e, 8 TB/s at 1,400W for B300 (Blackwell Ultra). Vera Rubin is 50 PFLOPS NVFP4, 288GB HBM4, 22 TB/s per GPU, with Vera Rubin NVL72 at 3,600 PFLOPS NVFP4 and 20.7TB HBM4 Nvidia Vera Rubin NVL72.

At system level, Nvidia says Vera Rubin NVL72 delivers up to 10x more tokens per megawatt and one tenth the cost per million tokens versus GB200 NVL72 at similar interactivity, with all optimizations on including NVFP4, multi token prediction and disaggregated prefill/decode Nvidia Vera Rubin NVL72. CoreWeave's first silicon bring up of Vera Rubin NVL72 measured the same 10x tokens per second per megawatt versus GB200 NVL72 on DeepSeek R1 CoreWeave. Nvidia's Vera Rubin POD overview details the rack scale co design behind those gains Nvidia Developer Blog.

The math check: does Jalapeno beat Blackwell and Rubin?

On raw silicon efficiency, the claim versus Blackwell verifies cleanly. On tokens per megawatt versus Rubin, the early lab result is plausible but not yet apples to apples.

FP4 compute per watt (TFLOPS/W) — published specs, Aug 25, 2026
Chip FP4 per watt Bandwidth per watt
Jalapeno 13.4 PFLOPS / 700W 19.1 TFLOPS/W 22.0 GB/s/W
B200 9 PFLOPS / 1,000W 9.0 TFLOPS/W 8.0 GB/s/W
B300 15 PFLOPS / 1,400W 10.7 TFLOPS/W 5.7 GB/s/W

Jalapeno is about 2.1x B200 and 1.8x B300 on compute per watt, and 2.8 to 3.9x on bandwidth per watt. That matches the "substantially better than current state of the art" characterization.

Versus Rubin, raw FLOPS per watt is not the right comparison because Rubin TDP is not disclosed and system level tokens per joule is what data centers monetize. If Rubin is 10x GB200 on tokens per MW, Jalapeno's 53 to 104x versus GB200 class implies about 5 to 10x versus Rubin at the tested 8k input, 1k output point. SemiAnalysis, which was invited to OpenAI's lab to run InferenceX, found Jalapeno's single token prediction throughput per MW surpassing Vera Rubin's multi token prediction results from July and far exceeding GB200 SemiAnalysis.

What SemiAnalysis verified, and what it flagged

SemiAnalysis confirms Jalapeno is real, reticle sized, HBM4 based, and industry leading on the tests they could run, beating every Nvidia, AMD and Google chip they have tested on multiple open models. They also stress the limits: results were provided by OpenAI and verified in lab, not a full InferenceX suite, no AgentX long context multi turn test which they consider the better real world proxy, models were not the largest frontier, and both Jalapeno and Rubin are immature and will improve. Jalapeno is still engineering samples while Rubin is shipping to customers now SemiAnalysis.

On TCO, SemiAnalysis puts Vera Rubin and Jalapeno head to head on tokens per dollar, but notes Jalapeno achieves its lead without speculative decoding while Rubin's number includes it. When Jalapeno adds speculative decoding, its TCO lead would widen. Part of the edge also comes from trading Nvidia's high margins for Broadcom's lower custom silicon margins SemiAnalysis.

This matters because the broader buildout is already power constrained. Recent GPU rental trends show B200 holding firm while H100 softens, and the 10GW contracting wave plus the supply side capex surge frame why throughput per watt is now revenue per watt.

What it means for Nvidia inference moat

Nvidia's moat is not broken, but it is dented where it hurts most. Inference at scale is power limited, and a first gen inference only ASIC with extreme hardware software co design beating a general purpose GPU on perf per watt and latency simultaneously is a meaningful proof point. It validates the path Google, Amazon and Meta are also on with TPUs, Trainium and MTIA.

Nvidia's remaining moat is system level, not just FLOPS. NVLink 6, Spectrum X and Quantum X800 networking, Dynamo and TensorRT LLM software, and the ability to run both training and inference on one platform still differentiate. Rubin is also 10x Blackwell and Jalapeno does not train. As the shortage spreads sideways to CPUs, HBM and fab tools, system integration and supply chain scale matter more than a single chip's efficiency.

Pricing pressure is the nearer risk. If hyperscalers can serve inference at a fraction of the power of Blackwell class systems, merchant inference pricing must fall. That is good for token consumers and for GPU price dynamics, but it squeezes GPU rental spreads. Nvidia captures part of the value back via higher Rubin ASPs and full NVL72 system sales.

What it means for OpenAI path to independence

Jalapeno gives OpenAI inference independence, not full independence. OpenAI stresses it is a generalized inference chip, it runs GPT OSS, DeepSeek R1, Kimi K2.5 and even Doom via Codex prompts, which increases odds Microsoft adopts it via its IP share. But OpenAI still needs Nvidia for training and for general purpose flexibility, and it still depends on Broadcom and TSMC for design and manufacturing. The reported 10GW Vera Rubin commitment to Nvidia for training underscores that diversification, not decoupling, is the strategy.

The leverage is real. Owning the inference stack gives OpenAI control over power, cost and roadmap, with Gen2 and Gen3 already in flight, and negotiating power on Nvidia pricing and allocation. It also lets OpenAI monetize efficiency via Jevons paradox, where lower cost per token expands total consumption.

Execution risk remains. Jalapeno must pass production qualification, mature its software stack, and prove AgentX scale performance at gigawatt scale with Microsoft and other partners starting late 2026. First gen chips rarely stay ahead without continuous software and networking co optimization.

What to watch

Sources


More research at bargo.ai/research.

Get Bargo research in your inbox
One email when we publish. No spam, unsubscribe anytime.
Get Bargo research in your inbox
One email when we publish. No spam.