The Jalapeno Moment: How Nvidia's Dominance Might End

OpenAI's first-gen Jalapeno just beat Nvidia's Blackwell and even Rubin. First-gen chips are never supposed to be competitive. This one is. That's the moment Nvidia's absolute dominance starts to crack.

Share
The Jalapeno Moment: How Nvidia's Dominance Might End

This screenshot from Dylan Patel stopped me cold:

"OpenAI Jalapeno is spicy. Usually first generation chips aren't competitive, but OpenAI is beating Nvidia Blackwell and even Rubin. This is huge news!"

SemiAnalysis got into OpenAI's lab and the numbers are absurd for a Gen 1 ASIC. This is not just another custom chip announcement. This is the Silicon Divorce going public.

1. The Economics Have Flipped

Nvidia built its moat on being the universal GPU. You could train anything, run anything, anywhere.

That was perfect for the training boom.

Inference is different. Inference is about TCO, watts, and tokens per second at scale. And there, custom wins.

On SemiAnalysis's InferenceX benchmark, Jalapeno delivered 1.5x to 1.9x more AI work per watt at peak throughput and 1.7x to 3.6x lower end-to-end latency than the best NVIDIA GB200/GB300 results that are available right now, tested on GPT-OSS 120B, DeepSeek R1 670B, and Kimi K2.5 1T. 

SemiAnalysis's own summary: "Jalapeño beats Blackwell... across almost all scenarios without being tuned for any specific point in the curve" and it "beats every Nvidia, AMD, and Google chip we have been able to test on multiple top open source models" on tokens per megawatt. 

The specs explain why: Jalapeno is a reticle-sized ASIC built with Broadcom on TSMC N3P, 13.4 PFLOPs of MXFP4 at 700W (vs Rubin's 900-1,150W), 15.4 TB/s HBM4 bandwidth, taped out in just 16 months.

OpenAI says it will deliver performance per watt substantially better than current state-of-the-art. 

Early deployment is 100MW this year, with a second-gen already approaching tape-out.

When your biggest customer can get nearly double the work per megawatt at half the latency, the "Nvidia tax" stops being tolerable.

2. The Customers Are Becoming Competitors

This is the pattern: JPMorgan projected that chips from Google, Amazon, Meta, and OpenAI will make up 45% of the AI-chip market. 

Jay Goldberg at Seaport put it bluntly: hyperscalers are building custom silicon because "they don't want to be stuck behind an Nvidia monopoly" and now "Nvidia now has to compete with its customers". 

Google has TPU v6 / Ironwood. Amazon has Trainium 3. Meta has MTIA. OpenAI now has Jalapeno. ByteDance is building custom CPUs too.

The reason is simple: by designing chips specifically for their own internal workloads, such as LLM inference, hyperscalers are achieving performance-per-watt efficiencies that general-purpose GPUs struggle to match. 

The result is what some analysts call The Great Decoupling: "The NVIDIA tax, the 70-80% margins the company once commanded, is being eroded" as hyperscalers offer their own silicon at 30-50% discount. 

Nvidia's stock is flat in 2026 (+2%) after 3 years powering the entire market rally for exactly this fear.

3. This Doesn't Mean Nvidia Collapses. It Means Nvidia Decouples.

This is the part everyone gets wrong. Nvidia won't die. It will decouple.

The old model: You buy Nvidia chips, Nvidia switches, Nvidia networking, CUDA stack. One throat to choke.

The new model Nvidia itself is pushing: NVLink Fusion. Let custom CPUs and ASICs plug into NVLink fabric. 3.6 TB/s cache-coherent chiplet link.

Keep ownership of the interconnect, even if you lose ownership of the compute.

Because the real moat was never just the chip. It was CUDA + NVLink. Jalapeno proves you can beat Blackwell on raw efficiency, but can you beat the entire platform without rebuilding 15 years of software?

That's the fight. OpenAI designed Jalapeno using Codex AI-assisted design, optimized for its own models, not for PyTorch in general.

It's vertical integration at its purest.

So Nvidia's dominance doesn't end with a bang. It ends with a fragmentation:

  • Training stays more Nvidia-centric (for now)
  • Frontier inference shifts to custom ASICs for cost
  • The interconnect becomes the new battleground

The era of "everyone just buys H100s/Blackwells" is over. The era of "every hyperscaler has its own silicon and Nvidia is the fabric" has begun.

Jalapeno is spicy indeed. And Nvidia is starting to sweat, just like in that SemiAnalysis meme.