Executive Summary
The AI chip market is splitting. General-purpose GPUs still win the flexible work, but the predictable, high-volume inference jobs are moving to silicon built for one customer and one workload. OpenAI’s Jalapeno is the clearest sign of that split.
Jalapeno came out of a joint program with Broadcom, and the benchmarks are real. The chip posts 1.5x to 1.9x more inference throughput than Nvidia’s GB200 and GB300 racks. Latency runs 1.7x to 3.6x lower, and each rack packs 128 accelerators. It also does inference only. It cannot train a model, and training is where the largest compute spend lives. OpenAI left out system-level power, so the result measures throughput per rack rather than per watt. Broadcom’s own numbers show where the volume is heading. AI semiconductor revenue hit $16.7 billion in the third quarter of 2026, up 221 percent. Custom accelerators drove 73 percent of it. Watch efficiency disclosures, not peak throughput.
Jalapeno is real and it is fast
OpenAI and Broadcom unveiled Jalapeno, OpenAI’s first custom AI chip, and the benchmarks are striking. On inference, OpenAI reports 1.5x to 1.9x more throughput than Nvidia‘s Blackwell-generation systems, with 1.7x to 3.6x lower end-to-end latency. Each rack packs 128 accelerators, 1.7 exaFLOPS of 4-bit compute, and 27.5 terabytes of HBM4.
It taped out in nine months, which OpenAI calls the fastest ASIC development cycle ever achieved in high-performance semis. Engineering samples are running models in the lab at production frequency and power. This is not a slide deck. It is a working chip, co-designed with Broadcom on silicon and networking and with Celestica on system integration.
The design goal matters too. Jalapeno is architected around OpenAI’s view of where LLM inference is going, and the flexibility to work with all models rather than just one. The roadmap goes beyond a single chip. OpenAI and Broadcom are building a multi-generation compute platform with initial deployment by the end of 2026 and expansion after that.
Why this is not the end of Nvidia
Read the fine print and the picture changes. Jalapeno does inference only. It cannot train a model. Nvidia’s training dominance is untouched, and training is where the biggest compute spend lives. OpenAI also compared its chip against GB200 and GB300 racks, which launched in 2024 and 2025, not against the upcoming Vera Rubin platform. The comparison is real but it is not against Nvidia’s newest hardware.
Even the throughput number has a caveat. OpenAI did not disclose system-level power. That means Jalapeno measures throughput per rack, not throughput per watt. Efficiency is the metric that decides what actually fills a data center, and that number is still hidden.
There is also the software question. Jalapeno does not support CUDA, the software layer that keeps developers anchored to Nvidia. Custom accelerators have to prove they can carry a workload before a team will port to them, and that switching cost is real. For most companies, the path of least resistance stays with the chip that runs the software they already have.
The economics are the real story
Broadcom’s custom-silicon business is the larger shift. AI semiconductor revenue hit $16.7 billion in the third quarter of 2026, up 221 percent year over year, and custom accelerators drove 73 percent of it. Broadcom expects AI chip revenue to top $100 billion in fiscal 2027, about double this year. Google, Meta, OpenAI, and Anthropic are all tied to Broadcom custom silicon programs.
The distinction that matters is workload type. Nvidia sells general-purpose GPUs that do anything. Broadcom designs application-specific chips around a single customer’s needs. A chip optimized for one inference or recommendation workload can beat a general-purpose chip on that workload while being cheaper and more power-efficient at it. Meta is moving its MTIA accelerator toward production. Google has worked with Broadcom across multiple generations of its TPUs. OpenAI is already preparing a second-generation chip and planning a third.
Custom silicon is taking over the predictable, high-volume work. Nvidia keeps the flexible work. That is a split, not a replacement.
What this means for the data center. For anyone running AI infrastructure, the takeaway is that the single-vendor assumption is over. The next few years will look like a mix. Nvidia GPUs for training and general compute, Broadcom-style ASICs for high-volume inference, and networking that ties both together. Broadcom’s networking position means it gets paid whether you buy custom chips or Nvidia GPUs.
This is not a replacement. It is a specialization. Nvidia still sells more AI hardware in a quarter than Broadcom expects from its chip business in a year. But the trend is toward chips built for a specific job, and specialization is exactly what a system designed for one workload is good at.
The headline is Jalapeno beating Blackwell. The story is that the AI chip market is finally wide enough for more than one winner, and that the winners will divide the workload rather than fight over all of it.
Every announced commitment in this space is tracked with its source in our AI data centre power commitments record.
Related reading. The Next Big AI Win Is Cutting Power Per Workload, Not Adding More Chips. The Cost of AI Is Finally Falling. The Cost of Using It Is Not.. Nvidia Is Letting Rivals Into Its Rack. That Is the Point.. AI Infrastructure Runs on Four Layers. Most Break Below the Model..
Get the next one before it is old news
Independent analysis of cloud-native infrastructure, Kubernetes and data centre economics. No vendor spin.

[…] Read the full analysis. […]