Executive Summary
The AI hardware fight has moved off the chip and onto the rack. Whoever owns the interconnect keeps earning whether or not the accelerator inside sells. That reframes every deal a custom silicon shop signs.
d-Matrix joined Nvidia’s NVLink Fusion partner list and will build its Raptor inference chips into Nvidia’s racks, switches, and supply chain. d-Matrix expects systems with up to 144 accelerators on a single all-to-all NVLink fabric by the end of next year. Nvidia earns even when it sells no GPU, because it supplies the Vera CPUs, NVSwitch appliances, BlueField and ConnectX networking, and Spectrum-X Ethernet underneath. It is paying rivals to join. Nvidia invested $3.5 billion in MediaTek and $2 billion in Marvell. Watch the lock-in behind the openness. Standardizing on one rack simplifies deployment and makes leaving the orbit harder.
What d-Matrix just signed up for
AI inference startup d-Matrix joined the growing list of chipmakers licensing Nvidia‘s NVLink Fusion interconnect and its MGX rack designs. The deal puts d-Matrix’s upcoming Raptor inference chips into the same racks, the same switches, and the same supply chain that Nvidia’s own GPUs use.
The specifics are striking. d-Matrix expects to offer systems with up to 144 Raptor accelerators connected by a single all-to-all NVLink fabric by the end of next year. Each Raptor card carries 32 gigabytes of 3D-stacked memory with about 100 terabytes per second of memory bandwidth, which the company says is roughly 4.5 times the memory bandwidth of Nvidia’s upcoming Rubin GPU. d-Matrix also pairs its chips with Nvidia’s Vera CPUs, NVSwitch appliances, BlueField and ConnectX networking, and Spectrum-X Ethernet.
The strategy hiding in plain sight
Read that list again. d-Matrix designs a competing accelerator, and it just agreed to build it on top of Nvidia’s switches, network cards, CPUs, and rack standard. Nvidia stands to make money on every one of those parts even when it does not sell a single GPU.
That is the whole play. Nvidia spent years building NVLink and its rack architecture into the default way AI clusters are wired. Now it is licensing that plumbing to everyone else. d-Matrix joins Qualcomm, Arm, Marvell, Amazon, Fujitsu, MediaTek, and others on the NVLink Fusion partner list. For each of them, the math is the same. Building your own scale-up network from scratch is slow and expensive. Plugging into Nvidia’s takes the problem off your plate.
Nvidia is even paying to bring partners in. It invested $3.5 billion in MediaTek and $2 billion in Marvell as part of those deals. When you spend billions to get competitors to adopt your interconnect, you are not trying to sell interconnects. You are trying to make your platform the place where the AI industry builds, no matter whose chips are in the rack.
It is worth being precise about what is and is not being built here. d-Matrix is not shipping a general-purpose GPU to take on Nvidia across the board. It is building a specialized inference chip for latency-sensitive work like coding assistants, chat, and voice agents, where customers will pay a premium for speed. That is a narrower target, and it is one where a purpose-built chip can actually win on cost per token. Nvidia is fine with that, because a d-Matrix rack still means Nvidia switches, Nvidia networking, and Nvidia’s supply chain underneath it.
What this means for the AI hardware market
This reframes the competition. The story of the last two years was that custom silicon and rivals would eat into Nvidia’s GPU dominance. Broadcom builds custom accelerators for Google, Meta, and OpenAI. Startups like d-Matrix target inference specifically, where latency and cost per token decide winners. That pressure is real.
But Nvidia is not just defending its chips. It is defending the layer underneath them. If every accelerator, whether Nvidia or not, plugs into NVLink switches and MGX racks, then Nvidia owns the toll road even as the cars change. The competitive fight shifts from the chip to the architecture around it.
For buyers, that cuts both ways. Standardizing on one interconnect and rack design simplifies deployment and lets you mix accelerators for different workloads, which is exactly what d-Matrix is promising. It also deepens lock-in. The more your racks and networking are Nvidia parts, the harder it is to leave the orbit entirely, even if you buy someone else’s silicon.
The takeaway is that the AI infrastructure race is being decided at the rack level, not the chip level. Nvidia seems to understand that better than anyone. It is happy to share the compute as long as it keeps the roads.
Related reading. The Next Big AI Win Is Cutting Power Per Workload, Not Adding More Chips. OpenAI and Broadcom Built a Chip That Beats Blackwell. Nvidia Is Not Worried.. Kubernetes Was Built to Schedule Pods. GPUs Broke the Model.. AI Infrastructure Runs on Four Layers. Most Break Below the Model..
Get the next one before it is old news
Independent analysis of cloud-native infrastructure, Kubernetes and data centre economics. No vendor spin.

[…] Why this quietly deepens lock-in even as it adds choice. Read the full analysis. […]