Executive Summary
NVIDIA published Topograph on 22 September, an open source toolkit that reads the physical network inside a GPU cluster and republishes it in whatever shape the workload manager reads. Kubernetes gets node labels. Slurm gets generated topology configuration. KAI Scheduler and Kueue get what they need for gang placement. The finding is smaller than the announcement. The blocker was never scheduling logic. It was that no scheduler held a current map of the hardware underneath it.
The pattern is map drift. Clusters change shape constantly. Nodes join, switches relink, cloud instances land in different locality domains, and the topology view a scheduler was initialized with stops matching reality. Every placement decision after that is made blind. Topograph regenerates its view on request and on watched cluster changes, with a typical 15 second aggregation delay, and it ships providers for Google Cloud, Lambda, Nebius, Nscale, Oracle Cloud Infrastructure and Crusoe, plus on premise InfiniBand, Spectrum-X and Multi-Node NVLink.
Topology-aware GPU scheduling sounds like a scheduler feature. It is not. The default Kubernetes scheduler has no idea how the hardware is wired, and it never has. NVIDIA’s answer is to stop asking the scheduler to discover the fabric and to hand it a published map instead.
The bandwidth math explains the urgency. A fifth generation NVLink fabric moves 1.8 terabytes per second per GPU on Blackwell parts, and the sixth generation on Vera Rubin doubles that to 3.6. An NVIDIA Quantum InfiniBand port runs at 800 gigabits per second. Spread a tightly coupled job across locality domains and those numbers stop being yours. Traffic crosses shared links, latency climbs, and GPUs burn provisioned power waiting on data.
Placement problems compound at scale and surface as network congestion. That is an infrastructure sentence, not a marketing one.
The scheduler was never the hard part
Slurm and Kubernetes both already support topology-aware allocation, and both have for years. The catch hides in a phrase that is easy to skim past. A scheduler can only act on the topology it observes.
Topograph splits that observation problem into two concepts. Providers discover fabric data from cloud APIs or on premise systems and normalize it into a canonical model. Engines translate the model into the format a specific manager expects, which today means Kubernetes node labels, Node Feature Discovery resources, Slinky ConfigMaps, Slurm topology configuration, or instance-oriented JSON.

Five components keep the view current. An API server validates requests and dispatches discovery. A node observer watches Kubernetes node and pod changes. A node data broker stores per-node attributes as annotations. The provider converts raw fabric data, and the engine writes out the result. The API exposes five endpoints, including a generate call that returns an accepted status while it works, a lookup that serves the cached result, and a metrics endpoint for Prometheus.
The view refreshes on watched cluster changes, not on a cron job a human maintains. The aggregation delay is deliberate, 15 seconds typical.
Five clouds and three on premise fabrics, one model
The provider list decides who can use it on day one. Working integrations cover Google Cloud, Lambda, Nebius, Nscale, Oracle Cloud Infrastructure and Crusoe. On premise, Topograph reads InfiniBand through the standard discovery tooling, or NetQ for Spectrum-X and Multi-Node NVLink domains.
KAI Scheduler and Kueue are optional but they are where the interesting placement happens, because gang scheduling is the workload type that suffers most when a job straddles two domains. Topograph itself ships as a Helm chart and needs Kubernetes 1.27 or later plus a supported provider. That means the integration work lands on the platform team, not the fabric vendor.
NVIDIA also owns the other half of this problem. It bought SchedMD, the company behind Slurm, in December 2025, and ships a Slinky engine that runs Slurm on Kubernetes and writes topology to a ConfigMap. If one map can drive Kubernetes labels and Slurm configuration, the case for running two placement regimes gets weaker each release.
What this changes for the people paying the power bill
AI factories are power-limited systems. NVIDIA’s own framing is that they deliver maximum value when fully optimized, and it defines the waste in the currency operators actually feel. Poor placement leaves GPUs consuming provisioned power while waiting on data. That is the story we have watched all year as the constraint moved from chips to substations.
The near term effect is less about the vendor and more about the schedule. Nothing here requires new hardware. It requires a platform team to treat topology as configuration they version, rather than a snapshot from a migration nobody revisited.
Three questions worth asking about your own cluster. Does your scheduler hold a topology view that predates your last hardware change? When a node joins or a switch relinks, what regenerates the map, a person or a controller? And can you state tokens per watt for a distributed job, or only GPU utilization?
Related reading. Kubernetes Was Built to Schedule Pods. GPUs Broke the Model. The GPU scheduling bottleneck. Kubernetes 1.37 Is Rebuilding the Scheduler for Accelerators here. Round-Robin Routing Wastes GPUs. AWS Just Made the Open Fix an EKS Add-On here.
Topograph is on GitHub, and the announcement has the deployment detail. Bandwidth figures come from NVIDIA NVLink.
Get the next one before it is old news
Independent analysis of cloud-native infrastructure, Kubernetes and data center economics. No vendor spin.

[…] Read the full analysis. Why the scheduler was never the hard part. […]
[…] reading. NVIDIA Topograph targets the same fleet problem one layer up, by mapping GPU topology so schedulers stop placing work across the wrong links. Enterprise AI […]
[…] reading. We broke down why the topology of a GPU cluster decides what the scheduler can do, and why token prices keep moving once caching enters the picture. The full detail is in […]