NVIDIA open sourced a toolkit that fixes the part of GPU scheduling nobody wanted to own. Here is what changed.
- Topograph discovers a cluster’s network topology from cloud APIs or on-premises fabric systems and normalizes it into one model.
- It republishes that model in the format each workload manager already reads, which means Kubernetes node labels, Node Feature Discovery resources, Slinky ConfigMaps, Slurm topology configuration or topology JSON.
- Providers exist for Google Cloud, Lambda, Nebius, Nscale, Oracle Cloud Infrastructure and Crusoe. On premises it reads InfiniBand, Spectrum-X and Multi-Node NVLink domains.
- The minimum Kubernetes version is 1.27, the typical aggregation delay is 15 seconds, and KAI Scheduler or Kueue handle topology-aware gang scheduling.
- NVIDIA frames the waste as provisioned power sitting on GPUs that are waiting for data.
Read the full analysis. Why the scheduler was never the hard part.
Get the next one before it is old news
Independent analysis of cloud-native infrastructure, Kubernetes and data center economics. No vendor spin.
