Photo by Kirill Sh on Unsplash. Source: https://unsplash.com/photos/fiber-optic-cables-in-network-switch-eVWWr6nmDf8 (Unsplash License).

NVIDIA open sourced a toolkit that fixes the part of GPU scheduling nobody wanted to own. Here is what changed.

  • Topograph discovers a cluster’s network topology from cloud APIs or on-premises fabric systems and normalizes it into one model.
  • It republishes that model in the format each workload manager already reads, which means Kubernetes node labels, Node Feature Discovery resources, Slinky ConfigMaps, Slurm topology configuration or topology JSON.
  • Providers exist for Google Cloud, Lambda, Nebius, Nscale, Oracle Cloud Infrastructure and Crusoe. On premises it reads InfiniBand, Spectrum-X and Multi-Node NVLink domains.
  • The minimum Kubernetes version is 1.27, the typical aggregation delay is 15 seconds, and KAI Scheduler or Kueue handle topology-aware gang scheduling.
  • NVIDIA frames the waste as provisioned power sitting on GPUs that are waiting for data.

Read the full analysis. Why the scheduler was never the hard part.

By Ivan Tarin

Ivan Tarin is a Principal Product Marketing Manager at SUSE, where he owns go-to-market strategy and positioning for a seven-product cloud-native portfolio spanning Kubernetes, virtualization, storage, security, and observability. A former full-stack developer who shipped production code for enterprise and public-sector clients including U.S. national laboratories, Ivan translates complex infrastructure and AI technology into messaging that lands with developers, platform teams, and enterprise buyers. He has presented at KubeCon, SUSECON, and AWS Developer Week, and is currently pursuing an MS in Artificial Intelligence at the University of Colorado Boulder.

Leave a Reply

Your email address will not be published. Required fields are marked *

Get the next one before it is old news

Independent analysis of cloud-native infrastructure, Kubernetes and data center economics. No vendor spin.