Photo by JESHOOTS.COM on Unsplash. Source: https://unsplash.com/photos/hand-installing-a-computer-processor-onto-a-motherboard-DQBHliTwIaU (Unsplash License).

More than 2,000 GPU servers are serving their monitoring metrics to anyone with a browser. One of the services they expose carries a high-severity flaw that lets an unauthenticated attacker crash it. Datacenter security firm Lava published the finding on 8 October, alongside a scan of the whole internet. The Register reported the scope the same day.

NVIDIA DCGM Exporter reads telemetry straight from the GPUs on a host, covering temperature, utilization, memory, power draw, and error events. It publishes those metrics over plain HTTP, usually on port 9400, so Prometheus can scrape them. Run one on every GPU node and you have a detailed picture of the hardware.

The exposure is the reconnaissance

Lava scanned the internet four times between March and May 2026. It found roughly 2,100 hosts serving NVIDIA DCGM Exporter metrics with no authentication. Those hosts reported more than 12,000 unique GPU IDs, an estimated $100 million of hardware. Half were datacenter accelerators such as H100s and H200s. The rest were consumer cards, mostly RTX 4090s and 5090s. The scans also caught 312 B200s and 32 B300s.

An open metrics endpoint hands an attacker the blueprint. It shows what hardware you run, how hard you run it, and when. Repeated reads reveal workload schedules. Where Kubernetes labels are exposed, pod and namespace names can point to the project using a specific GPU. Across the scans, 44 percent of the exposed GPUs sat in the United States, 17 percent in Romania, and 16 percent in China.

One endpoint can be turned against itself

The flaw is CVE-2026-47483, rated 8.2 high. It lives in the Go profiling endpoints that some exporters serve next to their metrics. Concurrent unauthenticated requests to those endpoints exhaust memory until the exporter crashes. Lava reproduced the crash with NVIDIA’s own DCGM Exporter container, not a misconfigured fork. NVIDIA assigned the CVE and published an advisory after Lava reported it, and the record sits on the NVD.

When the exporter dies, operators lose their view of GPU health. The same memory and CPU pressure can slow training or inference running on the same host, especially without strict resource limits. A monitoring tool has become a way to hurt the workload it watches.

The fix is network hygiene, not just a patch

Lava’s guidance is direct. Bind DCGM Exporter to a loopback or private interface and let only your monitoring stack reach it through firewall rules or security groups. Apply the same boundary to Prometheus, whose query API and target pages often stay open too. Lava points operators to DCGM Exporter 4.8.2 or later, where the profiling endpoint is opt-in rather than on by default.

Every exposed exporter is a decision someone made to skip a firewall rule. Three questions for your fleet. Can anything outside your network reach port 9400 on a GPU node? Does your exporter build postdate the fix, and is profiling switched off? And if your GPU telemetry went dark for an hour, would you notice before a customer did?

Related reading. This is the same class of mistake we covered in the Vercel sandbox escape, where an exposed control surface did more than leak data. For how agent workloads get isolated, see why the sandbox is moving into the operating system.

By Ivan Tarin

Ivan Tarin is a Principal Product Marketing Manager at SUSE, where he owns go-to-market strategy and positioning for a seven-product cloud-native portfolio spanning Kubernetes, virtualization, storage, security, and observability. A former full-stack developer who shipped production code for enterprise and public-sector clients including U.S. national laboratories, Ivan translates complex infrastructure and AI technology into messaging that lands with developers, platform teams, and enterprise buyers. He has presented at KubeCon, SUSECON, and AWS Developer Week, and is currently pursuing an MS in Artificial Intelligence at the University of Colorado Boulder.

One thought on “2,100 GPU Servers Expose Their Metrics, and One Bug Can Switch Them Off”

Leave a Reply

Your email address will not be published. Required fields are marked *

Get the next one before it is old news

Independent analysis of cloud-native infrastructure, Kubernetes and data center economics. No vendor spin.