Electrician testing an electrical panel with a multimeter. Photo by Toolmash Expo on Unsplash. Source: https://unsplash.com/photos/electrician-testing-electrical-panel-with-multimeter-PkHf7BUWbtk (Unsplash License).

Executive Summary

Two moves in one week point the same way. NVIDIA opened a program that qualifies batteries and cooling gear for AI factories, and figures published for its largest single site showed a campus heading past a million accelerators. The bottleneck has moved off the GPU and onto the power connection.

The scale is the story. A Memphis campus runs 720 Tesla Megapacks, described by an xAI developer as 3.3 GWh of storage and by the local utility chief as roughly 2,000 MW behind the meter. NVIDIA qualified three battery vendors on September 21. Capacity is no longer gated by chips alone. It is gated by whether a site can hold a steady load while a training job swings hundreds of megawatts at once.

Diagram of the AI factory power path. Grid interconnection feeds a battery buffer, which feeds AI factory racks.
Power delivery now gates AI capacity the way chip supply did in 2024.

Every AI capacity plan used to end with the same question. How many GPUs can we get. Chip supply is no longer the hard part. The question that sets the schedule now is how much steady power the site can hold, and how fast that load can move.

The constraint moved from chips to the grid

GPU allocation eased through 2026. Grid interconnection did not. Transmission capacity, generation queues and substation lead times now sit on the critical path for every large AI campus, and they run on their own clock.

NVIDIA treats this as a control problem rather than a megawatt problem. A cluster can change its draw faster than a thermal plant can respond, so the models of supply that held for a decade of steady data centers no longer fit. Batteries become the shock absorber, and they move the constraint into procurement.

The Colossus 2 site near Memphis shows the shape. Elon Musk said on September 25 that the campus holds 110,000 GB200 and 440,000 GB300 accelerators, with another 220,000 GB300 due within a week, another 220,000 in November, and a further 220,000 by late December if the schedule holds. That would push the site past a million accelerators on its own.

Independent tracking is more conservative. Epoch AI put Colossus 2 at roughly 110,000 GB200 and 330,000 GB300 as of September 23, supported by about 946 MW of IT power. Treat Musk’s number as a claim and the tracker as a floor.

NVIDIA turned power and cooling into a qualification list

NVIDIA launched DSX Ready on September 21. The program qualifies partner products against its DSX AI factory reference design. The first two categories are battery energy storage systems and coolant distribution units.

The qualified batteries came from Hitachi Energy, LG Energy Solution and Tesla. The qualified cooling units came from LG Electronics, LiquidStack and Vertiv, as Energy-Storage.news reported. LG Energy Solution offers 2.5 MW and 5.1 MWh grid-forming blocks. Vertiv’s CoolChip unit is rated at 2.3 MW.

The list matters because it is a menu. A builder can specify parts that already passed NVIDIA’s review instead of proving each one on site. NVIDIA is careful to say the qualification does not replace site engineering or guarantee site stability. That caveat is the honest part. A qualified battery does not fix a weak grid.

The battery is what closes the swing

Memphis is the reference deployment. Canary Media counted 720 Tesla Megapack containers in satellite imagery from July 11, worth about 2.8 GWh on their own. An xAI developer told the Tennessee Valley Authority board on August 20 that the site held 3.3 GWh. The local utility chief described about 2,000 MW behind the meter on September 2.

The reason is the swing. A synchronized training run can drop or spike hundreds of megawatts in milliseconds. Batteries absorb that and let the site draw a smoother baseline from the utility. For scale, the largest known grid battery before this sat near 3,287 MWh at a California solar site.

So the procurement order has flipped. Batteries and cooling units now gate deployment as much as GPUs do, and they arrive on longer lead times. Three questions decide the next capacity plan. Where does the site sit in the interconnection queue. How much swing can the local grid absorb on its own. And who owns the batteries that cover the rest.

Related reading. We covered Alibaba’s twenty gigawatt plan and the Texas permit freeze that stalled a 445.8 GW queue. Both are the same forcing function from different angles.

By Ivan Tarin

Ivan Tarin is a Principal Product Marketing Manager at SUSE, where he owns go-to-market strategy and positioning for a seven-product cloud-native portfolio spanning Kubernetes, virtualization, storage, security, and observability. A former full-stack developer who shipped production code for enterprise and public-sector clients including U.S. national laboratories, Ivan translates complex infrastructure and AI technology into messaging that lands with developers, platform teams, and enterprise buyers. He has presented at KubeCon, SUSECON, and AWS Developer Week, and is currently pursuing an MS in Artificial Intelligence at the University of Colorado Boulder.

2 thoughts on “AI Factories Now Buy the Battery Before the Chip”

Leave a Reply

Your email address will not be published. Required fields are marked *

Get the next one before it is old news

Independent analysis of cloud-native infrastructure, Kubernetes and data center economics. No vendor spin.