Skip to content
Saptarshi Joshi
← Research
ongoing2025 – present

Dynamic system modeling of liquid-cooled AI data centers

How a liquid-cooled AI data center behaves in time: under changing workloads, changing weather, and at increasing scale, with single-phase and two-phase direct-to-chip cooling.

Motivation

Data centers consumed about 415 TWh of electricity in 2024, and the International Energy Agency projects that figure to more than double by 2030. Rack power has grown from under 10 kW to well over 100 kW with AI hardware, and megawatt racks are on the roadmap. At these densities the industry is moving from air cooling to direct-to-chip liquid cooling.

Most analysis of these cooling systems is done at steady state. In operation, though, the load is not steady: AI training and inference workloads change from second to second and from hour to hour, the outdoor air that finally receives the heat changes with the weather, and grid operators increasingly ask data centers to adjust their load. Understanding how the cooling system responds in time matters for sizing, for control, for energy use and for reliability.

What we are building

We have developed a dynamic model of a direct-to-chip liquid-cooled AI rack that follows the heat from the chip to the outdoor air: the chip packages and cold plates, the coolant distribution units, the facility loop, the heat rejection equipment and the controls. Around it we are developing a framework for analyzing transient performance.

What we are studying

  • How time-varying workloads and weather together determine chip temperature and cooling energy.
  • How far chiller-free, water-free heat rejection with dry coolers can be pushed in hot climates, and what it means for usable rack capacity.
  • Energy efficiency, with a total usage effectiveness (TUE) near 1.01 as the target and the question of how much further it can go.
  • When a steady-state analysis is sufficient and when the dynamics have to be modeled.
  • The effect of thermal cycling on chip reliability.
  • Operation under demand response, and sizing a rack and its cooling for a given climate.

Where it is going

  • Scaling the rack model to 1 MW racks.
  • Scaling from the rack to rows, halls and whole sites, toward hyperscale and exascale facilities, using distributed and reduced-order modeling.
  • Comparing heat rejection and coolant distribution architectures.
  • Extending the model to pumped two-phase direct-to-chip cooling.

Status

Ongoing. A first-author manuscript is in preparation. The work has been presented at ACRC meetings, most recently in Fall 2026.