How Much Water Do AI Data Centers Consume? The Real Cost
One billion AI prompts sounds enormous. Yet at Google’s published estimate of 0.26 milliliters per median Gemini text prompt, that workload would consume roughly 260,000 liters—about 69,000 gallons.
That’s less than many people expect from a billion requests. It’s also only one slice of the problem. A physical AI facility can consume millions of gallons a year, depending on its cooling system, climate, electricity supply, and the accounting boundary used.
So the honest answer is that AI data centers may use almost no routine cooling water—or substantial amounts. The hardware alone doesn’t determine the result.
Start with the accounting
“Water use” can describe several different things.
- Withdrawal is water taken from a river, aquifer, municipal system, or reclaimed-water network.
- Discharge is water returned after use.
- Consumption is water not returned locally, usually because it evaporates.
A cooling tower may withdraw a large volume, circulate much of it repeatedly, and consume the remainder through evaporation and blowdown. A sealed liquid-cooling loop may be filled once and recirculate coolant for years, requiring only occasional maintenance.
Consumption is often the most useful number for environmental comparisons because it reflects water that has effectively left the local watershed. Withdrawals still matter, however. A facility can strain a municipal or river system through high peak demand even when its annual net consumption appears modest.
There’s also a second boundary: whether the estimate includes only water used at the facility or the water associated with generating its electricity. Those are separate burdens and should be reported separately.
The Lawrence Berkeley National Laboratory’s United States Data Center Energy Usage Report, published in 2024, estimated that all U.S. data centers—not AI facilities alone—directly consumed 66 billion liters, or 17.4 billion gallons, in 2023. Hyperscale and colocation facilities accounted for about 84% of that total.
The report estimated another 800 billion liters, or approximately 211 billion gallons, of indirect water consumption associated with electricity generation. That figure is about twelve times the direct total, but it isn’t a universal ratio. The result changes sharply with the regional power mix, generation technology, facility location, and the accounting method used.
Selected benchmarks
| Measure | Estimate | Boundary |
|---|---|---|
| U.S. direct consumption, 2023 | 66 billion liters / 17.4 billion gallons | All U.S. data centers |
| U.S. indirect consumption, 2023 | 800 billion liters / 211 billion gallons | Water associated with electricity generation |
| U.S. hyperscale projection, 2028 | 60–124 billion liters / 16–33 billion gallons annually | Hyperscale facilities only |
| Google median Gemini text prompt | 0.26 milliliters | Google’s May 2025 full-stack estimate |
| AWS and Microsoft WUE disclosures | 0.12 and 0.27 L/kWh | AWS global operations and Microsoft’s owned fleet, as reported in 2025 |
LBNL’s 2028 projection covers U.S. hyperscale facilities and carries a wide range because AI adoption, rack density, climate, construction pace, and cooling design remain uncertain.
Why AI hardware makes cooling harder
Every watt consumed by a processor eventually becomes heat. The facility has to move that heat away from the silicon and reject it outdoors.
Conventional data centers cool servers with air. Fans move air through the racks, while chillers, cooling towers, or outside-air systems carry heat away from the room. That approach works well at moderate rack densities.
AI clusters push much more power into the same footprint. LBNL estimated that U.S. electricity use by GPU-accelerated AI servers grew from less than 2 terawatt-hours in 2017 to more than 40 TWh in 2023. High-density racks create hot spots and leave less room for airflow mistakes or future hardware upgrades.
Direct-to-chip liquid cooling addresses the transport problem. Cold plates sit on CPUs and GPUs, carrying heat into a liquid loop. Liquid moves heat far more efficiently than air, allowing higher rack densities with less fan power.
NVIDIA’s GB200 NVL72 is a prominent example. Its compute trays and manifolds use liquid cooling, while some networking and storage components remain air-cooled.
That does not mean the system consumes water continuously. A sealed loop can circulate water or another coolant indefinitely. The crucial decision comes after the coolant absorbs heat.
A facility may reject that heat through:
- Dry coolers, which use air-cooled radiators and fans.
- Evaporative cooling towers, which consume water to improve heat rejection.
- Hybrid or adiabatic coolers, which use water during especially hot conditions.
- Immersion cooling, which places hardware in a specialized dielectric fluid and still requires a separate heat-rejection system.
A liquid-cooled rack connected to a dry cooler may have almost zero routine operational cooling-water consumption. The same rack connected to an evaporative tower can consume significant water. “Liquid cooling” describes the method of moving heat inside the building; it does not tell you how the facility disposes of that heat.
WUE helps, but it isn’t the whole answer
Operators commonly report Water Usage Effectiveness, or WUE:
WUE = liters of site water consumed ÷ kilowatt-hours of IT energy
AWS reported a global WUE of 0.12 liters per kilowatt-hour in sustainability reporting published in 2025. Microsoft reported 0.27 L/kWh for its owned fleet for its 2025 reporting period. These figures are useful, but they aren’t directly interchangeable: they cover different corporate footprints and operating boundaries.
WUE also doesn’t show whether the water was potable or reclaimed, how stressed the local basin is, how much water the site withdraws at peak times, or how much water is associated with its electricity. It may also exclude or separately account for sanitation, landscaping, humidification, and construction.
A low WUE in a water-rich region may present less local risk than a higher WUE in a drought-prone basin. Operators need both the efficiency number and the geography behind it.
What does one AI prompt consume?
Google’s May 2025 assessment of a median Gemini text prompt estimated:
- 0.24 watt-hours of electricity
- 0.03 grams of carbon-dioxide equivalent
- 0.26 milliliters of water
A narrower calculation that counted active accelerator consumption rather than broader infrastructure overhead produced an estimate of 0.12 milliliters.
Google’s figure is company-specific, not an industry benchmark. It reflects Google’s model, hardware, data-center operations, and accounting choices. It shouldn’t be presented as the water cost of an arbitrary AI prompt.
The number can change with input and output length, model architecture, hardware generation, utilization, batching, cooling design, climate, and workload type. Image and video generation generally require more computation than a short text response. Training, long-context reasoning, tool calls, and agentic systems create different profiles again.
That makes “water per prompt” a convenient public metric but a poor planning number. For infrastructure decisions, measure the workload on the actual model and hardware stack, then report the boundary used.
The trade-offs operators have to manage
Dry heat rejection is attractive in water-constrained regions because it minimizes evaporation. It can require more electricity, fans, or equipment during hot weather, though. If that extra electricity comes from a water-intensive generation system, some of the burden shifts to the grid rather than disappearing.
Evaporative cooling often delivers better thermal efficiency, especially in dry climates. It can reduce electricity use and operating cost, but it consumes water during the same hot periods when drought and municipal demand may be most severe.
Climate changes the balance. A cool site may use outside air for much of the year. A hot or humid site may need mechanical cooling and water-assisted heat rejection for longer. Two identical AI clusters can therefore have very different annual water profiles.
Providers are responding with a mix of designs. AWS has reported contracts for reclaimed-water use at 130 data centers across six countries, alongside air-cooled and selective evaporative systems. Microsoft reported that about 90% of its owned fleet used low- or zero-water cooling systems in 2025. Its newer AI-optimized design uses direct-to-chip cooling in a closed loop and is designed for zero routine cooling-water consumption during normal operation.
That last qualification matters. Such a facility still needs an initial coolant fill, maintenance and sanitation water, and potentially water during emergency or peak-condition operations. Zero cooling-water consumption is not the same as zero water use by the building.
Efficiency can also trigger a rebound effect. If cheaper inference leads to dramatically more AI usage, water per task may fall while total demand rises. Facility planning has to model both.
A practical checklist
For operators, the most useful questions are straightforward:
- Report withdrawal and consumption separately. Include annual totals, peak demand, water source, discharge, and drought restrictions.
- Match cooling to the site. Use direct-to-chip cooling for dense AI racks, then select dry, hybrid, or evaporative heat rejection based on water availability, climate, and power constraints.
- Use reclaimed water where it is reliable. Treatment requirements, quality, and long-term supply matter as much as the headline volume.
- Calculate indirect water use locally. Use the actual electricity mix and disclose the generation and accounting assumptions.
- Measure the entire workload. Include training, inference, networking, storage, cooling overhead, and idle capacity.
- Look beyond the individual building. Several moderate facilities in one basin can create a major cumulative demand.
Cloud customers should ask whether a provider’s number represents withdrawal or consumption, whether it includes power-related water, which model and hardware were measured, and whether the result covers training, inference, or both.
The best design depends on the constraint. Closed-loop liquid cooling with dry heat rejection is a strong choice where water scarcity dominates. Hybrid or evaporative systems may make sense where electricity, land, and thermal efficiency are more pressing and the water supply is resilient or reclaimed. Either way, judge the complete energy-and-water system—not just the pipe connected to the data center.
Frequently Asked Questions
How much water do AI data centers use?
There is no universal figure: LBNL estimated direct consumption by all U.S. data centers at 66 billion liters in 2023, while Google estimated 0.26 milliliters for a median Gemini text prompt.
Does liquid cooling use more water than air cooling?
Not necessarily; a sealed liquid loop can use little replacement water, while total consumption depends mainly on the facility’s heat-rejection system.
What is the hidden water footprint of an AI data center?
It is the water associated with producing electricity, which LBNL estimated at 800 billion liters for U.S. data centers in 2023.
Can an AI data center operate with zero cooling-water consumption?
Yes, closed-loop direct-to-chip systems with dry heat rejection can reach zero routine operational cooling-water consumption, although the building still needs water for filling, maintenance, sanitation, and possible exceptional conditions.
Share this research breakdown
Help friends and peers stay ahead with autonomous AI insights.
This technical article was compiled using autonomous research pipelines and third-party foundation models (including OpenAI and web-retrieval systems) to analyze papers, documentation, and market data. Content is structured by EveeStatistic for informational exploration. Readers should independently verify critical benchmarks.