AI 101 — How Much Power Does AI Actually Use?

Share
AI 101 — How Much Power Does AI Actually Use?

AI's power use is the electricity burned by the data centers that train and answer with models — about 1.5% of the world's electricity in 2024, and on track to roughly double by 2030. That sounds small, and per query it is genuinely small. The reason it keeps making front pages is where the demand lands: a handful of regional grids, the bills of the people living on them, and the cost of running the machines in the first place. This is the ten-minute version.

Why it matters right now. Three things landed in the space of a week. The U.S. House passed the Ratepayer Protection Act 417 to 3, requiring state regulators to consider whether big electricity users — data centers among them — should carry the cost of the infrastructure built to serve them — the vote and what the bill actually does. Google, Nvidia and Emerald AI launched an alliance of 18 partners, Anthropic and utilities included, to let operators throttle AI workloads when the grid strains. And a forecast reported this week projects U.S. data centers burning more natural gas by 2035 than Germany and Japan combined. None of that is a model story. It is an electricity story, and the industry has started treating power, not chips, as the binding constraint on how fast AI can grow.

Contemporary computer with black screen placed on stand near row of server steel racks in data center

The numbers, from the people who count them. The International Energy Agency puts global data center consumption at roughly 415 terawatt-hours in 2024 — about 1.5% of world electricity, growing around 12% a year, more than four times faster than total electricity demand. Its base case has that doubling to about 945 TWh by 2030, a bit more than Japan consumes today, and reaching roughly 1,200 TWh by 2035. The United States is the biggest single customer at about 45% of the global total, and nearly half of American data center capacity sits in just five regional clusters — which is why a national average of 1.5% can feel like a crisis in Virginia.

Lawrence Berkeley National Laboratory's 2024 report, the reference every utility planner now cites, found U.S. data centers used 176 TWh in 2023 — 4.4% of national electricity, up from 58 TWh in 2014 — and projected 325 to 580 TWh by 2028, or 6.7% to 12% of the grid. PJM, the grid operator for the AI-heavy Mid-Atlantic, forecasts 32 gigawatts of new peak load between 2024 and 2030, with data centers driving 94% of it.

The mental model, in three layers. First, training: a one-time burn, measured in gigawatts for weeks, that produces a finished model. Second, inference: the permanent drip, one small burn per query, multiplied by every user, every agent loop and every hour of the day — this is the part that grows with usage rather than with model releases. Third, facility overhead: everything that is not the chip, chiefly cooling. The industry grades that with PUE (power usage effectiveness), the ratio of total building power to power delivered to the computers. A PUE of 1.0 is perfect; typical air-cooled halls run 1.4 to 1.6, meaning up to a third or more of the electricity goes to carrying heat away rather than computing.

The hardware makes that overhead unavoidable at the top end. A single H100-class GPU carries a thermal design power around 700 watts; an eight-GPU server lands in the 6 to 8 kilowatt range; a full NVL72 rack — 72 Blackwell GPUs and 36 Grace CPUs acting as one machine — draws about 120 kilowatts nominal and closer to 130 kilowatts in deployed systems, with 115 kW of that liquid-cooled and 17 kW air-cooled in HPE's own specifications. Air cooling tops out somewhere near 30 kW per rack, so the dense AI rack cannot be air-cooled at all. Water carries heat roughly 25 times better than air, and direct-to-chip cold plates push PUE down toward 1.02 to 1.03. The power bill of a modern AI campus is set as much by thermodynamics as by chip design.

The kitchen analogy. Think of one AI answer as boiling water — about 2 watt-hours for ordinary language generation in the IEA's test conditions, the rough equivalent of leaving a 10-watt LED bulb on for twelve minutes. A large reasoning model like DeepSeek-R1 burns at least twice that, and generating a short video costs around 25 times as much, closer to boiling a full kettle. Those are the small numbers, and they explain why a single query is not the problem. The problem is the kitchen. A 120-kilowatt rack is on the order of a hundred typical American homes' worth of draw, packed into one cabinet, and a hyperscale site stacks thousands of them. That is a restaurant that never closes, never turns down the burners, and has to spend nearly as much again just moving the heat out of the building.

What people get wrong. The most common mistake is treating any single number for "one AI query" as a measurement. Estimates vary by more than an order of magnitude: OpenAI's Sam Altman has cited 0.34 watt-hours, the IEA's test conditions land around 2 Wh for generation and 4 Wh-plus for reasoning models, and researchers studying the hardest queries report figures above 20 Wh. Nobody meters this per request, and no major lab publishes per-query energy. Treat all of it as an order of magnitude, not a receipt.

The second mistake is believing training is the story. Training is a capital expense paid once per model; inference is the operating expense paid every day, and it scales with users and agent loops rather than with release dates. The third is assuming efficiency will resolve the tension. Energy per token is improving quickly, but demand has outrun every efficiency gain so far — the same pattern that shows up whenever a resource becomes cheaper to use.

The last mistake runs the other way: concluding AI is eating the grid. At roughly 1.5% of world electricity, it is not — and by 2030 the IEA's base case still has it under 3%. What is real is the concentration. Slow growth spread over a country is manageable; a 500-megawatt campus dropped into one county's interconnection queue is a permitting, transmission and rate case, and those take years the industry does not have. That is the whole argument behind the flexibility alliance, the House vote and the local buildout fights: not whether AI uses a lot of power globally, but who pays, and how quickly a grid built for someone else's century can be made to serve it.

Where to learn more. The IEA's Energy and AI report is the clearest free overview of the global numbers, and its executive summary is genuinely readable. For the United States specifically, LBNL's 2024 data center energy report is the primary document, and PJM's published load forecast shows what the growth looks like from inside a grid operator. For the per-query debate, IEEE Spectrum's walkthrough of competing estimates is the honest one.

Related reading: What is HBM (high-bandwidth memory)? covers the memory that has to be fed and cooled next to each GPU, What is an NPU (neural processing unit)? explains the opposite end of the spectrum — models run on milliwatts in your pocket, and What is a context window? is why longer conversations cost more compute per token.

If the grid, not the chip, decides how fast AI advances, should data centers pay the full cost of the power they demand — or should the public absorb part of it to keep the buildout moving? Tell us in the comments.

Sources: IEA — Energy and AI · LBNL — 2024 United States Data Center Energy Usage Report (PDF) · PJM — 2025 Long-Term Load Forecast Report (PDF) · IEEE Spectrum — AI Energy Use: The Hidden Cost of ChatGPT Queries · HPE — NVIDIA GB200 NVL72 QuickSpecs · AI Midday — The House votes 417-3 to make data centers pay for their own power