Photo: Unsplash
The Grid Problem Nobody in AI Is Talking About Publicly
In January 2025, Microsoft announced a $80 billion capital expenditure plan for AI data centers in fiscal year 2025. In March, Meta announced $60-65 billion for the same year. Amazon, Google, and Oracle announced comparable plans. The combined announced data center investment from major tech companies in 2024-2025 exceeded $300 billion.
Announcing investment is easy. Getting the power to run the data centers you’re investing in is not.
The power demand numbers are staggering when made concrete. A single hyperscale data center campus of one gigawatt — a scale that several companies are now planning — draws as much electricity as a medium-sized city. San Jose, California, population one million, uses approximately 1.5 gigawatts at peak. The AI infrastructure being planned in the next five years would, if fully built and powered, represent a demand increase roughly equivalent to adding several new San Joses to the US electrical grid. The grid was not designed for this. The utilities running the grid have not planned for this. The regulatory processes that govern utility investment were not designed to accommodate this pace of change.
Here’s the physical constraint that none of the press releases mention: a new large power connection request — meaning a facility seeking 100 megawatts or more — takes an average of four to five years to work through the interconnection queue in the United States. That’s the average. Some projects wait longer. And since 2020, the total capacity of pending interconnection requests in the US has grown by approximately 700%, from roughly 450 gigawatts in the queue to over 3,000 gigawatts by 2024. The grid connection queue is not a problem that more money fixes. It’s a permitting, engineering, and physical construction problem that operates on decade-scale timelines.
The electrical grid is not a utility in the same sense as cloud computing. You cannot spin up more grid capacity with a credit card and a browser tab.
A new transmission line requires environmental review, route surveys, easement negotiations with potentially thousands of individual landowners across a multi-hundred-mile path, state and sometimes federal regulatory approval, and physical construction that takes years even once all approvals are in hand. The Grain Belt Express transmission project — a 800-mile high-voltage DC line from Kansas to Indiana — was first proposed in 2010. It received its final major regulatory approval in 2023. Construction began in 2024. It will deliver power in 2026 at the earliest. Sixteen years from concept to operation, for a project that is not especially large by the standards of what AI data center expansion requires.
The interconnection process specifically — the process by which a new large power consumer gets a dedicated connection to the transmission grid — involves a series of engineering studies that must be completed sequentially, each study depending on the results of the previous one, with the queue processed by regional transmission organizations (RTOs) that are operating at or beyond capacity themselves. PJM Interconnection, which manages the grid for 65 million people in the mid-Atlantic and Midwest, had a backlog of over 2,700 pending projects as of 2024. Clearing that backlog requires not just processing each project but redoing studies when new projects enter the queue and change the grid topology assumptions of earlier studies.
The gap between what AI companies are announcing and what the grid can deliver is not a political talking point. It’s an engineering fact that grid operators, energy regulators, and utilities are discussing with increasing alarm in trade publications that tech journalists generally don’t read.
The North American Electric Reliability Corporation (NERC) — the nonprofit that monitors grid reliability — flagged rapidly growing data center demand as a significant grid reliability concern in its 2023 and 2024 Long-Term Reliability Assessments. The specific concern is not just capacity but timing: data centers, especially AI training clusters, can have very large power ramps — going from minimal load to full load quickly, or vice versa — that create stability challenges for grid operators. A hyperscale cluster that starts a training run simultaneously engages hundreds of megawatts of load. Grid operators need to maintain reserves that can balance these swings, and the reserve margin math is getting harder as more variable demand appears on the grid.
Microsoft’s deal with Constellation Energy to restart the Three Mile Island nuclear plant (the reactor that remained operational, not the ones involved in the 1979 accident) received extensive press coverage in 2023 as evidence that AI companies were taking their power needs seriously. What received less coverage: the plant won’t deliver power to Microsoft’s data centers until 2028 at the earliest, after safety assessments and regulatory approvals. Google’s deal with Kairos Power for small modular reactor capacity: 2030 at the earliest. Amazon’s nuclear deals: similar timeframes.
Nuclear power is the right long-term answer to zero-carbon baseload power for AI data centers. It is not an answer to the 2025-2028 power constraint.
What is filling the gap is natural gas.
The AI electricity demand surge is being met, in practice, by natural gas peaker plants — facilities designed to run when demand exceeds the grid’s baseload capacity, burning methane for power. The amount of new natural gas generation being permitted, built, and planned in response to data center demand is substantial and is being quietly tracked by energy researchers who note the obvious contradiction with tech company sustainability pledges.
Google, Microsoft, and Amazon each have commitments to be carbon-free or carbon-neutral by various dates in the 2025-2040 range. These commitments were made when the companies’ electricity consumption was growing at a manageable pace and could be plausibly offset by renewable energy purchases. The AI investment surge has broken this math. Renewable energy capacity is growing rapidly, but not rapidly enough to cover the kind of demand growth that $300 billion in data center investment represents. The gap gets filled by gas.
The companies use a mix of accounting approaches — renewable energy certificates, power purchase agreements for wind and solar that may not be geographically or temporally matched to their actual consumption, carbon offsets — to maintain the appearance of sustainability commitments while their actual grid draw includes substantial fossil generation. This is not unique to AI companies; it’s standard corporate sustainability accounting. But the scale makes it harder to sustain.
The geographic concentration of AI data centers is creating localized grid stress that is already visible in utility filings. Loudoun County, Virginia — “Data Center Alley,” where 35% of the world’s internet traffic routes through data centers — is experiencing power constraints severe enough that Dominion Energy, the regional utility, has had to defer some data center connection requests and prioritize them by strategic importance. Northern Virginia’s data center density has pushed the local transmission infrastructure to its design limits in ways that require multi-billion-dollar transmission upgrades that take years.
Similar stress is visible in central Texas (ERCOT), in the Pacific Northwest (where hydropower-dependent utilities are approaching capacity limits), and in the Ohio Valley (where coal plant retirements are reducing the baseload cushion that data center growth was relying on).
The practical consequence for AI development timelines is this: announcements of large new training clusters are not the same as those clusters being operational. A company can acquire land, order GPUs, and begin construction immediately. They cannot accelerate the grid interconnection queue. The announced capacity of AI infrastructure in 2025 significantly exceeds the electrically-deliverable capacity in any near-term timeframe. Some of this infrastructure will sit partially idle waiting for power. Some of it will get power through interim diesel generation or gas connections that were not in the sustainability accounting. All of it will cost more and take longer than the press releases suggested.
The solution exists in principle. More transmission infrastructure, faster permitting (the Inflation Reduction Act included some interconnection reform provisions, though not enough), coordinated utility investment in grid upgrades ahead of demand rather than behind it, and ultimately the nuclear capacity that is being contracted now but won’t deliver for five to ten years.
Getting from here to there requires sustained policy attention and capital allocation to electrical infrastructure — not glamorous, not a product launch, not something that generates press coverage — for decades. The US has historically underinvested in transmission infrastructure relative to its needs, and the AI demand surge is revealing the extent of that underinvestment abruptly.
The constraint is not chips. Chip supply chains have been strained and are recovering. The constraint is not talent — the global ML research community is large and growing. The constraint is not capital, which is available in extraordinary quantities. The constraint is the physical infrastructure for delivering electricity, built over decades by regulated utilities working within rate structures that did not anticipate this demand profile, and not adjustable on the timescales that tech company investment announcements imply.
The geographic dimension of this matters. AI companies are not distributing data center capacity evenly — they’re clustering it where land, water for cooling, and existing grid infrastructure coincide favorably. Northern Virginia, the Pacific Northwest, central Texas, parts of the Midwest. These clusters create local grid stress that utility planning did not anticipate and cannot quickly remedy. The PJM queue problem is not a national average; it’s particularly acute in the regions where data center demand is most concentrated. Utilities serving these regions are simultaneously trying to retire aging fossil generation (driven by economics and regulation), interconnect large amounts of new renewable generation (which requires the same interconnection queue infrastructure), and accommodate unprecedented new industrial load. They’re doing three difficult things at once with engineering staff and regulatory processes that were designed for a much quieter environment.
Demand response is being piloted as a partial solution — contractual agreements where data centers agree to curtail load during grid stress events in exchange for lower electricity rates. This is sensible and some hyperscalers have signed such agreements. The limitation is that curtailing a training run mid-process is not always possible or desirable — stopping a multi-day training job costs compute time and potentially corrupts checkpoints. Inference workloads are more curtailable than training workloads. As the AI workload mix shifts toward inference at scale, demand response becomes more practical, but the shift takes time.
The grid problem is real and the industry mostly doesn’t say so publicly, because saying so would invite questions about the credibility of the expansion plans and the sustainability commitments simultaneously. Quieter to announce the investment and let the engineers work the problem. They’re working it. It will take a while. The gap between “we announced a $50 billion data center expansion” and “that data center is operational and drawing the power it needs” is measured in years, and the honest public conversation about AI infrastructure development should acknowledge that gap rather than treating the announcement and the achievement as equivalent.
One email a month: the upcoming live event + free recording access for subscribers. No spam, unsubscribe anytime.

