What the Substation Headroom Argument Leaves Out

The case for putting small data centers on spare distribution capacity is stronger than its critics allow. It also concedes more than its advocates have noticed.

Electrical distribution substation with transmission lines at dusk, representing spare grid capacity available to data centers

A substation, a location naturally suited to smaller data centers with flexible load

By Alex Marshall

There is a serious argument circulating that the data center industry does not need to build its own power. It deserves a serious answer, because the mathematics behind it is logical and the response to it has not been.

In February 2026, the Electric Power Research Institute, an independent non-profit research organization funded largely by the utility sector, announced a collaboration with Nvidia, the industrial property developer Prologis, and the prefabricated data center builder InfraPartners. The plan is to site small data centers, between 5 and 20 megawatts each, at or near electrical substations that already have unused capacity. The stated target was at least five pilot sites in development across the United States by the end of 2026. Recent reporting in IEEE Spectrum put the eventual fleet at around 25 sites across five utilities.

The logic is straightforward. Marc Spieler, who leads Nvidia's global energy industry work, has pointed out that there are roughly 55,000 substations in the United States, and that spare capacity of 5, 10, or 20MW at each of them adds up quickly. Ben Sooter, who runs EPRI's micro data center program, has said that individual substations typically have capacity in that range available on average.

This is not a marketing position as it rests on a real body of work. Duke University's Nicholas Institute for Energy, Environment and Sustainability published research in February 2025 finding that the largest 22 balancing authorities in the United States, the entities responsible for matching electricity supply and demand across a defined territory, could absorb around 76 gigawatts of new load if that load accepted curtailment averaging 0.25 percent of its annual consumption. At 0.5 percent the figure rose to 98 gigawatts, and at 1.0 percent to 126 gigawatts. The underlying reason is that the US power system was built to serve a peak which is rarely reaches. Grid operators use around 53 percent of their generation capacity on average, and the highest-demand hours occur for less than 200 hours a year.

So the theory that there is a large quantity of underutilized capacity sitting inside the existing distribution network, is correct, and the industry has spent three years talking about generation while under-using the flexible power assets that are installed. Anyone dismissing this as a distraction is wrong.

The problem is with what the argument is being used to prove.

The percentage and the hours are two different numbers

Sooter has said that compute workloads in the pilot would need to be moved to a different substation only about 0.1 percent of the time. That sounds trivial. It works out to roughly nine hours a year (or 97.61% availability - or zero nines in data center terms).

Compare that to the figure in the same reporting that peak demand periods run to just under 200 hours a year. Those two numbers are more than twenty times apart, and they are describing the same grid.

The Duke research shows why the gap is important. Tyler Norris, one of the report's authors, has explained that the headline curtailment rate is measured in energy, not in hours. A site being asked to shed 20 percent of its load for four hours registers as a very small energy percentage, but it is still a four-hour operational event. At the 0.25 percent energy figure, Duke found that some degree of curtailment would occur across about 85 hours a year on average. At 0.5 percent it was around 177 hours.

That is the number an operator needs, and it is not the number being quoted. The honest question for anyone signing a flexible interconnection agreement is not how often they will be fully relocated. It is how many hours a year they will be running at something less than full capacity, and by how much. Those two questions have very different answers, and only one of them is currently in circulation.

Headroom is a shared resource, and shared resources deplete

The 55,000 substation calculation describes a stock, not a flow. Each megawatt of spare distribution capacity can be allocated once.

The first tenant at a given substation faces an uncontested resource. The tenth tenant at the same substation faces a different situation entirely, and so does the first tenant once the other nine arrive. Every assumption behind the 0.1 percent figure is a first-mover assumption, calculated against a network where nobody else is doing this yet.

Nobody has published what that number looks like in steady state. If the model works, it will be copied, and the conditions that made it work will be consumed by the copying. This is the ordinary behavior of a common-pool resource, and it is the standard reason that early participants in a capacity market get terms that later participants do not.

For a developer, this is the decisive point. You are not evaluating whether spare capacity exists. You are evaluating whether it will still be there, on the same terms, at the end of a fifteen-year asset life. Those questions have almost nothing in common.

The whole model rests on peaks not coinciding

Distributing 25 sites across five utilities delivers resilience only if those five utilities experience stress at different times. Shift a workload out of a constrained substation and it has to arrive somewhere with headroom to spare, in that hour.

Weather does not cooperate with this. Winter Storm Uri in 2021 and Winter Storm Elliott in 2022 both produced simultaneous stress across multiple balancing authorities covering large parts of the country. Extreme heat behaves the same way. The events that actually threaten a data center's power supply are, by their nature, the events most likely to be correlated across a wide geography.

Peak non-coincidence across the fleet is the load-bearing assumption in this entire model. It has not been published, quantified, or stress-tested in public. Until somebody shows the correlation analysis, the flexibility on offer is only demonstrated for ordinary conditions, which are precisely the conditions where nobody needed it.

It addresses the part of the load that was never the constraint

Valerie Crafton of the modular data center company Mod42 has made the boundary explicit: inference is one of very few workloads that can be dynamically routed. Inference means running a trained model to produce an answer. Training means building the model in the first place, and it requires large numbers of processors interconnected tightly enough that splitting them across sites is not an option.

The pilot is designed for inference. That is a sensible engineering decision and I have no quarrel with it.

But the interconnection queues, the multi-year utility timelines, and the decade-long grid connection waits that gave rise to this whole debate are being driven by training campuses drawing hundreds of megawatts at a single location. The substation model does not touch them. It is a good answer to a question that was not the binding constraint.

This is not a criticism of the pilot. It is a criticism of how the pilot is being cited.

What the argument actually concedes

Here is the part that has gone unremarked.

The Duke research names three mechanisms by which a large load can deliver the flexibility that unlocks curtailment-enabled headroom: shifting the workload elsewhere, reducing operations, or running onsite generation. Norris has said directly that the growth of onsite generation and storage is one of the reasons this approach is now viable, because a curtailed facility does not have to go dark. It can transfer to its own supply.

The substation model does not remove onsite generation from the system. It changes what onsite generation is for.

In the conventional framing, generation on site exists because the grid cannot deliver enough power fast enough. In the flexible interconnection framing, generation on site exists because it is the instrument that makes the flexibility commitment credible. Without it, curtailment means lost revenue and broken service level agreements. With it, curtailment is a fuel switch. The asset moves from being the primary supply to being the thing that makes the contract signable, and it becomes more strategically important in that role, not less.

This is the pattern I have argued elsewhere as control following assets. The technical debate looks like it is about capacity. It is actually about which side of the meter the dispatch decision sits on. An operator taking spare substation capacity is accepting that somebody else decides when its load comes down, in exchange for connecting years earlier. That is a rational trade. It is also a real transfer of control, and it should be priced as one rather than treated as a free lunch because the megawatts were already sitting there.

What would change my mind

I would treat this differently if four things happened.

Publish the curtailment terms in hours and depth, not as an annual energy percentage. Publish the peak correlation analysis across the participating utilities. Show what the terms look like for the second and third tenant at the same substation rather than the first. And show whether the pilot sites are being built with onsite generation or storage behind them, and how much.

The first three would tell the market whether this is a durable model or a first-mover arbitrage. The fourth would settle the question this article is really about.

My expectation is that the sites get built, that they work, that they are genuinely useful for inference at the edge of the network, and that a meaningful proportion of them end up with generation or storage on site anyway. Not as a fallback. As the reason the interconnection agreement could be signed at all.

Spare grid capacity is real. Firmness is a separate product, and somebody still has to supply it.

 

Five Nines and Fast Power

Making Better Decisions in Infrastructure Investments

 

Questions and Answers

Can data centers use spare substation capacity instead of building onsite generation?

Not as a straight substitute. Spare substation capacity is real and substantial, but it is generally offered on flexible interconnection terms, meaning the operator agrees to reduce load when the grid is stressed. Delivering that reduction requires either shifting the workload elsewhere, cutting operations, or running onsite generation. In practice, onsite generation is often what makes the flexibility commitment credible, which means the substation model changes the role of onsite generation rather than removing it.

What is curtailment-enabled headroom?

Curtailment-enabled headroom is the amount of additional electrical load a power system can absorb using existing capacity, provided that load agrees to be temporarily reduced during periods of peak stress. The term was introduced by researchers at Duke University's Nicholas Institute for Energy, Environment and Sustainability in February 2025. It reframes grid capacity as a function of how flexible the new load is willing to be, rather than as a fixed quantity.

How much spare capacity does the United States grid actually have?

Research from Duke University's Nicholas Institute found that the 22 largest balancing authorities could absorb around 76 gigawatts of new load at an average annual curtailment rate of 0.25 percent, rising to 98 gigawatts at 0.5 percent and 126 gigawatts at 1.0 percent. Those balancing authorities cover roughly 95 percent of United States peak demand. The underlying reason is that grid operators use only about 53 percent of their generation capacity on average, because the system is built for peaks that occur for fewer than 200 hours a year.

What is the Nvidia and EPRI micro data center pilot?

It is a collaboration announced in February 2026 between the Electric Power Research Institute, Nvidia, Prologis and InfraPartners to site small data centers of 5 to 20 megawatts at or near electrical substations with unused capacity. The facilities are intended for distributed inference workloads rather than model training. The stated goal was at least five pilot sites in development across the United States by the end of 2026.

Why does the difference between curtailment percentage and curtailment hours matter?

Because they describe very different operational realities. A curtailment rate expressed as a percentage measures energy, not time. A facility asked to shed a fifth of its load for several hours registers as a very small energy percentage while still experiencing a multi-hour event. Duke University's research found that a 0.25 percent energy curtailment rate corresponded to some degree of curtailment across roughly 85 hours a year, and 0.5 percent to roughly 177 hours. Operators evaluating a flexible interconnection agreement need the hours figure and the depth of reduction, not the annual energy percentage.

Does flexible interconnection work for AI training or only for inference?

Primarily for inference. Inference means running a trained model to generate a response, and it can be distributed across sites and dynamically routed. Training requires large numbers of processors interconnected tightly enough that splitting the workload across geographically separate locations is impractical. The interconnection queues and multi-year utility timelines currently constraining the sector are driven largely by training campuses drawing hundreds of megawatts at a single site, which distributed inference models do not address.

What is the main risk in relying on spare substation capacity?

That it is a shared resource being valued on first-mover terms. Spare distribution capacity at a given substation can be allocated once. The conditions that make the arrangement attractive to an early participant, including the low expected frequency of curtailment, do not necessarily hold once the model is widely copied. A second risk is peak correlation. Distributing sites across multiple utilities only provides flexibility if those utilities experience stress at different times, and severe weather events tend to affect large geographic areas simultaneously.

Next
Next

The Grid Is Learning to Dispatch Demand