The Grid Is Learning to Dispatch Demand

The standard framing of the data center power problem is a supply story. Demand is surging, generation is not keeping pace, and the fix is to build more: more gas turbines, more nuclear, more transmission. That framing is not wrong, but it is incomplete, and the gap in it is starting to show up in how serious operators are actually behaving.

A Duke University analysis published by the Nicholas Institute for Energy, Environment and Sustainability in February 2025 modeled 22 of the largest balancing authorities in the country, covering roughly 95 percent of US peak power load. It found that if new large loads, including data centers, accept modest curtailment during the hours the grid is genuinely stressed, the existing system could absorb an additional 76 to 126 gigawatts with no new generation at all. The curtailment required to unlock that headroom was small: between a quarter and one percent of the year, averaging under two and a half hours at a time. For roughly 90 percent of the hours curtailment was needed, the load kept running at half capacity rather than shutting down entirely.

That is not a modest number. It is more new load than the entire existing global data center fleet consumes. And it does not require a single new power plant. It requires knowing when to ask a load to step back, for a couple of hours, a handful of times a year.

The industry is already testing this

The theory is now being built. Nvidia, EPRI, the data center developer InfraPartners, and the real estate operator Prologis are constructing a pilot fleet of roughly 25 small data centers, each in the 5 to 20 megawatt range, sited next to utility substations across five US utilities. The premise is straightforward: instead of requiring one very large interconnection at one location, which can mean a decade-long queue, place several smaller facilities where spare substation capacity already exists, and shift the compute between them as conditions change. If one substation is overloaded or a utility needs to shed load, the workload moves to a site with headroom rather than the facility going dark.

EPRI's projection is that these facilities will need to relocate workload to a different substation only about 0.1 percent of the time. That figure has circulated as if it were a measured result. It is not. Construction of the pilot fleet was not slated to begin until the end of 2026, so the number is a pre-launch estimate from the project's own sponsors, not an observed outcome. That matters less than it sounds. An ex ante projection of how rarely you need to exercise an option is still a statement about the option's value, and a rare trigger is exactly what you want from insurance. The claim to watch is not whether 0.1 percent holds up, but whether the model works at all once the fleet is live.

Where this does and does not apply

The honest caveat is that this only works for a subset of data center workloads. AI training runs need large numbers of GPUs tightly networked together in one place, because the job cannot be split across distant sites without destroying performance. Inference, the process of actually running a trained model to answer a query, has no such requirement. It is comparatively lightweight, does not need the same interconnect density, and can be dynamically routed to wherever power happens to be available. The pilot fleet is built for inference, not training, and that is not a limitation so much as a scope. The two workload types are becoming different infrastructure problems with different solutions, and treating them as one undifferentiated "AI power crunch" obscures that.

The Duke study also has its own honest limits. It did not model transmission constraints, only generation headroom, and it used peak demand levels rather than reserve margin, both of which could move the real number in either direction. Flexibility unlocks capacity on paper. Whether the wires can move it to where it is needed is a separate question the report does not answer.

The reframe

None of this changes how much power the country will eventually need to build. It changes the sequence in which the constraint actually binds. The industry has spent two years asking how fast new generation can come online. The more useful question, and the one this pilot is quietly answering, is how much of the demand curve can be served by routing existing capacity better before a single new turbine is ordered.

That is a sequencing and asset-classification problem, not a generation problem, and it belongs on the same axis as the rest of the asset classes that determine how a site is actually powered: firm generation, storage, and now flexible, dispatchable load itself as something to be sized, sited, and optioned like any other asset on the balance sheet. Treating dispatchability as a genuine asset class, rather than a footnote to a generation story, is the shift this pilot represents. Whether that shift shows up in how developers underwrite the next generation of sites is the thing worth watching over the next eighteen months.

Questions and answers

What did the Duke Nicholas Institute study actually find? Modeling 22 of the largest US balancing authorities, covering about 95 percent of peak power load, the February 2025 study found that 76 to 126 gigawatts of new load could be integrated into the existing grid with no new generation, provided that load accepts brief curtailment between 0.25 and 1 percent of the year, averaging under two and a half hours per event.

What is the Nvidia/EPRI pilot project? A fleet of roughly 25 small data centers, each sized between 5 and 20 megawatts, built next to utility substations across five US utilities. The partners are Nvidia, EPRI, developer InfraPartners, and real estate operator Prologis. The concept is to shift inference workloads between sites based on which substation has spare capacity at a given moment, rather than requiring one very large interconnection at a single location.

Is the 0.1 percent workload relocation figure real? It is a projection, not a measured result. Construction of the pilot fleet had not begun as of the estimate's publication, so the number reflects the sponsors' modeling assumptions rather than operating data. It should be read as an ex ante claim about how rarely the option needs to be exercised, which is a reasonable way to think about the value of optionality, not as evidence the model already works.

Does this solve the AI power shortage? No, and it is not designed to. It applies to inference workloads, which can be routed to wherever power is available. Training workloads require tightly networked GPU clusters in one location and cannot be split this way. The pilot addresses one workload type, not the aggregate demand curve.

Why does this matter for how data centers get financed and sited? Because it reframes flexible load from an operational nicety into a genuine asset class, alongside generation and storage, that can be sized, sited, and optioned as part of a site's power strategy from the start rather than retrofitted later. Developers and financiers who treat dispatchability as a first-order design decision will have more sequencing options than those who treat it as a backup plan.

Previous
Previous

What the Substation Headroom Argument Leaves Out

Next
Next

Why Is The Structured Transition Model New?