A data center can have spare megawatts and still have nowhere to install the next rack.
That sounds contradictory only if capacity is treated as one pool of electricity. In a real facility, capacity has a location, a redundancy basis, a cooling requirement and an electrical path. Five unused megawatts at campus level do not help a row whose busway is full, a hall with no liquid-cooling loop, or a building whose remaining utility allowance is structurally reserved for facility systems.
This is why capacity planning should move through several layers rather than starting with a simple “MW ÷ kW per rack” calculation. The arithmetic is useful, but it is only the first filter.
What the site can receive and support under its design basis.
Capacity that exists but is not directly available to IT compute.
The capacity that can actually be allocated to IT under the intended design.
Row, busway, PDU, cooling and floor constraints decide whether the capacity is usable.
The first mistake is treating utility MW and IT MW as the same number
The grid or on-site generation sees the entire facility. Servers see only the share delivered to IT equipment. Cooling systems, pumps, fans, electrical conversion losses, controls and auxiliary building loads all sit between those two numbers.
Uptime Institute's May 2026 work on AI-era capacity allocation makes this distinction explicit. PUE tells us how efficiently a facility operates once energized, but it does not answer how much of the site's provisioned electrical envelope is structurally available to IT under the declared redundancy design. Uptime argues that this allocation question is becoming increasingly important as grid capacity tightens.
Imagine a project described as a “20 MW site.” Before translating that into racks, I would want to know whether 20 MW means utility service, total facility demand, IT capacity or a marketing shorthand for eventual build-out. Each definition produces a different rack plan.
This is also why our guide to data center power use separates design capacity, operating load and total facility demand. Capacity planning breaks down quickly when those terms are allowed to drift.
A real 5 MW AI design shows why average rack density can be almost useless
Vertiv's current AI Reference Design #028 is a particularly good example because it publishes the capacity allocation in enough detail to inspect.
The design is described as a 5 MW solution using NVIDIA DGX GB300 NVL72 architecture. Its stated total IT load is 4,992 kW across 64 racks. Divide those numbers and the apparent average is exactly 78 kW per rack.
If that were the only number in a planning spreadsheet, the spreadsheet would be badly misleading.
The actual deployment contains 32 compute racks at 142 kW each and 32 support racks at 14 kW each. Half the racks operate at more than ten times the density of the other half. Vertiv also specifies a cooling split of roughly 78% liquid and 22% air. The reference design explicitly tells planners to align AI clusters with data center “capacity blocks” to avoid stranded power.
The total still reconciles perfectly: 4,544 kW plus 448 kW equals 4,992 kW. But the facility cannot be designed around an imaginary row of identical 78 kW racks. The compute positions need the electrical and thermal infrastructure for 142 kW; the support positions do not.
Whenever I see “average rack density” in a capacity document, the next number I want is the distribution. An average can be statistically correct and operationally useless.
Capacity planning works better in blocks than in averages
The Vertiv design organizes the 5 MW solution as two pods of 2,496 kW each. Within each pod, the high-density compute racks and lower-density support racks form a repeatable deployment unit.
That is the idea behind a capacity block: instead of asking whether the campus has another 2.5 MW somewhere, ask whether it has another complete 2.5 MW block with the required power topology, cooling mix, network adjacency and physical rack positions.
The distinction becomes increasingly important with rack-scale AI. A 142 kW GB300 rack cannot simply occupy any empty cabinet position. Its location needs access to the required branch capacity, liquid-cooling infrastructure and residual air cooling. The entire cluster may also need to remain physically organized around its network topology.
This changes expansion planning. A facility with 3 MW of scattered spare capacity can be less useful to an AI tenant than another facility with one contiguous, pre-engineered 2.5 MW pod.
Stranded power is capacity you own but cannot productively use
Data centers have always had stranded capacity. AI makes it more visible because rack density is so uneven.
Suppose a hall has 1 MW of unused electrical capacity. On paper, ten 100 kW racks fit exactly. But imagine the remaining capacity is distributed as 20 kW across fifty rack positions. The total is still 1 MW. There is nowhere to place even one 100 kW rack without redesigning the local distribution.
Cooling can strand capacity in the same way. A row may have enough electrical kW but not enough heat-removal capacity. Or liquid-cooled racks may need to sit near a CDU loop that has already reached its usable design load.
Physical space can invert the problem. A facility may have technically available power and cooling but no adjacent rack positions for the cluster. Moving the workload into multiple disconnected areas can be operationally or architecturally unacceptable.
1,000 kW available.
50 separate positions with only 20 kW each.
10 contiguous positions at 100 kW each.
Redundancy changes how much provisioned capacity becomes usable IT
A capacity plan also needs to distinguish installed equipment capacity from sustainable IT allocation under the facility's failure design.
Vertiv's 5 MW reference uses a four-to-make-three power configuration and N+1 cooling. The purpose of installing redundant infrastructure is precisely that not every component has to be available simultaneously for the IT load to remain supported.
This means adding the nameplate ratings of every UPS, generator or cooling unit and calling the result “capacity” will overstate what the facility can sell or sustainably allocate to IT. The usable figure is the load supported after applying the declared redundancy basis and design constraints.
Uptime's 2026 capacity-allocation work frames this as a structural question: given the site's permitted power envelope and its redundancy architecture, how much power is sustainably allocatable to IT?
PCE is an interesting new metric — but I would not treat it as an industry standard yet
Uptime's paper discusses Power and Compute Effectiveness (PCE), a metric developed by cooling-system provider Airsys. PCE is defined as provisioned IT compute power allocation divided by total provisioned site electrical capacity.
The attraction is obvious. PUE tells us about operating overhead. PCE tries to expose how effectively a constrained electrical envelope is allocated before workloads are even energized.
If a hypothetical 20 MW provisioned site can structurally allocate 16 MW to IT under its chosen redundancy and facility design, its PCE under that definition would be 80%. The remaining 4 MW of the envelope is not necessarily “waste”; it can represent cooling, electrical losses, auxiliary loads and the consequences of the reliability architecture.
Uptime is appropriately cautious. It says the value of PCE will depend on whether operators adopt it and whether broader industry data demonstrates that it provides meaningful insight. So I would use the concept today as a planning lens, not present PCE as a settled industry benchmark.
PUE and capacity allocation answer different questions
This distinction is easy to miss.
PUE asks how much total energy the facility uses relative to IT energy while operating. A site with 10 MW of IT load and PUE 1.20 is using roughly 12 MW at facility level under a simplified steady-state example.
Capacity allocation asks a different question: how much of the site's provisioned electrical envelope can be assigned to IT under the intended design and redundancy conditions?
A facility can therefore have an excellent PUE and still make poor use of a scarce utility connection if large amounts of provisioned capacity are stranded by topology, over-allocation, cooling constraints or conservative infrastructure design. Conversely, aggressively maximizing allocatable IT capacity without preserving the required redundancy would be equally misguided.
The goal is not to maximize one percentage. It is to understand where the power went and whether that allocation matches the business and reliability requirements.
Rack-space planning needs peak density, not just average density
Return to the Vertiv example. Sixty-four racks averaging 78 kW sounds like a relatively uniform high-density room. In reality, half the positions are at 142 kW.
That changes aisle planning, liquid distribution, branch circuits, maintenance procedures and the physical consequence of a failed cooling segment. It also changes how future capacity can be filled. Replacing a 14 kW support rack with another 142 kW compute rack is not a simple cabinet swap unless the local infrastructure was designed for it.
For mixed-density halls, the capacity model should therefore preserve rack classes rather than collapsing everything into one average. “32 × 142 kW + 32 × 14 kW” contains operational information that “64 racks × 78 kW average” throws away.
The rack count formula is still useful — after the design density is chosen
Once the relevant rack class is known, the simple formula regains its value:
A 10 MW homogeneous IT deployment at 20 kW/rack would imply about 500 racks. At 50 kW, about 200. At 100 kW, about 100.
The formula should be treated as a first-pass capacity count, not a floor plan. Network racks, storage, support equipment, spare positions, maintenance clearances and mixed workload densities can all increase the actual cabinet requirement.
Our 20 kW vs 50 kW vs 100 kW analysis explores why the infrastructure around those three deployments changes even when total IT MW remains identical.
Deployment phasing can matter as much as final MW
A 50 MW campus that eventually supports 50 MW of IT is not necessarily a 50 MW operating load in year one.
Capacity may arrive in utility phases. Buildings may be commissioned sequentially. Customers may deploy racks over months or years. AI hardware can also arrive in large batches tied to product cycles rather than in a smooth monthly ramp.
A good capacity plan therefore has at least two time axes: when infrastructure becomes available and when IT load is expected to consume it. The gap between those dates represents capacity that exists financially and physically but is not yet producing compute or revenue.
That gap can be deliberate. Building ahead of demand can protect future expansion rights. The mistake is hiding it by presenting ultimate campus MW as if it were current usable load.
AI power management adds a new layer between installed capacity and operating load
NVIDIA's current Power Reservation Steering documentation illustrates how software is beginning to participate in capacity allocation. For GB200 NVL72, NVIDIA lists a designed rack power of 120 kW; for GB300, 135 kW. Its software can manage the dynamic portion of node power within a defined power domain rather than assuming every device draws its maximum simultaneously.
NVIDIA's default utilization threshold in the current documentation is 93%, preserving headroom for runtime fluctuations before new jobs are admitted to a power domain. The documentation also distinguishes static unmanaged power from dynamic GPU/CPU power that the scheduler can control.
This does not mean electrical infrastructure can simply be undersized because software exists. It means AI introduces a controllable layer that conventional capacity models often did not have. Any design credit taken for software-managed diversity needs to be explicit, validated and consistent with the facility's reliability objectives.
Capacity is also a cooling problem
Uptime's capacity-allocation analysis specifically includes capacity structurally reserved for cooling at peak design load because cooling systems draw from the same provisioned site envelope as everything else.
Vertiv's 5 MW design makes the interdependence concrete. The cooling topology is 78% liquid and 22% air, with dedicated liquid-cooling, air-cooling and heat-rejection equipment. The compute capacity cannot be increased independently of those systems just because an electrical meter says spare kW remain.
At high density, cooling can become the local constraint before site power does. A CDU loop can be full while the upstream utility connection has megawatts remaining. That is why liquid-cooling planning increasingly happens at the pod level rather than by isolated rack.
Network topology can strand physical rack space too
Capacity planning is usually discussed as power plus cooling, but large AI clusters add a third constraint: connectivity.
Rack-scale and pod-scale architectures depend on specific high-bandwidth network relationships. An empty rack position on the opposite side of a building may have enough power and cooling but still be a poor place to expand a tightly coupled cluster if doing so changes cable reach, network topology or operational boundaries.
I would therefore avoid treating every powered rack position as fungible capacity in an AI facility. Some positions have more strategic value because they complete a contiguous cluster block.
A useful capacity ledger should reconcile five numbers
I would want a planning document to be able to reconcile the site from the top down without changing definitions halfway through.
Provisioned site capacity. What is permitted and available from grid and on-site sources under the defined operating basis?
Facility reserve and infrastructure load. What portion is structurally required for cooling, losses, auxiliaries and reliability?
Allocatable IT capacity. How many MW can sustainably reach IT under the design?
Local block capacity. Where can those MW actually be delivered at the required rack density and cooling type?
Deployed and operating load. How much of the available IT capacity is installed, energized and actively used today?
When those five numbers reconcile, a site-level MW headline becomes useful. When they do not, “20 MW available” can mean almost anything.
The utilization percentage can hide two very different problems
Suppose a facility has 10 MW of allocatable IT capacity and only 6 MW is operating. Calling it 60% utilized does not tell us why the remaining 4 MW is idle.
One possibility is commercial: the capacity is genuinely available and simply has not been leased or deployed yet.
Another is structural: parts of the 4 MW cannot be used by the next intended workload because the capacity is fragmented across the wrong halls, cooling systems or rack densities.
Those situations look identical in a site-wide utilization percentage and have very different economic implications. The first is inventory. The second is stranded infrastructure.
Capacity planning should end at a rack position, not at a campus number
For a conventional low-density estate, planners could often work comfortably from hall-level or room-level averages. AI makes that increasingly risky.
A useful plan should eventually be able to point to the exact positions where the next deployment goes and answer four questions at once: Is the electrical capacity there? Is the cooling capacity there? Does the redundancy architecture still hold? Does the workload fit the physical and network layout?
If the answer depends on moving capacity from somewhere else, that movement should be part of the project plan rather than assumed to be free.
MW tells you the scale of a data center. kW tells you how that capacity is distributed. Rack space tells you where it can physically be used. Capacity planning is the discipline of making all three statements true at the same time — under the actual cooling, redundancy and workload constraints of the facility.
