A GPU rack can be described as 120 kW, 130 kW or 142 kW without anyone necessarily being wrong.
That sounds like a bad way to begin a sizing exercise, but it is exactly the problem infrastructure teams face in 2026. NVIDIA's current Mission Control documentation lists a designed rack power of 120 kW for GB200 NVL72 and 135 kW for GB300 NVL72. Vertiv, meanwhile, publishes infrastructure reference designs around 130 kW for GB200 and 142 kW for GB300.
The numbers are not interchangeable because they answer different questions. NVIDIA is describing the rack platform's designed power envelope. Vertiv is designing the facility infrastructure that has to support deployments built around those systems. The second number can therefore include a different design basis, margin or reference-architecture assumption.
This is the first rule of GPU rack planning: never use a rack-power number until you know what boundary it describes.
NVIDIA designed rack power for GB200 / GB300 in current Mission Control guidance.
Vertiv reference-design rack density for GB200 / GB300 deployments.
Total facility demand also includes cooling and other non-IT loads.
NVIDIA's current Mission Control power-reservation guidance is particularly useful because it makes the product-level distinction explicit. It lists 120 kW designed rack power for GB200 and 135 kW for GB300, then explains how that budget is divided across the 18 nodes in each rack.
Vertiv's 2026 AI infrastructure reference designs, by contrast, show deployments engineered around 130 kW GB200 racks and 142 kW GB300 racks. I would not “correct” one source with the other. I would ask which number the electrical designer intends to guarantee.
GB200 is already a 120 kW rack-scale computer
NVIDIA GB200 NVL72 puts 72 Blackwell GPUs and 36 Grace CPUs into a single liquid-cooled rack-scale system connected through NVLink. NVIDIA describes the rack as an exascale computer rather than a collection of independent servers, which is an important physical change as well as a compute change.
In its power-management documentation, NVIDIA uses approximately 120 kW at full load for the complete GB200 rack, including nodes and rack components such as switches and interconnects. The same guidance says a fully provisioned infrastructure power budget should be at least 120 kW for that product-level design envelope.
That is roughly the electrical demand of a small commercial building concentrated into one rack footprint.
Those totals are simple multiplication of designed rack power, not a prediction of simultaneous operating draw. That distinction becomes important once an installation contains hundreds of racks.
GB300 moves the designed rack envelope to 135 kW
NVIDIA's latest Mission Control documentation lists 135 kW designed rack power for GB300 NVL72. Each rack still contains 18 nodes, but the per-GPU maximum rises from 1.2 kW in the GB200 entry to 1.4 kW for GB300, increasing the managed compute envelope.
NVIDIA's DGX SuperPOD GB300 reference architecture also shows how much power hardware now sits inside the rack: eight power shelves, each capable of delivering up to 33 kW, are used in the rack system. The presence of substantially more installed power-supply capability than normal operating demand is not evidence that the rack continuously consumes that sum. Redundancy, transient behavior and power delivery architecture all matter.
That is another reason “add up the PSU labels” is not a sound way to estimate facility load.
Why Vertiv designs around 142 kW when NVIDIA says 135 kW
This apparent disagreement is actually one of the most useful public examples of how facility planning works.
NVIDIA says GB300 has a 135 kW designed rack power. Vertiv's current AI Hub lists GB300 new-build reference designs at 142 kW rack density. Vertiv also shows GB200 infrastructure designs around 130 kW even though NVIDIA's product-level design figure is 120 kW.
The difference is around 5% for GB300 and about 8% for GB200. It would be tempting to call that “the required headroom.” I would not do that without the design documents saying so. Reference-architecture values can reflect vendor-specific assumptions, ancillary loads, practical sizing choices and the particular system boundary being modeled.
What the comparison does prove is that a facility designer should not assume the product's headline rack power is automatically the final branch-circuit, busway or room-density design value.
NVIDIA: “What is the designed rack power envelope?”
Vertiv: “What rack density should this infrastructure reference design support?”
Do not average the numbers. Resolve the boundary.Rack power is not the same as GPU power
A rack contains much more than GPUs. CPUs, NVLink switches, network interfaces, storage, management hardware and power-conversion losses all contribute to the rack's demand.
NVIDIA's current GB300 power-budget example makes this visible. It models four GPUs per node at up to 1.4 kW each, plus CPU power and a static allowance for the rest of the node and rack infrastructure. The resulting rack-level design number is therefore more useful for facility planning than multiplying GPU TDP by GPU count.
The same principle applies outside NVIDIA. A server specification may publish accelerator power, but a data center has to power the complete server, the rack networking and the supporting electrical path.
If a capacity model starts with “number of GPUs × GPU watts” and jumps directly to facility MW, I would treat it as unfinished rather than wrong. It has skipped too many layers to be trusted yet.
Power density and energy consumption are different questions
A 135 kW rack rating tells the facility how much power it may need to deliver. It does not say that every GB300 rack will draw exactly 135 kW for 8,760 hours each year.
Uptime Institute's May 2026 work on AI-era capacity planning draws an important distinction between workloads. It says AI training clusters often operate near their peak power envelope for extended periods, while inference workloads can behave more like conventional compute and show more variation in utilization. That difference affects how much of the reserved electrical capacity becomes actual energy consumption.
Long periods near a high power envelope can make reserved capacity heavily utilized.
More variation can create a larger gap between provisioned power and average consumption.
Those sketches are conceptual rather than measured traces. The point is that infrastructure must survive the required peak behavior, while the energy bill follows the load across time.
Power management is starting to attack the gap between provisioned and used capacity
This gap is financially important because scarce data center power is often reserved based on conservative peak assumptions. If every rack is provisioned for theoretical maximum draw but rarely reaches it at the same moment, expensive infrastructure can sit underutilized.
NVIDIA is now explicitly addressing that problem in software. Its Dynamic Power Software and Power Reservation Steering tools coordinate power budgets across GPU infrastructure rather than treating each accelerator as an independent fixed load.
In one NVIDIA test using Megatron training, energy-storage-enhanced GB300 power shelves reduced the peak demand seen at the AC input by 30% while smoothing rapid rack-level fluctuations. That is a measured result for a specific test, not a promise that every data center can cut utility capacity by 30%.
The broader direction is nevertheless important. Future AI data centers may extract more compute from the same electrical envelope not only through more efficient chips, but by orchestrating when thousands of chips are allowed to consume peak power.
A 100-rack deployment is already a multi-megawatt electrical system
Consider 100 GB300 racks. Using NVIDIA's 135 kW designed rack power gives a nominal rack-level envelope of 13.5 MW. Using Vertiv's 142 kW infrastructure reference-design density gives 14.2 MW.
The 700 kW difference is not evidence that one calculation is wrong. It shows what happens when different design boundaries are scaled across a large deployment.
Then facility overhead has to be added. At a hypothetical annual PUE of 1.20, 13.5 MW of steady IT/rack load corresponds to about 16.2 MW total facility load. At PUE 1.30, it would be about 17.55 MW.
A planning document therefore needs to say explicitly whether “14 MW AI pod” means rack-level IT capacity or total facility demand. The difference can be several megawatts.
At 1,000 racks, a small assumption becomes utility-scale
The same boundary problem becomes much larger at campus scale.
One thousand GB300 racks at NVIDIA's 135 kW product-level design figure equal 135 MW. At Vertiv's 142 kW reference density, the infrastructure planning figure is 142 MW. That is a 7 MW difference before applying PUE.
Seven megawatts is larger than many entire legacy enterprise data centers. This is why hyperscale AI planning cannot tolerate ambiguous definitions even when the percentage difference looks small.
The next generation is already pushing beyond 200 kW per rack
Uptime Institute said in July 2026 that major AI hardware roadmaps are expected to push rack power above 200 kW. Its conclusion is that the industry's planning question has shifted from whether direct liquid cooling will be needed to how much liquid-cooled capacity will be needed over time.
NVIDIA ecosystem material presented at GTC 2026 describes Vera Rubin NVL72 as a 100% liquid-cooled rack above 200 kW, with 45°C coolant and a move toward in-row rather than in-rack CDUs because of the space and thermal requirements. That is a roadmap-era design, not a reason to size every 2026 deployment at 200 kW.
The more useful lesson is architectural. A facility intended to host several generations of AI hardware may have to decide whether today's 135–142 kW rack is the endpoint or simply the first stage of a path toward 200 kW and beyond.
This is why liquid cooling is now part of the power discussion
Once rack demand reaches triple digits, electrical and cooling design stop being separable conversations. Nearly all of the power consumed by the rack becomes heat that has to be removed.
GB200 and GB300 NVL72 are liquid-cooled rack-scale systems. Vertiv's 130 kW and 142 kW reference designs use liquid plus air, reflecting the fact that not every watt necessarily leaves the rack through the liquid loop.
A facility capable of delivering 142 kW electrically but removing only 80 kW thermally is not a 142 kW-ready facility in any useful sense. Power availability, coolant capacity, residual air cooling and heat rejection all have to be sized against the same deployment.
That relationship is covered in more detail in our liquid cooling cost guide, where the economic unit shifts from an isolated rack toward the complete pod and thermal path.
Do not design only for average rack density
A mixed data center might average 30 kW per rack while containing an AI zone at 140 kW. The building average can therefore look comfortable even when one busway, row or cooling loop is at its limit.
This is a spatial problem. Capacity has to exist where the high-density racks physically sit. Ten spare megawatts elsewhere on campus do not necessarily help a hall whose electrical distribution and cooling infrastructure cannot carry another 500 kW.
Uptime's 2026 work on AI-era capacity allocation argues that traditional aggregate metrics are becoming less informative for exactly this reason. High-density workloads can create local constraints that are hidden by site-wide averages.
A “150 kW-ready” claim needs at least four definitions
The phrase sounds precise. It is not precise until the provider explains what it means.
Electrical: can the rack position receive 150 kW continuously under the promised redundancy architecture?
Thermal: can the cooling system remove the corresponding heat load under design conditions?
Distribution: can the row, busway, PDU and upstream path support multiple racks at that density simultaneously?
Operational: is 150 kW available after required margins, maintenance states and failure scenarios?
If only the first question has an answer, the rack position is not fully 150 kW-ready yet.
The redundancy architecture can change how much upstream capacity a rack requires
Two racks with identical 135 kW compute demand can sit inside different electrical architectures. One facility may reserve redundant paths so either side can carry the required load after a failure. Another may use a different topology or accept a lower redundancy level for a training cluster that can tolerate interruption.
That is why upstream installed capacity cannot be inferred by multiplying rack power by rack count and then simply doubling it. The correct result depends on the redundancy topology, utilization assumptions and how loads are shared during normal and failure conditions.
For AI training, the economics can be especially different from traditional enterprise IT. A training job may be able to checkpoint and restart, which can justify infrastructure choices that would be unacceptable for a payment system or real-time application.
Power per rack is becoming a software problem too
Historically, data center electrical design was mostly about supplying the peak load hardware might request. AI introduces enough synchronized compute that the behavior of the workload itself can become part of capacity management.
NVIDIA's current Dynamic Power Software documentation describes policies that manage power from individual racks through installations with thousands of GPUs. The objective is not simply to throttle hardware; it is to maximize workload throughput inside a defined electrical envelope.
That creates an important distinction between hardware-capable rack power and operationally allocated rack power. A 135 kW rack can be installed in a system whose orchestration deliberately limits some workloads below maximum under certain conditions.
For developers and operators, this may eventually increase the amount of sellable or usable compute that fits inside a constrained utility connection. For electrical engineers, it also creates a dependency on control software that has to be understood before theoretical diversity is converted into reduced physical infrastructure.
What number should go into a capacity model?
If the model is checking whether a rack product can be supported, start with the manufacturer's designed rack power and the exact system configuration.
If the model is sizing busway, PDUs, cooling or a room, use the facility design basis agreed with the infrastructure engineering team. That may be higher than the product headline, as the Vertiv references demonstrate.
If the model is estimating annual electricity cost, use an expected operating-load profile rather than assuming every rack sits permanently at design maximum.
And if the model is sizing the utility service, move one level higher again: include the facility overhead, redundancy assumptions, deployment phasing and the degree to which power-management software can be relied upon in the final design.
Manufacturer rack envelope
Engineering design density
Expected load over time
Total facility demand + resilience + phasing
The number I would write next to a GB300 rack in 2026
I would write two numbers, not one.
135 kW beside “NVIDIA designed rack power,” because that is what NVIDIA's current Mission Control documentation publishes for GB300 NVL72. Then I would write the project's actual facility design density beside it — which might be around the 142 kW used in Vertiv's current reference architectures, or another value justified by the project's engineering basis.
Keeping both numbers visible prevents a product specification from silently becoming an electrical design assumption.
GPU rack power has stopped being a specification that can live only on a server datasheet. At 120, 135 and soon 200+ kW, it determines the shape of the room, the cooling system, the electrical topology and ultimately how many megawatts of AI compute a site can deploy. The difficult part is no longer finding a number. It is making sure everyone is using the same boundary when they say it.
