Skip to content
Uniqcli

NVIDIA GB300 NVL72 by HPE

Rack-scale AI power for models above 1 trillion parameters

Overview

The NVIDIA GB300 NVL72 by HPE is a liquid-cooled, rack-scale AI system built to train and serve models above 1 trillion parameters, and it's the right fit for organizations that have outgrown standard 8-GPU servers and need a full rack that behaves as one accelerator. It packages 72 NVIDIA Blackwell Ultra GPUs and 36 Grace CPUs into a single NVLink domain, delivered by HPE as an integrated system rather than parts you assemble yourself.

Inside that domain, fifth-generation NVLink ties every GPU together at 130 TB/s, and each Blackwell Ultra GPU carries 288 GB of HBM3e memory, about 50% more than the prior generation, for roughly 37 TB of combined fast memory across the rack. NVIDIA rates the platform at 1,080 PFLOPS of dense FP4 Tensor Core throughput, tuned for the long-context, token-heavy demands of reasoning models and real-time generative inference at production volume. Networking runs over NVIDIA ConnectX-8 SuperNICs at 800 Gb/s per GPU, with a choice of Quantum-X800 InfiniBand or Spectrum-X Ethernet depending on how your operations team scales out.

HPE wraps that NVIDIA hardware with direct liquid cooling engineered for the rack's roughly 132 kW thermal load, plus deployment services, HPE GreenLake consumption options, and support that carries the system from facility planning through production operations. GB300 NVL72 began shipping in December 2025 and is now running in hyperscale production, including Vultr's 2026 global AI infrastructure buildout, so this is proven, supportable capacity rather than a future roadmap promise. As an authorized HPE partner, Uniqcli scopes, sources, and supports the full deployment for federal, SLED, healthcare, and enterprise buyers.

Request a quote
NVIDIA GB300 NVL72 by HPE

Why NVIDIA GB300 NVL72 by HPE

The NVIDIA GB300 NVL72 by HPE is the current flagship rack-scale platform for training and serving AI models beyond 1 trillion parameters, and buyers need to know it's shipping and proven, not a future promise. Since shipping began in December 2025, it's already running in hyperscale production, giving federal, SLED, healthcare, and enterprise buyers a validated path to frontier-scale AI capacity backed by an authorized HPE partner that can scope facility readiness, financing, and support before a 132 kW rack ever lands on the floor.

One 72-GPU NVLink domain per rack

72 NVIDIA Blackwell Ultra GPUs and 36 Grace CPUs connect over fifth-generation NVLink at 130 TB/s, so the rack functions as a single large accelerator instead of a cluster of separate servers.

288 GB HBM3e per GPU, ~37 TB fast memory total

Blackwell Ultra's 12-high HBM3e stacks push memory to 288 GB per GPU, roughly 50% more than the prior generation, giving the rack a combined pool of GPU and Grace CPU memory large enough for trillion-parameter model partitions.

1,080 PFLOPS FP4, built for reasoning-scale inference

NVIDIA rates dense FP4 Tensor Core throughput at 1,080.2 PFLOPS (about 1.1 exaFLOPS), engineered specifically for the token-hungry, long-context inference patterns of frontier reasoning models.

800 Gb/s per GPU with fabric choice

Every GPU gets a dedicated ConnectX-8 SuperNIC path at 800 Gb/s, and the rack supports either NVIDIA Quantum-X800 InfiniBand or Spectrum-X Ethernet, so you scale out on the fabric your operations team already runs.

HPE direct liquid cooling at rack scale

Cold plates on every compute and NVLink switch tray capture roughly 90% of the rack's heat to liquid, letting a 132 kW-class rack run at sustained density that air cooling alone can't touch.

Shipping and running in production today

GB300 NVL72 began shipping in December 2025 and is now deployed at hyperscale, including a 2026 buildout for Vultr's global AI infrastructure, evidence this is a supportable platform, not a future roadmap item.

Backed by the full HPE AI Factory stack

HPE wraps the NVIDIA hardware with integrated networking, software, deployment services, and HPE GreenLake consumption options, so your team isn't integrating a rack-scale system from parts alone.

What it does

72-GPU single NVLink domain

Fifth-generation NVIDIA NVLink connects all 72 Blackwell Ultra GPUs into one shared-memory domain with 130 TB/s of aggregate NVLink bandwidth, so the rack behaves as a single accelerator instead of 72 separate GPUs stitched together over a network.

288 GB HBM3e per GPU

Each Blackwell Ultra GPU carries 288 GB of HBM3e memory (12-high stacks, up from 8-high on the prior generation), giving the full rack roughly 20 TB of fast GPU memory plus 17 TB of Grace CPU LPDDR5X, for a combined ~37 TB fast-memory pool inside one NVLink domain.

1,080 PFLOPS of FP4 inference compute

NVIDIA rates the rack at 1,080.2 PFLOPS dense FP4 Tensor Core throughput (roughly 1.1 exaFLOPS) and 720 PFLOPS at FP8/FP6, sized for reasoning models and long-context inference at production scale.

800 Gb/s per-GPU networking

Each GPU gets its own NVIDIA ConnectX-8 SuperNIC path at 800 Gb/s, with the rack supporting either NVIDIA Quantum-X800 InfiniBand or Spectrum-X Ethernet for scale-out to multi-rack clusters, so the network fabric choice matches your existing data center standard.

HPE direct liquid cooling at rack density

Compute and NVLink switch trays use cold plates cooled by HPE direct liquid cooling; at the rack level roughly 90% of heat is captured to liquid and 10% to air, with a nominal 132 kW thermal design power and up to ~155 kW peak electrical draw per rack.

18 compute trays, 2 GB300 boards each

The rack packages 18 liquid-cooled compute trays (2 GB300 boards per tray) plus liquid-cooled NVLink switch trays and 8 power shelves feeding a DC bus bar, engineered as one delivered unit rather than assembled on-site component by component.

Part of NVIDIA AI Computing by HPE

The GB300 NVL72 slots into the broader HPE AI Factory portfolio alongside HPE services, software, and HPE GreenLake, so training, real-time inference, and agentic AI workloads run on infrastructure HPE integrates and supports end to end.

Production-proven at hyperscale

Beyond single-rack deployments, the platform is running in cloud-scale environments, including Vultr's global infrastructure buildout announced in 2026 using GB300 NVL72 by HPE with NVIDIA Spectrum-X Ethernet, evidence this is shipping, supportable infrastructure rather than a paper launch.

The NVIDIA GB300 NVL72 by HPE lineup

NVIDIA GB300 NVL72 by HPE

The full rack configuration for training and serving trillion-parameter-class models on one HPE-integrated system.

Liquid-cooled rack: 72 Blackwell Ultra GPUs, 36 Grace CPUs, single NVLink domain, 1,080 PFLOPS FP4

GB300 NVL72 with Quantum-X800 InfiniBand

Best where your cluster standard is InfiniBand and you're scaling multiple racks into one large training fabric.

Rack fabric option using NVIDIA Quantum-X800 InfiniBand and ConnectX-8 SuperNICs at 800 Gb/s per GPU

GB300 NVL72 with Spectrum-X Ethernet

Best for teams standardizing on Ethernet operations or building cloud-style, multi-tenant AI infrastructure.

Rack fabric option using NVIDIA Spectrum-X Ethernet, proven at hyperscale (Vultr's 2026 global AI buildout)

HPE AI Factory for sovereigns

Best for regulated industries, government agencies, and national AI programs that require in-country control of data and models.

GB300 NVL72 delivered with data sovereignty, in-country lifecycle control, and compliance alignment

HPE AI Factory at-scale

Best for service providers, model builders, and hyperscale operators needing cloud-like, multi-tenant operations across many racks.

Multi-rack GB300 NVL72 deployments networked together for hundreds to tens of thousands of GPUs

At a glance

GPUs
72 NVIDIA Blackwell Ultra GPUs
CPUs
36 NVIDIA Grace CPUs (2,592 Arm Neoverse V2 cores)
GPU memory
288 GB HBM3e per GPU, ~20 TB total
GPU interconnect
5th-gen NVIDIA NVLink, 130 TB/s aggregate
Compute performance
1,080.2 PFLOPS dense FP4 Tensor Core; 720 PFLOPS FP8/FP6
Networking
ConnectX-8 SuperNIC, 800 Gb/s per GPU; Quantum-X800 InfiniBand or Spectrum-X Ethernet
Rack layout
18 compute trays (2 GB300 boards each) + NVLink switch trays
Power
132 kW nominal TDP per rack, ~155 kW peak EDPp
Cooling
Direct liquid cooling, ~90% heat to liquid / 10% to air
Availability
Shipping since December 2025

How to buy NVIDIA GB300 NVL72 by HPE

The GB300 NVL72 is sold as a complete HPE-integrated rack, not a build-your-own GPU server, so procurement runs through a scoped configuration rather than a catalog SKU pick.

Direct capital purchase

Order the rack-scale system outright through an authorized HPE partner. We handle the bill of materials (compute trays, NVLink switch trays, power shelves, networking, services) and align it to TAA-compliant configurations where required.

HPE GreenLake consumption

Consume the GB300 NVL72 as a metered, pay-per-use service under HPE GreenLake instead of a capital purchase, useful if you want AI Factory economics without carrying the full asset on your books.

Federal and public-sector contract vehicles

GSA MAS (application in progress), SAP/FAR channels, and GPC direct are common paths for agencies buying HPE AI infrastructure; GPC (Government Purchase Card) can cover smaller add-on components and services within threshold.

HPE support and lifecycle services attach

Pointnext Tech Care or Complete Care support tiers, deployment services, and ongoing operations are quoted alongside the hardware since a 132 kW liquid-cooled rack is not a self-install product.

Multi-rack AI Factory programs

Sovereign, service-provider, and hyperscale buyers typically negotiate multi-rack or multi-year programs rather than a single-unit purchase; we help structure phased delivery against facility readiness.

Where it fits

Training and fine-tuning frontier and reasoning models above 1 trillion parameters
Real-time, high-throughput generative AI inference at production scale
Sovereign and national AI factory programs requiring in-country data control
Cloud and hyperscale AI infrastructure buildouts (multi-tenant, GPU-as-a-service)
Converged HPC and AI research workloads in federal and higher-education labs
Multi-rack AI clusters scaling from a single NVLink domain to tens of thousands of GPUs

Frequently asked

What exactly ships in a GB300 NVL72 rack from HPE?

One integrated rack: 18 liquid-cooled compute trays (36 NVIDIA Grace CPUs and 72 Blackwell Ultra GPUs across 2 GB300 boards per tray), liquid-cooled NVLink switch trays for the 72-GPU NVLink domain, 8 power shelves on a DC bus bar, and your choice of NVIDIA Quantum-X800 InfiniBand or Spectrum-X Ethernet networking with ConnectX-8 SuperNICs. HPE integrates it as a single delivered unit rather than a parts list you assemble on-site.

How do we procure the GB300 NVL72 through Uniqcli?

As an authorized HPE partner, we scope and quote the GB300 NVL72 rack and the surrounding AI Factory configuration for you, including networking, services, and support. We can source TAA-compliant configurations and support purchases through GSA MAS (application in progress), SAP/FAR channels, and GPC direct, plus GPC for smaller line items, alongside HPE GreenLake consumption if you prefer an as-a-service model over a capital purchase.

What are the power and cooling requirements for one rack?

Plan for roughly 132 kW nominal thermal design power per rack, with electrical design power peaking around 155 kW and HPE recommending bus way provisioning up to 192 kW EDPp. Cooling is direct liquid cooling at the rack, capturing about 90% of heat to liquid and 10% to air. This is not a retrofit into a standard air-cooled row: your facility needs liquid cooling distribution (CDUs) and power delivery sized for this density before the rack lands.

How does GB300 NVL72 compare to the prior GB200 NVL72 generation?

GB300 swaps in Blackwell Ultra GPUs with 288 GB of HBM3e memory per GPU (up from 192 GB), delivering roughly 1.5x the dense FP4 compute and materially higher per-GPU token throughput on large reasoning models. Each compute tray also shifts to 4 GPUs plus 2 Grace CPUs, versus 2 GPUs per tray on GB200. If you already run GB200 NVL72, GB300 is a memory and inference-throughput upgrade on the same NVLink rack-scale architecture, not a different platform to relearn.

How do we size a GB300 NVL72 deployment for our workload?

Sizing starts with your workloads: model sizes, training versus inference mix, context length, dataset volumes, and concurrent user counts. From there we map GPU count, whether a single 72-GPU NVLink domain covers it or you need multiple racks networked together, and which fabric (InfiniBand or Ethernet) fits your operations team. We can arrange access to HPE and NVIDIA technical resources to validate the design before you commit budget.

Is this suitable for sovereign, federal, or other regulated environments?

Yes. HPE positions the GB300 NVL72 within its AI Factory portfolio for service providers, sovereign entities, and large enterprises, with configurations aimed at regulated industries and national AI programs that need in-country data control. We help align the specific configuration, contract vehicle, and support terms to your regulatory and security requirements as an authorized HPE partner.

Should we buy GB300 NVL72 now or wait for the next NVIDIA platform?

GB300 NVL72 is shipping today and running in production at hyperscale (Vultr's global buildout is one public example), while HPE has already announced its successor, the Vera Rubin NVL72 platform, for December 2026 availability. If you need trillion-parameter-class capacity in 2026, GB300 is the proven, supportable choice now; if your timeline allows waiting a year and you want the newest silicon, it's worth discussing both paths with us before you commit.

Where do we get a quote for a GB300 NVL72 deployment?

Submit your workload details and target timeline through our quote request, and we'll come back with a scoped configuration covering the rack, networking fabric, facility prerequisites, and support tier, along with applicable contract-vehicle pricing for federal, SLED, healthcare, or enterprise procurement.

Works with

The GB300 NVL72 rarely stands alone. It anchors an AI Factory buildout that pairs with HPE's broader compute, storage, and networking portfolio, plus the services and comparisons that help you scope the right tier before committing.

Build your HPE bill of materials.

Send us the requirement, the project, or an existing quote to beat. We come back with a validated, TAA-compliant HPE configuration and a real price, often below list.

connect [at] getuniqcli.com · Chicago, IL