GB200 NVL72 Rack-Scale Architecture
The NVIDIA GB200 and GB300 NVL72 represents a paradigm-shifting rack-scale computing platform, treating 72 Blackwell GPUs and 36 Grace CPUs as a single,
Source: mortalapps.com- The NVIDIA GB200 and GB300 NVL72 represents a paradigm-shifting rack-scale computing platform, treating 72 Blackwell GPUs and 36 Grace CPUs as a single, unified computational entity.
- The architecture utilizes nine dedicated NVLink Switch Trays and a highly advanced passive copper backplane to generate 130 TB/s of fully non-blocking, all-to-all bandwidth.
- It enforces strict network isolation, implementing separate high-speed fabrics for East/West compute traffic (ConnectX-8 SuperNICs) and North/South storage and management traffic (BlueField-3 DPUs).
- Unprecedented power demands, peaking at approximately 142 kW per rack, necessitate a fully integrated, modular liquid-cooling design built upon the MGX architecture.
Why This Matters
As advanced generative models scale aggressively past the one trillion parameter threshold, the traditional 8-GPU server node boundary introduces severe pipeline and tensor parallel fragmentation. The NVL72 architecture shatters this limitation by pushing the scale-up domain from 8 GPUs to 72 GPUs, enabling real-time inference for trillion-parameter models at speeds 30x faster than previous architectures. It fundamentally shifts the basic unit of infrastructure provisioning from the individual server chassis to the entire integrated rack.
Core Intuition
Instead of constructing a cluster by manually wiring discrete servers together via fragile and complex optical transceivers, the NVL72 architecture conceptually transforms the entire physical rack into a single, massive motherboard. The copper backplane of the rack functions directly as the NVLink circuit board. Every individual compute tray plugs straight into this spine. Because electrical signals travel mere inches over dense copper instead of traversing meters over optical fiber, the system achieves unprecedented throughput and exceptional energy efficiency by eliminating the need for signal retimers and power-hungry optical conversions within the L1 compute domain.
Technical Deep Dive
The NVL72 compute rack houses eighteen 1RU compute trays. Each tray is populated with two NVIDIA Grace CPUs (each boasting 72 ARM Neoverse V2 cores and linked via NVLink C2C) paired with four Blackwell B300 GPUs, yielding an aggregate,152 GB of HBM3 memory per tray.
The networking architecture is strictly bifurcated. The East/West Compute Network features four dual-port ConnectX-8 SuperNICs integrated directly onto the compute baseboard, guaranteeing a mathematically perfect 1:1 GPU-to-NIC ratio. This provides up to 800 Gb/s Ethernet or InfiniBand throughput per GPU for massive scale-out communication. Conversely, the North/South Fabric relies on a single dual-port QSFP112 BlueField-3 B3240 DPU per tray, offloading storage orchestration, Out-of-Band (OOB) management, and enforcing zero-trust security parameters independently of the host CPU.
Nine NVLink Switch Trays (containing two NVSwitch ASICs each) govern the internal fabric. Every one of the 72 GPUs pushes 18 physical NVLink connections directly into the copper backplane, routing exactly one link to each of the 18 NVSwitches to construct the 130 TB/s non-blocking domain.
| Component | Quantity per NVL72 Rack |
|---|---|
| Key Specification | Compute Trays |
| 18 (1RU) | 2 Grace CPUs, 4 Blackwell GPUs per tray |
| NVLink Switch Trays | 9 (1RU) |
| 2 NVSwitch ASICs per tray, 130 TB/s total BW | ConnectX-8 SuperNICs |
| 800 Gb/s per NIC (1:1 GPU ratio) | BlueField-3 DPUs |
| Dual-port QSFP112, Storage/Management | Power Delivery |
| 8 Shelves | ~142 kW total draw, liquid-cooled 26 |