AI fabric planning
NVIDIA AI Server Networking Requirements: Fabric and Storage Planner
NVIDIA AI server networking requirements depend on what crosses the wire: GPU collectives, parameter or expert traffic, model and dataset reads, checkpoints, client requests and management. Link speed alone is not an architecture. The design needs traffic classes, topology, oversubscription targets, NIC placement, switching, cabling and failure domains that match the actual workload.
Quick answer
What to size before you buy
Separate scale-out compute traffic, storage traffic, client or north-south traffic and management. Size each path from measured demand, then verify NIC, PCIe, switch and cable/optic support as one end-to-end system.
Current Amazon listings
Supporting hardware matched into separate catalogue classes
Live product cards are discovery aids for supporting infrastructure. They do not imply NVIDIA, OEM or facility certification. Exact model, condition, interface, warranty and compatibility must be verified before purchase.
Technical decision
Turn the requirement into a measurable decision
Move beyond 25/100GbE when measured traffic or future node count justifies it, not because a faster port exists. Distributed training and rack-scale NVIDIA platforms can need very high-bandwidth fabrics, while many inference servers are constrained elsewhere.
Interactive planning tool
NVIDIA AI Network Bandwidth Planner
Use this as a screening calculation. It does not certify a server, predict benchmark performance, design high-voltage electrical work, or replace the current OEM and facility documentation.
Before you buy
Four checks that keep planning estimates in context
Start with current documentation
Use the exact platform or OEM system guide as the source of truth for supported configurations and limits.
Keep assumptions visible
Every calculator input is an assumption until it is replaced by a measurement, vendor limit or facility design value.
Separate nameplate from application performance
Port speed, SSD peak rate, GPU memory and power ratings do not guarantee end-to-end workload results.
Escalate facility decisions
High-voltage distribution, rack electrical work, cooling design and liquid loops require qualified professionals and current codes.
Classify traffic before choosing ports
AI clusters carry several traffic types with different latency, loss and burst characteristics. Treating them as one stream makes capacity calculations misleading. This boundary belongs in the NVIDIA AI Server Networking acceptance plan.
For NVIDIA AI Server Networking, list collective, storage, inference ingress, management, backup and telemetry flows with source, destination and peak rate. Recheck it after material changes. A pass/fail note for classify traffic before choosing ports belongs in the NVIDIA AI Server Networking commissioning record.
Understand scale-up versus scale-out
NVLink or other local interconnect handles tightly coupled accelerator communication inside supported systems; Ethernet or InfiniBand connects nodes and racks. Those domains solve different problems. This boundary belongs in the NVIDIA AI Server Networking acceptance plan.
For NVIDIA AI Server Networking, draw the boundary of each scale-up island and calculate what actually leaves it during the workload. Recheck it after material changes. A pass/fail note for understand scale-up versus scale-out belongs in the NVIDIA AI Server Networking commissioning record.
Estimate east-west demand from the parallelism strategy
Data parallel, tensor parallel, pipeline parallel and expert parallel patterns move different amounts of data at different synchronization points. This boundary belongs in the NVIDIA AI Server Networking acceptance plan.
For NVIDIA AI Server Networking, profile the framework with representative model size and batch settings rather than deriving fabric bandwidth solely from GPU count. Recheck it after material changes. A pass/fail note for estimate east-west demand from the parallelism strategy belongs in the NVIDIA AI Server Networking commissioning record.
Protect storage traffic from collective bursts
Dataset reads and checkpoints can overlap with compute communication. Shared links may become the hidden throttle even when average utilization looks low. This boundary belongs in the NVIDIA AI Server Networking acceptance plan.
For NVIDIA AI Server Networking, capture peak concurrent traffic and decide whether separate storage interfaces, VLAN/QoS policy or a separate fabric is warranted. Recheck it after material changes. A pass/fail note for protect storage traffic from collective bursts belongs in the NVIDIA AI Server Networking commissioning record.
Check NIC-to-CPU and NIC-to-GPU locality
A high-speed NIC can lose efficiency when it sits behind a constrained PCIe switch or crosses NUMA boundaries unexpectedly. GPUDirect paths also depend on supported topology and software. This boundary belongs in the NVIDIA AI Server Networking acceptance plan.
For NVIDIA AI Server Networking, map NIC and GPU slots to PCIe roots, confirm ACS/IOMMU behavior where relevant, and benchmark the actual placement. Recheck it after material changes. A pass/fail note for check nic-to-cpu and nic-to-gpu locality belongs in the NVIDIA AI Server Networking commissioning record.
Treat RDMA as an engineered feature
RDMA can reduce CPU overhead and latency, but it introduces requirements around NIC support, drivers, fabric configuration and operational discipline. This boundary belongs in the NVIDIA AI Server Networking acceptance plan.
For NVIDIA AI Server Networking, use the vendor’s supported configuration and validate lossless or congestion-control behavior appropriate to the selected technology before production. Recheck it after material changes. A pass/fail note for treat rdma as an engineered feature belongs in the NVIDIA AI Server Networking commissioning record.
Calculate oversubscription explicitly
Leaf-spine or multi-plane fabrics can have enough edge ports yet insufficient uplink capacity when many nodes communicate simultaneously. This boundary belongs in the NVIDIA AI Server Networking acceptance plan.
For NVIDIA AI Server Networking, compute worst-case edge-to-uplink ratios per failure domain and verify that the intended collective pattern fits them. Recheck it after material changes. A pass/fail note for calculate oversubscription explicitly belongs in the NVIDIA AI Server Networking commissioning record.
Count ports, optics and cables early
High-speed designs can be constrained by switch radix, OSFP/QSFP choices, breakout support, fiber type, DAC reach and power or thermal limits of optics. This boundary belongs in the NVIDIA AI Server Networking acceptance plan.
For NVIDIA AI Server Networking, build a port-level cable schedule before ordering so the logical topology can actually be wired and serviced. Recheck it after material changes. A pass/fail note for count ports, optics and cables early belongs in the NVIDIA AI Server Networking commissioning record.
Keep management independent enough to recover
When the main fabric fails, operators still need BMC, switch and orchestration access. Putting every control path on the same failure domain can make repair much harder. This boundary belongs in the NVIDIA AI Server Networking acceptance plan.
For NVIDIA AI Server Networking, design an out-of-band management path and test recovery with the production fabric intentionally unavailable. Recheck it after material changes. A pass/fail note for keep management independent enough to recover belongs in the NVIDIA AI Server Networking commissioning record.
Plan for fabric growth without disruptive recabling
Adding racks can consume spare spine capacity or require a new topology stage. The initial design should state its supported expansion point. This boundary belongs in the NVIDIA AI Server Networking acceptance plan.
For NVIDIA AI Server Networking, reserve ports, power and physical pathways for the next increment, and document the trigger for adding another fabric plane or switch tier. Recheck it after material changes. A pass/fail note for plan for fabric growth without disruptive recabling belongs in the NVIDIA AI Server Networking commissioning record.
Monitor errors, not only utilization
CRC errors, drops, retransmits, congestion and pause behavior can harm distributed jobs long before link utilization reaches 100 percent. This boundary belongs in the NVIDIA AI Server Networking acceptance plan.
For NVIDIA AI Server Networking, collect NIC and switch counters with job telemetry so performance regressions can be tied to network events. Recheck it after material changes. A pass/fail note for monitor errors, not only utilization belongs in the NVIDIA AI Server Networking commissioning record.
Benchmark the communication pattern
iperf-style throughput confirms a link but does not represent all-reduce, all-to-all or storage I/O behavior. This boundary belongs in the NVIDIA AI Server Networking acceptance plan.
For NVIDIA AI Server Networking, use NCCL tests and application-level runs at the intended node count, then compare results against a documented acceptance envelope. Recheck it after material changes. A pass/fail note for benchmark the communication pattern belongs in the NVIDIA AI Server Networking commissioning record.
Methodology and official references
The network planner uses requested node count and per-node traffic to expose aggregate demand and oversubscription. It does not convert an Ethernet or InfiniBand port rate into guaranteed NCCL or application throughput. Current NVIDIA platform and fabric documentation remain authoritative.
- NVIDIA Vera Rubin NVL72
- NVIDIA GB300 NVL72
- NVIDIA GB200 NVL72
- NVIDIA NVL72 AI Factory reference architecture
- NVIDIA NVL72 node configurations
- NVIDIA NVL72 logical network architecture
- NVIDIA Rubin platform
As an Amazon Associate, Cloudzat may earn from qualifying purchases. Marketplace listings are supporting-hardware discovery, not certification. Product revisions, firmware, software, electrical limits, thermals, topology and workload behavior can change results; verify the exact hardware and current vendor documentation before purchase.
Frequently asked questions
What should I know about “Classify traffic before choosing ports”?
AI clusters carry several traffic types with different latency, loss and burst characteristics. Treating them as one stream makes capacity calculations misleading. To address “Classify traffic before choosing ports”, list collective, storage, inference ingress, management, backup and telemetry flows with source, destination and peak rate. Test that result on NVIDIA AI Server Networking.
How should I validate “Understand scale-up versus scale-out”?
NVLink or other local interconnect handles tightly coupled accelerator communication inside supported systems; Ethernet or InfiniBand connects nodes and racks. Those domains solve different problems. To address “Understand scale-up versus scale-out”, draw the boundary of each scale-up island and calculate what actually leaves it during the workload. Test that result on NVIDIA AI Server Networking.
Why does “Estimate east-west demand from the parallelism strategy” affect the final design?
Data parallel, tensor parallel, pipeline parallel and expert parallel patterns move different amounts of data at different synchronization points. To address “Estimate east-west demand from the parallelism strategy”, profile the framework with representative model size and batch settings rather than deriving fabric bandwidth solely from GPU count. Test that result on NVIDIA AI Server Networking.
Which measurement matters most for “Protect storage traffic from collective bursts”?
Dataset reads and checkpoints can overlap with compute communication. Shared links may become the hidden throttle even when average utilization looks low. To address “Protect storage traffic from collective bursts”, capture peak concurrent traffic and decide whether separate storage interfaces, VLAN/QoS policy or a separate fabric is warranted. Test that result on NVIDIA AI Server Networking.
When can “Check NIC-to-CPU and NIC-to-GPU locality” become a bottleneck?
A high-speed NIC can lose efficiency when it sits behind a constrained PCIe switch or crosses NUMA boundaries unexpectedly. GPUDirect paths also depend on supported topology and software. To address “Check NIC-to-CPU and NIC-to-GPU locality”, map NIC and GPU slots to PCIe roots, confirm ACS/IOMMU behavior where relevant, and benchmark the actual placement. Test that result on NVIDIA AI Server Networking.
How much reserve is appropriate for “Treat RDMA as an engineered feature”?
RDMA can reduce CPU overhead and latency, but it introduces requirements around NIC support, drivers, fabric configuration and operational discipline. To address “Treat RDMA as an engineered feature”, use the vendor’s supported configuration and validate lossless or congestion-control behavior appropriate to the selected technology before production. Test that result on NVIDIA AI Server Networking.
Can extra hardware solve “Calculate oversubscription explicitly” by itself?
Leaf-spine or multi-plane fabrics can have enough edge ports yet insufficient uplink capacity when many nodes communicate simultaneously. To address “Calculate oversubscription explicitly”, compute worst-case edge-to-uplink ratios per failure domain and verify that the intended collective pattern fits them. Test that result on NVIDIA AI Server Networking.
What should be documented for “Count ports, optics and cables early”?
High-speed designs can be constrained by switch radix, OSFP/QSFP choices, breakout support, fiber type, DAC reach and power or thermal limits of optics. To address “Count ports, optics and cables early”, build a port-level cable schedule before ordering so the logical topology can actually be wired and serviced. Test that result on NVIDIA AI Server Networking.
How should “Keep management independent enough to recover” be tested before production?
When the main fabric fails, operators still need BMC, switch and orchestration access. Putting every control path on the same failure domain can make repair much harder. To address “Keep management independent enough to recover”, design an out-of-band management path and test recovery with the production fabric intentionally unavailable. Test that result on NVIDIA AI Server Networking.
How does growth change the plan for “Plan for fabric growth without disruptive recabling”?
Adding racks can consume spare spine capacity or require a new topology stage. The initial design should state its supported expansion point. To address “Plan for fabric growth without disruptive recabling”, reserve ports, power and physical pathways for the next increment, and document the trigger for adding another fabric plane or switch tier. Test that result on NVIDIA AI Server Networking.