NVIDIA HGX AI Factory
NVIDIA HGX AI Factory: Reference Architecture and Sizing Guide
Size the overall planning boundary in NVIDIA HGX AI Factory from the HGX node outward. Multiply HGX server and GPU count across the target node count, then check whether per-server network fabric capacity scales at the same rate.
Surface server-generation mismatch early because HGX server topology, NIC count, PCIe layout, power, and cooling can differ by generation and OEM. Build a repeatable node profile before estimating the whole cluster. Validate the node profile through current NVIDIA HGX reference architecture.
For enterprise HGX cluster architects, the profile should include GPU count, network rails, local and shared storage paths, firmware baseline, rack power, and thermal constraints. Expansion is safer when new racks repeat a verified unit and the team knows exactly which shared services require additional capacity. Avoid assuming that doubling servers automatically doubles usable AI throughput.
Quick answer
What NVIDIA HGX AI Factory should settle first
Size the first decision gate in NVIDIA HGX AI Factory from the HGX node outward. Multiply HGX server and GPU count across the target node count, then check whether rack power and storage throughput scales at the same rate.
Surface PCIe or fabric topology variation early because HGX server topology, NIC count, PCIe layout, power, and cooling can differ by generation and OEM.
Current Amazon listings
Supporting hardware for nvidia ai factory & dsx
Live product cards are discovery aids for the planning workflow. They do not certify a complete architecture. Verify exact model, condition, interface, warranty, firmware, compatibility and seller details before purchase.
Technical decision
Turn NVIDIA HGX AI Factory into a verified design
Validate the node profile through network rail topology. For enterprise HGX cluster architects, the profile should include GPU count, network rails, local and shared storage paths, firmware baseline, rack power, and thermal constraints. Expansion is safer when new racks repeat a verified unit and the team knows exactly which shared services require additional capacity.
Decision table
NVIDIA HGX AI Factory planning inputs and verification
| Planning item | Why it matters | Verify with |
|---|---|---|
| Hgx server and gpu count | Controls the capacity boundary and can expose server-generation mismatch. | current NVIDIA HGX reference architecture |
| Per-server network fabric capacity | Controls the throughput boundary and can expose PCIe or fabric topology variation. | selected certified server datasheet |
| Rack power and storage throughput | Controls the fit boundary and can expose network rail oversubscription. | network rail topology |
| Hgx server and gpu count | Controls the resilience boundary and can expose rack density mismatch. | storage performance validation |
| Per-server network fabric capacity | Controls the facility boundary and can expose storage not scaling with GPU count. | rack electrical and thermal design |
Interactive planning tool
NVIDIA HGX AI Factory Sizing Screen
Use this as a screening calculation. It does not certify a design, guarantee benchmark performance, replace a provider quote, or override current OEM, software, network or facility documentation.
Before you buy
Four checks that keep planning estimates in context
Start with current documentation
Use the exact platform or OEM system guide as the source of truth for supported configurations and limits.
Keep assumptions visible
Every calculator input is an assumption until it is replaced by a measurement, vendor limit or facility design value.
Separate nameplate from application performance
Port speed, SSD peak rate, GPU memory and power ratings do not guarantee end-to-end workload results.
Escalate facility decisions
High-voltage distribution, rack electrical work, cooling design and liquid loops require qualified professionals and current codes.
Define the HGX server generation
Size define the hgx server generation in NVIDIA HGX AI Factory from the HGX node outward. Multiply HGX server and GPU count across the target node count, then check whether per-server network fabric capacity scales at the same rate.
Surface server-generation mismatch early because HGX server topology, NIC count, PCIe layout, power, and cooling can differ by generation and OEM. Build a repeatable node profile before estimating the whole cluster.
Validate the node profile through current NVIDIA HGX reference architecture. For enterprise HGX cluster architects, the profile should include GPU count, network rails, local and shared storage paths, firmware baseline, rack power, and thermal constraints.
Expansion is safer when new racks repeat a verified unit and the team knows exactly which shared services require additional capacity. Avoid assuming that doubling servers automatically doubles usable AI throughput.
Choose a validated certified system
Size choose a validated certified system in NVIDIA HGX AI Factory from the HGX node outward. Multiply per-server network fabric capacity across the target node count, then check whether rack power and storage throughput scales at the same rate.
Surface PCIe or fabric topology variation early because HGX server topology, NIC count, PCIe layout, power, and cooling can differ by generation and OEM. Build a repeatable node profile before estimating the whole cluster.
Validate the node profile through selected certified server datasheet. For enterprise HGX cluster architects, the profile should include GPU count, network rails, local and shared storage paths, firmware baseline, rack power, and thermal constraints. Expansion is safer when new racks repeat a verified unit and the team knows exactly which shared services require additional capacity.
Avoid assuming that doubling servers automatically doubles usable AI throughput.
Map GPUs and local CPU resources
Size map gpus and local cpu resources in NVIDIA HGX AI Factory from the HGX node outward. Multiply rack power and storage throughput across the target node count, then check whether HGX server and GPU count scales at the same rate.
Surface network rail oversubscription early because HGX server topology, NIC count, PCIe layout, power, and cooling can differ by generation and OEM. Build a repeatable node profile before estimating the whole cluster.
Validate the node profile through network rail topology. For enterprise HGX cluster architects, the profile should include GPU count, network rails, local and shared storage paths, firmware baseline, rack power, and thermal constraints. Expansion is safer when new racks repeat a verified unit and the team knows exactly which shared services require additional capacity.
Avoid assuming that doubling servers automatically doubles usable AI throughput.
Design scale-out networking
Size design scale-out networking in NVIDIA HGX AI Factory from the HGX node outward. Multiply HGX server and GPU count across the target node count, then check whether per-server network fabric capacity scales at the same rate.
Surface rack density mismatch early because HGX server topology, NIC count, PCIe layout, power, and cooling can differ by generation and OEM. Build a repeatable node profile before estimating the whole cluster.
Validate the node profile through storage performance validation. For enterprise HGX cluster architects, the profile should include GPU count, network rails, local and shared storage paths, firmware baseline, rack power, and thermal constraints. Expansion is safer when new racks repeat a verified unit and the team knows exactly which shared services require additional capacity.
Avoid assuming that doubling servers automatically doubles usable AI throughput.
Plan local and shared storage
Size plan local and shared storage in NVIDIA HGX AI Factory from the HGX node outward. Multiply per-server network fabric capacity across the target node count, then check whether rack power and storage throughput scales at the same rate.
Surface storage not scaling with GPU count early because HGX server topology, NIC count, PCIe layout, power, and cooling can differ by generation and OEM. Build a repeatable node profile before estimating the whole cluster.
Validate the node profile through rack electrical and thermal design. For enterprise HGX cluster architects, the profile should include GPU count, network rails, local and shared storage paths, firmware baseline, rack power, and thermal constraints.
Expansion is safer when new racks repeat a verified unit and the team knows exactly which shared services require additional capacity. Avoid assuming that doubling servers automatically doubles usable AI throughput.
Translate nodes into rack density
Size translate nodes into rack density in NVIDIA HGX AI Factory from the HGX node outward. Multiply rack power and storage throughput across the target node count, then check whether HGX server and GPU count scales at the same rate.
Surface server-generation mismatch early because HGX server topology, NIC count, PCIe layout, power, and cooling can differ by generation and OEM. Build a repeatable node profile before estimating the whole cluster.
Validate the node profile through current NVIDIA HGX reference architecture. For enterprise HGX cluster architects, the profile should include GPU count, network rails, local and shared storage paths, firmware baseline, rack power, and thermal constraints.
Expansion is safer when new racks repeat a verified unit and the team knows exactly which shared services require additional capacity. Avoid assuming that doubling servers automatically doubles usable AI throughput.
Plan power and thermal headroom
Size plan power and thermal headroom in NVIDIA HGX AI Factory from the HGX node outward. Multiply HGX server and GPU count across the target node count, then check whether per-server network fabric capacity scales at the same rate.
Surface PCIe or fabric topology variation early because HGX server topology, NIC count, PCIe layout, power, and cooling can differ by generation and OEM. Build a repeatable node profile before estimating the whole cluster.
Validate the node profile through selected certified server datasheet. For enterprise HGX cluster architects, the profile should include GPU count, network rails, local and shared storage paths, firmware baseline, rack power, and thermal constraints. Expansion is safer when new racks repeat a verified unit and the team knows exactly which shared services require additional capacity.
Avoid assuming that doubling servers automatically doubles usable AI throughput.
Design failure domains and spares
Size design failure domains and spares in NVIDIA HGX AI Factory from the HGX node outward. Multiply per-server network fabric capacity across the target node count, then check whether rack power and storage throughput scales at the same rate.
Surface network rail oversubscription early because HGX server topology, NIC count, PCIe layout, power, and cooling can differ by generation and OEM. Build a repeatable node profile before estimating the whole cluster.
Validate the node profile through network rail topology. For enterprise HGX cluster architects, the profile should include GPU count, network rails, local and shared storage paths, firmware baseline, rack power, and thermal constraints. Expansion is safer when new racks repeat a verified unit and the team knows exactly which shared services require additional capacity.
Avoid assuming that doubling servers automatically doubles usable AI throughput.
Standardize firmware and software
Size standardize firmware and software in NVIDIA HGX AI Factory from the HGX node outward. Multiply rack power and storage throughput across the target node count, then check whether HGX server and GPU count scales at the same rate.
Surface rack density mismatch early because HGX server topology, NIC count, PCIe layout, power, and cooling can differ by generation and OEM. Build a repeatable node profile before estimating the whole cluster.
Validate the node profile through storage performance validation. For enterprise HGX cluster architects, the profile should include GPU count, network rails, local and shared storage paths, firmware baseline, rack power, and thermal constraints. Expansion is safer when new racks repeat a verified unit and the team knows exactly which shared services require additional capacity.
Avoid assuming that doubling servers automatically doubles usable AI throughput.
Validate cluster management
Size validate cluster management in NVIDIA HGX AI Factory from the HGX node outward. Multiply HGX server and GPU count across the target node count, then check whether per-server network fabric capacity scales at the same rate.
Surface storage not scaling with GPU count early because HGX server topology, NIC count, PCIe layout, power, and cooling can differ by generation and OEM. Build a repeatable node profile before estimating the whole cluster.
Validate the node profile through rack electrical and thermal design. For enterprise HGX cluster architects, the profile should include GPU count, network rails, local and shared storage paths, firmware baseline, rack power, and thermal constraints.
Expansion is safer when new racks repeat a verified unit and the team knows exactly which shared services require additional capacity. Avoid assuming that doubling servers automatically doubles usable AI throughput.
Scale by repeatable units
Size scale by repeatable units in NVIDIA HGX AI Factory from the HGX node outward. Multiply per-server network fabric capacity across the target node count, then check whether rack power and storage throughput scales at the same rate.
Surface server-generation mismatch early because HGX server topology, NIC count, PCIe layout, power, and cooling can differ by generation and OEM. Build a repeatable node profile before estimating the whole cluster.
Validate the node profile through current NVIDIA HGX reference architecture. For enterprise HGX cluster architects, the profile should include GPU count, network rails, local and shared storage paths, firmware baseline, rack power, and thermal constraints.
Expansion is safer when new racks repeat a verified unit and the team knows exactly which shared services require additional capacity. Avoid assuming that doubling servers automatically doubles usable AI throughput.
Create the HGX AI factory baseline
Size create the hgx ai factory baseline in NVIDIA HGX AI Factory from the HGX node outward. Multiply rack power and storage throughput across the target node count, then check whether HGX server and GPU count scales at the same rate.
Surface PCIe or fabric topology variation early because HGX server topology, NIC count, PCIe layout, power, and cooling can differ by generation and OEM. Build a repeatable node profile before estimating the whole cluster.
Validate the node profile through selected certified server datasheet. For enterprise HGX cluster architects, the profile should include GPU count, network rails, local and shared storage paths, firmware baseline, rack power, and thermal constraints. Expansion is safer when new racks repeat a verified unit and the team knows exactly which shared services require additional capacity.
Avoid assuming that doubling servers automatically doubles usable AI throughput.
Methodology and official references
Size the validation method in NVIDIA HGX AI Factory from the HGX node outward. Multiply rack power and storage throughput across the target node count, then check whether HGX server and GPU count scales at the same rate.
Surface rack density mismatch early because HGX server topology, NIC count, PCIe layout, power, and cooling can differ by generation and OEM. Build a repeatable node profile before estimating the whole cluster. Validate the node profile through storage performance validation.
For enterprise HGX cluster architects, the profile should include GPU count, network rails, local and shared storage paths, firmware baseline, rack power, and thermal constraints. Expansion is safer when new racks repeat a verified unit and the team knows exactly which shared services require additional capacity. Avoid assuming that doubling servers automatically doubles usable AI throughput.
As an Amazon Associate, Cloudzat may earn from qualifying purchases. Marketplace listings are supporting-hardware discovery, not certification. Product revisions, firmware, software, electrical limits, thermals, topology and workload behavior can change results; verify the exact hardware and current vendor documentation before purchase.
Frequently asked questions
What should I verify first for NVIDIA HGX AI factory?
Size FAQ checkpoint 1 for NVIDIA HGX AI Factory in NVIDIA HGX AI Factory from the HGX node outward. Multiply HGX server and GPU count across the target node count, then check whether per-server network fabric capacity scales at the same rate.
Surface network rail oversubscription early because HGX server topology, NIC count, PCIe layout, power, and cooling can differ by generation and OEM. Build a repeatable node profile before estimating the whole cluster. NVIDIA HGX AI Factory checkpoint 1 retains selected certified server datasheet; the following NVIDIA HGX AI Factory review tracks rack density mismatch.
Which NVIDIA HGX AI factory values should be treated as NVIDIA-published facts?
Validate the node profile through rack electrical and thermal design. For enterprise HGX cluster architects, the profile should include GPU count, network rails, local and shared storage paths, firmware baseline, rack power, and thermal constraints.
Expansion is safer when new racks repeat a verified unit and the team knows exactly which shared services require additional capacity. Avoid assuming that doubling servers automatically doubles usable AI throughput. NVIDIA HGX AI Factory checkpoint 2 retains network rail topology; the following NVIDIA HGX AI Factory review tracks storage not scaling with GPU count.
How should I use the NVIDIA HGX AI Factory calculator?
Size FAQ checkpoint 3 for NVIDIA HGX AI Factory in NVIDIA HGX AI Factory from the HGX node outward. Multiply rack power and storage throughput across the target node count, then check whether HGX server and GPU count scales at the same rate.
Surface storage not scaling with GPU count early because HGX server topology, NIC count, PCIe layout, power, and cooling can differ by generation and OEM. Build a repeatable node profile before estimating the whole cluster. NVIDIA HGX AI Factory checkpoint 3 retains storage performance validation; the following NVIDIA HGX AI Factory review tracks server-generation mismatch.
What is the most common sizing mistake for NVIDIA HGX AI Factory?
Validate the node profile through selected certified server datasheet. For enterprise HGX cluster architects, the profile should include GPU count, network rails, local and shared storage paths, firmware baseline, rack power, and thermal constraints. Expansion is safer when new racks repeat a verified unit and the team knows exactly which shared services require additional capacity.
Avoid assuming that doubling servers automatically doubles usable AI throughput. NVIDIA HGX AI Factory checkpoint 4 retains rack electrical and thermal design; the following NVIDIA HGX AI Factory review tracks PCIe or fabric topology variation.
How should networking be validated for NVIDIA HGX AI Factory?
Size FAQ checkpoint 5 for NVIDIA HGX AI Factory in NVIDIA HGX AI Factory from the HGX node outward. Multiply per-server network fabric capacity across the target node count, then check whether rack power and storage throughput scales at the same rate.
Surface PCIe or fabric topology variation early because HGX server topology, NIC count, PCIe layout, power, and cooling can differ by generation and OEM. Build a repeatable node profile before estimating the whole cluster.
NVIDIA HGX AI Factory checkpoint 5 retains current NVIDIA HGX reference architecture; the following NVIDIA HGX AI Factory review tracks network rail oversubscription.
How should storage and memory headroom be planned for NVIDIA HGX AI Factory?
Validate the node profile through storage performance validation. For enterprise HGX cluster architects, the profile should include GPU count, network rails, local and shared storage paths, firmware baseline, rack power, and thermal constraints. Expansion is safer when new racks repeat a verified unit and the team knows exactly which shared services require additional capacity.
Avoid assuming that doubling servers automatically doubles usable AI throughput. NVIDIA HGX AI Factory checkpoint 6 retains selected certified server datasheet; the following NVIDIA HGX AI Factory review tracks rack density mismatch.
How should power and cooling be handled for NVIDIA HGX AI Factory?
Size FAQ checkpoint 7 for NVIDIA HGX AI Factory in NVIDIA HGX AI Factory from the HGX node outward. Multiply HGX server and GPU count across the target node count, then check whether per-server network fabric capacity scales at the same rate.
Surface rack density mismatch early because HGX server topology, NIC count, PCIe layout, power, and cooling can differ by generation and OEM. Build a repeatable node profile before estimating the whole cluster.
NVIDIA HGX AI Factory checkpoint 7 retains network rail topology; the following NVIDIA HGX AI Factory review tracks storage not scaling with GPU count.
When does a NVIDIA HGX AI factory plan need to be recalculated?
Validate the node profile through current NVIDIA HGX reference architecture. For enterprise HGX cluster architects, the profile should include GPU count, network rails, local and shared storage paths, firmware baseline, rack power, and thermal constraints.
Expansion is safer when new racks repeat a verified unit and the team knows exactly which shared services require additional capacity. Avoid assuming that doubling servers automatically doubles usable AI throughput. NVIDIA HGX AI Factory checkpoint 8 retains storage performance validation; the following NVIDIA HGX AI Factory review tracks server-generation mismatch.
How much reserve should NVIDIA HGX AI Factory include?
Size FAQ checkpoint 9 for NVIDIA HGX AI Factory in NVIDIA HGX AI Factory from the HGX node outward. Multiply rack power and storage throughput across the target node count, then check whether HGX server and GPU count scales at the same rate.
Surface server-generation mismatch early because HGX server topology, NIC count, PCIe layout, power, and cooling can differ by generation and OEM. Build a repeatable node profile before estimating the whole cluster.
NVIDIA HGX AI Factory checkpoint 9 retains rack electrical and thermal design; the following NVIDIA HGX AI Factory review tracks PCIe or fabric topology variation.
What should be documented before buying hardware for NVIDIA HGX AI Factory?
Validate the node profile through network rail topology. For enterprise HGX cluster architects, the profile should include GPU count, network rails, local and shared storage paths, firmware baseline, rack power, and thermal constraints. Expansion is safer when new racks repeat a verified unit and the team knows exactly which shared services require additional capacity.
Avoid assuming that doubling servers automatically doubles usable AI throughput. NVIDIA HGX AI Factory checkpoint 10 retains current NVIDIA HGX reference architecture; the following NVIDIA HGX AI Factory review tracks network rail oversubscription.