The fifth pillar of NVIDIA AI networking
NVIDIA Scale-In Networking: BlueField-4 and AI Factory Guide
NVIDIA introduced Scale-In as a dedicated infrastructure domain for the traffic that enters, leaves and services an AI factory. Instead of letting host CPUs absorb policy enforcement, storage access, security and telemetry, the model uses BlueField-4 DPUs and DOCA services connected through Spectrum-X Ethernet. NVIDIA describes BlueField-4 with up to 800 Gb/s throughput for this role.
Interactive calculator
BlueField-4 Scale-In Capacity Planner
Enter your own topology and bandwidth assumptions. These planning calculators estimate raw links, capacity and ratios; they do not certify fabric goodput, rail mapping, cable reach, firmware interoperability, electrical design, cooling or OEM compatibility.
Live Amazon supporting hardware
Current supporting components and Price Options
Compare current Amazon listings relevant to this guide. Product availability and prices can change.
Quick answer
What is NVIDIA Scale-In?
Scale-In is NVIDIA’s name for the accelerated north-south infrastructure layer of an AI factory. BlueField-4 handles networking, storage, security, tenant isolation and observability services independently of the host CPU, while Spectrum-X Ethernet connects those services at high bandwidth. NVIDIA calls it the fifth pillar of its AI networking architecture.
Use Scale-In when shared infrastructure services are becoming a bottleneck
The architecture is most relevant to multi-tenant or agentic AI factories where data access, storage, security policy and operational traffic can consume substantial host resources. Size from actual north-south traffic and service requirements, then decide whether BlueField offload improves isolation and predictability enough to justify the DPU and operational stack.
Where Scale-In fits
Scale-In is distinct from the fabrics that synchronize GPUs, even though all of them can use high-speed networking.
| Network domain | Primary traffic | NVIDIA building block | Planning question |
|---|---|---|---|
| Scale-up | Inside rack / tightly coupled accelerators | NVLink / NVLink Switch | How do accelerators behave as one system? |
| Scale-out | GPU cluster communication | Spectrum-X or Quantum networking | How do racks exchange collective traffic? |
| Scale-across | Between AI data centers | Spectrum-XGS | How do sites act as one larger system? |
| Scale-In | Users, services, storage and operations entering AI factory | BlueField-4 + DOCA + Spectrum-X | How are infrastructure services accelerated and isolated? |
| Management / telemetry | Control and observability | DOCA / platform tooling | Can operations remain independent of host workload? |
Before you use the result for procurement
Draw the physical topology
Map every server-facing port, leaf uplink, spine link and network plane. Aggregate bandwidth alone can hide impossible port or lane assumptions.
Verify exact endpoints
Confirm NIC form factor, PCIe generation, host lane budget, port speed, connector and firmware support on the exact server platform.
Qualify optics and cables
Match OSFP/QSFP form factor, lane rate, breakout, reach, fiber type and both endpoint qualification lists. Do not treat equal headline speed as automatic compatibility.
Test failure and congestion behavior
Validate oversubscription, ECMP or multiplane path behavior, switch failure domains and recovery under the traffic patterns the AI workload will actually generate.
Scale-In focuses on traffic entering the AI factory
Training and inference clusters are usually discussed in terms of accelerator-to-accelerator communication, but production systems also move prompts, datasets, checkpoints, storage traffic, API calls, security events and telemetry. NVIDIA uses Scale-In to describe the infrastructure domain that handles those flows as they cross into the accelerated environment.
This distinction is useful because north-south services have different policy and reliability requirements from collective traffic. They may need tenant isolation, encryption, storage virtualization or observability even while the GPU fabric is optimized for synchronized east-west communication.
BlueField-4 is the processing engine for the new domain
NVIDIA positions BlueField-4 DPU as the host-independent processor that runs networking, storage and security services for Scale-In. The company cites up to 800 Gb/s throughput and emphasizes that infrastructure work can continue without consuming the host CPU cycles needed by AI applications.
A DPU should not be purchased merely because 800G sounds large. Identify the services that will actually be offloaded and confirm they are supported in the current DOCA release. The value comes from predictable isolation and acceleration of real infrastructure work.
DOCA turns DPU hardware into deployable services
The hardware by itself does not create a policy plane. NVIDIA DOCA supplies APIs, libraries and services for networking, storage, security, telemetry and lifecycle management on BlueField. Scale-In therefore needs software operations as much as NIC or switch engineering.
Teams should include DOCA versioning, deployment automation, monitoring and incident response in the architecture review. Moving services off the host changes where logs, certificates, policy updates and failure diagnostics live.
Host-independent processing protects GPU-serving CPUs
Agentic AI factories can place heavy orchestration and data-processing demands on CPUs. If those same CPUs also terminate storage, networking and security workloads, unpredictable infrastructure traffic can interfere with application latency. BlueField provides a separate processing domain.
The economic case is strongest when host CPU contention is measurable. Profile CPU time spent in network, storage and security paths before estimating savings. Offload should free resources or improve isolation in a way that can be observed, not simply relocate work to another processor.
Tenant isolation belongs close to the network edge
Shared AI factories need strong boundaries between customers, projects and agents. NVIDIA describes BlueField and DOCA as enforcing policy independently of the tenant host, reducing the opportunity for a compromised workload to bypass controls running in the same CPU domain.
The security design still needs identity, key management, segmentation policy and auditability. A DPU is an enforcement point, not a complete zero-trust program. Include the DPU firmware and management plane in threat models and patch procedures.
Storage traffic is a first-class Scale-In workload
Modern AI services repeatedly fetch model data, embeddings, vector indexes, context and checkpoints. Storage can therefore consume hundreds of gigabits per server and compete with application traffic. Scale-In treats data access as an accelerated infrastructure service instead of a background CPU task.
Measure storage flows separately from user-facing API traffic. The calculator asks for both because their peaks may overlap. If storage and application traffic share the same DPU or switch paths, the combined demand and overhead must fit within the planned utilization target.
Telemetry should not become a hidden tax on the host
Large clusters emit enormous volumes of network counters, security events and performance telemetry. Collecting and processing that information on the host can consume resources and become unreliable when the host is already overloaded.
A host-independent telemetry path can improve observability during failures because the DPU remains a separate operational point. Confirm which metrics are available and how they feed the existing monitoring stack before relying on Scale-In as an operational improvement.
Spectrum-X provides the high-bandwidth transport
Scale-In is not a standalone appliance. NVIDIA connects BlueField services over Spectrum-X Ethernet so policy, storage and application traffic can move at the rates expected by modern AI factories. This creates an end-to-end infrastructure path from DPU through the Ethernet fabric.
Network designers should decide whether Scale-In traffic shares switches with other domains or uses dedicated capacity. Shared hardware can improve utilization, while separate paths can simplify fault isolation and performance guarantees.
Redundancy should be modeled as service capacity
If a DPU or path fails, the remaining devices need enough headroom to absorb the traffic. Simply buying two devices does not create high availability if both normally run near line rate. The calculator therefore reserves capacity before adding redundant units.
Test failover with active storage, security and API load. The key metric is whether policy and data services continue within latency targets, not whether a standby device eventually becomes reachable.
Scale-In is distinct from Spectrum-X Multiplane
Multiplane addresses scale-out fabric topology and resilience across GPU-serving networks. Scale-In addresses the infrastructure services crossing into the AI factory. Both can use Spectrum-X and NVIDIA endpoint technology, but they solve different bottlenecks.
Keeping the terms separate helps procurement. A team evaluating BlueField for storage and security offload should not accidentally assume it has also solved the collective-communication topology needed by thousands of GPUs.
The best sizing input is measured traffic by service
A generic “800G server” label hides whether bandwidth is used by prompts, storage reads, checkpoint writes, security inspection or telemetry. Capture service-level peaks and concurrency, then combine them with realistic overlap assumptions.
The calculator uses a simple sum plus overhead because that is transparent. Sophisticated deployments may use time-series percentiles or traffic matrices, but the underlying principle remains: size from observed service demand rather than a marketing maximum.
Adoption requires platform and operations ownership
Introducing DPUs adds firmware, drivers, DOCA services, provisioning and a management surface. The architecture can reduce host complexity while increasing infrastructure specialization. Assign clear ownership for upgrades and troubleshooting before production rollout.
A successful proof of concept should show not only throughput but operational benefits: lower host CPU use, stable tail latency, enforceable isolation, faster incident visibility or simpler storage virtualization. Those outcomes make the Scale-In investment defensible.
Methodology and sources
Cloudzat uses NVIDIA’s August 2026 BlueField-4 technical material for the definition of Scale-In and its cited 800 Gb/s DPU throughput. The capacity tool combines user-entered service traffic with utilization and redundancy assumptions; it does not estimate application acceleration or security effectiveness.
- NVIDIA: BlueField-4 and Scale-In networking
- NVIDIA Spectrum-X Ethernet platform
- NVIDIA ConnectX-9 SuperNIC introduction
- NVIDIA: Spectrum-6 arrives in gigascale AI factories
- NVIDIA Vera Rubin, LPX and Spectrum-X update
As an Amazon Associate, Cloudzat may earn from qualifying purchases. Marketplace listings on these pages are supporting networking hardware such as NICs, switches, optics and high-speed cables. A marketplace row is not represented as a Spectrum-6 switch, ConnectX-9 SuperNIC, Thor Ultra NIC or qualified NVIDIA fabric unless the exact listing evidence supports that identity. Verify model, speed, connector, firmware, warranty and OEM qualification before purchase.
Frequently asked questions
Why does NVIDIA call Scale-In a fifth networking pillar?
NVIDIA separates it from scale-up, scale-out and other AI fabric roles because it focuses on infrastructure traffic and services entering the AI factory, accelerated by BlueField-4 and DOCA.
What hardware powers Scale-In?
NVIDIA’s current architecture centers on BlueField-4 DPUs connected through Spectrum-X Ethernet.
How fast is BlueField-4?
NVIDIA cites up to 800 Gb/s throughput for BlueField-4 in the Scale-In context. Exact performance depends on workload and enabled services.
What services can be offloaded?
NVIDIA highlights networking, storage access, security, policy enforcement, tenant isolation, telemetry and related infrastructure functions.
Is Scale-In the same as scale-out networking?
No. Scale-out moves distributed accelerator traffic across the cluster. Scale-In handles north-south application, data, storage and infrastructure services.
Does Scale-In replace host CPUs?
No. It offloads infrastructure work so CPUs can focus more on application and orchestration tasks. The host still performs its normal compute role.
Do I need DOCA?
DOCA is the software platform NVIDIA uses to develop and operate services on BlueField, so it is central to the Scale-In model.
How should I size DPU count?
Combine application and storage bandwidth, add service overhead, limit normal utilization and reserve enough capacity for failures. Then validate actual service performance.
Can Scale-In improve security?
It can provide host-independent enforcement and isolation, but it must be integrated into a broader identity, key-management, patching and audit program.
What should a proof of concept measure?
Measure host CPU reduction, DPU utilization, storage throughput, API latency, policy enforcement, failover behavior and telemetry continuity under realistic load.