NVIDIA Groq 3 LPX: Architecture, Specs and Use Cases

NVIDIA Vera Rubin inference architecture

NVIDIA Groq 3 LPX: Architecture, Specs and Use Cases

NVIDIA Groq 3 LPX is a rack-scale inference accelerator designed to make token generation highly responsive inside the Vera Rubin platform. It is not a replacement for Rubin GPUs. The architecture combines LPUs with Vera Rubin NVL72 so large-context GPU processing and latency-sensitive generation can be assigned to the parts of the system built to handle them best.

Interactive calculator

Groq 3 LPX Deployment Planner

Enter your own workload or capacity assumptions. Public LPX benchmarks are workload-specific, and this calculator does not certify application performance, electrical design, cooling or OEM compatibility.

Live Amazon supporting hardware

Current supporting components and Price Options

Compare current Amazon listings relevant to this guide. Product availability and prices can change.

Loading current Amazon listings...

Quick answer

What is NVIDIA Groq 3 LPX?

Groq 3 LPX is NVIDIA's interactive AI inference accelerator for Vera Rubin. A full LPX rack contains 256 Groq 3 LPU accelerators, 128 GB of aggregate SRAM, 40 PB/s of on-chip SRAM bandwidth and 640 TB/s of scale-up bandwidth. NVIDIA positions it as a low-latency complement to Rubin GPUs for agentic and long-context inference.

The key buying question is interactivity, not peak FLOPS alone

LPX makes the most sense when individual requests must generate tokens quickly and predictably while the broader Vera Rubin platform continues to carry context-heavy work. Teams should size it from measured application interactivity, concurrency, context length and token demand rather than comparing one headline benchmark with unrelated GPU specifications.

NVIDIA Groq 3 LPX rack specifications at a glance

These are NVIDIA-published rack-scale figures. Application performance still depends on model, context, serving method and the way LPX is paired with Vera Rubin NVL72.

NVIDIA-published Groq 3 LPX rack specifications and practical interpretation.
ResourceLPX rack figureWhy it mattersPlanning note
Groq 3 LPU accelerators256Parallel low-latency inference engineA rack-scale resource, not 256 independent retail cards
Aggregate SRAM128 GBKeeps latency-sensitive working data close to computeSRAM is distributed across the LPUs
SRAM bandwidth40 PB/sFeeds deterministic token-generation executionDo not compare directly with GPU HBM bandwidth without workload context
Scale-up bandwidth640 TB/sConnects LPUs across the rackDirect chip-to-chip communication is part of the architecture
DDR5 memory12 TBExtends capacity for large models and workloadsDifferent role from the on-chip SRAM
AI inference compute315 PFLOPS FP8Rack-level compute capability published by NVIDIAPrecision and workload must match before using the number

Before you use the result for procurement

Benchmark the exact workload

Keep the model, context length, output length, concurrency, precision and serving-software version beside every TPS result.

Separate context from generation

LPX is positioned around low-latency token generation; Rubin and Rubin CPX handle different context-heavy work. Measure both phases.

Treat rack specs as architecture

LPUs, SRAM and internal bandwidth explain the system design but do not by themselves predict application throughput.

Verify the current OEM design

Use qualified system documentation for rack power, cooling, networking, serviceability and supported configuration before procurement.

LPX is an inference accelerator inside the Vera Rubin system, not a stand-alone GPU substitute

The cleanest way to understand LPX is as a specialized lane inside a larger inference factory. Vera Rubin NVL72 supplies Rubin GPUs, high-bandwidth memory and the general accelerated-computing environment. LPX adds a rack of language processing units optimized around low-latency, deterministic token generation. NVIDIA describes the two as a co-designed system because the work can be split and coordinated rather than forcing every inference phase onto one processor type.

That distinction matters when planning infrastructure. A buyer should not ask whether 256 LPUs can replace a certain number of Rubin GPUs. The more useful question is which phases of the serving loop are limiting user-perceived responsiveness and whether moving latency-sensitive decode work onto LPX improves the complete service. The answer depends on the model, context length, request shape, batching strategy and software stack.

Why agentic AI makes token-generation latency more important

A normal chat response may involve one prompt and one answer. An agent can run hundreds of reasoning, tool-use and verification steps, with generated text from each step feeding later steps. Small delays therefore accumulate. If every action waits on slow token generation, the entire job feels sluggish even when the infrastructure can process a large number of total tokens across all users.

LPX targets that interactivity problem. NVIDIA's August 2026 announcement emphasizes output-token speed and describes the platform as especially relevant to coding and other latency-sensitive agentic work. The important operational metric is not just aggregate tokens per second. Teams should also watch time to first token, inter-token latency, request completion time and the percentage of requests that stay inside the desired latency objective.

The 256-LPU rack is designed to behave as a tightly coordinated inference engine

NVIDIA publishes 256 Groq 3 LPU accelerators per LPX rack. Those chips are interconnected through direct chip-to-chip links and are compiler scheduled so communication can be planned alongside computation. The architectural goal is to reduce the arbitration and jitter that can appear when data movement must be decided dynamically while an inference step is already running.

For capacity planning, think of the rack as one coordinated resource rather than a pile of unrelated add-in cards. Per-chip arithmetic is useful for understanding memory and bandwidth, but production sizing should be done at the deployable rack and service level. A benchmark measured on a particular configuration cannot automatically be divided by 256 to produce a trustworthy per-chip service number.

SRAM is central to LPX latency behavior

Each Groq 3 LPU provides 500 MB of SRAM and 150 TB/s of SRAM bandwidth according to NVIDIA. Across 256 chips, the published rack figures become 128 GB of aggregate SRAM and 40 PB/s of on-chip SRAM bandwidth. SRAM offers very fast, predictable access, which fits the architecture's emphasis on deterministic execution and rapid token generation.

The capacity is modest compared with the multi-terabyte memory footprints associated with large GPU systems, so the design does not attempt to keep every model state exclusively in SRAM. LPX is paired with other memory tiers and with Rubin GPUs. This is another reason not to compare memory-capacity numbers in isolation: the platform is deliberately heterogeneous, with different memory technologies serving different phases of inference.

DDR5 extends capacity while SRAM protects the latency-sensitive path

NVIDIA lists 12 TB of DDR5 memory per LPX rack in addition to the 128 GB of SRAM. That combination is sometimes described as a fusion memory approach: very high-bandwidth SRAM is available where deterministic low-latency processing matters, while a much larger DDR5 pool supports models and workloads that cannot fit entirely in the on-chip memory.

The DDR5 inside an LPX deployment should not be confused with an ordinary user-upgradeable server memory shopping decision. System configuration is an OEM and platform matter. The live Amazon rows on this page are therefore limited to adjacent infrastructure such as NICs, enterprise NVMe and managed power equipment; they are not presented as internal LPX replacement parts.

Direct chip-to-chip links provide the rack-scale fabric

NVIDIA states that each Groq 3 LPU has 2.5 TB/s of scale-up bandwidth and the full rack provides 640 TB/s. The long-context technical description also explains that the compiler has visibility into the communication resources and can plan transfers before execution. This tight coupling is part of how the platform tries to keep latency stable when work is distributed over many processors.

When evaluating a design, separate this internal LPX scale-up fabric from the external data-center network. ConnectX and Spectrum-X are relevant to communication between servers and racks, while the LPX chip-to-chip network is part of the internal rack architecture. Mixing those bandwidth figures in one comparison table can produce misleading conclusions.

Rubin GPUs still carry critical inference work

The LPX architecture does not remove the GPU from the serving path. NVIDIA describes heterogeneous decode in which Rubin GPUs continue handling work such as prefill and attention while LPUs can accelerate latency-sensitive feed-forward or mixture-of-experts execution. Other serving configurations can divide work differently depending on the model and software stack.

This makes the platform more similar to a team of specialized processors than to a single universal accelerator. A workload with no severe interactivity requirement may remain well served by a GPU-centric configuration. A workload with large context and tight latency targets can justify a more specialized split. Production traces, not architecture diagrams alone, should decide where the bottleneck actually sits.

Deterministic compiler scheduling is a defining difference

Groq's execution model is built around planning work before it runs. NVIDIA's technical material says the compiler can schedule compute and communication down to fine-grained timing, using its knowledge of the LPU resources and chip-to-chip links. That approach reduces the need for real-time scheduling decisions during token generation and is intended to make latency more predictable.

Predictability can be valuable for premium inference services because a fast average is less useful when tail latency remains poor. Still, deterministic hardware scheduling does not eliminate every source of application jitter. Queueing, networking, model routing, storage, tool calls and external APIs can dominate an end-to-end agent session. Measure the complete request path.

NVIDIA Dynamo is part of the serving story

NVIDIA positions Dynamo as the software layer that can orchestrate disaggregated inference and route work between processors. In an LPX plus Vera Rubin deployment, software has to decide how requests move through prefill, attention, feed-forward and other phases while preserving the service objective. The hardware benefit is only realized when the serving stack can use the heterogeneous resources effectively.

Operators should therefore include software readiness in the deployment plan. Model support, compiler maturity, routing behavior, observability and upgrade procedures can be as important as raw processor specifications. A rack that benchmarks well in a controlled demonstration may still require extensive service engineering before it can carry a production multi-tenant inference API.

Third-party benchmark results are useful, but they are workload specific

Artificial Analysis measured 3,431 output tokens per second on Groq 3 LPX with Gemma 4 31B at 100,000-token context, and NVIDIA's press release rounds the result to 3,400. That is strong evidence that the platform can deliver high interactivity on that specific benchmark. It is not a universal capacity rating for every model, prompt length, batch size or serving tier.

Use the result as a reference point, then benchmark the models that matter to your business. Record context length, output length, concurrency, quantization or precision, quality settings and the exact software release. Without those details, two tokens-per-second figures can look comparable while measuring very different operating conditions.

Nebius adoption is an important production signal

NVIDIA announced Nebius as the first AI cloud to adopt Groq 3 LPX, with plans to bring the hardware into Nebius Token Factory. That matters because it moves the discussion beyond a laboratory architecture and toward a cloud service that has to expose predictable performance through an API, schedule real customers and operate the systems continuously.

An early cloud deployment does not automatically tell every enterprise to buy racks. Many teams may be better served consuming LPX-backed inference from a provider until their sustained demand, data requirements or economics justify dedicated infrastructure. Compare cloud access, reserved capacity and owned-rack economics using the same workload traces.

Plan networking, storage, power and cooling as separate workstreams

LPX is a liquid-cooled rack-scale platform aligned with the broader Vera Rubin MGX infrastructure. Data-center deployment therefore involves more than ordering compute. External network topology, storage paths, rack power, facility distribution, cooling loops and operational access all have to be engineered. Those layers have their own failure modes and procurement lead times.

The Amazon table above is useful only for price discovery on support components that commonly appear in retail channels. It should never be used as a substitute for an OEM bill of materials for LPX. Exact rack power, CDU requirements, supported optics and validated networking should come from the system vendor and current NVIDIA platform documentation.

Use a measured rollout before committing the whole inference fleet

A sensible adoption path is to define one or two production workloads, capture their baseline latency and throughput on the existing platform, then validate the same service on LPX-backed Vera Rubin infrastructure. The comparison should include cost per completed task, tail latency, model quality, utilization and operational overhead rather than only peak token speed.

Once the service behavior is understood, capacity can be expanded in rack increments with explicit redundancy and headroom. The calculator at the top of this page intentionally asks for measured sustained TPS per rack so it can be replaced with your real number. That keeps the planning model useful even as software and model performance evolve.

Methodology and sources

Cloudzat separates NVIDIA-published architectural specifications from workload-dependent performance. Rack figures come from the current NVIDIA Groq 3 LPX product and technical pages; benchmark figures are labeled with their model and context. Calculator capacity is driven by user-entered sustained performance rather than an assumed universal LPX rating.

As an Amazon Associate, Cloudzat may earn from qualifying purchases. Marketplace listings on these pages cover adjacent networking, storage, memory or power hardware only. They are not represented as LP30 chips, LPX trays or qualified Vera Rubin rack components. Verify exact models, warranty, interfaces and current OEM documentation before purchase.

Frequently asked questions

Is NVIDIA Groq 3 LPX a GPU?

No. LPX is a rack-scale inference platform built around Groq 3 language processing units. NVIDIA pairs it with Rubin GPUs inside the Vera Rubin architecture rather than describing it as a general-purpose GPU replacement.

How many LPUs are in one Groq 3 LPX rack?

NVIDIA lists 256 Groq 3 LPU accelerators per LPX rack. The system is designed as a tightly connected rack-scale resource.

How much SRAM does an LPX rack have?

The published aggregate is 128 GB of SRAM, derived from 500 MB per LPU across 256 accelerators.

What is the LPX rack SRAM bandwidth?

NVIDIA publishes 40 PB/s of aggregate on-chip SRAM bandwidth per rack. That figure describes the internal SRAM path and should not be compared directly with external network bandwidth.

Does LPX replace Vera Rubin NVL72?

No. NVIDIA describes LPX as an extension of Vera Rubin. Rubin GPUs and LPUs can cooperate on inference, with each processor type handling work suited to its architecture.

What workloads are best suited to LPX?

The strongest fit is latency-sensitive agentic inference where rapid output-token generation and predictable response time matter, especially alongside large context handled by the broader Vera Rubin system.

What benchmark has been published for Groq 3 LPX?

Artificial Analysis measured 3,431 output tokens per second on Gemma 4 31B with a 100K-token context. Treat that as a specific benchmark result, not a universal performance guarantee.

Is Groq 3 LPX in production?

NVIDIA announced on August 24, 2026 that Groq 3 LPX is in full production and named Nebius as the first AI cloud adopter.

Can I buy Groq 3 LPX on Amazon?

Cloudzat does not present Amazon listings as a source for LPX racks. The marketplace table on this page covers adjacent infrastructure only. Complete systems should be sourced through qualified NVIDIA/OEM channels.

How should I size an LPX deployment?

Start with measured sustained TPS for your exact model and serving stack, apply realistic utilization and headroom, add redundancy, and then convert the requirement into whole racks. Revalidate after software or model changes.

Scroll to Top