LPX capacity and redundancy sizing
NVIDIA Groq 3 LPX Rack Calculator: Size Inference Capacity
Rack sizing should start with sustained performance from the model and service you actually plan to run. This calculator converts a target output-token rate into whole LPX racks after accounting for usable utilization, capacity headroom and explicit redundancy. It also shows the resulting LPU, SRAM and internal scale-up resources.
Interactive calculator
Groq 3 LPX Rack Sizing Calculator
Enter your own workload or capacity assumptions. Public LPX benchmarks are workload-specific, and this calculator does not certify application performance, electrical design, cooling or OEM compatibility.
Live Amazon supporting hardware
Current supporting components and Price Options
Compare current Amazon listings relevant to this guide. Product availability and prices can change.
Quick answer
Do not size LPX from a headline benchmark alone
Enter measured sustained TPS per rack, not a theoretical peak. The tool reduces that number by your utilization target, adds demand headroom, rounds up to whole racks and then adds the redundancy count you choose. It does not assume perfect linear scaling between racks.
Size for the failure and maintenance case, not only normal operation
If one rack must be offline for maintenance or a fault, the remaining fleet still needs enough capacity to meet the service objective. N+1 is a common starting concept, but the right redundancy model depends on cluster topology, traffic routing and business availability requirements.
Inputs that control LPX rack count
The biggest sizing errors come from treating peak benchmark performance as sellable production capacity.
| Sizing input | What it represents | Too low causes | Too high causes |
|---|---|---|---|
| Measured rack TPS | Sustained output rate for your workload | Underestimated rack count | Over-purchase if measurement is pessimistic |
| Utilization ceiling | Capacity you are willing to sell during normal operation | Queueing and poor tail latency | Idle specialized capacity |
| Demand headroom | Buffer above current peak demand | Frequent emergency expansion | Capital sits unused |
| Redundancy | Capacity available during failures or maintenance | Availability risk | Higher CAPEX |
| Growth allowance | Expected future traffic increase | Early re-procurement | Excess capacity if demand misses forecast |
| Model mix | Variation around the measured baseline | One slow model can saturate the pool | Over-segmentation can waste capacity |
Before you use the result for procurement
Benchmark the exact workload
Keep the model, context length, output length, concurrency, precision and serving-software version beside every TPS result.
Separate context from generation
LPX is positioned around low-latency token generation; Rubin and Rubin CPX handle different context-heavy work. Measure both phases.
Treat rack specs as architecture
LPUs, SRAM and internal bandwidth explain the system design but do not by themselves predict application throughput.
Verify the current OEM design
Use qualified system documentation for rack power, cooling, networking, serviceability and supported configuration before procurement.
Begin with a sustained per-rack measurement
The most important input is the amount of output traffic one LPX rack can sustain on your exact service while meeting the latency objective. A short peak run is not enough. Test realistic context lengths, output lengths, concurrency and model mix for long enough to expose queueing, software or thermal behavior.
If you do not yet have hardware, a public benchmark can be a placeholder for scenario planning, but mark it clearly. Replace that placeholder as soon as a cloud trial, vendor proof of concept or acceptance test produces a workload-specific number.
Set a utilization ceiling below the point where latency becomes unstable
Accelerator utilization and user latency are related. Running a specialized inference rack at 100% theoretical capacity leaves little room for bursts, uneven requests or software overhead. Queueing can rise sharply before a system reaches its absolute throughput limit.
The calculator therefore applies a utilization ceiling to the measured rack TPS. A 70% setting means only 70% of the measured rate is treated as normally sellable capacity. Use load tests to find the threshold where p95 and p99 latency begin to degrade.
Headroom protects against demand variance
Traffic forecasts are never exact. Product launches, model changes or a single large customer can create demand above the current peak. Headroom is a deliberate reserve for that uncertainty, applied to the target workload before rack count is calculated.
Keep growth and headroom conceptually separate. Headroom protects near-term operating variance; annual growth represents expected structural demand increase. Combining both into one unexplained buffer makes future capacity reviews harder.
Redundancy must be added after normal capacity is sized
If the service requires N+1 behavior, first calculate the racks required to carry demand and then add the standby or recoverable rack. That prevents the redundancy unit from being counted as normal sellable capacity. More complex designs may need enough spare capacity per availability zone or failure domain rather than one global spare.
The calculator uses a simple extra-rack count because topology is site-specific. It cannot determine whether one rack protects a multi-row cluster, whether maintenance overlaps failures or how traffic is redistributed between data centers.
Account for model mix instead of using one optimistic benchmark
A fleet rarely serves one model with one context length forever. Premium models may be slower, context windows grow, and new agent workflows can change output length. If the measured rack TPS comes from the easiest workload, the production pool can saturate when the mix shifts.
Build a weighted workload profile or size to the most demanding service tier that shares the pool. Another option is to isolate slow models into their own capacity group so a burst cannot degrade every customer.
Growth planning should be recalculated after each major model change
A 30% traffic forecast means little if a new model uses twice the compute per token. Capacity demand is the combination of business traffic and technical efficiency. Model upgrades, compiler releases and serving improvements can move the denominator as much as user growth moves the numerator.
Re-run the sizing model whenever a major model or software release changes sustained TPS. Keep the previous measurement so finance can distinguish demand-driven expansion from performance-driven expansion.
Whole-rack rounding creates stepwise capacity economics
Rack-scale platforms do not let a buyer purchase 0.3 of a production rack simply because the arithmetic says so. Capacity therefore arrives in large steps. A service near the threshold for another rack can show a sudden increase in marginal cost even if demand grows only slightly.
That step function is important for pricing and cloud-versus-own comparisons. Small workloads may be cheaper to consume as a service until they can utilize the next rack efficiently.
The output also shows installed LPUs and memory resources
Once the rack count is known, the tool multiplies it by 256 LPUs, 128 GB SRAM, 12 TB DDR5 and 640 TB/s internal scale-up bandwidth. These totals help describe the installed platform footprint but should not be used as direct performance multipliers.
For example, two racks provide twice the installed LPU count, but delivered application throughput can be lower than 2x if external networking, routing or model partitioning becomes the bottleneck.
External network capacity must scale with the serving architecture
LPX racks participate in a broader Vera Rubin AI factory. Requests, KV state, storage traffic or disaggregated inference data may cross the scale-out network depending on software topology. A rack-sizing model that ignores the network can buy compute faster than the fabric can feed it.
Create a separate network model using actual bytes moved per request and the communication path. High-speed NIC listings above are only market references; production fabrics should use qualified networking and optics.
Power and cooling reserve can constrain how many racks fit in a site
Even when inference demand justifies more racks, the data hall may not have electrical and liquid-cooling capacity ready. High-density rack-scale systems require coordinated facility expansion, and lead time for switchgear, cooling distribution or construction can exceed server delivery time.
Capacity planning should therefore have both a compute forecast and a facility envelope. The tool does not invent LPX rack power because those details must come from current OEM engineering documentation.
Use phased acceptance tests for multi-rack deployment
Instead of energizing an entire large order and discovering a software issue everywhere, validate a small number of racks against the defined performance and availability tests. Confirm model support, network behavior, failover and observability before scaling the same configuration.
The measured results from the first phase should then replace planning assumptions for the next phase. This turns capacity expansion into a feedback loop rather than a fixed spreadsheet built months before production.
Revisit the calculator when service-level objectives change
A business can reduce rack needs by accepting slower responses, or require more capacity when it launches a premium low-latency tier. The same hardware can therefore have different sellable capacity depending on the latency promise.
Save the utilization, headroom and redundancy assumptions beside the commercial SLA. That makes it possible to explain why two services with identical token volume need different infrastructure.
Methodology and sources
The rack calculator uses user-entered sustained TPS, an explicit utilization ceiling, demand headroom and whole-rack redundancy. It then reports installed resources using NVIDIA's published 256-LPU, 128-GB SRAM, 12-TB DDR5 and 640-TB/s scale-up figures per rack.
- NVIDIA Groq 3 LPX product page
- NVIDIA technical blog: inside Groq 3 LPX
- NVIDIA technical blog: LPX long-context interactivity
- NVIDIA Groq 3 LPX full-production announcement
- NVIDIA Vera Rubin NVL72
- NVIDIA technical blog: Vera Rubin POD architecture
As an Amazon Associate, Cloudzat may earn from qualifying purchases. Marketplace listings on these pages cover adjacent networking, storage, memory or power hardware only. They are not represented as LP30 chips, LPX trays or qualified Vera Rubin rack components. Verify exact models, warranty, interfaces and current OEM documentation before purchase.
Frequently asked questions
What TPS value should I enter per rack?
Use sustained performance from your exact model, context, concurrency and serving software while the service still meets its latency objective.
Can I use 3,431 TPS as the rack value?
Only as an early scenario if you understand it comes from the specific Gemma 4 31B 100K-context benchmark. Replace it with your own sustained result.
Why does the calculator reduce TPS by utilization?
Because operating continuously at a benchmark peak leaves no room for bursts, request variance or maintenance and can degrade tail latency.
What does headroom mean?
Headroom is extra capacity above expected demand for near-term variance and forecast error.
Should redundancy be counted in normal capacity?
Not if it is reserved to survive a rack outage. The calculator adds redundancy after sizing the normal load.
How many LPUs are installed for each rack?
256 per rack.
How much SRAM does each added rack contribute?
128 GB aggregate SRAM.
Does rack throughput scale linearly?
Not necessarily. External network, routing, software and model behavior can prevent perfect scaling.
Should I include annual growth and headroom?
They represent different risks. Growth models expected demand increase; headroom protects short-term variance. Use both only when they are justified.
What else must be checked before ordering racks?
Validate facility power and cooling, network topology, software support, OEM lead time, failover and production acceptance criteria.