Advertisement


Home Server Accelerators NVIDIA’s Groq 3 LPU Accelerators for Heterogeneous AI Compute at Hot Chips...

NVIDIA’s Groq 3 LPU Accelerators for Heterogeneous AI Compute at Hot Chips 2026

0
NVIDIA Groq 3 LPX Rack
NVIDIA Groq 3 LPX Rack

The second AI presentation of the afternoon comes from NVIDIA, who besides doing talks on the Vera CPU and Rubin GPU, are also presenting a talk on the use of language processing units (LPUs) in their Vera Rubin racks.

This talk is arguably a bit of an unusual one, because NVIDIA does not currently produce their own LPUs (though they are under development). The LPUs being used in Vera Rubin generation racks – and specifically the dedicated LPX racks – come from Groq, whom NVIDIA is buying the chips from as part of a broader acquihire of the company. None the less, with the bulk of Groq’s talent now at NVIDIA, it falls to NVIDIA to promote the chips.

Groq’s LPUs are designed to fill a weak spot in NVIDIA’s Vera Rubin stack. GPUs are great for pre-fill, and depending on how the design of the chip is optimized, so-so at decode. Whereas LPUs – essentially purpose-built chips with large amounts of on-die SRAM to keep model data as close to the compute hardware as possible. Fittingly, the biggest rationale for the use of LPUs in a Vera Rubin server cluster is to boost performance at low latencies, taking advantage of that SRAM to get results back ASAP so that AI models can move on to the next token.

NVIDIA Vera Hot Chips 2026 Groq LPX3
NVIDIA Vera Hot Chips 2026 Groq LPX3

This article is being written live from the presentation, so please excuse any typos.

NVIDIA’s Groq 3 LPU Accelerators for Heterogeneous AI Compute at Hot Chips 2026

Yesterday NVIDIA released the first third-party benchmarks of an LPX rack, so the timing of that and this talk is not coincidental. Nor is the fact that NVIDIA has been promoting the use of LPUs in both their Vera and Rubin talks. Today’s talk is reiterating parts of that for the Hot Chips crowd, as well as laying out the case for using LPUs with Vera Rubin and how they fit in to the broader ecosystem.

NVIDIA Groq 3 LPU Hot Chips 2026 Agentic AI
NVIDIA Groq 3 LPU Hot Chips 2026 Agentic AI

Recapping a common refrain through all of NVIDIA’s talks at Hot Chips, the company considers agenetic AI to be the most complex computing workload in history. And to that end, it has required new approaches to hardware development – as well as a whole lot more hardware.

NVIDIA Groq 3 LPU Hot Chips 2026 Vera Rubin Stack
NVIDIA Groq 3 LPU Hot Chips 2026 Vera Rubin Stack

LPUs are the building block of the LPX rack, which is one of several racks that make up the larger Vera Rubin hardware ecosystem.

NVIDIA Groq 3 LPU Hot Chips 2026 Groq 3 LPX
NVIDIA Groq 3 LPU Hot Chips 2026 Groq 3 LPX

And once again showing off NVIDIA’s performance curves/frontiers for the Vera Rubin ecosystem, plotting tokens/sec/watt versus tokens/sec/user. The Rubin GPU hardware has its own performance curve, but overall performance can drop pretty hard if you try to push higher user interactivity (i.e. lowering token latency). This is where the LPX rack comes in, adding a second layer of hardware to significantly boost the per-user token rate and reduces the latency accordingly. Or as the Groq team puts it: this is ludicrous mode.

NVIDIA Groq 3 LPU Hot Chips 2026 Tokens and Latency
NVIDIA Groq 3 LPU Hot Chips 2026 Tokens and Latency

Agentic AI systems need a lot of processing time due to the large amounts of context tokens in play. Decode is the largest portion of the compute workload in term of the amount of time taken. And that is where the LPU comes in to handle this portion of the inference task more efficiently and quickly.

NVIDIA Groq 3 LPU Hot Chips 2026 Groq 3 LPX Engineering
NVIDIA Groq 3 LPU Hot Chips 2026 Groq 3 LPX Engineering

Presenting at Hot Chips is apparently a bit of a “pinch me” moment for the Groq team members who moved over to NVIDIA.

NVIDIA Groq 3 LPU Hot Chips 2026 LPX Performance
NVIDIA Groq 3 LPU Hot Chips 2026 LPX Performance

Diving into the LPX rack, a single rack can decode 11,000 tokens per second on a 31B model (Gemma 4).

NVIDIA Groq 3 LPU Hot Chips 2026 LPX Long-Context Decode
NVIDIA Groq 3 LPU Hot Chips 2026 LPX Long-Context Decode

NVIDIA this week published third-party benchmarks, which were conducted by Artificial Analysis. The LPX rack offed a 4x higher token output rate than the next-closest public competitor, which NVIDIA doesn’t name but looks to be a Cerebas CS-3.

NVIDIA Groq 3 LPU Hot Chips 2026 Near-SRAM Compute
NVIDIA Groq 3 LPU Hot Chips 2026 Near-SRAM Compute

Diving into the LPU architecture, one of the key elements of the chip design is close connection between memory and compute. Groq has put quite a bit of SRAM very close to their compute elements.

This is a fully deterministic chip. There is no branching or other sources of non-determinism.

NVIDIA Groq 3 LPU Hot Chips 2026 Software Scheduling
NVIDIA Groq 3 LPU Hot Chips 2026 Software Scheduling

Because the chip is deterministic, the instruction scheduling is handled in software as well. Which means there’s no need for significant scheduling hardware on the chip.

NVIDIA Groq 3 LPU Hot Chips 2026 LPU Core SIMDs
NVIDIA Groq 3 LPU Hot Chips 2026 LPU Core SIMDs

And here is a step-by-step example of how a thread is scheduled and executed over an LPU, with an emphasis on how threads can be handled concurrently.

NVIDIA Groq 3 LPU Hot Chips 2026 LPU Core Schedule
NVIDIA Groq 3 LPU Hot Chips 2026 LPU Core Schedule

 

NVIDIA Groq 3 LPU Hot Chips 2026 LPU Core Schedule 2
NVIDIA Groq 3 LPU Hot Chips 2026 LPU Core Schedule 2

 

NVIDIA Groq 3 LPU Hot Chips 2026 LPU Core Schedule 3
NVIDIA Groq 3 LPU Hot Chips 2026 LPU Core Schedule 3

 

NVIDIA Groq 3 LPU Hot Chips 2026 LPU Core Schedule 4
NVIDIA Groq 3 LPU Hot Chips 2026 LPU Core Schedule 4

 

NVIDIA Groq 3 LPU Hot Chips 2026 LPU Core Schedule 5
NVIDIA Groq 3 LPU Hot Chips 2026 LPU Core Schedule 5

 

NVIDIA Groq 3 LPU Hot Chips 2026 LPU Core Schedule 6
NVIDIA Groq 3 LPU Hot Chips 2026 LPU Core Schedule 6

 

NVIDIA Groq 3 LPU Hot Chips 2026 LPU Core Schedule 7
NVIDIA Groq 3 LPU Hot Chips 2026 LPU Core Schedule 7

 

NVIDIA Groq 3 LPU Hot Chips 2026 LPU Network
NVIDIA Groq 3 LPU Hot Chips 2026 LPU Network

Ultimately an LPU isn’t going to be working alone. It’s going to be thousands of chips working together, effectively acting as a single massive LPU core.

NVIDIA Groq 3 LPU Hot Chips 2026 Low Overhead Network
NVIDIA Groq 3 LPU Hot Chips 2026 Low Overhead Network

Looking at the LPU networking, each chip acts as both a processor and a router. Again owing to the determinism, the compiler schedules all resources, both processing and networking.

NVIDIA Groq 3 LPU Hot Chips 2026 LPU Packet Overhead
NVIDIA Groq 3 LPU Hot Chips 2026 LPU Packet Overhead

Interestingly, there is no hardware flow control or virtual channels here. Groq keeps things very bare and basic.

NVIDIA Groq 3 LPU Hot Chips 2026 Scale-Up Interconnect
NVIDIA Groq 3 LPU Hot Chips 2026 Scale-Up Interconnect

Getting back to the physical world, there are 256 LPUs in a single LPX rack. This combines to make for a 128GB of SRAM for memory, offering a total of 40PB/second of aggregate SRAM bandwidth. Or in terms of compute, there are 315 PFLOPS of FP8 compute performance.

NVIDIA Groq 3 LPU Hot Chips 2026 Mass Production
NVIDIA Groq 3 LPU Hot Chips 2026 Mass Production

One focus has been on ensuring the LPX rack and trays are optimized for mass production and are reliable once assembled. The LPU team is using NVIDIA’s MGX rack architecture here, which means using a standardized and well-tested piece of hardware across NVIDIA’s ecosystem.

NVIDIA Groq 3 LPU Hot Chips 2026 Determinism
NVIDIA Groq 3 LPU Hot Chips 2026 Determinism
NVIDIA Groq 3 LPU Hot Chips 2026 Rack-Scale Smoothing
NVIDIA Groq 3 LPU Hot Chips 2026 Rack-Scale Smoothing

The emphasis on determinism also means that power consumption is very deterministic. That means that the LPX can employ look-ahead techniques to smooth power consumption, and even order current from power regulators ahead of time so that it can be available when it’s needed in a moment. The net effect is that it cuts down on both droop (60%) and overshoot, flattening out the overall curve.

(And the audience just applauded at this)

NVIDIA Groq 3 LPU Hot Chips 2026 Determinism TDP
NVIDIA Groq 3 LPU Hot Chips 2026 Determinism TDP

The software backing the LPX also takes into consideration thermal considerations to better spread out the workload over the chip to avoid thermal hot spots and thermal throttling. This being another benefit of deterministic execution.

And that’s the LPU hardware in a nutshell.

NVIDIA Groq 3 LPU Hot Chips 2026 GPU-LPU Co-Design
NVIDIA Groq 3 LPU Hot Chips 2026 GPU-LPU Co-Design

Looking at the bigger picture, LPX is but one part of the larger Vera Rubin hardware ecosystem. NVIDIA will be using the LPXes in conjunction with their GPU hardware to maximize overall performance. Which means co-designing the GPU and LPU.

NVIDIA Groq 3 LPU Hot Chips 2026 LPU Sync Domain
NVIDIA Groq 3 LPU Hot Chips 2026 LPU Sync Domain

Getting the LPUs and GPUs talking to each other and efficiently interacting is trickier than it may first appear. The LPUs are in a synchronous domain, but talking to the GPUs or going to an external KV cache is an async operation. So NVIDIA employs a FPGA to function as an async bridge between the two worlds.

NVIDIA Groq 3 LPU Hot Chips 2026 Static Scheduling
NVIDIA Groq 3 LPU Hot Chips 2026 Static Scheduling

Larger mixed clusters also have to account for the fact that not everything within LPX’s domain is perfectly deterministic. So there is an need to handle static scheduling of dynamic workloads.

NVIDIA Groq 3 LPU Hot Chips 2026 Speculative Decode
NVIDIA Groq 3 LPU Hot Chips 2026 Speculative Decode

LPU and GPU racks keep their own KV caches. Only draft tokens are exchanged between the two racks.

NVIDIA Groq 3 LPU Hot Chips 2026 Disaggregated Decode
NVIDIA Groq 3 LPU Hot Chips 2026 Disaggregated Decode

NVIDIA offloads the attention portion of the decode process back to the GPUs, making decode a disaggregated process.

NVIDIA Groq 3 LPU Hot Chips 2026 Disaggregated Prefill
NVIDIA Groq 3 LPU Hot Chips 2026 Disaggregated Prefill

And, of course, prefill and decode are disaggreated as well, with prefill taking place in the GPUs while most of the decode process takes place in the LPUs.

NVIDIA Groq 3 LPU Hot Chips 2026 Max Utilization
NVIDIA Groq 3 LPU Hot Chips 2026 Max Utilization

NVIDIA uses micro-batches of workloads that are executed concurrently on the GPUs and LPUs. This allows NVIDIA to overlap these workloads and their communication to help keep up the utilization of the hardware.

NVIDIA Groq 3 LPU Hot Chips 2026 CUDA LPU Support
NVIDIA Groq 3 LPU Hot Chips 2026 CUDA LPU Support

The integration of LPUs into the NVIDIA hardware ecosystem means that the CUDA software platform needs to be similarly updated. NVIDIA has adding LPU support to CUDA ecosystem, making them a fully-supported target for CUDA.

NVIDIA Groq 3 LPU Hot Chips 2026 Performance Curves
NVIDIA Groq 3 LPU Hot Chips 2026 Performance Curves

Looking at the culmination of the LPU hardware, the rest of the Vera Rubin hardware, and NVIDIA’s software changes significantly alters the performance curve/pareto frontier. Vera Rubin can push a lot of tokens overall at low interactivity, but tapers off quickly at higher interactivity rates. Combining this with the LPUs extends these curves further, and the more that is offloaded to the LPUs the higher the interactivity rates gets, up to a 5x improvement at the highest token/user/second rate. Though the trade-off is that total throughput efficiency is dropping the more the LPUs are used; GPUs are still the king of efficiency when total throughput is all that matters and high latencies are acceptable.

NVIDIA Groq 3 LPU Hot Chips 2026 Key Takeaways
NVIDIA Groq 3 LPU Hot Chips 2026 Key Takeaways

And that is the Groq 3 LPU, and NVIDIA’s Vera Rubin LPX rack. The disaggregation unlocks new performance possibilities for the NVIDIA ecosystem, and gives the platform the tools needed to offer much faster low-latency inference than what GPUs can provide on their own.

LEAVE A REPLY

Please enter your comment!
Please enter your name here

This site uses Akismet to reduce spam. Learn how your comment data is processed.