Qualcomm Elite Gen 6 NPU: Adding Element Accelerator for Transformers
Rounding out our look at the compute blocks of the Elite Gen 6 family, let us talk about the Hexagon NPU.
At a high level, the NPU on the Elite Gen 6 is architecturally very similar to the Elite Gen 5. Qualcomm’s collection of matrix, scalar, and vector cores does not have any new major features, and the lowest supported precisions remain unchanged at INT2 and FP8. Even the core counts are the same, with 1 matrix core, 8 vector cores, and 12 scalar cores.

What has changed, however, is that Qualcomm has added a fourth compute/accelerator block to the Hexagon NPU: what they are calling an element accelerator. According to the company, the purpose of this block is to specifically accelerate transformers, which in the last few years have become a key technology in AI inference. Specifically, this block seems to be designed to accelerate attestation, which is something of a weakness in more basic NPU designs.

Qualcomm is not quoting a specific performance improvement figure from the element accelerator alone, but overall the company is touting a 14% improvement in NPU performance for the vanilla Elite Gen 8, with a 20% improvement in performance-per-watt. Notably, all versions of the Elite Gen 6 get the updated NPU with the element accelerator, so this is not restricted to the Extreme.
What is restricted to the Extreme, however, is a larger shared memory pool for the NPU. Qualcomm is not disclosing the precise size of this shared memory for either SKU, but the Extreme gets a 50% larger cache. And, accordingly, it is able to deliver higher performance: Qualcomm says that the Extreme gets 35% better NPU performance than the Elite Gen 5, or a 33% improvement in per-per-watt.
Qualcomm is also framing the shared memory pool as a key constraint on overall model size. Specifically, the company is saying that the Extreme is fast enough to run 30B parameter MoE models – but only because of its larger shared memory.
Combined with GPU matrix cores being restricted to the Extreme SKU, there is a clear pattern of Qualcomm restricting its newest AI hardware features (and the performance they enable) to the Extreme SKU. So rather than the chips differing significantly by CPU performance or even GPU performance, the biggest difference between the two of them looks to be their AI inference performance.
Qualcomm Elite Gen 6 Memory: Adding LPDDR6 Support
Feeding the beast that is the rest of the Snapdragon 8 Elite Gen 6 family, Qualcomm is also using this generation to upgrade the memory capabilities of their flagship SoCs.
Brand-new to Gen 6, Qualcomm is introducing LPDDR6 memory support. The next-gen low-power memory is just now becoming available, and over time will be replacing LPDDR5(X) as the cutting-edge memory technology for energy-efficient devices (be it smartphones or servers).

LPDDR6 brings its own slew of changes, but the most important improvements are all focused around memory bandwidth. And the standard accomplishes this in two ways:
First and foremost, LPDDR6 is getting a wider memory bus. Whereas LPDDR5(X) is typically configured with one (or more) 16-bit-wide memory channels, LPDDR6 will have an effective width of 24 bits, comprised of two 12-bit-wide subchannels. As a result, a 4-channel design, such as a flagship Qualcomm chip, will feature a 96-bit wide (24×4) memory bus when using LPDDR6, a 50% wider bus than LPDDR5(X).

The second big bandwidth improvement from LPDDR6 comes from support for higher clock speeds. Over time, the memory standard is expected to reach as high as 14,400 MT/sec, whereas LPDDR5X today is topping out at 10,600 MT/sec. However, this latter aspect will not be applicable to Elite Gen 6 phones, as the memory controller is not qualified for those higher memory speeds.
Ultimately, LPDDR6 support will sit alongside LPDDR5X support for the Elite Gen 6 chips. So device vendors can pair the SoC with either memory type, depending on cost considerations and memory availability. The LPDDR5X performance of the Elite Gen 6 is unchanged from the Elite Gen 5, with the chip supporting 4×16-bit channels of LPDDR5-10600, for around 84.8GB/second of memory bandwidth. For LPDDR6, on the other hand, we are looking at 4×24-bit channels of LPDDR6-10600, for a peak memory bandwidth of 127.2GB/second.
The full performance impact of LPDDR6 on the Elite Gen 6 series remains to be seen. But on paper, at least, LPDDR6 promises a good deal more bandwidth than what LPDDR5X can deliver.


