Advertisement


Home Networking NVIDIA Spectrum-X Ethernet Multiplane Network Architecture at Hot Chips 2026

NVIDIA Spectrum-X Ethernet Multiplane Network Architecture at Hot Chips 2026

0
Hot Chips 2026 NVIDIA Spectrum-X Multiplane Network Architecture Slide 16 Spectrum-X Ethernet Multiplane Extends AI-Factory Scale by 64X
Hot Chips 2026 NVIDIA Spectrum-X Multiplane Network Architecture Slide 16 Spectrum-X Ethernet Multiplane Extends AI-Factory Scale by 64X

NVIDIA is presenting the Spectrum-X Ethernet Multiplane Network Architecture here at Hot Chips 2026. This talk explains how NVIDIA plans to scale AI factory networking from thousands of GPUs toward half a million.

This article is being written live from the presentation, so please excuse any typos.

NVIDIA Spectrum-X Ethernet Multiplane Network Architecture at Hot Chips 2026

NVIDIA is showing the same slide for the 4th time at Hot Chips 2026. I think it is done for framing Agentic AI, but it is also the 4th time (at least) we have seen this.

Hot Chips 2026 NVIDIA Spectrum-X Multiplane Network Architecture Slide 2 Agentic AI Is the Most Complex
Hot Chips 2026 NVIDIA Spectrum-X Multiplane Network Architecture Slide 2 Agentic AI Is the Most Complex

Again, we get another look at the AI Factory platform. Gilad is saying this is because this is what NVIDIA is building.

Hot Chips 2026 NVIDIA Spectrum-X Multiplane Network Architecture Slide 3 NVIDIA Vera Rubin Is a Full-Stack AI Factory Platform
Hot Chips 2026 NVIDIA Spectrum-X Multiplane Network Architecture Slide 3 NVIDIA Vera Rubin Is a Full-Stack AI Factory Platform

Here NVIDIA divides the AI factory into five purpose-built networks: scale-across, scale-in, scale-out, scale-up, and the AI context scale. That scale-in is the new one today, and it was a bit awkward watching Broadcom’s Thor Ultra presentation today knowing that scale-in was going to become a thing in the subsequent presentation. A single general-purpose fabric cannot serve all of these well, and that is the load-bearing claim behind the rest of the talk.

Hot Chips 2026 NVIDIA Spectrum-X Multiplane Network Architecture Slide 4 AI Factories Need Five Purpose-Built Networks
Hot Chips 2026 NVIDIA Spectrum-X Multiplane Network Architecture Slide 4 AI Factories Need Five Purpose-Built Networks

A quick comment that NVIDIA co-packaged optics solutions are in production. We are also talking about why NVIDIA needs Astra, scale-in, and BlueField-4.

NVIDIA’s scale-up networking runs on NVLink. NVLink’s NVL72 rack pairs an NVLink spine with switch trays to form a 72-GPU scale-up domain, which NVIDIA credits with leading tokens per megawatt.

Hot Chips 2026 NVIDIA Spectrum-X Multiplane Network Architecture Slide 5 NVIDIA NVLink Scale-up Networking Fabric for AI Factories
Hot Chips 2026 NVIDIA Spectrum-X Multiplane Network Architecture Slide 5 NVIDIA NVLink Scale-up Networking Fabric for AI Factories

Enterprise, hyperscale, and service provider networks each have their own spine designs, and NVIDIA argues that AI factories need a purpose-built Ethernet rather than reusing general-purpose topologies.

Hot Chips 2026 NVIDIA Spectrum-X Multiplane Network Architecture Slide 6 The Different Ethernet Architectures
Hot Chips 2026 NVIDIA Spectrum-X Multiplane Network Architecture Slide 6 The Different Ethernet Architectures

Spectrum-X is NVIDIA’s giga-scale answer for scale-out Ethernet. NVIDIA claims 1.6x higher RDMA bandwidth, 2.2x better multi-tenancy, and 1.3x lower bandwidth jitter around 102.4T switch systems and 1.6T SuperNICs. Our earlier look at the MRC RDMA transport protocol digs into the transport underneath Spectrum-X. NVIDIA is focused on solving jitter not just with a single device but with an entire end-to-end system.

Hot Chips 2026 NVIDIA Spectrum-X Multiplane Network Architecture Slide 7 Spectrum-X:
Hot Chips 2026 NVIDIA Spectrum-X Multiplane Network Architecture Slide 7 Spectrum-X:

NVIDIA’s headline here is 1.9x higher training performance in a multi-tenant AI factory, drawn from a DeepSeek V3 multi-job training run. In that noisy shared environment, step times for off-the-shelf Ethernet climb well above the Spectrum-X curves, with the gap widening across the 80-plus steps shown. NVIDIA attributes the separation to low-jitter communications and tenant noise isolation.

Hot Chips 2026 NVIDIA Spectrum-X Multiplane Network Architecture Slide 8 1.9x Higher Training Performance in Multi-Tenant AI Factory
Hot Chips 2026 NVIDIA Spectrum-X Multiplane Network Architecture Slide 8 1.9x Higher Training Performance in Multi-Tenant AI Factory

Extreme co-design is most evident in NCCL performance. For the main NCCL operations in Nemotron Ultra pre-training, gradient all-reduce leads at 14x, token dispatch sits at 3.5x, and both gradient reduce-scatter and parallel matrix multiply land at 2x, all relative to off-the-shelf Ethernet.

Hot Chips 2026 NVIDIA Spectrum-X Multiplane Network Architecture Slide 9 Up to 14x Higher NCCL Performance
Hot Chips 2026 NVIDIA Spectrum-X Multiplane Network Architecture Slide 9 Up to 14x Higher NCCL Performance

Scale-out AI leans on optics to a degree traditional clouds do not. NVIDIA puts the optical share in an AI factory at 10% of the compute power, well above what a conventional cloud data center typically carries.

Hot Chips 2026 NVIDIA Spectrum-X Multiplane Network Architecture Slide 10 Scale-Out AI Depends on Optics
Hot Chips 2026 NVIDIA Spectrum-X Multiplane Network Architecture Slide 10 Scale-Out AI Depends on Optics

Spectrum-X Ethernet photonics is the production play behind that. NVIDIA has a co-packaged optics chip with micro ring modulators in production and a 3D-stacked silicon photonics engine on the TSMC Coupe process, and it claims 4x fewer lasers, lower power, and 10x lower mean time between interruptions.

Hot Chips 2026 NVIDIA Spectrum-X Multiplane Network Architecture Slide 11 Spectrum-X Ethernet
Hot Chips 2026 NVIDIA Spectrum-X Multiplane Network Architecture Slide 11 Spectrum-X Ethernet

“100G SerDes is last year’s stuff” just after Broadcom presents 100G SerDes Thor Ultra at Hot Chips. Awesome.

Here is the chip:

NVIDIA CES 2026 Keynote Spectrum X Co Packaged Optics
NVIDIA CES 2026 Keynote Spectrum X Co-Packaged Optics

Today’s hyperscale cloud commonly starts from a top-of-rack switch topology. This 2k-scale baseline is the conventional layout NVIDIA wants to move beyond for AI factories.

Hot Chips 2026 NVIDIA Spectrum-X Multiplane Network Architecture Slide 13 Traditional Hyperscale Cloud Top-of-Rack Switch Topology
Hot Chips 2026 NVIDIA Spectrum-X Multiplane Network Architecture Slide 13 Traditional Hyperscale Cloud Top-of-Rack Switch Topology

Multi-rail is the first step up. NVIDIA shows 8k Rubin GPUs at 1.6T per-GPU scale-out bandwidth, with 100T switches carrying 64 1.6T ports and splitting traffic across four rails. In multi-rail, the full bandwidth of the NIC goes to the switch port.

Hot Chips 2026 NVIDIA Spectrum-X Multiplane Network Architecture Slide 14 AI Factory With Spectrum-X Ethernet Multi-Rail Topology
Hot Chips 2026 NVIDIA Spectrum-X Multiplane Network Architecture Slide 14 AI Factory With Spectrum-X Ethernet Multi-Rail Topology

Now NVIDIA introduces the multiplane topology. Multiplane separates the fabric into planes that can share switch hardware, rather than dedicating full rails, and that is the core idea of this talk. This is where the NIC connects to multiple switches. Instead of 1.6Tbps to one switch, this is 8x 200Gbps, with each link going to a different switch.

Hot Chips 2026 NVIDIA Spectrum-X Multiplane Network Architecture Slide 15 Introducing Spectrum-X Ethernet Multiplane Topology
Hot Chips 2026 NVIDIA Spectrum-X Multiplane Network Architecture Slide 15 Introducing Spectrum-X Ethernet Multiplane Topology

Multiplane extends AI factory scale by 64x over multi-rail. NVIDIA lands at 512k Rubin GPUs, still at 1.6T scale-out bandwidth per GPU, using 100T switches with 512 ports of 200G laid out as eight planes across four rails.

Hot Chips 2026 NVIDIA Spectrum-X Multiplane Network Architecture Slide 16 Spectrum-X Ethernet Multiplane Extends AI-Factory Scale by 64X
Hot Chips 2026 NVIDIA Spectrum-X Multiplane Network Architecture Slide 16 Spectrum-X Ethernet Multiplane Extends AI-Factory Scale by 64X

Multiplane also trims the physical footprint. NVIDIA claims 1.7x fewer scale-out switches than a traditional multi-tier single-rail topology, reducing power, rack space, and costs.

Hot Chips 2026 NVIDIA Spectrum-X Multiplane Network Architecture Slide 17 Spectrum-X Ethernet Multiplane Saves 1.7x Scale-Out Switches
Hot Chips 2026 NVIDIA Spectrum-X Multiplane Network Architecture Slide 17 Spectrum-X Ethernet Multiplane Saves 1.7x Scale-Out Switches

Fault behavior is where multiplane is meant to pay off. Under partial bandwidth loss, the Spectrum-X multiplane topology holds 90% bandwidth with a 2.68ms detection time, while a traditional multi-tier topology drops to 0% bandwidth.

Hot Chips 2026 NVIDIA Spectrum-X Multiplane Network Architecture Slide 18 Spectrum-X Ethernet Multiplane Maintains 90% Bandwidth
Hot Chips 2026 NVIDIA Spectrum-X Multiplane Network Architecture Slide 18 Spectrum-X Ethernet Multiplane Maintains 90% Bandwidth

NVIDIA converts that into a 1.6x goodput advantage over off-the-shelf Ethernet multiplane. Spectrum-X detects faults in 2.68ms, roughly 400x faster, and recovers in 100ms, about 11x faster, holding 90% bandwidth, whereas the OTS topology suffers full bandwidth loss with 1080ms detection and recovery.

Hot Chips 2026 NVIDIA Spectrum-X Multiplane Network Architecture Slide 19 Spectrum-X Ethernet Multiplane Delivers 1.6X Higher
Hot Chips 2026 NVIDIA Spectrum-X Multiplane Network Architecture Slide 19 Spectrum-X Ethernet Multiplane Delivers 1.6X Higher

Spectrum-XGS reaches across multiple AI factories. With 800Gb/s per port on 102.4Tb/s switches and ConnectX-9 SuperNICs, NVIDIA claims up to 1.9x lower multi-site latency and doubled scale-across performance for distributed AI operations.

Hot Chips 2026 NVIDIA Spectrum-X Multiplane Network Architecture Slide 20 Spectrum-XGS Connects AI Factories
Hot Chips 2026 NVIDIA Spectrum-X Multiplane Network Architecture Slide 20 Spectrum-XGS Connects AI Factories

NVLink Fusion brings third-party XPUs onto the NVIDIA AI platform. NVIDIA shows 3.6 TB/s all-to-all bandwidth per XPU connecting 72 XPUs in a single domain, wrapped in a cableless MGX rack rated for 45C inlet and 100% liquid cooling, with 3x lower latency.

Hot Chips 2026 NVIDIA Spectrum-X Multiplane Network Architecture Slide 21 NVLink Fusion Connects XPUs to the NVIDIA AI Platform
Hot Chips 2026 NVIDIA Spectrum-X Multiplane Network Architecture Slide 21 NVLink Fusion Connects XPUs to the NVIDIA AI Platform

Software must keep pace with that hardware. DSX is NVIDIA’s AI factory platform, spanning power optimization and infrastructure software plus platform software across DSX OS, DSX Sim, DSX MaxLPS, and DSX Flex.

Hot Chips 2026 NVIDIA Spectrum-X Multiplane Network Architecture Slide 22 NVIDIA DSX AI Factory Platform
Hot Chips 2026 NVIDIA Spectrum-X Multiplane Network Architecture Slide 22 NVIDIA DSX AI Factory Platform

NVIDIA argues AI factories need simulation and validation at scale before construction. This is more important when you have increasingly large network topologies because you also need to understand the cabling needs.

Hot Chips 2026 NVIDIA Spectrum-X Multiplane Network Architecture Slide 23 AI Factory Requires Simulation and Validation at Scale
Hot Chips 2026 NVIDIA Spectrum-X Multiplane Network Architecture Slide 23 AI Factory Requires Simulation and Validation at Scale

DSX Air compresses that deployment timeline from months to days. NVIDIA shows infrastructure bring-up falling from six months to one week, infrastructure software from three months to one week, and deployment from two weeks to one day.

Hot Chips 2026 NVIDIA Spectrum-X Multiplane Network Architecture Slide 24 DSX Air Compresses Orchestrator
Hot Chips 2026 NVIDIA Spectrum-X Multiplane Network Architecture Slide 24 DSX Air Compresses Orchestrator

NVIDIA closes with the five networking infrastructures of the AI factory. Scale-in, scale-up, scale-out, scale-across, and context scale each get purpose-built silicon, and the figure piles on multipliers from 18x down across bandwidth, packet rate, latency, and jitter.

Hot Chips 2026 NVIDIA Spectrum-X Multiplane Network Architecture Slide 25 The Five Networking Infrastructures of NVIDIA AI Factory
Hot Chips 2026 NVIDIA Spectrum-X Multiplane Network Architecture Slide 25 The Five Networking Infrastructures of NVIDIA AI Factory

Every networking dimension shown here is purpose-built and co-designed with the platform rather than borrowed from a general-purpose fabric. That is the through-line of the entire talk.

Final Words

NVIDIA is making the argument that Ethernet can handle AI factory scale-out, but only when the hardware is designed specifically for AI. This multiplane topology is the headline, expanding the addressable scale from 8k GPUs to 512k GPUs while claiming fewer switches and better fault behavior. Whether those figures hold in live deployments is the real test, and the rest of Hot Chips should show how the Vera Rubin software-and-optics story comes together.

LEAVE A REPLY

Please enter your comment!
Please enter your name here

This site uses Akismet to reduce spam. Learn how your comment data is processed.