While we are out building fun clusters, Gigabyte is doing something really interesting. It has a 40-node cluster with 320 cores 1.28TB of memory, 80x SSDs, and 40 (integrated) GPUs in a single 1U chassis. We encountered Gigabyte’s hardware on display at Computex 2026, and it is quite neat.
A 40-Node 1U Cluster Gigabyte R1C7-K0A-AS1 at Computex 2026
Here are the quick specs of the cluster. There are 40 Intel Core Ultra 7 258V (Lunar Lake) nodes, each with 32GB of LPDDR5X memory and two M.2 slots.

To understand the server, Gigabyte took almost GPU-looking cartridges and installed five of them in the server with three in front and two in the second row.

Taking one of these cartridges out, you can see that there are eight nodes.

The nodes all come on a small board with relatively few connections to the overall system. Under one heatsink we have the Lunar Lake CPU. Under the other, we have two PCIe Gen5 x2 M.2 slots.

Moving to the back of the system, this is where things are really neat. There is not a lot of information on this one, but it looks like each cartridge is getting two MCIO 8i connectors. Again, this is all speculation at this point since this is not in the specs shown. The big heatsink at the rear is likely how all of this is consolidated to two QSFP28 ports at the rear.

Inside, there is a chassis management controller, so there is a rear management port. We also get two 3.2kW Titanium rated power supplies. Most interesting, however, are the two QSFP28 ports.

Nobody we talked to knew how this was connected internally. If you look at the cartridge connectors, they do not have a huge number of pins for power and data. If those are MCIO connections, then with 16 lanes per cartridge, that is x2 per node. Alternatively, they could be carrying Ethernet, which would make this system really easy, since the chip under the big heatsink at the rear of the chassis would then be an Ethernet switch, and the topology would be extremely easy to manage.
Final Words
Overall, this is a neat system. It is also one that we could certainly see in the future. This is using Lunar Lake, so Patrick’s take on this:
“Companies have used Intel iGPU nodes in clusters for years, like the old Intel Visual Compute Accelerator, which was popular with a certain Indian Broadcaster, from what I remember. Intel Quick Sync transcoding is super easy to integrate, and you also get x86 cores for your pipeline. Alternatively, these designs have been used for physical desktop nodes in the data center and cloud.”
Gigabyte’s R1C7-K0A-AS1 is not yet released, so hopefully, we get to see more of it in the future. One fun and interesting note: given the density here, with 4 P cores and 4 E cores per node (8 total), there are a total of 320 cores per U. Which over 40U of rack space adds up to 12800 CPU cores, 3200 SSDs, 1600 iGPUs, and 51.2TB of LPDDR5X memory that can be connected via 80 power connections, 80 100GbE ports, and 40 management ports. That is a higher density than solutions like the upcoming Arm AGI CPU, which makes it a neat stat on just how dense these systems are.




How much do you think it will cost?
I sincerely hope these nodes can be packaged individually. Pair them with an inexpensive mini-ITX or micro-ATX carrier board, they could be funnelled into the small home server market or rescued from the trash bin at the next hardware refresh. Perfect for Project TinyMiniMicro!
These are a decent option if you have limited rack space and need high-density hosting, like for a CDN. A few drawbacks: the drives aren’t hot-swappable, so you must shut down the node to replace media. They draw a lot of power — roughly two of these units per 40U rack on 208V/30A circuits (3200 W ÷ 208 V ? 15.4 A). Instead of having physical drives per node, I would suggest using ‘network storage’ which you would then have your hot-swappable media.
Um… Why? Look at how much of that memory is wasted with read-only operating system, library, and actual data? This kind of system design makes little sense to me. They should be mostly OS-less devices connected to a very, very smart NIC / storage interface. With some local scratch storage.
Yes, software needs to change to use this kind of density. But the overall waste of *not* changing needs to be in the accounting.
No ECC RAM, right?
No ECC. This is Lunar Lake, so the only memory is on-package, and Intel did not offer ECC configurations.
In previous companies we had the Moonshot for video transcoding because they did an Xeon E3 with QSV. It was very powerful and quite efficient for the time, we had it running probably 500-700 transcodes in parallel. I replaced that with a handful of higher core count machines maxed out with Intel DC Flex GPUs, but now Intel doesn’t care about supplying the market with GPUs, even enterprise. So I might take a look at this in the future.