NVIDIA unveils Rubin at CES to cut inference token costs
NVIDIA just dropped its latest AI supercomputing platform, Rubin, at CES 2026. The big news? Rubin is designed to make running huge AI models way more affordable.
NVIDIA says it can cut inference token costs by up to 10 times and needs only one-fourth of the graphics cards compared to its last-generation Blackwell system for training mixture-of-experts (MoE) models.
It's all thanks to a fresh "extreme codesign" setup where six different chips team up as one powerful supercomputer.
Cloud providers line up for Rubin
Rubin packs some serious hardware: an 88-core Vera CPU, a next-generation GPU with massive processing power (50 petaflops), and speedy networking technology for smoother data flow.
It'll be ready for action in the second half of 2026, and big players like Amazon Web Services, Google Cloud, and Microsoft are already lining up to use it for their own AI projects.
So if you're into tech or curious about where AI is headed next, Rubin is definitely one to watch.