NVIDIA unveils Rubin promising up to 10x inference cost reduction
NVIDIA just introduced Rubin, its latest AI supercomputing platform, at CES 2026.
Rubin is built to make developing and running huge AI models easier and much cheaper: think up to 10 times lower inference token costs and requiring only one-fourth of the graphics cards compared to NVIDIA's last-generation Blackwell system to train mixture-of-experts (MoE) models.
Rubin packs 6 chips, 50 petaflops
Rubin packs six integrated chips into one powerful machine, featuring an 88-core Vera CPU and a GPU that can hit up to 50 petaflops of performance.
With smarter design and faster connections, Rubin is all about slashing infrastructure costs so more companies (and cloud giants like AWS, Google Cloud, and Microsoft) can use advanced AI without breaking the bank.
It's set to roll out to partners sometime in the second half of 2026, so expect even bigger things from AI soon.