The H100 does about 989 trillion 16-bit operations every second. Rubin has 896 Tensor Cores to the H100’s 528.
Side by side
Rubin and the H100
Memory80 GB
288 GB · 3.6×
Bandwidth3.35 TB/s
22 TB/s · 6.6×
H100 · 989 TFLOPS16-bit
Rubin · 50 PFLOPS4-bit NVFP4
Different number formats. Those two do not compare directly.
The catch
Operations per byte
Rubin needs · ~2,270 ops per byte
One user, 4-bit weights · ~4
One user, 16-bit weights · ~1
50 PFLOPS ÷ 22 TB/s. A chatbot answering one person does about one. So most of the time, even NVIDIA’s most powerful chip is waiting.
Waiting, in time
8B model, 4-bit weights, one user
Read the weights4 GB ÷ 22 TB/s≈ 180 µs
Do the math15 B ops ÷ 50 PFLOPS≈ 0.3 µs
Not waiting for math. Waiting for memory.
Keeping it fed
Do more work with every byte
Read each weight once↓
User 1
User 2
…
User 500
~500× the work per byte, limited by KV-cache memory
16-bit weights16 GB
4-bit weights4 GB
Blackwell8 TB/s
Rubin22 TB/s · 2.8×
Beyond the chip
From one GPU to a rack
Rubin GPU · 288 GB↓ ×2 + 1 Vera CPUVera Rubin Superchip · 576 GB↓ ×36 over NVLink 6NVL72 · 72 GPUs · 20.7 TB · 3.6 exaFLOPS NVFP4↓ InfiniBand or EthernetMany racks → an AI factory
Remember
Keep the Tensor Cores fed
One small job: multiply and add↓896 Tensor Cores · 224 SMs · 8 clusters↓Two dies · eight HBM4 stacks · 288 GB · 22 TB/s↓Needs 2,000+ ops per byte↓Batching and small formats keep it busy