← Back to the lesson
1 / 1

Week 3

Inside NVIDIA Vera Rubin

NVIDIA’s newest AI chip, taken apart one layer at a time.

Click or press → to reveal each idea. Use the buttons on a slide to explore.

Meet Rubin

336 billion transistors

Vera Rubin — the platform
VeraThe CPU · 88 Arm cores
RubinThe GPU
H10080 B
Blackwell208 B
Rubin336 B

Nearly all of those transistors exist to do one small job, very fast.

One small job

Multiply two numbers. Add to a running total.

total= total+ a× b
8 billion parameters
~0.5 BEmbedding · looked up
~7.5 BMultiply-add
~7.5 billion multiply-adds per token

Five layers

Take it apart

5 · LidProtects and spreads heat
HBM4HBM4 Compute die 1Compute die 2 HBM4HBM4 HBM4HBM4 HBM4HBM4

3 · Two compute dies  ·  4 · Eight memory stacks

2 · Silicon interposerFine wiring between dies and memory
1 · BoardPower in, signals out

Why two dies and not one big one?

Why two dies

No chip can be bigger than one field

One Rubin die
fills almost all of it

One exposure field · 26 × 33 mm

Die 1
Die 2

Joined to work as one GPU

Chips are printed one field at a time. So NVIDIA builds two dies and joins them.

The memory

Eight stacks, 288 GB

8stacks
×
12DRAM chips each
→
36 GBper stack
=
288 GBup to 22 TB/s

This is where a model’s weights live.

Inside a stack

Down, then sideways

Base die
↓ then →through the interposer
GPU die
Silicon interposer

Copper channels run straight down through every chip, so the whole stack shares one very wide connection.

A wider highway

384 lanes, or 16,384?

Gaming GPU 384-bit bus
Rubin 8 × 2,048 = 16,384-bit

HBM4 doubles the width of each stack. It only works because the memory sits millimetres from the compute.

From above

Links, clusters, and SMs

NVLink 6 · other GPUs
C2C · Vera CPU
GPC28 SMs
GPC28 SMs
GPC28 SMs
GPC28 SMs
Shared L2 cacheGigaThread Engine
GPC28 SMs
GPC28 SMs
GPC28 SMs
GPC28 SMs
PCIe Gen 6 · host
Memory controllers · HBM4
NVLink 6Other GPUs · 3.6 TB/s
NVLink-C2CVera CPU · 1.8 TB/s
PCIe Gen 6Host · 256 GB/s
Tensor Core
Tensor Core
Tensor Core
Tensor Core
CUDA cores · schedulers · registers · shared memory / L1

Count them

896 Tensor Cores

8clusters
×
28SMs each
=
224SMs
×
4Tensor Cores per SM
=
896Tensor Cores

GPU → GPC → SM → Tensor Core.

The Tensor Core

Counting with the H100

512multiply-adds per tick
×
4per SM
×
132SMs
×
2ops each
×
~1.83 Bticks / s
=
~989 Tops / s

The H100 does about 989 trillion 16-bit operations every second. Rubin has 896 Tensor Cores to the H100’s 528.

Side by side

Rubin and the H100

Memory80 GB
288 GB · 3.6×
Bandwidth3.35 TB/s
22 TB/s · 6.6×
H100 · 989 TFLOPS16-bit
Rubin · 50 PFLOPS4-bit NVFP4

Different number formats. Those two do not compare directly.

The catch

Operations per byte

Rubin needs · ~2,270 ops per byte
One user, 4-bit weights · ~4
One user, 16-bit weights · ~1

50 PFLOPS ÷ 22 TB/s. A chatbot answering one person does about one. So most of the time, even NVIDIA’s most powerful chip is waiting.

Waiting, in time

8B model, 4-bit weights, one user

Read the weights 4 GB ÷ 22 TB/s ≈ 180 µs
Do the math 15 B ops ÷ 50 PFLOPS ≈ 0.3 µs

Not waiting for math. Waiting for memory.

Keeping it fed

Do more work with every byte

Read each weight once
User 1
User 2
…
User 500
~500× the work per byte, limited by KV-cache memory
16-bit weights16 GB
4-bit weights4 GB
Blackwell8 TB/s
Rubin22 TB/s · 2.8×

Beyond the chip

From one GPU to a rack

Rubin GPU · 288 GB Vera Rubin Superchip · 576 GB NVL72 · 72 GPUs · 20.7 TB · 3.6 exaFLOPS NVFP4 Many racks → an AI factory

Remember

Keep the Tensor Cores fed

One small job: multiply and add 896 Tensor Cores · 224 SMs · 8 clusters Two dies · eight HBM4 stacks · 288 GB · 22 TB/s Needs 2,000+ ops per byte Batching and small formats keep it busy

Open the full lesson →

Click or → to reveal