Inside NVIDIA Vera Rubin

NVIDIA’s newest AI chip, taken apart one layer at a time.

Open slides →

We’ll start with the one small job it does ↓

Chapter 1

Meet Rubin

In the Introduction to GPU lesson, we looked at the parts found inside any modern GPU: transistors, SMs, Tensor Cores, memory, NVLink, and cooling.

Now let’s look at a real one, and a very new one.

This is Rubin, NVIDIA’s newest chip for artificial intelligence.

It holds 336 billion transistors, and nearly all of them exist to do one small job, very fast.

First, a quick note on the name, because it gets mixed up often:

Vera Rubin — the platform
VeraThe CPU · 88 Arm cores
RubinThe GPU · this lesson
Vera Rubin Superchip = 1 Vera CPU + 2 Rubin GPUs Vera Rubin NVL72 = 36 Vera CPUs + 72 Rubin GPUs in one rack
Vera is the CPU. Rubin is the GPU. Vera Rubin is the platform that puts them together. Both are named after the astronomer Vera Rubin.

Most of this lesson is about the Rubin GPU. We’ll come back to the full platform at the end.

To get a sense of scale, here is how the transistor count has grown across NVIDIA’s data-center GPUs:

Transistor counts of recent NVIDIA data-center GPUs
GPUTransistorsCompute dies
H100 (Hopper)80 billion1
B200 (Blackwell)208 billion2
Rubin336 billion2
Rubin has about 4× the transistors of the H100 and about 1.6× Blackwell.
Chapter 2

One small job

So what is the one small job?

Multiply two numbers, and add the result to a running total.

Then do it again. And again.

total = total + a × b
A multiply-add, also called a fused multiply-add or FMA. This is the operation at the heart of every neural network.

Why this operation? Because a neural network is mostly matrix multiplication, and matrix multiplication is nothing more than this multiply-add repeated a very large number of times.

Every weight in the model gets multiplied by some input value, and the results get added together.

To write a single token, a small piece of a word, a model with 8 billion parameters does this about 7.5 billion times.

Why 7.5 and not 8?

Because roughly half a billion of those parameters are the embedding table. The model simply looks up a row in that table for each token. It does not multiply through it. Nearly every other weight is used once per token, in one multiply-add.

8 billion parameters
~0.5 BEmbedding table · looked up
~7.5 BUsed in a multiply-add
~7.5 billion multiply-adds per token
Attention over earlier tokens adds a little more work, and that part grows as the conversation gets longer.

And that is for one token. A single answer may contain hundreds of tokens, and a real service is answering many users at the same time.

AI hardware = a machine for doing multiply-adds as fast as possible
Chapter 3

Five layers

Take Rubin apart and there are five layers.

5LidProtects the chip and spreads heat to the cooling
3Two compute dies4Eight memory stacks
HBM4 HBM4 Compute die 1 Compute die 2 HBM4 HBM4 HBM4 HBM4 HBM4 HBM4
2Silicon interposerExtremely fine wiring between the dies and the memory
1Board (package substrate)Carries power in and signals out
A board, a silicon interposer, two compute dies, eight stacks of memory, and a lid. A simplified view; the real layout is more detailed.

From the bottom up:

  1. A board
    The package substrate. It carries power into the chip and connects it to the rest of the system.
  2. A silicon interposer
    A thin layer with extremely fine wiring. The dies and the memory all sit on it, so they can talk to each other over very short, very wide connections.
  3. Two compute dies
    This is where the transistors, SMs, and Tensor Cores live.
  4. Eight stacks of memory
    HBM4, sitting right next to the compute dies.
  5. A lid
    It protects the silicon and helps move heat into the cooling system.

Remember the GPU anatomy picture from the Introduction to GPU lesson? This is the same idea, zoomed in on the package that sits at the center of the card.

One question stands out: why two dies, and not one big one?

Chapter 4

Why two dies

Chips are printed onto silicon wafers, one small field at a time.

A lithography machine projects the chip’s pattern onto the wafer through a mask called a reticle. It exposes one rectangle, steps over, exposes the next, and so on across the wafer.

No chip can be bigger than one field: 26 × 33 mm, about 858 mm².

This is called the reticle limit. It comes from the optics of the machine itself. The lens can only keep a pattern that size sharp and undistorted.

One Rubin die
fills almost all of it

One exposure field · 26 × 33 mm

Die 1
Die 2

Two dies, joined to work as one GPU

Each Rubin die is built right up to the reticle limit. To get more transistors, NVIDIA builds two and joins them.

Each Rubin die fills almost all of that field.

So NVIDIA builds two, and joins them to work as one GPU.

The connection between them is called NV-HBI, the NVIDIA High-Bandwidth Interface. It is fast enough that software sees one GPU, not two.

There is a second reason this approach helps: yield.

Every wafer has a few tiny defects scattered across it. A defect usually ruins the die it lands on. The bigger the die, the more likely it is to be hit. Two reticle-sized dies, tested separately before they are joined, waste less silicon than one impossibly large die would.

Blackwell used the same two-die idea. Rubin keeps it and puts more into each die.

Can’t print a bigger chip → print two chips → join them so they behave like one
Chapter 5

The memory

Around the dies stand eight memory stacks, holding 288 GB.

This is where a model’s weights live.

In the Introduction to GPU lesson, we called this VRAM. On data-center GPUs, it is built from a special kind of memory called HBM, or High Bandwidth Memory. Rubin uses the newest version, HBM4.

How Rubin reaches 288 GB of memory
Memory stacks8
DRAM chips in each stack12 (a “12-high” stack)
Capacity per stack36 GB
Total8 × 36 GB = 288 GB
Peak bandwidthup to 22 TB/s
Eight 12-high stacks of HBM4 give Rubin 288 GB of memory.

Inside one stack

Here is how a stack is built. The one pictured in most teardowns is from an earlier generation, but the idea is the same.

Memory chips are piled on top of a base chip, right beside the compute.

Copper channels run straight down through every chip. These are called through-silicon vias, or TSVs. Because they pass through every layer, the whole stack shares one very wide connection.

The data flows down, then sideways into the GPU.

Base die
↓ then →through the interposer
GPU die
Silicon interposer
Twelve DRAM chips on a base die. Copper TSVs (teal) run straight down through every layer, then the data moves sideways through the interposer into the GPU.

Why this design matters

Remember the highway analogy from the bandwidth section of the Introduction to GPU lesson?

A gaming GPU might have a 384-bit memory bus.

HBM4 gives each stack a 2,048-bit interface, twice as wide as HBM3e. With eight stacks, Rubin has a 16,384-bit highway to its memory.

Gaming GPU 384-bit bus A few lanes
Rubin 8 × 2,048 = 16,384-bit A very wide highway
Stacking memory right next to the compute is what makes such a wide connection possible.

You could not run that many wires across a normal circuit board. It only works because the memory sits on the interposer, millimetres away from the compute dies.

Stack the memory → put it next to the compute → connect it with a very wide road
Chapter 6

Links, clusters, and SMs

Now look at the compute side from above.

NVLink 6 · to other GPUs · 3.6 TB/s
NVLink-C2C · to Vera CPU
GPC
28 SMs
GPC
28 SMs
GPC
28 SMs
GPC
28 SMs
Shared L2 cacheGigaThread Engine
GPC
28 SMs
GPC
28 SMs
GPC
28 SMs
GPC
28 SMs
PCIe Gen 6 · to the host
Memory controllers · to the eight HBM4 stacks
A simplified floor plan. Links along the edges, the shared L2 cache and the work scheduler in the middle, and eight clusters of 28 SMs around them.

Along the edges: links to the outside world

Along the edges are its links to the outside world: to the CPU, to other GPUs, and to everything else.

Rubin’s external links
LinkConnects toBandwidth
NVLink 6Other Rubin GPUs, through NVLink switches3.6 TB/s
NVLink-C2CThe Vera CPU, with shared, coherent memory1.8 TB/s
PCIe Gen 6 (x16)The host system and other devicesup to 256 GB/s
NVLink 6 doubles the per-GPU GPU-to-GPU bandwidth of the previous generation.

This is the same picture as the NVLink section of the Introduction to GPU lesson: PCIe to the host, NVLink between GPUs. Rubin adds a very fast, direct link to its own CPU.

In the middle: a shared cache and a work scheduler

In the middle sits a shared cache and an engine that hands out the work.

  • The L2 cache is the shared pantry from the memory lesson. Every SM can reach it before making the longer trip to HBM.
  • The GigaThread Engine is the dispatcher. It takes the work the program launches and hands it out to the SMs.

Around them: eight clusters, 224 SMs

Around them are eight clusters holding 224 streaming multiprocessors, or SMs.

That is 28 in each cluster. Each cluster is a GPC, the Graphics Processing Cluster from the hierarchy we learned earlier:

GPU → GPC → SM → Warp → Threads

And inside every SM, four Tensor Cores.

One SM

Tensor Core
Tensor Core
Tensor Core
Tensor Core
CUDA cores · schedulers · registers · shared memory / L1
224 SMs × 4 Tensor Cores = 896 Tensor Cores.

Multiply it out:

8 clusters × 28 SMs = 224 SMs
224 SMs × 4 Tensor Cores = 896 Tensor Cores
Chapter 7

The Tensor Core

So how much work does one Tensor Core actually do?

To count it, let’s borrow an older chip that many of us rent in the cloud: the H100.

Remember our small job: multiply and add.

One H100 Tensor Core does 512 of them in a single tick of its clock, working on 16-bit numbers.

Four Tensor Cores make an SM, and the H100 has 132 SMs.

Now multiply it all out:

512multiply-adds per Tensor Core per tick
×
4Tensor Cores per SM
×
132SMs
×
2operations per multiply-add
×
~1.83 billionticks per second
=
~989 trillionoperations per second
A multiply-add counts as two operations: one multiply and one add. This is the H100’s dense 16-bit Tensor Core peak.

So the H100 does about 989 trillion operations every second. That is 989 TFLOPS, the number you see on its datasheet.

Rubin has 896 Tensor Cores, compared with the H100’s 528. Each one is a newer generation, and it is especially fast on very small number formats.

Datasheet tip

NVIDIA datasheets often show two numbers for the same chip: one “with sparsity” and one without. The sparsity number is twice as high and assumes half the weights are zero in a special pattern. Most models don’t use that pattern, so the dense number is the one to plan with.

Chapter 8

Rubin and the H100

Put them side by side.

Rubin compared with the H100
H100 (SXM)RubinChange
Transistors80 B336 B~4.2×
Compute dies12
SMs132224~1.7×
Tensor Cores528896~1.7×
Memory80 GB HBM3288 GB HBM43.6×
Memory bandwidth3.35 TB/s22 TB/s6.6×
GPU-to-GPU (NVLink)900 GB/s3.6 TB/s4×
Peak Tensor Core math989 TFLOPS (16-bit)50 PFLOPS (4-bit NVFP4)Not comparable
Rubin figures are NVIDIA’s published “up to” numbers. The two peak-math rows use different number formats.

Rubin holds more than three and a half times the H100’s memory, and reads it more than six times as fast.

Their peak math uses different number formats, so those two numbers do not compare directly.

The H100’s 989 TFLOPS is for 16-bit numbers. Rubin’s 50 PFLOPS is for NVFP4, NVIDIA’s 4-bit format. A 4-bit number is a quarter of the size of a 16-bit one, so the hardware can move and multiply far more of them per second.

Recall the Tensor Core section of the Introduction to GPU lesson: slightly lower precision in exchange for much higher speed. NVFP4 pushes that trade-off about as far as it goes today. Rubin’s Transformer Engine manages the precision so the model keeps its accuracy.

Always check the number format before comparing two FLOPS numbers.
Chapter 9

The catch

Rubin can do 50 PFLOPS of 4-bit math, fed by 22 terabytes a second.

Divide one by the other:

50,000 trillion operations ÷ 22 trillion bytes ≈ 2,270 operations per byte

To keep that math busy, Rubin needs to do more than 2,000 operations for every byte that arrives from memory.

Now think about a chatbot answering one person.

For every token, it reads every weight once and does one multiply-add with it. For our 8B model stored in 16-bit numbers, that is:

  • about 15 billion operations (7.5 billion multiply-adds × 2)
  • about 16 GB of weights read from memory

That works out to about one operation per byte.

What Rubin needs to stay busy · ~2,270 operations per byte
A chatbot answering one person, 4-bit weights · ~4 operations per byte
A chatbot answering one person, 16-bit weights · ~1 operation per byte
Even with 4-bit weights, one user gives the Tensor Cores a tiny fraction of the work they could do.

So most of the time, even NVIDIA’s most powerful chip is waiting.

Not for math. For memory.

What waiting looks like in time

Take the same 8B model with 4-bit weights, about 4 GB:

Reading the weights 4 GB ÷ 22 TB/s ≈ 180 microseconds
Doing the math 15 billion ÷ 50,000 trillion ≈ 0.3 microseconds
For one user, the Tensor Cores finish in a fraction of a microsecond, then wait hundreds of times longer for the next weights to arrive.

This is the idea from the bandwidth section, made concrete:

Fast cores + not enough data per byte = cores waiting

How the industry fights back

If the problem is too little work per byte, the fix is to do more work with every byte you read.

  • Batching
    Serve many users at once. Each weight is read once, then used for every request in the batch. With 500 users in a batch, each byte does roughly 500 times the work. This is exactly why continuous batching matters so much in the inference lesson.
  • Smaller number formats
    4-bit weights are a quarter of the bytes of 16-bit weights, so the same bandwidth delivers four times as many.
  • More bandwidth
    HBM4’s wider interface is Rubin’s biggest single jump over Blackwell, about 2.8×.

But batching has a limit too. Every user in the batch needs their own KV cache, and that cache has to fit in the 288 GB. This is where the GPU memory lesson and the inference lesson meet.

Beyond the chip

From one GPU to a rack

NVIDIA does not really sell Rubin as a single chip. It sells it as part of a platform.

Rubin GPU 288 GB HBM4 · 22 TB/s · 50 PFLOPS NVFP4 Vera Rubin Superchip 576 GB HBM4 · 88 CPU cores · 1.5 TB LPDDR5X Vera Rubin NVL72 72 GPUs · 36 CPUs · 20.7 TB HBM4 · 260 TB/s NVLink · 3.6 exaFLOPS NVFP4 Many racks → an AI factory
The GPU, the superchip, and the rack. NVLink joins GPUs inside the rack; InfiniBand or Ethernet joins racks together.
  • Vera CPU
    88 custom Arm-compatible cores, up to 1.5 TB of LPDDR5X memory, and a 1.8 TB/s coherent link to its two Rubin GPUs.
  • NVL72 rack
    72 Rubin GPUs that can all talk to each other through NVLink 6 switches. To a model, the rack behaves like one enormous GPU with 20.7 TB of fast memory.
  • Liquid cooling
    The whole rack is liquid cooled. Remember the cooling section: more activity means more heat, and at this density air cooling is not enough.

This matches the NVLink picture from the Introduction to GPU lesson:

Inside the rack → NVLink
Between racks → InfiniBand or Ethernet
Bringing It Together

What to remember about Rubin

One small job: multiply and add 896 Tensor Cores in 224 SMs across 8 clusters Two reticle-sized dies joined as one GPU Eight HBM4 stacks · 288 GB · 22 TB/s Needs ~2,000+ operations per byte to stay busy Batching and small formats keep it fed
From the one small job to the reason the chip waits.
  • Rubin has 336 billion transistors, and nearly all of them exist to multiply and add.
  • A die cannot be bigger than one lithography field, so Rubin is two dies working as one GPU.
  • Eight stacks of HBM4 hold 288 GB, the place where a model’s weights live.
  • Eight clusters hold 224 SMs, and each SM holds four Tensor Cores, 896 in total.
  • Compared with the H100, it has 3.6× the memory and 6.6× the bandwidth. Peak math numbers need the same number format to be compared.
  • To keep 50 PFLOPS busy on 22 TB/s, the chip needs more than 2,000 operations per byte. A chatbot answering one person does about one.

That is the key idea:

A faster chip does not help if it spends its time waiting for data. The job of the software around it is to keep the Tensor Cores fed.

And what that waiting looks like, from the moment you press send, is exactly the journey we followed in the inference lesson.