Inside NVIDIA Vera Rubin
NVIDIA’s newest AI chip, taken apart one layer at a time.
We’ll start with the one small job it does ↓
Meet Rubin
In the Introduction to GPU lesson, we looked at the parts found inside any modern GPU: transistors, SMs, Tensor Cores, memory, NVLink, and cooling.
Now let’s look at a real one, and a very new one.
This is Rubin, NVIDIA’s newest chip for artificial intelligence.
It holds 336 billion transistors, and nearly all of them exist to do one small job, very fast.
First, a quick note on the name, because it gets mixed up often:
Most of this lesson is about the Rubin GPU. We’ll come back to the full platform at the end.
To get a sense of scale, here is how the transistor count has grown across NVIDIA’s data-center GPUs:
| GPU | Transistors | Compute dies |
|---|---|---|
| H100 (Hopper) | 80 billion | 1 |
| B200 (Blackwell) | 208 billion | 2 |
| Rubin | 336 billion | 2 |
One small job
So what is the one small job?
Multiply two numbers, and add the result to a running total.
Then do it again. And again.
Why this operation? Because a neural network is mostly matrix multiplication, and matrix multiplication is nothing more than this multiply-add repeated a very large number of times.
Every weight in the model gets multiplied by some input value, and the results get added together.
To write a single token, a small piece of a word, a model with 8 billion parameters does this about 7.5 billion times.
Why 7.5 and not 8?
Because roughly half a billion of those parameters are the embedding table. The model simply looks up a row in that table for each token. It does not multiply through it. Nearly every other weight is used once per token, in one multiply-add.
And that is for one token. A single answer may contain hundreds of tokens, and a real service is answering many users at the same time.
AI hardware = a machine for doing multiply-adds as fast as possible
Five layers
Take Rubin apart and there are five layers.
From the bottom up:
- A board
The package substrate. It carries power into the chip and connects it to the rest of the system. - A silicon interposer
A thin layer with extremely fine wiring. The dies and the memory all sit on it, so they can talk to each other over very short, very wide connections. - Two compute dies
This is where the transistors, SMs, and Tensor Cores live. - Eight stacks of memory
HBM4, sitting right next to the compute dies. - A lid
It protects the silicon and helps move heat into the cooling system.
Remember the GPU anatomy picture from the Introduction to GPU lesson? This is the same idea, zoomed in on the package that sits at the center of the card.
One question stands out: why two dies, and not one big one?
Why two dies
Chips are printed onto silicon wafers, one small field at a time.
A lithography machine projects the chip’s pattern onto the wafer through a mask called a reticle. It exposes one rectangle, steps over, exposes the next, and so on across the wafer.
No chip can be bigger than one field: 26 × 33 mm, about 858 mm².
This is called the reticle limit. It comes from the optics of the machine itself. The lens can only keep a pattern that size sharp and undistorted.
fills almost all of it
One exposure field · 26 × 33 mm
Two dies, joined to work as one GPU
Each Rubin die fills almost all of that field.
So NVIDIA builds two, and joins them to work as one GPU.
The connection between them is called NV-HBI, the NVIDIA High-Bandwidth Interface. It is fast enough that software sees one GPU, not two.
There is a second reason this approach helps: yield.
Every wafer has a few tiny defects scattered across it. A defect usually ruins the die it lands on. The bigger the die, the more likely it is to be hit. Two reticle-sized dies, tested separately before they are joined, waste less silicon than one impossibly large die would.
Blackwell used the same two-die idea. Rubin keeps it and puts more into each die.
Can’t print a bigger chip → print two chips → join them so they behave like one
The memory
Around the dies stand eight memory stacks, holding 288 GB.
This is where a model’s weights live.
In the Introduction to GPU lesson, we called this VRAM. On data-center GPUs, it is built from a special kind of memory called HBM, or High Bandwidth Memory. Rubin uses the newest version, HBM4.
| Memory stacks | 8 |
|---|---|
| DRAM chips in each stack | 12 (a “12-high” stack) |
| Capacity per stack | 36 GB |
| Total | 8 × 36 GB = 288 GB |
| Peak bandwidth | up to 22 TB/s |
Inside one stack
Here is how a stack is built. The one pictured in most teardowns is from an earlier generation, but the idea is the same.
Memory chips are piled on top of a base chip, right beside the compute.
Copper channels run straight down through every chip. These are called through-silicon vias, or TSVs. Because they pass through every layer, the whole stack shares one very wide connection.
The data flows down, then sideways into the GPU.
Why this design matters
Remember the highway analogy from the bandwidth section of the Introduction to GPU lesson?
A gaming GPU might have a 384-bit memory bus.
HBM4 gives each stack a 2,048-bit interface, twice as wide as HBM3e. With eight stacks, Rubin has a 16,384-bit highway to its memory.
You could not run that many wires across a normal circuit board. It only works because the memory sits on the interposer, millimetres away from the compute dies.
Stack the memory → put it next to the compute → connect it with a very wide road
Links, clusters, and SMs
Now look at the compute side from above.
Along the edges: links to the outside world
Along the edges are its links to the outside world: to the CPU, to other GPUs, and to everything else.
| Link | Connects to | Bandwidth |
|---|---|---|
| NVLink 6 | Other Rubin GPUs, through NVLink switches | 3.6 TB/s |
| NVLink-C2C | The Vera CPU, with shared, coherent memory | 1.8 TB/s |
| PCIe Gen 6 (x16) | The host system and other devices | up to 256 GB/s |
This is the same picture as the NVLink section of the Introduction to GPU lesson: PCIe to the host, NVLink between GPUs. Rubin adds a very fast, direct link to its own CPU.
In the middle: a shared cache and a work scheduler
In the middle sits a shared cache and an engine that hands out the work.
- The L2 cache is the shared pantry from the memory lesson. Every SM can reach it before making the longer trip to HBM.
- The GigaThread Engine is the dispatcher. It takes the work the program launches and hands it out to the SMs.
Around them: eight clusters, 224 SMs
Around them are eight clusters holding 224 streaming multiprocessors, or SMs.
That is 28 in each cluster. Each cluster is a GPC, the Graphics Processing Cluster from the hierarchy we learned earlier:
GPU → GPC → SM → Warp → Threads
And inside every SM, four Tensor Cores.
One SM
Multiply it out:
8 clusters × 28 SMs = 224 SMs
224 SMs × 4 Tensor Cores = 896 Tensor Cores
The Tensor Core
So how much work does one Tensor Core actually do?
To count it, let’s borrow an older chip that many of us rent in the cloud: the H100.
Remember our small job: multiply and add.
One H100 Tensor Core does 512 of them in a single tick of its clock, working on 16-bit numbers.
Four Tensor Cores make an SM, and the H100 has 132 SMs.
Now multiply it all out:
So the H100 does about 989 trillion operations every second. That is 989 TFLOPS, the number you see on its datasheet.
Rubin has 896 Tensor Cores, compared with the H100’s 528. Each one is a newer generation, and it is especially fast on very small number formats.
NVIDIA datasheets often show two numbers for the same chip: one “with sparsity” and one without. The sparsity number is twice as high and assumes half the weights are zero in a special pattern. Most models don’t use that pattern, so the dense number is the one to plan with.
Rubin and the H100
Put them side by side.
| H100 (SXM) | Rubin | Change | |
|---|---|---|---|
| Transistors | 80 B | 336 B | ~4.2× |
| Compute dies | 1 | 2 | |
| SMs | 132 | 224 | ~1.7× |
| Tensor Cores | 528 | 896 | ~1.7× |
| Memory | 80 GB HBM3 | 288 GB HBM4 | 3.6× |
| Memory bandwidth | 3.35 TB/s | 22 TB/s | 6.6× |
| GPU-to-GPU (NVLink) | 900 GB/s | 3.6 TB/s | 4× |
| Peak Tensor Core math | 989 TFLOPS (16-bit) | 50 PFLOPS (4-bit NVFP4) | Not comparable |
Rubin holds more than three and a half times the H100’s memory, and reads it more than six times as fast.
Their peak math uses different number formats, so those two numbers do not compare directly.
The H100’s 989 TFLOPS is for 16-bit numbers. Rubin’s 50 PFLOPS is for NVFP4, NVIDIA’s 4-bit format. A 4-bit number is a quarter of the size of a 16-bit one, so the hardware can move and multiply far more of them per second.
Recall the Tensor Core section of the Introduction to GPU lesson: slightly lower precision in exchange for much higher speed. NVFP4 pushes that trade-off about as far as it goes today. Rubin’s Transformer Engine manages the precision so the model keeps its accuracy.
Always check the number format before comparing two FLOPS numbers.
The catch
Rubin can do 50 PFLOPS of 4-bit math, fed by 22 terabytes a second.
Divide one by the other:
50,000 trillion operations ÷ 22 trillion bytes ≈ 2,270 operations per byte
To keep that math busy, Rubin needs to do more than 2,000 operations for every byte that arrives from memory.
Now think about a chatbot answering one person.
For every token, it reads every weight once and does one multiply-add with it. For our 8B model stored in 16-bit numbers, that is:
- about 15 billion operations (7.5 billion multiply-adds × 2)
- about 16 GB of weights read from memory
That works out to about one operation per byte.
So most of the time, even NVIDIA’s most powerful chip is waiting.
Not for math. For memory.
What waiting looks like in time
Take the same 8B model with 4-bit weights, about 4 GB:
This is the idea from the bandwidth section, made concrete:
Fast cores + not enough data per byte = cores waiting
How the industry fights back
If the problem is too little work per byte, the fix is to do more work with every byte you read.
- Batching
Serve many users at once. Each weight is read once, then used for every request in the batch. With 500 users in a batch, each byte does roughly 500 times the work. This is exactly why continuous batching matters so much in the inference lesson. - Smaller number formats
4-bit weights are a quarter of the bytes of 16-bit weights, so the same bandwidth delivers four times as many. - More bandwidth
HBM4’s wider interface is Rubin’s biggest single jump over Blackwell, about 2.8×.
But batching has a limit too. Every user in the batch needs their own KV cache, and that cache has to fit in the 288 GB. This is where the GPU memory lesson and the inference lesson meet.
From one GPU to a rack
NVIDIA does not really sell Rubin as a single chip. It sells it as part of a platform.
- Vera CPU
88 custom Arm-compatible cores, up to 1.5 TB of LPDDR5X memory, and a 1.8 TB/s coherent link to its two Rubin GPUs. - NVL72 rack
72 Rubin GPUs that can all talk to each other through NVLink 6 switches. To a model, the rack behaves like one enormous GPU with 20.7 TB of fast memory. - Liquid cooling
The whole rack is liquid cooled. Remember the cooling section: more activity means more heat, and at this density air cooling is not enough.
This matches the NVLink picture from the Introduction to GPU lesson:
Inside the rack → NVLink
Between racks → InfiniBand or Ethernet
What to remember about Rubin
- Rubin has 336 billion transistors, and nearly all of them exist to multiply and add.
- A die cannot be bigger than one lithography field, so Rubin is two dies working as one GPU.
- Eight stacks of HBM4 hold 288 GB, the place where a model’s weights live.
- Eight clusters hold 224 SMs, and each SM holds four Tensor Cores, 896 in total.
- Compared with the H100, it has 3.6× the memory and 6.6× the bandwidth. Peak math numbers need the same number format to be compared.
- To keep 50 PFLOPS busy on 22 TB/s, the chip needs more than 2,000 operations per byte. A chatbot answering one person does about one.
That is the key idea:
A faster chip does not help if it spends its time waiting for data. The job of the software around it is to keep the Tensor Cores fed.
And what that waiting looks like, from the moment you press send, is exactly the journey we followed in the inference lesson.