Week 3 · GPU
GPUs are everywhere in AI conversations today.
Click or press → to reveal each idea. Use the buttons on a slide to explore.
The question
What exactly does a GPU do that a CPU cannot do as efficiently?
And how did hardware originally designed to make video games look better become the foundation of modern AI?
A common misconception
GPU stands for Graphics Processing Unit, and its original job was much simpler: help computers draw graphics for video games.
Instead of a few complicated tasks, the computer needed huge numbers of similar calculations at the same time.
CPU vs GPU
A few highly skilled workers handling different and complicated jobs. Optimized for latency: how quickly can one task be completed?
A very large team of workers doing similar jobs in parallel. Optimized for throughput: how much work can be completed at the same time?
The key idea
CPU = optimize one task for speed
GPU = optimize many similar tasks for parallel execution
Neural networks perform huge numbers of similar mathematical operations, especially matrix calculations. That is why GPUs became so important for AI.
The journey, in five moments
GPUs did not suddenly become useful for AI. Their design evolved over many years.
Five moments
1990s — GPUs were created for video games. Early GPUs were built to draw graphics quickly: polygons, textures, lighting, and pixels.
2000s — Graphics became more realistic. GPUs gained more processing units and became much better at parallel computation.
2006 — NVIDIA introduced CUDA. Developers could write programs that used the GPU for general-purpose computation, not just drawing graphics.
2012 — AlexNet. It showed the AI community that GPUs could dramatically speed up deep-learning training.
Today — GPUs are a major part of modern AI infrastructure for training and running models.
Inside the GPU
You can think of the entire GPU card as a small specialized computer.
Around the die
Transistors
A transistor is a tiny electronic switch. At the simplest level, it can be thought of as having two states:
ON or OFF
Think of a single transistor like one LEGO brick. Billions of carefully arranged bricks can create an incredibly complex machine.
Transistors
Performance also depends on how those transistors are organized, memory bandwidth, cache design, clock speed, power limits, and software.
Streaming Multiprocessor
The hierarchy
The entire processor. It contains many GPCs and runs massive parallel workloads.
A Graphics Processing Cluster: a large processing block inside the GPU that contains multiple SMs.
A Streaming Multiprocessor is a small compute unit. It has CUDA cores, Tensor Cores, schedulers, registers, shared memory, and cache.
On NVIDIA GPUs, threads are grouped into sets of 32, called a warp. The threads in a warp execute the same instruction on different pieces of data.
A thread is one small unit of work. Example: Thread 1 adds A[0] + B[0]. Thread 2 adds A[1] + B[1].
1 warp = 32 threads
GPU Cores
The general-purpose workers. They handle addition, subtraction, multiplication, and division again and again.
The AI specialists. Designed for matrix multiplication. CUDA Core = a worker using hand tools. Tensor Core = a machine built for one high-speed job.
The graphics specialists. RT stands for ray tracing. They are not a major part of most AI workloads.
GPU Memory
The answer is simple: speed and capacity are a trade-off.
Registers → Shared Memory / L1 Cache → L2 Cache → VRAM
Keep data close
The data currently in your hands. Extremely fast, but there is very little of it.
Shared memory is a whiteboard used by the whole team. Software decides what goes there. L1 cache is the chef's countertop: hardware manages it automatically.
L2 is the pantry shared by the entire kitchen. Larger than L1, shared across the GPU. Path: SM → L1 → L2 → VRAM.
The GPU's main memory. For an AI model that can include model weights, activations, and the KV cache during LLM inference.
A fast GPU is not useful if its cores are constantly waiting for data.
Bus width & bandwidth
Bus width = how wide is the road?
Bandwidth = how much data can travel on that road every second?
Fast cores + slow memory = cores waiting. Fast cores + high memory bandwidth = cores stay busy.
NVLink
PCIe connects the GPU to the host system.
NVLink connects GPUs to other GPUs at very high speed.
InfiniBand or Ethernet connects larger groups of systems across the network.
Cooling System
If the GPU gets too hot: temperature rises → clock speed drops → performance decreases.
More GPU activity → more heat. Better cooling → more stable performance.
Bringing it together
Imagine you type: “What is Artificial Intelligence?”
The flow
The big picture
That is the core idea behind a GPU and the reason it became such an important part of modern AI.