← Back to the lesson
1 / 1

Week 3 · Kubernetes

Kubernetes on GPU

A beginner-friendly guide to discovering, scheduling, sharing, and operating GPUs in Kubernetes.

Click or press → to reveal each idea. Use the buttons on a slide to explore.

The gap

Kubernetes already knows CPU and memory

GPUs are specialized devices. A few extra components have to discover them, expose them as resources, and make them usable inside containers.

A GPU sitting in a server is not a Kubernetes GPU.

The path

Keep this picture in mind

Everything else in this lesson explains one step of this flow.

Physical NVIDIA GPU
        ↓
Linux / NVIDIA Driver
        ↓
GPU discovery (NFD / GFD)
        ↓
NVIDIA Device Plugin
        ↓
Kubernetes sees nvidia.com/gpu
        ↓
Scheduler chooses a node
        ↓
Container runtime exposes the GPU
        ↓
Pod runs the AI workload

Why GPUs

Generative AI is a lot of parallel math

Large language models perform a huge number of similar calculations when they process prompts and generate tokens. GPUs were built for graphics, and that same parallel design fits training and inference.

Having a GPU installed does not mean Kubernetes knows how to use it.

The default cluster

Kubernetes does not ship one vendor’s GPU stack

  • CPU, memory, storage, and network are built in.
  • NVIDIA, AMD, and Intel each have their own drivers, libraries, and runtimes.
  • The core scheduler stays vendor-neutral. The GPU path is added on.

A mixed cluster

A GPU pod can only land where a GPU exists

  • Worker 1 — CPU only. Cannot run a GPU workload.
  • Worker 2 — NVIDIA T4. Potentially yes.
  • Worker 3 — NVIDIA A100. Potentially yes.

“Potentially” is the whole lesson. The card has to be discovered, advertised, scheduled, and exposed to the container.

Before scheduling

Questions Kubernetes has to answer

  • Which nodes have GPUs?
  • Which model, and how many?
  • How much GPU memory?
  • How does the container reach the device?
  • Can several apps share it?
  • What if the model needs more than one GPU?

Discovery

First find the hardware

NFD discovers general node hardware. GFD adds NVIDIA-GPU-specific details.

Node Feature Discovery scans each node and adds labels: CPU features, PCI devices, network, kernel, and GPUs. It runs as a DaemonSet — one copy on every applicable node.

GPU Feature Discovery says what kind of NVIDIA GPU it is: product, memory, family, count, and whether MIG is supported. A T4, an A100, and an H100 are not interchangeable.

NFD

This node contains NVIDIA GPU hardware

feature.node.kubernetes.io/pci-0302_10de.present: "true"
  • 0302 — PCI class for a 3D controller.
  • 10de — NVIDIA’s PCI vendor ID.
  • true — that hardware is present.

Vendor IDs

The same label pattern, different vendors

  • 10de — NVIDIA
  • 1002 — AMD
  • 8086 — Intel

NFD can see the card. It still does not make the GPU allocatable to a pod.

GFD

Labels that describe the device

nvidia.com/gpu.count: "4"
nvidia.com/gpu.product: "A100-SXM4-40GB"
nvidia.com/gpu.memory: "40537"
nvidia.com/gpu.family: "ampere"
nvidia.com/mig.capable: "true"

An application can then target one model:

nodeSelector:
  nvidia.com/gpu.product: A100-SXM4-40GB

Device Plugin

Make the GPU a Kubernetes resource

Physical NVIDIA GPU
        ↓
NVIDIA Device Plugin
        ↓
Kubelet
        ↓
nvidia.com/gpu: 1

It runs on GPU nodes as a DaemonSet and talks to the kubelet.

Four jobs

What the Device Plugin actually does

  1. Discover the GPUs on the node.
  2. Report capacity to the kubelet.
  3. Allocate a free GPU when a pod asks for one.
  4. Track health so a bad GPU is not offered again.

Four healthy GPUs show up as nvidia.com/gpu: 4.

The request

“My container needs one NVIDIA GPU.”

resources:
  limits:
    nvidia.com/gpu: 1

The scheduler can place that pod on any node with a free GPU.

By itself, this does not say T4, A100, or H100. One GPU is one GPU.

Two different problems

Scheduled is not the same as usable

Can Kubernetes place the pod on a GPU node?

The Device Plugin helps with allocation.

Can the process inside the container talk to that GPU?

The Container Toolkit and the NVIDIA runtime make the selected GPU visible.

A node can advertise nvidia.com/gpu correctly and the pod can still fail if the runtime path is missing.

Scheduling

Each mechanism answers a different question

  • Any NVIDIA GPU — resource request.
  • A particular model — resource request plus nodeSelector.
  • Required and preferred traits — resource request plus Node Affinity.
  • Keep ordinary pods off expensive nodes — taints and tolerations.
  • Describe the device itself — Dynamic Resource Allocation.

Production clusters usually combine several of these.

Pick a need

Which mechanism fits?

Resource request. nvidia.com/gpu: 1 is enough when every GPU node is equivalent for the workload. A 16 GB T4 and an 80 GB H100 both count as one GPU, so this is the wrong tool when memory size matters.

nodeSelector is strict. nvidia.com/gpu: 1 allocates one GPU. The selector only considers nodes with that exact label. If none are free, the pod stays Pending. It does not quietly move to a different GPU.

Node Affinity. Required means the node must match — for example GPU memory greater than 40,000 MiB. Preferred means, among the nodes that qualify, favor a Hopper GPU. Preferred never overrides required.

Taints and tolerations. A taint answers “who may use this node?” A toleration is permission, not placement. The pod still needs a GPU request, and usually a selector or affinity, to actually land there.

nodeSelector

Allocate one GPU, and only on this model

spec:
  nodeSelector:
    nvidia.com/gpu.product: Tesla-T4
  containers:
  - name: llm
    resources:
      limits:
        nvidia.com/gpu: 1

Two jobs: the limit allocates a device. The selector refuses every other GPU type.

Protect the node

A toleration does not pull the pod onto the GPU

kubectl taint nodes gpu-node \
  nvidia.com/gpu=true:NoSchedule
tolerations:
- key: "nvidia.com/gpu"
  operator: "Exists"
  effect: "NoSchedule"

Ordinary pods stay off the expensive node. GPU pods are allowed in, then still have to request a GPU.

DRA

Ask for a device, not a node label

Traditional

“Run me on a node with this GPU label.”

Dynamic Resource Allocation

“I need a GPU with these characteristics. Please find it.”

A vendor DRA driver publishes the real devices. Kubernetes matches the claim.

DRA objects

Five names, one request form

What category of device can be requested.

What devices are actually available right now.

A reusable GPU request: class, count, and how to allocate.

The actual device request for this workload.

The vendor-specific bridge to the real hardware. Attribute names come from the driver, so check them rather than guessing.

The claim

Describe the GPU. Don’t name the card.

  • Type: NVIDIA GPU
  • Count: 1
  • Memory: at least 40 GB
  • Optional: a particular model or family
exactly:
  deviceClassName: gpu.nvidia.com
  allocationMode: ExactCount
  count: 1

Memorize the intent: which class, how to allocate, how many. The field names follow the installed driver.

Wiring it up

The pod asks. One container receives the device.

resourceClaims:
- name: high-memory-gpu
  resourceClaimTemplateName: gpu-claim-template

containers:
- name: model-runner
  resources:
    claims:
    - name: high-memory-gpu

A sidecar in the same pod does not have to receive GPU access.

Device Plugin vs DRA

Simple counts, or a described device

  • Traditional: ask for nvidia.com/gpu: 1, often plus node labels.
  • DRA: select through device attributes.
  • Traditional thinks about which node has the GPU.
  • DRA thinks about what the application needs.

nvidia.com/gpu stays the straightforward default. DRA pays off when requirements get specific or the cluster is mixed.

MIG

One large GPU, several isolated slices

Without MIG

A request for one GPU reserves the whole device.

With MIG

On an A100 or H100, several workloads each get their own memory and compute slice of the same card.

Useful when smaller jobs should not occupy an entire high-end GPU.

GPU Operator

One installer for the whole stack

  • NVIDIA Driver — Linux talks to the GPU.
  • NFD and GFD — discover the hardware and label it.
  • Device Plugin — expose nvidia.com/gpu.
  • Container Toolkit — let the container use the allocated GPU.
  • MIG Manager — configure slices on supported cards.
  • DCGM Exporter — GPU telemetry for monitoring.

Hands-on

Install the Operator, then prove the GPU is advertised

On the node, nvidia-smi has to work first. If the host cannot see the GPU, Kubernetes cannot either.

helm repo add nvidia https://helm.ngc.nvidia.com/nvidia
helm repo update
kubectl create namespace gpu-operator
helm install gpu-operator nvidia/gpu-operator \
  --namespace gpu-operator

Then check pods in gpu-operator, and confirm allocatable GPUs on the node.

The first pod

If nvidia-smi prints inside the container, the path works

containers:
- name: cuda-container
  image: nvidia/cuda:12.9.0-base-ubuntu24.04
  command: ["nvidia-smi"]
  resources:
    limits:
      nvidia.com/gpu: 1

kubectl apply, then kubectl logs gpu-test. The NVIDIA-SMI table means allocation and the runtime both succeeded.

Pending

Pending often means “no node fits,” not “the app is broken”

  • No GPU on the nodes that match.
  • Every GPU is already allocated.
  • nodeSelector or affinity is too tight.
  • A taint blocks the pod, and nothing tolerates it.
  • A DRA claim cannot find a matching device.

Read events with kubectl describe pod before debugging the model.

Troubleshooting

Start at the host. End at the model.

Run nvidia-smi on the worker. Can Linux see the GPU?

Are the GPU Operator pods healthy?

Do the expected GPU labels exist on the node?

Does the node show allocatable nvidia.com/gpu?

Does the pod request a GPU, and do selectors, affinity, and taints allow placement?

Can a tiny CUDA pod run nvidia-smi?

Only now debug the model or the inference server.

Do not debug the application while the platform path is already broken.

Choosing a method

Start with the smallest rule that fits

  • Any GPU — nvidia.com/gpu.
  • A specific model — add nodeSelector.
  • Must-have plus nice-to-have — Node Affinity.
  • Reserve GPU nodes — taints and tolerations.
  • Describe the device — DRA.

Remember

Five things

  1. GPUs need vendor-specific integration. CPU and memory do not.
  2. NFD finds hardware. GFD describes the NVIDIA GPU.
  3. The Device Plugin exposes capacity as nvidia.com/gpu.
  4. Requests, selectors, affinity, taints, and DRA each solve a different placement problem.
  5. The GPU Operator manages that software stack across the cluster.

A physical GPU is a Kubernetes GPU only when discovery, the Device Plugin, the scheduler, and the container runtime all agree.

Open the full lesson →

Click or → to reveal