GenAI on Kubernetes: teaching the cluster about GPUs

Kubernetes is great at scheduling containers across machines — but it was built around CPUs and memory first. GPU support came later, through extensions. If you want Ollama, vLLM, PyTorch, or TensorFlow pods on GPUs, you need that bridge installed on purpose.

Start with why GPUs are invisible by default ↓

Prefer to watch this one instead of reading it? This lesson also exists as a short video, covering the same ground: why Kubernetes needs the GPU Operator, Device Plugin, and how to run your first GPU pod.

Video thumbnail for the Day 6 lesson: GenAI on Kubernetes
GenAI on Kubernetes (the video version) Watch on YouTube ↗
The default gap

Why Kubernetes cannot use GPUs out of the box

A fresh Kubernetes install already understands everyday resources:

  • CPU
  • Memory
  • Storage
  • Network interfaces

Those are standard OS and Kubernetes territory. GPUs are different. Every vendor ships its own drivers, libraries, and tools — NVIDIA, AMD, Intel. Because each stack is different, Kubernetes does not bake GPU support into the core.

Picture plugging in a new printer. The OS may notice something is connected, but until the right driver lands, you still cannot print. GPUs work the same way. Kubernetes knows the machine exists, but not whether a GPU is installed, how many there are, which pods should get them, or how apps reach them. That bridge is what NVIDIA provides with the GPU Operator and the Device Plugin.

Stack from AI pod through Kubernetes scheduler, Device Plugin, container toolkit, and NVIDIA drivers down to GPU hardware
From pod request to silicon: scheduler, Device Plugin, container runtime, then drivers.
NVIDIA GPU Operator

One installer for a messy pile of GPU software

Doing this by hand means installing drivers, the NVIDIA Container Toolkit, the Device Plugin, monitoring pieces, GPU Feature Discovery, and sometimes MIG Manager — then keeping all of that healthy across dozens of nodes. That gets painful fast.

The GPU Operator automates it. You install the operator; it deploys and manages the GPU stack for you. Think of it as the installation and lifecycle manager for everything NVIDIA-GPU-related inside the cluster. After install, it keeps watching that those components stay healthy and current.

Device Plugin

This is how Kubernetes finally sees nvidia.com/gpu

Drivers alone are not enough. The Device Plugin is a small component on every GPU node. It detects GPUs, reports them to Kubernetes, lets containers request them, and tracks allocation.

Without it, the cluster only sees CPU and memory. With it, a node might advertise:

nvidia.com/gpu: 4

That means four allocatable NVIDIA GPUs. Apps can request GPUs the same way they request CPU or memory — as first-class Kubernetes resources.

RuntimeClass

Having a GPU on the node is not the same as a container seeing it

By default, Kubernetes starts containers with a normal runtime built for ordinary apps. That path does not automatically expose GPUs.

The NVIDIA Container Runtime adds GPU support at container start time. A RuntimeClass tells Kubernetes which runtime to use for a pod. For GPU workloads, you want the NVIDIA runtime so the process inside the container can talk to the card. Wrong runtime, and the app may behave as if no GPU exists — even when the node has one.

Multi-Instance GPU

One expensive card, several isolated slices

On GPUs like the A100 and H100, MIG can split one physical GPU into smaller, isolated GPUs. Each slice gets its own memory, its own compute, and hardware isolation from the others.

Without MIG, one app often owns the whole device — even if it only needs a little. With MIG, several AI apps can share the same physical GPU more safely, which usually means better utilization and lower cost.

Comparison of one app owning a full GPU versus MIG slicing the GPU into isolated partitions
MIG turns “one app, whole card” into “several isolated slices on one card.”
GPU scheduling

Pending is not always broken — sometimes you are just out of GPUs

Scheduling is Kubernetes deciding where pods run. For normal apps, that is mostly CPU and memory. For AI workloads, GPU availability matters too.

Imagine three workers: Node A has no GPU, Node B has one, Node C has four. A pod that asks for one GPU will only land on B or C. If none are free, it stays Pending until capacity appears. That is intentional — GPU work should not quietly land on a CPU-only box.

Three nodes with zero, one, and four GPUs; GPU pods only schedule onto nodes with free GPUs
Request nvidia.com/gpu: 1 and the scheduler skips nodes that cannot satisfy it.

Unlike CPU, GPUs are usually allocated as whole devices. Asking for one GPU typically reserves the entire GPU for that workload — unless you are using something like MIG to share more finely.

Hands-on

Deploy the NVIDIA GPU Operator

Once the operator is in, it pulls in drivers (when applicable), Device Plugin, Container Toolkit, Feature Discovery, and friends. After that, the cluster is much closer to running GPU AI workloads such as Ollama, vLLM, TensorFlow, and PyTorch.

To access the code for this lab — including the setup scripts and the sample gpu-pod.yaml — use the GitHub repo: ideaweaver-ai/kubernetes-genai.

Prerequisites

  • A Kubernetes cluster
  • A worker node with an NVIDIA GPU
  • Drivers already present, or installable by the operator
  • kubectl pointed at the cluster
  • helm on your machine

Step 1 — Prove the OS can see the GPU

On the GPU node:

nvidia-smi

You should see something like an L40 (or your card), driver version, and memory. If this fails, fix the host before blaming Kubernetes.

Steps 2–4 — Helm install

helm repo add nvidia https://helm.ngc.nvidia.com/nvidia helm repo update kubectl create namespace gpu-operator helm install gpu-operator nvidia/gpu-operator \ --namespace gpu-operator

Give it a few minutes. One chart deploys a lot of moving parts.

Step 5 — Pods healthy?

kubectl get pods -n gpu-operator

Expect Running pods for feature discovery, container toolkit, device plugin, driver (when managed), and the operator itself. Exact names vary by version.

Step 6 — Does Kubernetes advertise the GPU?

kubectl describe node <gpu-node-name> # or kubectl get nodes \ -o custom-columns=NAME:.metadata.name,GPU:.status.allocatable.nvidia\.com/gpu

In Capacity / Allocatable you want a line like nvidia.com/gpu: 1 (or more). Example output:

NAME GPU worker-01 1 worker-02 4
Six GPU Operator components: driver, container toolkit, device plugin, feature discovery, MIG manager, DCGM exporter
One Helm chart quietly installs this whole toolbox.
ComponentPurpose
NVIDIA DriverLets Linux talk to the GPU
Container ToolkitLets containers access the GPU
Device PluginAdvertises nvidia.com/gpu to Kubernetes
GPU Feature DiscoveryDetects features and labels the node
MIG ManagerConfigures MIG on supported GPUs
DCGM ExporterExposes GPU metrics for monitoring
Hands-on

Deploy your first GPU-enabled pod

Goal: run nvidia-smi inside a container. If it works, Kubernetes allocated a GPU, the container can reach it, and the NVIDIA runtime path is healthy.

Create gpu-pod.yaml (or grab it from the repo):

apiVersion: v1 kind: Pod metadata: name: gpu-test spec: restartPolicy: Never containers: - name: cuda-container image: nvidia/cuda:12.9.0-base-ubuntu24.04 command: ["nvidia-smi"] resources: limits: nvidia.com/gpu: 1

What matters most:

  • Image — NVIDIA CUDA base image already has nvidia-smi.
  • Command — print GPU info, then exit (not a long-running server).
  • nvidia.com/gpu: 1 — “only schedule me where a free NVIDIA GPU exists.” Without this line, Kubernetes treats the pod as a normal CPU workload and will not allocate a GPU.
kubectl apply -f gpu-pod.yaml kubectl logs gpu-test

If you see the familiar NVIDIA-SMI table in the logs, congratulations — your cluster is no longer GPU-blind.

What to remember

Five things worth carrying forward

  1. Kubernetes knows CPU and memory by default — not GPUs.
  2. GPU Operator installs and manages the NVIDIA stack; Device Plugin advertises nvidia.com/gpu.
  3. RuntimeClass / NVIDIA runtime is how containers actually reach the card.
  4. MIG can slice big GPUs; without it, one GPU request usually means the whole device.
  5. Always verify host nvidia-smi, then node allocatable GPUs, then a tiny GPU pod.

Say it back to me

Tap each card. Can you remember what it means, before you flip it?

Host sees GPU → Operator + Device Plugin → node advertises nvidia.com/gpu → pod requests it → runtime exposes it.