Kubernetes on GPU
A beginner-friendly guide to discovering, scheduling, sharing, and operating GPUs in Kubernetes.
Start with the end-to-end picture ↓
Kubernetes is excellent at scheduling CPU and memory, but GPUs are specialized devices. We need a few extra components to help Kubernetes discover GPUs, expose them as resources, and make them usable inside containers.
Keep this path in mind
Before learning individual components, keep this end-to-end picture in mind. Everything else in this lesson is simply explaining one part of this flow.
Why GPUs matter for Generative AI
Large language models perform a huge number of mathematical operations when they process prompts and generate tokens. CPUs can perform those calculations, but GPUs are designed to perform many calculations in parallel, which makes them especially useful for AI workloads.
GPUs were originally built for graphics and gaming, but the same parallel-computing design is a strong fit for model training and inference.
Having a GPU physically installed in a server does not automatically mean Kubernetes knows how to use it.
Why Kubernetes does not use GPUs automatically
A fresh Kubernetes cluster already understands common resources such as:
- CPU
- Memory
- Storage
- Network interfaces
GPUs are different. NVIDIA, AMD, Intel, and other vendors have different drivers, libraries, runtimes, and management tools. Kubernetes therefore does not build one vendor-specific GPU implementation directly into its core.
Imagine a cluster with three workers:
| Node | Hardware | Can it run a GPU workload? |
|---|---|---|
| Worker 1 | CPU only | No |
| Worker 2 | NVIDIA T4 | Potentially yes |
| Worker 3 | NVIDIA A100 | Potentially yes |
For a GPU workload, Kubernetes may need to answer:
- Which nodes have GPUs?
- Which GPU model is installed?
- How many GPUs are available?
- How much GPU memory does the device have?
- How does the container access the GPU?
- Can multiple applications share the device?
- What happens when a model needs more than one GPU?
That leads us to the first major topic: GPU discovery.
First find the hardware
Before Kubernetes can make a good scheduling decision, it needs information about the hardware available on each node.
NFD discovers general node hardware. GFD adds NVIDIA-GPU-specific details.
4.1 Node Feature Discovery (NFD)
Node Feature Discovery is like a hardware scanner for Kubernetes nodes. It inspects each node and adds labels that describe useful hardware and software features.
NFD can discover information such as:
- CPU features
- PCI devices
- Network devices
- Kernel features
- GPUs and other hardware
NFD typically runs as a DaemonSet. If you are new to Kubernetes, a DaemonSet simply means: run one copy of this component on every applicable node.
One way to install NFD is with Kustomize:
NFD_REPO=https://github.com/kubernetes-sigs/node-feature-discovery
kubectl apply -k \
$NFD_REPO/deployment/overlays/default
After NFD starts, you can inspect labels on a node:
kubectl get node <node-name> -o yaml
You may see a label similar to:
feature.node.kubernetes.io/pci-0302_10de.present: "true"
In plain English:
0302is a PCI class associated with a 3D controller.10deis NVIDIA’s PCI vendor ID.truemeans that hardware is present on the node.
So this label is essentially telling us: this node contains NVIDIA GPU hardware.
| Vendor ID | Vendor |
|---|---|
| 10de | NVIDIA |
| 1002 | AMD |
| 8086 | Intel |
4.2 GPU Feature Discovery (GFD)
NFD can tell us that NVIDIA hardware exists, but AI workloads often need more detail. A T4, A100, and H100 are all GPUs, but they differ significantly in memory and capabilities.
GPU Feature Discovery focuses on NVIDIA GPUs and can add labels describing the device in more detail.
GFD may publish labels such as:
nvidia.com/gpu.count: "4"
nvidia.com/gpu.product: "A100-SXM4-40GB"
nvidia.com/gpu.memory: "40537"
nvidia.com/gpu.family: "ampere"
nvidia.com/mig.capable: "true"
| Label | Meaning |
|---|---|
nvidia.com/gpu.count | How many GPUs are present |
nvidia.com/gpu.product | The NVIDIA GPU model |
nvidia.com/gpu.memory | GPU memory reported by discovery |
nvidia.com/gpu.family | GPU architecture family |
nvidia.com/mig.capable | Whether the GPU supports MIG |
This becomes useful when a cluster contains different GPU models. For example, an application can target a specific GPU label:
nodeSelector:
nvidia.com/gpu.product: A100-SXM4-40GB
Discovery describes the GPU, but discovery alone does not make the GPU allocatable to a pod.
Make the GPU a Kubernetes resource
The NVIDIA Device Plugin is the bridge between the physical NVIDIA GPU and Kubernetes resource management.
The plug-in normally runs on GPU nodes as a DaemonSet and communicates with the kubelet.
Its main responsibilities are:
- Device discovery — detect GPUs available on the node.
- Report capacity — tell kubelet how many GPU resources are available.
- Allocation — help assign an available GPU when a pod requests one.
- Health tracking — stop unhealthy GPUs from being offered to new workloads.
For example, a node with four healthy GPUs may advertise:
nvidia.com/gpu: 4
5.1 Requesting a GPU from a pod
Once the Device Plugin is working, requesting a GPU is simple:
resources:
limits:
nvidia.com/gpu: 1
This means: “My container needs one NVIDIA GPU.”
If the cluster has Node 1 with no GPUs, Node 2 with two GPUs, and Node 3 with four GPUs, the scheduler can consider Node 2 or Node 3 as long as at least one GPU is free.
nvidia.com/gpu: 1 asks for one GPU. By itself, it does not say whether that GPU should be a T4, A100, H100, or another model.
The container still needs a path to the GPU
There are two separate questions:
Helps with allocation
Makes the selected GPU visible inside the container
A node can advertise nvidia.com/gpu correctly and a pod can still fail to use the GPU if the container runtime path is not configured correctly.
How Kubernetes decides where the pod runs
GPU scheduling starts with the same scheduler Kubernetes already uses. The difference is that GPU workloads add device requirements.
| Need | Useful Kubernetes mechanism |
|---|---|
| Any available NVIDIA GPU | Resource request |
| A particular GPU model | Resource request + nodeSelector |
| Required and preferred GPU characteristics | Resource request + Node Affinity |
| Protect expensive GPU nodes | Taints + Tolerations |
| Describe device requirements dynamically | Dynamic Resource Allocation (DRA) |
7.1 Resource-based scheduling
The simplest case is when any NVIDIA GPU is acceptable:
resources:
limits:
nvidia.com/gpu: 1
This works well when all GPU nodes are equivalent for the workload.
But if your model needs 35 GB of GPU memory, a 16 GB T4 and an 80 GB H100 cannot be treated as interchangeable even though both count as one GPU.
7.2 nodeSelector: choose an exact GPU type
If the workload needs a particular GPU model, use a label-based scheduling rule.
spec:
nodeSelector:
nvidia.com/gpu.product: Tesla-T4
containers:
- name: llm
image: my-llm:latest
resources:
limits:
nvidia.com/gpu: 1
These two settings do different jobs:
nvidia.com/gpu: 1means: allocate one GPU.nodeSelectormeans: only consider nodes with this exact GPU label.
nodeSelector is strict: if no matching node is available, the pod remains Pending rather than silently moving to a different GPU type.
7.3 Node Affinity: required plus preferred
Node Affinity is useful when the scheduling rule is more flexible.
For example: the model must have more than 40,000 MiB of GPU memory, and we prefer a Hopper-family GPU.
affinity:
nodeAffinity:
requiredDuringSchedulingIgnoredDuringExecution:
nodeSelectorTerms:
- matchExpressions:
- key: nvidia.com/gpu.memory
operator: Gt
values:
- "40000"
preferredDuringSchedulingIgnoredDuringExecution:
- weight: 1
preference:
matchExpressions:
- key: nvidia.com/gpu.family
operator: In
values:
- hopper
The YAML looks complex, but the idea is simple:
- Required: the node must satisfy this condition.
- Preferred: if several nodes satisfy the requirement, prefer this kind of node.
7.4 Taints and Tolerations: protect expensive GPU nodes
nodeSelector and affinity answer: where should this pod run? Taints and tolerations answer a different question: which pods are allowed onto this node?
GPU nodes can be expensive, so you may want to keep ordinary workloads away from them.
Taint the GPU node:
kubectl taint nodes gpu-node \
nvidia.com/gpu=true:NoSchedule
Then allow a GPU workload to tolerate that taint:
tolerations:
- key: "nvidia.com/gpu"
operator: "Exists"
effect: "NoSchedule"
A toleration does not force a pod onto a GPU node. It only gives the pod permission to run there. You still normally combine it with a GPU resource request, nodeSelector, or affinity.
Ask for a device, not a node label
Traditional GPU scheduling often makes us think about nodes: find the node with the right labels, then schedule the pod there.
Dynamic Resource Allocation changes the mental model:
The core DRA framework is a Kubernetes mechanism for describing and allocating specialized devices. A compatible vendor DRA driver publishes the real devices and their attributes.
A simple way to remember the main objects:
| Object | Beginner-friendly meaning |
|---|---|
| DeviceClass | What category of device can be requested |
| ResourceSlice | What devices are available |
| ResourceClaimTemplate | A reusable GPU request template |
| ResourceClaim | The actual device request |
| DRA driver | The vendor-specific bridge to the real hardware |
8.1 ResourceClaim: think of it as a GPU request form
Instead of naming a specific physical GPU, the workload describes what it needs.
8.2 A simplified ResourceClaimTemplate
A simplified example uses the Kubernetes resource API like this:
apiVersion: resource.k8s.io/v1
kind: ResourceClaimTemplate
metadata:
name: gpu-claim-template
spec:
spec:
devices:
requests:
- name: high-memory-gpu
exactly:
deviceClassName: gpu.nvidia.com
allocationMode: ExactCount
count: 1
The key idea is not to memorize every API field. Focus on the intent:
deviceClassNamesays which category of device we want.allocationModesays how allocation should work.countsays how many devices are needed.
DRA can also use driver-published attributes and selectors to describe more specific device requirements. Exact attribute names depend on the installed DRA driver, so they should be checked rather than guessed.
8.3 Connecting the claim to an application
apiVersion: apps/v1
kind: Deployment
metadata:
name: inference-server
spec:
replicas: 1
selector:
matchLabels:
app: inference-server
template:
metadata:
labels:
app: inference-server
spec:
resourceClaims:
- name: high-memory-gpu
resourceClaimTemplateName: gpu-claim-template
containers:
- name: model-runner
image: myorg/llm-inference:latest
resources:
claims:
- name: high-memory-gpu
The pod-level resourceClaims section creates or associates the device request. The container-level resources.claims section says which container should receive access to that allocated device.
This matters because a pod may contain a GPU-using model container and another sidecar container that does not need GPU access.
8.4 Traditional Device Plugin vs DRA
| Traditional approach | DRA approach |
|---|---|
Ask for nvidia.com/gpu: 1 | Describe a device requirement |
| Often combine with node labels | Select through device attributes |
| Think about which node has the GPU | Think about what device the application needs |
| Simple and widely understood | More expressive for complex device requirements |
Traditional nvidia.com/gpu requests remain useful and straightforward. DRA becomes especially interesting as device requirements become more specific or heterogeneous.
Split one large GPU into smaller isolated instances
On supported GPUs such as A100 and H100, Multi-Instance GPU (MIG) can divide one physical GPU into smaller isolated GPU instances.
Without MIG, a request for one GPU often reserves the whole device. With MIG, multiple workloads can use isolated slices of the same physical GPU, each with its own assigned memory and compute resources.
Why it matters: MIG can improve utilization when several smaller workloads do not need an entire high-end GPU.
Manage the GPU software stack
Manually preparing every GPU node can become painful. You may need drivers, the Container Toolkit, the Device Plugin, discovery components, monitoring, and MIG management.
The NVIDIA GPU Operator acts as an installation and lifecycle manager for this stack.
| Component | Purpose |
|---|---|
| NVIDIA Driver | Lets Linux communicate with the GPU |
| Node Feature Discovery (NFD) | Discovers general node hardware |
| GPU Feature Discovery (GFD) | Adds detailed NVIDIA GPU labels |
| NVIDIA Device Plugin | Exposes GPU resources to Kubernetes |
| NVIDIA Container Toolkit | Lets containers access allocated GPUs |
| MIG Manager | Configures MIG on supported GPUs |
| DCGM Exporter | Exposes GPU telemetry for monitoring |
This is one of the main reasons the GPU Operator is useful: instead of managing several pieces individually across many nodes, you install the Operator and let it manage the NVIDIA GPU stack.
Deploy the NVIDIA GPU Operator
Prerequisites
- A Kubernetes cluster
- At least one worker node with an NVIDIA GPU
- NVIDIA drivers already present, or installable by the Operator
kubectlconfigured for the cluster- Helm installed on your machine
Step 1 — Confirm that Linux can see the GPU
Run this directly on the GPU node:
nvidia-smi
You should see the GPU model, driver version, GPU memory, and other device information.
If nvidia-smi fails on the host, fix the host-level GPU problem before debugging Kubernetes.
Steps 2–4 — Install the GPU Operator with Helm
helm repo add nvidia https://helm.ngc.nvidia.com/nvidia
helm repo update
kubectl create namespace gpu-operator
helm install gpu-operator nvidia/gpu-operator \
--namespace gpu-operator
The chart deploys several components, so give the cluster time to create and initialize the pods.
Step 5 — Check that the Operator pods are healthy
kubectl get pods -n gpu-operator
You should expect components related to the operator, device plugin, container toolkit, discovery, monitoring, and drivers when driver management is enabled. Exact pod names can vary by version.
Step 6 — Confirm that Kubernetes advertises the GPU
kubectl describe node <gpu-node-name>
# or
kubectl get nodes \
-o custom-columns=NAME:.metadata.name,GPU:.status.allocatable.nvidia\.com/gpu
You want to see allocatable GPU capacity, for example:
NAME GPU
worker-01 1
worker-02 4
Deploy your first GPU-enabled pod
The simplest end-to-end test is to run nvidia-smi inside a container. If it works, Kubernetes allocated the GPU and the container can access it.
apiVersion: v1
kind: Pod
metadata:
name: gpu-test
spec:
restartPolicy: Never
containers:
- name: cuda-container
image: nvidia/cuda:12.9.0-base-ubuntu24.04
command: ["nvidia-smi"]
resources:
limits:
nvidia.com/gpu: 1
What the important lines mean:
- Image: the NVIDIA CUDA base image includes
nvidia-smi. - Command: run
nvidia-smiand then exit. nvidia.com/gpu: 1: schedule this pod only where one NVIDIA GPU can be allocated.
Apply the pod and read its logs:
kubectl apply -f gpu-pod.yaml
kubectl logs gpu-test
If the logs show the NVIDIA-SMI table, the basic GPU path is working end to end.
Why a GPU pod can remain Pending
Pending does not always mean something is broken. It may simply mean that no node currently satisfies the GPU request.
For example, a pod requesting one GPU can run only where an allocatable GPU is free. It should not quietly fall back to a CPU-only node.
- No GPU exists on matching nodes.
- All GPUs are already allocated.
- nodeSelector or affinity rules are too restrictive.
- A taint blocks the pod and no matching toleration is present.
- A DRA request cannot currently find a matching device.
Always inspect pod events before assuming the application itself is broken:
kubectl describe pod <pod-name>
A beginner-friendly troubleshooting order
- Host — run
nvidia-smion the worker. Can Linux see the GPU? - Operator — are the GPU Operator components healthy?
- Discovery — do expected GPU labels exist on the node?
- Resource advertisement — does the node show allocatable
nvidia.com/gpu? - Scheduling — does the pod request a GPU, and do its selectors, affinity, and taints allow placement?
- Runtime — can a tiny CUDA test pod run
nvidia-smisuccessfully? - Application — only after the platform path works, debug the model or inference server itself.
It moves from the lowest layer to the highest layer, so you do not waste time debugging an application when the host or Kubernetes GPU path is already broken.
Which scheduling method should I use?
| Requirement | Good starting point |
|---|---|
| I just need any GPU | nvidia.com/gpu resource request |
| I need a specific GPU model | Resource request + nodeSelector |
| I need required and preferred characteristics | Resource request + Node Affinity |
| I want to reserve GPU nodes for GPU workloads | Taints + Tolerations |
| I want to describe device characteristics dynamically | DRA |
These approaches are not mutually exclusive. Production environments often combine resource requests, labels or affinity, and taints/tolerations.
Put everything together
Five things to remember
- Kubernetes understands CPU and memory by default, but GPUs need vendor-specific integration.
- NFD discovers general hardware; GFD adds detailed NVIDIA GPU information.
- The NVIDIA Device Plugin exposes GPU capacity as resources such as
nvidia.com/gpu. - Scheduling can be simple or advanced: resource requests, selectors, affinity, taints/tolerations, and DRA each solve different problems.
- The GPU Operator simplifies lifecycle management of the NVIDIA GPU software stack across the cluster.
Quick glossary
| Term | Simple meaning |
|---|---|
| NFD | Discovers general hardware and node features |
| GFD | Adds NVIDIA-GPU-specific labels and characteristics |
| Device Plugin | Reports GPU capacity and supports allocation |
nvidia.com/gpu | The traditional Kubernetes extended resource used to request an NVIDIA GPU |
| Container Toolkit | Provides the container runtime path to the GPU |
| MIG | Splits a supported physical GPU into isolated GPU instances |
| nodeSelector | Hard requirement for a node label |
| Node Affinity | Flexible required/preferred node matching |
| Taint | Keeps pods away from a node unless tolerated |
| Toleration | Allows a pod onto a matching tainted node |
| DRA | Framework for requesting specialized devices by characteristics |
| GPU Operator | Installs and manages major NVIDIA GPU software components |
| DCGM Exporter | Exports NVIDIA GPU metrics for monitoring |