Kubernetes on GPU

A beginner-friendly guide to discovering, scheduling, sharing, and operating GPUs in Kubernetes.

Open slides →

Start with the end-to-end picture ↓

Kubernetes is excellent at scheduling CPU and memory, but GPUs are specialized devices. We need a few extra components to help Kubernetes discover GPUs, expose them as resources, and make them usable inside containers.

1. The simplest mental model

Keep this path in mind

Before learning individual components, keep this end-to-end picture in mind. Everything else in this lesson is simply explaining one part of this flow.

Physical NVIDIA GPU Linux / NVIDIA Driver GPU discovery (NFD / GFD) NVIDIA Device Plugin Kubernetes sees nvidia.com/gpu Scheduler chooses a suitable node Container runtime exposes the GPU Pod runs the AI workload
From the physical card to a running AI pod. Each later section explains one step.
2. Why GPUs matter

Why GPUs matter for Generative AI

Large language models perform a huge number of mathematical operations when they process prompts and generate tokens. CPUs can perform those calculations, but GPUs are designed to perform many calculations in parallel, which makes them especially useful for AI workloads.

GPUs were originally built for graphics and gaming, but the same parallel-computing design is a strong fit for model training and inference.

Important

Having a GPU physically installed in a server does not automatically mean Kubernetes knows how to use it.

3. The default gap

Why Kubernetes does not use GPUs automatically

A fresh Kubernetes cluster already understands common resources such as:

  • CPU
  • Memory
  • Storage
  • Network interfaces

GPUs are different. NVIDIA, AMD, Intel, and other vendors have different drivers, libraries, runtimes, and management tools. Kubernetes therefore does not build one vendor-specific GPU implementation directly into its core.

Imagine a cluster with three workers:

Worker 1 CPU only
Worker 2 NVIDIA T4
Worker 3 NVIDIA A100
A GPU pod can only land on a worker that actually has a usable GPU.
Which workers can run a GPU workload
NodeHardwareCan it run a GPU workload?
Worker 1CPU onlyNo
Worker 2NVIDIA T4Potentially yes
Worker 3NVIDIA A100Potentially yes

For a GPU workload, Kubernetes may need to answer:

  • Which nodes have GPUs?
  • Which GPU model is installed?
  • How many GPUs are available?
  • How much GPU memory does the device have?
  • How does the container access the GPU?
  • Can multiple applications share the device?
  • What happens when a model needs more than one GPU?

That leads us to the first major topic: GPU discovery.

4. GPU discovery

First find the hardware

Before Kubernetes can make a good scheduling decision, it needs information about the hardware available on each node.

Easy way to remember

NFD discovers general node hardware. GFD adds NVIDIA-GPU-specific details.

4.1 Node Feature Discovery (NFD)

Node Feature Discovery is like a hardware scanner for Kubernetes nodes. It inspects each node and adds labels that describe useful hardware and software features.

NFD can discover information such as:

  • CPU features
  • PCI devices
  • Network devices
  • Kernel features
  • GPUs and other hardware

NFD typically runs as a DaemonSet. If you are new to Kubernetes, a DaemonSet simply means: run one copy of this component on every applicable node.

Node 1 NFD worker Scan hardware
Node 2 NFD worker Scan hardware
Node 3 NFD worker Scan hardware
A DaemonSet runs one NFD worker on each node so every machine gets scanned.

One way to install NFD is with Kustomize:

NFD_REPO=https://github.com/kubernetes-sigs/node-feature-discovery

kubectl apply -k \
  $NFD_REPO/deployment/overlays/default

After NFD starts, you can inspect labels on a node:

kubectl get node <node-name> -o yaml

You may see a label similar to:

feature.node.kubernetes.io/pci-0302_10de.present: "true"

In plain English:

  • 0302 is a PCI class associated with a 3D controller.
  • 10de is NVIDIA’s PCI vendor ID.
  • true means that hardware is present on the node.

So this label is essentially telling us: this node contains NVIDIA GPU hardware.

PCI vendor IDs
Vendor IDVendor
10deNVIDIA
1002AMD
8086Intel

4.2 GPU Feature Discovery (GFD)

NFD can tell us that NVIDIA hardware exists, but AI workloads often need more detail. A T4, A100, and H100 are all GPUs, but they differ significantly in memory and capabilities.

GPU Feature Discovery focuses on NVIDIA GPUs and can add labels describing the device in more detail.

NFD “This node has NVIDIA GPU hardware.” GFD “A100, about 40 GB, supports MIG.”
NFD says the hardware exists. GFD says what kind of NVIDIA GPU it is.

GFD may publish labels such as:

nvidia.com/gpu.count: "4"
nvidia.com/gpu.product: "A100-SXM4-40GB"
nvidia.com/gpu.memory: "40537"
nvidia.com/gpu.family: "ampere"
nvidia.com/mig.capable: "true"
Common GFD labels
LabelMeaning
nvidia.com/gpu.countHow many GPUs are present
nvidia.com/gpu.productThe NVIDIA GPU model
nvidia.com/gpu.memoryGPU memory reported by discovery
nvidia.com/gpu.familyGPU architecture family
nvidia.com/mig.capableWhether the GPU supports MIG

This becomes useful when a cluster contains different GPU models. For example, an application can target a specific GPU label:

nodeSelector:
  nvidia.com/gpu.product: A100-SXM4-40GB
Remember

Discovery describes the GPU, but discovery alone does not make the GPU allocatable to a pod.

5. Device Plugin

Make the GPU a Kubernetes resource

The NVIDIA Device Plugin is the bridge between the physical NVIDIA GPU and Kubernetes resource management.

Physical NVIDIA GPU NVIDIA Device Plugin Kubelet Kubernetes sees nvidia.com/gpu: 1
The Device Plugin reports GPU capacity so Kubernetes can allocate it like any other resource.

The plug-in normally runs on GPU nodes as a DaemonSet and communicates with the kubelet.

Its main responsibilities are:

  1. Device discovery — detect GPUs available on the node.
  2. Report capacity — tell kubelet how many GPU resources are available.
  3. Allocation — help assign an available GPU when a pod requests one.
  4. Health tracking — stop unhealthy GPUs from being offered to new workloads.

For example, a node with four healthy GPUs may advertise:

nvidia.com/gpu: 4

5.1 Requesting a GPU from a pod

Once the Device Plugin is working, requesting a GPU is simple:

resources:
  limits:
    nvidia.com/gpu: 1

This means: “My container needs one NVIDIA GPU.”

If the cluster has Node 1 with no GPUs, Node 2 with two GPUs, and Node 3 with four GPUs, the scheduler can consider Node 2 or Node 3 as long as at least one GPU is free.

Important limitation

nvidia.com/gpu: 1 asks for one GPU. By itself, it does not say whether that GPU should be a T4, A100, H100, or another model.

6. The container path

The container still needs a path to the GPU

There are two separate questions:

1. Can Kubernetes schedule this pod onto a node that has a GPU? Device Plugin

Helps with allocation

2. Can the process inside the container talk to that GPU? Container Toolkit + NVIDIA runtime

Makes the selected GPU visible inside the container

Scheduling and container access are different problems. Both have to work.
Common confusion

A node can advertise nvidia.com/gpu correctly and a pod can still fail to use the GPU if the container runtime path is not configured correctly.

7. GPU scheduling

How Kubernetes decides where the pod runs

GPU scheduling starts with the same scheduler Kubernetes already uses. The difference is that GPU workloads add device requirements.

Pod requests a GPU Scheduler
├──Resource request — any NVIDIA GPU
├──nodeSelector — a particular model
├──Node Affinity — required and preferred traits
├──Taints + Tolerations — who may use the node
└──DRA — describe the device itself
Each mechanism answers a different scheduling question. They are often combined.
Scheduling needs and the Kubernetes mechanism that fits
NeedUseful Kubernetes mechanism
Any available NVIDIA GPUResource request
A particular GPU modelResource request + nodeSelector
Required and preferred GPU characteristicsResource request + Node Affinity
Protect expensive GPU nodesTaints + Tolerations
Describe device requirements dynamicallyDynamic Resource Allocation (DRA)

7.1 Resource-based scheduling

The simplest case is when any NVIDIA GPU is acceptable:

resources:
  limits:
    nvidia.com/gpu: 1

This works well when all GPU nodes are equivalent for the workload.

But if your model needs 35 GB of GPU memory, a 16 GB T4 and an 80 GB H100 cannot be treated as interchangeable even though both count as one GPU.

7.2 nodeSelector: choose an exact GPU type

If the workload needs a particular GPU model, use a label-based scheduling rule.

spec:
  nodeSelector:
    nvidia.com/gpu.product: Tesla-T4

  containers:
  - name: llm
    image: my-llm:latest
    resources:
      limits:
        nvidia.com/gpu: 1

These two settings do different jobs:

  • nvidia.com/gpu: 1 means: allocate one GPU.
  • nodeSelector means: only consider nodes with this exact GPU label.

nodeSelector is strict: if no matching node is available, the pod remains Pending rather than silently moving to a different GPU type.

7.3 Node Affinity: required plus preferred

Node Affinity is useful when the scheduling rule is more flexible.

For example: the model must have more than 40,000 MiB of GPU memory, and we prefer a Hopper-family GPU.

affinity:
  nodeAffinity:
    requiredDuringSchedulingIgnoredDuringExecution:
      nodeSelectorTerms:
      - matchExpressions:
        - key: nvidia.com/gpu.memory
          operator: Gt
          values:
          - "40000"

    preferredDuringSchedulingIgnoredDuringExecution:
    - weight: 1
      preference:
        matchExpressions:
        - key: nvidia.com/gpu.family
          operator: In
          values:
          - hopper

The YAML looks complex, but the idea is simple:

  • Required: the node must satisfy this condition.
  • Preferred: if several nodes satisfy the requirement, prefer this kind of node.

7.4 Taints and Tolerations: protect expensive GPU nodes

nodeSelector and affinity answer: where should this pod run? Taints and tolerations answer a different question: which pods are allowed onto this node?

GPU nodes can be expensive, so you may want to keep ordinary workloads away from them.

Taint the GPU node:

kubectl taint nodes gpu-node \
  nvidia.com/gpu=true:NoSchedule

Then allow a GPU workload to tolerate that taint:

tolerations:
- key: "nvidia.com/gpu"
  operator: "Exists"
  effect: "NoSchedule"
Very important

A toleration does not force a pod onto a GPU node. It only gives the pod permission to run there. You still normally combine it with a GPU resource request, nodeSelector, or affinity.

8. Dynamic Resource Allocation

Ask for a device, not a node label

Traditional GPU scheduling often makes us think about nodes: find the node with the right labels, then schedule the pod there.

Dynamic Resource Allocation changes the mental model:

Traditional “Run me on a node with this GPU label.”
DRA “I need a GPU with these characteristics. Please find it.”
DRA shifts the request from “which node” to “what device.”

The core DRA framework is a Kubernetes mechanism for describing and allocating specialized devices. A compatible vendor DRA driver publishes the real devices and their attributes.

A simple way to remember the main objects:

DRA objects
ObjectBeginner-friendly meaning
DeviceClassWhat category of device can be requested
ResourceSliceWhat devices are available
ResourceClaimTemplateA reusable GPU request template
ResourceClaimThe actual device request
DRA driverThe vendor-specific bridge to the real hardware

8.1 ResourceClaim: think of it as a GPU request form

Instead of naming a specific physical GPU, the workload describes what it needs.

GPU request
├──Type: NVIDIA GPU
├──Count: 1
├──Memory: at least 40 GB
└──Optional: a particular model or family
Scheduler + DRA driver find a matching device
The claim describes the device. Kubernetes and the DRA driver try to match it.

8.2 A simplified ResourceClaimTemplate

A simplified example uses the Kubernetes resource API like this:

apiVersion: resource.k8s.io/v1
kind: ResourceClaimTemplate
metadata:
  name: gpu-claim-template
spec:
  spec:
    devices:
      requests:
      - name: high-memory-gpu
        exactly:
          deviceClassName: gpu.nvidia.com
          allocationMode: ExactCount
          count: 1

The key idea is not to memorize every API field. Focus on the intent:

  • deviceClassName says which category of device we want.
  • allocationMode says how allocation should work.
  • count says how many devices are needed.

DRA can also use driver-published attributes and selectors to describe more specific device requirements. Exact attribute names depend on the installed DRA driver, so they should be checked rather than guessed.

8.3 Connecting the claim to an application

apiVersion: apps/v1
kind: Deployment
metadata:
  name: inference-server
spec:
  replicas: 1
  selector:
    matchLabels:
      app: inference-server
  template:
    metadata:
      labels:
        app: inference-server
    spec:
      resourceClaims:
      - name: high-memory-gpu
        resourceClaimTemplateName: gpu-claim-template

      containers:
      - name: model-runner
        image: myorg/llm-inference:latest
        resources:
          claims:
          - name: high-memory-gpu

The pod-level resourceClaims section creates or associates the device request. The container-level resources.claims section says which container should receive access to that allocated device.

This matters because a pod may contain a GPU-using model container and another sidecar container that does not need GPU access.

8.4 Traditional Device Plugin vs DRA

Traditional device plugin compared with DRA
Traditional approachDRA approach
Ask for nvidia.com/gpu: 1Describe a device requirement
Often combine with node labelsSelect through device attributes
Think about which node has the GPUThink about what device the application needs
Simple and widely understoodMore expressive for complex device requirements
Practical takeaway

Traditional nvidia.com/gpu requests remain useful and straightforward. DRA becomes especially interesting as device requirements become more specific or heterogeneous.

9. Multi-Instance GPU

Split one large GPU into smaller isolated instances

On supported GPUs such as A100 and H100, Multi-Instance GPU (MIG) can divide one physical GPU into smaller isolated GPU instances.

Without MIG, a request for one GPU often reserves the whole device. With MIG, multiple workloads can use isolated slices of the same physical GPU, each with its own assigned memory and compute resources.

Without MIG Application A → whole GPU
With MIG Application A → GPU slice 1 Application B → GPU slice 2 Application C → GPU slice 3
MIG turns one exclusive card into several isolated slices.

Why it matters: MIG can improve utilization when several smaller workloads do not need an entire high-end GPU.

10. NVIDIA GPU Operator

Manage the GPU software stack

Manually preparing every GPU node can become painful. You may need drivers, the Container Toolkit, the Device Plugin, discovery components, monitoring, and MIG management.

The NVIDIA GPU Operator acts as an installation and lifecycle manager for this stack.

GPU Operator components
ComponentPurpose
NVIDIA DriverLets Linux communicate with the GPU
Node Feature Discovery (NFD)Discovers general node hardware
GPU Feature Discovery (GFD)Adds detailed NVIDIA GPU labels
NVIDIA Device PluginExposes GPU resources to Kubernetes
NVIDIA Container ToolkitLets containers access allocated GPUs
MIG ManagerConfigures MIG on supported GPUs
DCGM ExporterExposes GPU telemetry for monitoring

This is one of the main reasons the GPU Operator is useful: instead of managing several pieces individually across many nodes, you install the Operator and let it manage the NVIDIA GPU stack.

11. Hands-on

Deploy the NVIDIA GPU Operator

Prerequisites

  • A Kubernetes cluster
  • At least one worker node with an NVIDIA GPU
  • NVIDIA drivers already present, or installable by the Operator
  • kubectl configured for the cluster
  • Helm installed on your machine

Step 1 — Confirm that Linux can see the GPU

Run this directly on the GPU node:

nvidia-smi

You should see the GPU model, driver version, GPU memory, and other device information.

Debugging rule

If nvidia-smi fails on the host, fix the host-level GPU problem before debugging Kubernetes.

Steps 2–4 — Install the GPU Operator with Helm

helm repo add nvidia https://helm.ngc.nvidia.com/nvidia
helm repo update

kubectl create namespace gpu-operator

helm install gpu-operator nvidia/gpu-operator \
  --namespace gpu-operator

The chart deploys several components, so give the cluster time to create and initialize the pods.

Step 5 — Check that the Operator pods are healthy

kubectl get pods -n gpu-operator

You should expect components related to the operator, device plugin, container toolkit, discovery, monitoring, and drivers when driver management is enabled. Exact pod names can vary by version.

Step 6 — Confirm that Kubernetes advertises the GPU

kubectl describe node <gpu-node-name>

# or

kubectl get nodes \
  -o custom-columns=NAME:.metadata.name,GPU:.status.allocatable.nvidia\.com/gpu

You want to see allocatable GPU capacity, for example:

NAME        GPU
worker-01   1
worker-02   4
12. Hands-on

Deploy your first GPU-enabled pod

The simplest end-to-end test is to run nvidia-smi inside a container. If it works, Kubernetes allocated the GPU and the container can access it.

apiVersion: v1
kind: Pod
metadata:
  name: gpu-test
spec:
  restartPolicy: Never
  containers:
  - name: cuda-container
    image: nvidia/cuda:12.9.0-base-ubuntu24.04
    command: ["nvidia-smi"]
    resources:
      limits:
        nvidia.com/gpu: 1

What the important lines mean:

  • Image: the NVIDIA CUDA base image includes nvidia-smi.
  • Command: run nvidia-smi and then exit.
  • nvidia.com/gpu: 1: schedule this pod only where one NVIDIA GPU can be allocated.

Apply the pod and read its logs:

kubectl apply -f gpu-pod.yaml
kubectl logs gpu-test

If the logs show the NVIDIA-SMI table, the basic GPU path is working end to end.

13. Pending pods

Why a GPU pod can remain Pending

Pending does not always mean something is broken. It may simply mean that no node currently satisfies the GPU request.

For example, a pod requesting one GPU can run only where an allocatable GPU is free. It should not quietly fall back to a CPU-only node.

  • No GPU exists on matching nodes.
  • All GPUs are already allocated.
  • nodeSelector or affinity rules are too restrictive.
  • A taint blocks the pod and no matching toleration is present.
  • A DRA request cannot currently find a matching device.

Always inspect pod events before assuming the application itself is broken:

kubectl describe pod <pod-name>
14. Troubleshooting

A beginner-friendly troubleshooting order

1. Host — can Linux see the GPU with nvidia-smi? 2. Operator — are the GPU Operator components healthy? 3. Discovery — do expected GPU labels exist on the node? 4. Resource advertisement — is nvidia.com/gpu allocatable? 5. Scheduling — request, selectors, affinity, and taints 6. Runtime — can a tiny CUDA pod run nvidia-smi? 7. Application — only then debug the model
Move from the lowest layer to the highest so you do not debug the app while the platform path is broken.
  1. Host — run nvidia-smi on the worker. Can Linux see the GPU?
  2. Operator — are the GPU Operator components healthy?
  3. Discovery — do expected GPU labels exist on the node?
  4. Resource advertisement — does the node show allocatable nvidia.com/gpu?
  5. Scheduling — does the pod request a GPU, and do its selectors, affinity, and taints allow placement?
  6. Runtime — can a tiny CUDA test pod run nvidia-smi successfully?
  7. Application — only after the platform path works, debug the model or inference server itself.
Why this order works

It moves from the lowest layer to the highest layer, so you do not waste time debugging an application when the host or Kubernetes GPU path is already broken.

15. Choosing a method

Which scheduling method should I use?

Which scheduling method to start with
RequirementGood starting point
I just need any GPUnvidia.com/gpu resource request
I need a specific GPU modelResource request + nodeSelector
I need required and preferred characteristicsResource request + Node Affinity
I want to reserve GPU nodes for GPU workloadsTaints + Tolerations
I want to describe device characteristics dynamicallyDRA

These approaches are not mutually exclusive. Production environments often combine resource requests, labels or affinity, and taints/tolerations.

16. The full path

Put everything together

Physical GPU NVIDIA Driver NFD — “Hardware exists” GFD — detailed NVIDIA GPU characteristics NVIDIA Device Plugin — expose GPU capacity nvidia.com/gpu Pod requests a GPU Scheduler
├──Resource availability
├──nodeSelector / affinity
├──Taints / tolerations
└──DRA, when used
Kubelet + NVIDIA container runtime / toolkit Container receives GPU access LLM / AI workload runs
Discovery, allocation, scheduling, and the container runtime all have to line up before the model runs.
17. What to remember

Five things to remember

  1. Kubernetes understands CPU and memory by default, but GPUs need vendor-specific integration.
  2. NFD discovers general hardware; GFD adds detailed NVIDIA GPU information.
  3. The NVIDIA Device Plugin exposes GPU capacity as resources such as nvidia.com/gpu.
  4. Scheduling can be simple or advanced: resource requests, selectors, affinity, taints/tolerations, and DRA each solve different problems.
  5. The GPU Operator simplifies lifecycle management of the NVIDIA GPU software stack across the cluster.
Glossary

Quick glossary

Glossary of Kubernetes GPU terms
TermSimple meaning
NFDDiscovers general hardware and node features
GFDAdds NVIDIA-GPU-specific labels and characteristics
Device PluginReports GPU capacity and supports allocation
nvidia.com/gpuThe traditional Kubernetes extended resource used to request an NVIDIA GPU
Container ToolkitProvides the container runtime path to the GPU
MIGSplits a supported physical GPU into isolated GPU instances
nodeSelectorHard requirement for a node label
Node AffinityFlexible required/preferred node matching
TaintKeeps pods away from a node unless tolerated
TolerationAllows a pod onto a matching tainted node
DRAFramework for requesting specialized devices by characteristics
GPU OperatorInstalls and manages major NVIDIA GPU software components
DCGM ExporterExports NVIDIA GPU metrics for monitoring