Process Scheduler

The CPU can execute only a limited number of tasks at any moment. Hundreds or even thousands of processes and threads may be ready to run. The scheduler is how Linux chooses among them.

Start with what the scheduler actually decides ↓

Inside the kernel

The Process Scheduler is one of the most important parts of Linux

The Process Scheduler is one of the most important parts of the Linux kernel because the CPU can execute only a limited number of tasks at any moment. On a system, however, hundreds or even thousands of processes and threads may be ready to run. The scheduler decides which task gets CPU time, on which CPU core, and for how long. Its goal is to keep the system responsive, share CPU time fairly, and make good use of all available cores.

A simple way to visualize it is:

Four processes feed into the Process Scheduler, which selects one task to run on a CPU core.
The scheduler sits between ready processes and the CPU. Only a few tasks can actually run at once.
Process A Process B Process C Process D
Process Scheduler CPU Core

If several processes are ready at the same time, the scheduler chooses one of them to run. After that process has used some CPU time, becomes blocked, or is preempted, the scheduler may select another process. The switch from one running task to another is called a context switch.

For example:

CPU Core 0 Process A Process B Process A Process C
These switches can happen very quickly, which creates the impression that many applications are running simultaneously.

The scheduler also works closely with process states. A process may be:

Running Ready / Runnable Sleeping / Waiting Running again

For example, suppose a process is waiting for data from disk or the network. There is no point in keeping it on the CPU while it waits, so the kernel can put it to sleep and allow another runnable process to use the CPU.

On a multi-core system, the scheduler also decides which core should execute a task:

Linux Scheduler
CPU 0 Process A
CPU 1 Process B
CPU 2 Process C

It may also move tasks between cores to keep the workload balanced. This is known as load balancing.

The Linux process scheduler decides which process or thread gets the CPU, when it gets the CPU, and for how long.

Now that you understand the basic role of the Process Scheduler, the next step is to understand what exactly the scheduler is scheduling and how we can observe it on a Linux system.

Process and Thread

Before going further, we need to understand two words

A process is basically a running program.

For example:

python app.py

creates a Python process.

A process can contain one or more threads. Threads are smaller units of execution inside that process.

Python Process
├── Thread 1
├── Thread 2
├── Thread 3
└── Thread 4
These threads share many resources belonging to the same process, such as memory and open files.

On a multi-core system, different threads can potentially run on different CPU cores:

Thread 1 CPU 0
Thread 2 CPU 1
Thread 3 CPU 2
Thread 4 CPU 3

Linux scheduling is actually performed at the task/thread level, which is why multithreaded applications can make use of multiple CPU cores.

Process states

A process is not always running

One of the most important things to understand is that a process does not spend its entire lifetime executing on the CPU.

Sometimes it is running. Sometimes it is ready to run but waiting for CPU time. Sometimes it is waiting for disk or network activity. Sometimes it has already finished.

Linux represents these situations using process states. The most important ones are:

  • R = Running / Runnable
  • S = Sleeping
  • D = Uninterruptible Sleep
  • T = Stopped
  • Z = Zombie

These states are extremely important during Linux troubleshooting and interviews.

R

Running or Runnable

R means the process is either currently using a CPU or is ready to use one.

Imagine four processes but only one available CPU:

CPU Process A — currently running

Waiting for CPU

Process B Process C Process D

Processes B, C, and D are ready to run. The scheduler decides when each of them gets CPU time.

So R does not always mean “this process is currently executing.”

It can also mean “this process is ready and waiting for CPU.”

S

Sleeping

Most processes on a normal Linux server spend a lot of time sleeping.

Sleeping simply means: the process has nothing useful to do until something happens.

For example, Nginx may be waiting for a new request.

Nginx Waiting for network request S state

Once a request arrives:

Request arrives Process wakes up Runnable Scheduler gives it CPU

So seeing many processes in S state is completely normal.

D

Uninterruptible Sleep

D state is much more interesting.

A process usually enters D state when it is waiting inside the kernel for something such as:

NOTE: “Waiting inside the kernel” means the process made a request to the kernel and is blocked until that kernel operation, usually I/O, finishes.

  • Disk I/O
  • NFS
  • Storage
  • Filesystem
  • Device
  • Driver

For example:

Application Read file Kernel Waiting for storage D state

The process is alive, but it cannot continue until that kernel operation finishes.

This leads to one of the most famous Linux interview questions:

Why doesn't kill -9 kill a process stuck in D state?

Because the process is currently waiting inside the kernel.

Even if we send:

kill -9 1234

Linux may not terminate it immediately. The process first needs to return from that uninterruptible kernel operation.

That means if a process remains in D state for a long time, the real question is not “How do I kill this process?”

The better question is: What is this process waiting for?

It could be a slow disk, broken NFS mount, storage problem, driver issue, or another I/O problem.

To check processes currently in D state (uninterruptible sleep) on Linux, use:

ps -eo pid,ppid,state,comm,wchan:32 | awk '$3 == "D"'

You can also quickly check with:

ps aux | awk '$8 ~ /^D/'
Z

Zombie Process

A zombie is completely different from a process in D state.

A zombie process has already finished running.

Suppose a parent creates a child process:

Parent Child

The child finishes:

Parent Child → Finished

When the child exits, Linux keeps a very small amount of information about it so that the parent can check how it finished. The parent normally collects this information using something like:

wait()

After that, Linux completely removes the child.

But what happens if the parent does not collect the result?

Child finishes Parent doesn't collect result Zombie

That process appears as Z, or sometimes <defunct>, in commands such as ps.

Can we kill a zombie?

Another classic interview question:

Can we kill a zombie process using kill -9?

No. Why? Because the zombie is already dead. There is nothing left to kill.

The problem is actually with its parent process, which has not collected the child's exit information.

That gives us a very easy way to remember the difference:

D state

Still alive Waiting inside the kernel

Z state

Already finished Parent has not cleaned it up

To check for zombie processes on Linux, look for process state Z.

ps -eo pid,ppid,state,stat,comm | awk '$3 == "Z"'

A simpler command is:

ps aux | awk '$8 ~ /^Z/'

You can also use:

ps -el | grep ' Z '
Context Switching

Linux needs a bookmark when it changes books

Now suppose we have one CPU and three runnable processes: Process A, Process B, and Process C.

The scheduler may run them like this:

Time ─────────────────────────────→

A B A C B A

When Linux stops running Process A and starts running Process B, that is called a context switch.

Linux needs to remember where Process A stopped so that it can continue later.

Think of reading two books. You are reading Book A. Someone asks you to read Book B. You put a bookmark in Book A, switch to Book B, and later return to exactly where you stopped.

Linux does something conceptually similar.

Process A Save its state Switch Process B

Context switching is necessary, but it also has a cost. If the system switches between tasks excessively, CPU time is being spent switching instead of doing useful work.

Load average

Does a process in D state or Zombie state increase the load average?

Yes, but in very different ways.

A process in D state can increase the Linux load average, even though it may be using little or no CPU. Linux load average includes tasks that are either runnable (R) or stuck in uninterruptible sleep (D). So if many processes are waiting on disk, NFS, storage, or another kernel I/O operation, you can see a high load average while CPU utilization stays relatively low.

For example:

CPU usage:     20%
Load average:  15.0

That can happen when many processes are in D state waiting for I/O.

A zombie (Z) process, on the other hand, does not consume CPU and does not normally contribute to load average. It has already finished executing. It remains only as a small process-table entry because its parent has not yet collected its exit status.

A useful way to remember it:

State CPU usage Contributes to load average?
R Running / Runnable Yes / potentially Yes
D Uninterruptible sleep Usually no Yes
S Sleeping No No
Z Zombie No No

This is why high load average does not always mean high CPU usage. A server with many D-state processes can have a very high load while the CPUs are mostly idle.

CPU utilization and load average are not the same thing.

Think of it this way:

CPU usage

CPU usage How busy is the CPU right now?

Load average

Load average How many tasks are running / waiting for CPU (R), or waiting in uninterruptible I/O (D)?

So you can absolutely have:

CPU Usage:     15%
Load Average:  20

and that is not contradictory.

For example, imagine 20 processes are stuck waiting for slow storage:

Process 1 Process 2 Process 3 … Process 20
Disk / NFS / Storage is slow

State = D. The CPU might be mostly idle, but Linux still counts those tasks toward load average.

A very useful mental model is:

CPU utilization = how busy the CPUs are.
Load average = how much work is competing for CPU or stuck waiting on uninterruptible I/O.

And this leads to a great troubleshooting lesson:

High load plus high CPU likely means CPU pressure. High load plus low CPU means look for D-state processes and I/O problems. Low load plus high CPU means a small number of CPU-intensive processes.
High load is a queue. High CPU is busy cores. Those two facts can disagree, and that disagreement is a clue.
High Load + High CPU Likely CPU pressure
High Load + Low CPU Look for D-state processes / I/O problems
Low Load + High CPU A small number of CPU-intensive processes

This distinction is actually one of the most important things to understand when troubleshooting Linux performance.

Two kinds of switches

Voluntary and involuntary context switches

There are two useful types.

A voluntary context switch happens when a process cannot continue and gives up the CPU.

Process Needs data from disk Waits CPU given to another process

An involuntary context switch happens when the process could continue running, but Linux decides another task should get CPU time.

Process A running Scheduler interrupts A Process B runs

In simple terms:

Voluntary

“I cannot continue right now.”

Involuntary

“You have had enough CPU for now.”

Multiple CPU cores

What happens with more than one CPU?

Most modern machines have several CPU cores. Suppose we have four cores:

Scheduler
CPU 0 Task A
CPU 1 Task B
CPU 2 Task C
CPU 3 Task D

Linux tries to distribute work across those CPUs. If this happens:

CPU 0 = 100%
CPU 1 = 100%
CPU 2 = 5%
CPU 3 = 3%

the scheduler may move runnable tasks so that the workload is distributed more efficiently. This is known as load balancing.

CPU Affinity

Sometimes we restrict which CPU a task may use

Normally Linux is free to move a task between CPU cores.

Process A
CPU 0 CPU 1 CPU 2

But sometimes we want a process to run only on certain CPUs. This is called CPU affinity.

For example:

taskset -c 2 ./my_application

means: run this application on CPU 2.

CPU affinity is useful in areas such as performance testing, databases, low-latency applications, NUMA systems, and some high-performance or GPU workloads.

For a beginner, however, the most important idea is simply:

Normally the scheduler chooses the CPU. CPU affinity lets us restrict that choice.
How do we see it?

How do we see process states?

One of the easiest commands is:

ps -eo pid,ppid,stat,comm

You may see:

PID    PPID   STAT   COMMAND
1200     1      S    nginx
1400     1      R    python
1600     1      D    backup
1800  1500      Z    worker

Now the letters actually mean something:

  • R → Ready / running on CPU
  • S → Waiting normally
  • D → Waiting inside kernel, often for I/O
  • Z → Already finished, waiting for parent cleanup

That is much more useful than simply memorizing the letters.

top

The next command is:

top

It gives us a live view of processes and CPU usage. Some useful CPU values are:

  • us = CPU time running user applications
  • sy = CPU time running kernel code
  • id = CPU idle time
  • wa = CPU time associated with I/O wait

vmstat

Another very useful command is:

vmstat 1

Two scheduler-related fields worth watching are r and cs. r represents runnable tasks, while cs represents context switches.

If a machine has 4 CPUs but consistently has a large runnable queue, such as:

CPUs = 4
Runnable tasks = 25

many tasks are competing for CPU time. That can be an indication of CPU contention.

pidstat

To look at individual processes:

pidstat 1

And to examine context switches:

pidstat -w 1

This can show voluntary and involuntary context switches for each process.

mpstat

To look at individual CPU cores:

mpstat -P ALL 1

This helps answer questions like: is the whole machine busy, or is only one CPU core overloaded? That distinction can be very useful during troubleshooting.

Putting everything together

Suppose someone tells us: “The server is slow.”

Instead of immediately restarting something, we can start building a picture:

Is CPU busy? top / mpstat Are many processes waiting for CPU? vmstat Which process is consuming CPU? top / pidstat Are processes stuck in D state? ps Are there zombie processes? ps Are context switches unusually high? vmstat / pidstat

Now the scheduler is no longer just theory. It becomes part of our troubleshooting process.

Connect it to GenAI

And this still matters for GenAI

Even if an application is using GPUs, there is still significant CPU-side activity.

For example:

Incoming Request Python / vLLM / PyTorch Tokenization CPU Threads Linux Scheduler CPU Cores GPU Work Submitted GPU
The GPU may perform the heavy matrix calculations, but Linux is still scheduling CPU threads that handle requests, networking, tokenization, data movement, and interaction with GPU drivers.

So the key idea to remember is:

The Linux scheduler decides which task gets CPU time, where it runs, and when another task should take its place. Process states such as R, S, D, and Z help us understand what each process is doing—or waiting for—at any given moment.
A stronger toolkit

Once you move beyond ps and top

Linux gives you several much more powerful commands for understanding what a process is doing, what it is waiting for, which threads it has, which files it opened, how much memory it uses, and where CPU time is going.

A good advanced toolkit is:

What you want to inspect Command
Detailed process information ps
Parent/child relationship pstree
Per-process CPU usage pidstat
Threads inside a process ps -L, top -H
Open files / sockets lsof
Process memory map pmap
Kernel / process details /proc/<PID>/
System calls strace
CPU performance perf
CPU affinity taskset
Scheduling policy / priority chrt
What a blocked process is waiting on wchan, /proc/<PID>/stack
Before you go

If you remember only four things

1. The scheduler chooses who runs. It decides which process or thread gets the CPU, on which core, and for how long.

2. R, S, D, and Z are not trivia. They tell you whether a process is competing for CPU, waiting normally, stuck inside the kernel, or already finished.

3. Load average is not CPU usage. D-state tasks can raise load while the CPUs sit idle. Zombies do not.

4. GenAI still has a CPU side. Tokenization, networking, request handling, and GPU submission are still scheduled by Linux.