Linux memory management

When a process asks for memory, where does that memory actually come from, and how does Linux manage it safely?

Start between the application and the RAM chips ↓

The layer in the middle

Applications do not manage physical RAM themselves

When we start an application such as Python, Nginx, MySQL, or an AI workload, that application needs memory for its code, variables, heap, stack, libraries, buffers, and other data. The application does not directly manage physical RAM. Instead, the Linux kernel memory manager sits between applications and physical memory and decides how memory should be allocated, mapped, protected, shared, reclaimed, or moved to swap.

Application Virtual Memory Linux Memory Manager Physical RAM
The hardware helper

The CPU never walks RAM with a virtual address

When a process uses a memory address, that address is a logical (virtual) address. A piece of hardware called the MMU — Memory Management Unit — translates it into a physical address before RAM is touched.

The CPU issues a logical address to the MMU, which translates it into a physical address used to reach physical memory.
The MMU sits between the CPU and RAM. Virtual addresses become physical addresses before a byte is read or written.
CPU MMU Physical Memory
Physical Memory vs Virtual Memory

16 GB of RAM is not what the process thinks it has

Suppose your machine has 16 GB RAM. That is the actual physical memory installed in the system.

But applications do not normally work directly with those physical addresses. Each process gets its own virtual address space.

Process A and Process B each have a virtual address space with code, data, heap, mappings, and stack. Linux maps those to different physical RAM pages.
Both processes may think they have memory starting from similar addresses. Linux maps those virtual addresses to different locations in physical RAM.

Process A

Virtual Address Space Stack Memory Mapping Heap Data Code

Process B

Virtual Address Space Stack Memory Mapping Heap Data Code

Conceptually:

Process A Virtual Memory Linux Kernel Physical RAM Page 100
Process B Virtual Memory Linux Kernel Physical RAM Page 500

This provides isolation. Process A normally cannot simply access Process B's memory.

Virtual Memory

Virtual Memory in Linux

When an application runs on Linux, it normally does not access physical RAM directly. Instead, every process gets its own virtual address space, and the Linux kernel maps virtual addresses to physical memory as needed. This provides isolation between processes and lets Linux manage memory efficiently.

A simplified view is:

Six steps from left to right: Application, Virtual Address, Page Table, MMU, Physical Address, and RAM.
The MMU uses the kernel’s page table to translate each virtual address before RAM is touched.
Application Virtual Address Page Table MMU Physical Address RAM

The CPU's MMU (Memory Management Unit) translates virtual addresses into physical addresses using page tables maintained by the kernel. Linux manages memory in fixed-size blocks called pages, commonly 4 KB on x86 systems.

Virtual address space is not RAM

Suppose your machine has:

Physical RAM = 32 GB

A process may still have a potential virtual address space of many terabytes.

For example:

$ lscpu | grep "Address sizes"

Address sizes: 46 bits physical, 48 bits virtual

This tells us two different things:

46-bit physical addressing
48-bit virtual addressing

For virtual addresses:

2^48 bytes
= 281,474,976,710,656 bytes
= 256 TiB

Because:

2^40 bytes = 1 TiB

2^48
= 2^8 × 2^40
= 256 TiB

So a CPU supporting 48-bit virtual addresses has a theoretical virtual-address range of 256 TiB.

Why do we often say 128 TiB User Space + 128 TiB Kernel Space?

On x86-64 with 48-bit virtual addressing, addresses are divided into two canonical ranges.

Conceptually:

A 48-bit virtual address space of 256 TiB, split into a lower half of 128 TiB for user space and an upper half of 128 TiB for kernel space, with a non-canonical gap between them.
Each canonical half is 128 TiB. User space is the lower half. Kernel space is the upper half. A non-canonical gap sits between them.
48-bit Virtual Address Space · 256 TiB
Lower half User Space 128 TiB
Upper half Kernel Space 128 TiB

Each half is:

2^47 bytes = 128 TiB

The lower canonical range is approximately:

0x0000000000000000
        ↓
0x00007fffffffffff

and is normally used for user-space addresses.

The upper canonical range is approximately:

0xffff800000000000
        ↓
0xffffffffffffffff

and Linux uses the upper region for kernel-space mappings.

There is a large non-canonical gap between them.

Why doesn't a 64-bit CPU use all 64 bits for addresses?

A 64-bit CPU does not mean 64 address bits must be implemented.

Using all 64 address bits would provide:

2^64 bytes = 16 EiB

which is far beyond what today's systems normally need.

Supporting that entire range would also increase hardware complexity in areas such as:

  • MMU
  • Page tables
  • TLBs
  • Address translation
  • Cache structures

So x86-64 was designed to initially implement fewer address bits while leaving room for expansion.

For example:

Common systems
48-bit virtual
→ 256 TiB

Newer 5-level paging systems
57-bit virtual
→ 128 PiB

This is why bits above the implemented virtual-address width must follow the architecture's canonical-address rules.

What does 46-bit physical mean?

From:

46 bits physical

we get:

2^46 bytes
= 64 TiB

So the CPU can theoretically address up to 64 TiB of physical address space.

It does not mean your machine actually contains 64 TiB of RAM.

For example:

CPU physical addressing capability → 64 TiB

Actual installed RAM → 64 GB

Those are completely different things.

Useful Linux commands

To check the CPU's supported address sizes:

lscpu | grep "Address sizes"

To see how much virtual memory a process currently has mapped:

grep VmSize /proc/<PID>/status

For your current shell:

grep VmSize /proc/$$/status

To inspect the actual virtual-address mappings:

cat /proc/<PID>/maps

or:

pmap <PID>

The most important idea

Do not confuse these three things:

Three cards: Virtual Address Space is 48 bits, 256 TiB. Physical Address Space is 46 bits, 64 TiB. Installed RAM is 32 GB. A process may actually use only 2 GB.
Virtual address space, physical address space, and installed RAM are three different sizes. A process may use only a small part of any of them.
Virtual Address Space

How much address range a process can potentially use

Physical Address Space

How much physical memory the CPU can address

Installed RAM

How much memory is actually present in the machine

For example:

64-bit CPU

Virtual address size:
48 bits
→ 256 TiB total canonical virtual range

Physical address size:
46 bits
→ up to 64 TiB addressable

Installed RAM:
32 GB

And a process may actually use only:

2 GB

Linux maps only the virtual pages that are required to physical memory, often allocating physical pages on demand when they are first accessed.

Simple takeaway: Virtual memory gives every process the illusion of having its own large, private memory space, while Linux and the MMU handle the mapping to the much smaller amount of real physical RAM underneath.
Why virtual memory?

The biggest win is process isolation

Virtual memory gives Linux several important benefits. The biggest one is process isolation.

Imagine if every process directly accessed physical RAM:

Process A Process B Process C
Same Physical RAM
A bad pointer in one application could overwrite another application's memory.

Instead Linux gives every process its own virtual view:

Process A Virtual Address Space A
Process B Virtual Address Space B
Process C Virtual Address Space C

The kernel controls how those virtual addresses map to actual physical memory.

Pages

Memory is managed in pages, not one byte at a time

Linux does not normally manage memory one byte at a time. Memory is divided into blocks called pages.

On many Linux systems the normal page size is 4 KB. You can check it with:

getconf PAGE_SIZE

You may see:

4096

which means 4096 bytes = 4 KB.

So conceptually physical RAM looks like:

Physical Memory
Page 0 Page 1 Page 2 Page 3 Page 4 Page 5 …

Virtual memory is also divided into pages.

Page Tables

Linux has to remember which virtual page is which physical page

That information is stored in page tables.

For example:

Process Virtual Memory
Virtual Page 1 Physical Page 120
Virtual Page 2 Physical Page 450
Virtual Page 3 Physical Page 900

The CPU's MMU uses these page tables when translating virtual addresses into physical addresses.

So the simplified flow is:

Application Virtual Address Page Table Physical Address RAM

The Linux kernel is responsible for creating and maintaining these mappings.

Demand paging

Does Linux allocate RAM immediately?

This is where memory management becomes interesting.

Suppose a program asks for 1 GB of memory. That does not necessarily mean Linux immediately reserves and physically fills 1 GB of RAM.

Linux often uses demand paging. Memory may only be physically allocated when the process actually starts using it.

Application asks for memory Virtual address range created Application accesses memory Physical page allocated

This helps Linux use memory efficiently.

Page Fault

A page fault is not automatically a problem

Suppose a process accesses a virtual memory address. The CPU checks the page table. If the required page is not currently mapped as needed, the CPU triggers a page fault.

This sounds like an error, but a page fault can be completely normal.

Process accesses memory Page is not currently mapped Page Fault Kernel handles it Memory page becomes available Process continues

There are two terms worth knowing.

Minor page fault

The required data is already available in memory, but the process's page table needs to be updated. This is relatively cheap.

Major page fault

Linux must retrieve the required data from storage. This is much slower.

For example, a major page fault looks like:

Process Needs page Page not in RAM Read from disk RAM

This distinction becomes very important during performance troubleshooting.

Stack and Heap

Two important areas inside the virtual address space

The stack is commonly used for things such as function calls and local variables.

For example:

void test() {
    int x = 10;
}

x may live on the stack.

The heap is commonly used for dynamically allocated memory.

For example:

malloc(1024);

Conceptually:

Process memory layout from high addresses to low: Stack growing down, Memory Mapping Segment, Heap growing up, Data and BSS, Text Code, then Kernel Space.
A process’s virtual address space, from high addresses down to low. Stack grows down. Heap grows up. Kernel space sits below the user-space mapping.
Process Virtual Memory

High address

Stack Memory Mapping Heap Data / BSS Code

Low address

Same layout as boxes: high address at the top, code near the bottom of user space.
When RAM gets tight

What happens when RAM becomes full?

Suppose the machine has 8 GB RAM and applications keep requesting memory. Linux does not immediately crash. The kernel first tries to reclaim memory that is no longer urgently needed.

One important source is the page cache. Linux uses unused RAM to cache filesystem data.

So instead of leaving RAM empty, Linux may use unused RAM as page cache. This makes disk access much faster. If applications need the memory later, Linux can reclaim some of that cache.

Low “free” memory does not automatically mean a Linux server has a memory problem.

That is a very important Linux concept.

Page Cache

Linux deliberately uses available RAM for caching

Suppose an application reads /data/file.txt.

The first time:

Application Disk RAM / Page Cache

The next time the same data is needed:

Application Page Cache Much faster

So when you run:

free -h

you should not look only at free. The more useful number is usually available, because some cached memory can be reclaimed when applications need it.

Swap

If Linux is under memory pressure, it may move pages to disk

This area is called swap.

RAM Less frequently used pages Swap Disk

Later, if the application needs those pages:

Swap RAM Application

But disk is much slower than RAM. So heavy swapping can cause serious performance problems.

You may see this situation:

RAM nearly full Heavy swap activity Applications become slow
OOM Killer

Out Of Memory Killer

Eventually Linux may reach a point where it cannot reclaim enough memory.

For example: RAM full, plus swap full or unavailable, plus memory requests that continue. Now the kernel faces a serious problem.

Linux may invoke the OOM Killer — Out Of Memory Killer. The kernel selects one or more processes to terminate in order to free memory.

You may find messages like this in dmesg or journalctl -k. For example:

Out of memory
Killed process 1234 (python)

This is another very common Linux interview and troubleshooting topic.

Kernel Memory

Not all RAM belongs to applications

The Linux kernel itself needs memory. For example: kernel objects, filesystem metadata, network structures, process structures, and device structures.

This is where concepts such as the slab allocator come in. And this connects directly with the command slabtop.

So you can think of memory broadly as:

Physical RAM
├── Application Memory
├── Page Cache
├── Kernel Memory
└── Other System Usage
The simple picture

Keep this memory-management path in mind

Application Virtual Memory Pages Page Tables Linux Memory Manager Physical RAM If memory pressure increases Reclaim Cache / Swap If memory still unavailable OOM Killer
Linux gives every process its own virtual address space and maps that virtual memory to physical RAM using pages and page tables, while the kernel continuously manages allocation, caching, reclaim, swap, and memory protection.

The natural next step is to go into free, vmstat, /proc/meminfo, RSS vs VSZ, page cache, swap, major/minor page faults, OOM Killer, and slabtop from a troubleshooting perspective.

Connect it to GenAI

Memory is one of the biggest bottlenecks in AI systems

Linux memory management becomes much more interesting when we connect it to GenAI, because memory is one of the biggest bottlenecks in AI systems.

For GenAI, we are usually dealing with two different memory worlds: CPU memory (RAM) and GPU memory (VRAM). The Linux memory manager primarily manages the CPU-side virtual memory and physical RAM, while GPU memory is managed through the GPU driver and frameworks such as CUDA. But the two are tightly connected because model weights, input data, tokenization, networking buffers, pinned memory, and GPU transfers all start or interact with the CPU side.

GenAI Application Python / PyTorch / vLLM CPU Virtual Memory Linux Memory Manager Physical RAM CUDA / GPU Driver GPU VRAM

For example, when you load a model with PyTorch:

model = AutoModelForCausalLM.from_pretrained(...)

the model may initially be read from storage into RAM. Linux uses the page cache while reading those model files. Then, when you move the model to the GPU:

model.to("cuda")

the data must travel from CPU memory to GPU memory.

Model on Disk Linux Page Cache CPU RAM CUDA Driver GPU VRAM

This is why an AI server may need a large amount of system RAM even when most of the actual inference happens on GPUs.

Pinned memory

Some RAM is locked in place for faster GPU transfers

Normally, application memory can be moved or reclaimed by the operating system. But GPU frameworks sometimes use page-locked or pinned host memory so that transfers between CPU RAM and GPU VRAM can happen efficiently.

Normal RAM

CPU Memory

Pinned Memory

Faster DMA transfer GPU VRAM

In PyTorch, you may encounter pin_memory=True when using a DataLoader. This matters because excessive pinned memory can also put pressure on system RAM.

Virtual memory and workers

Each process sees its own address space

Suppose you are running multiple inference workers: vLLM Worker 1, vLLM Worker 2, Tokenizer, API Server, Monitoring Agent. Each process sees its own virtual address space, even though all of them ultimately share the same physical RAM.

vLLM Process → Virtual Memory A Tokenizer → Virtual Memory B
Linux Memory Manager Physical RAM
This isolation is extremely important because one AI worker should not simply be able to overwrite another worker's memory.
Page faults in inference

A few faults are normal. A storm of major faults is not.

Imagine an inference process accesses part of a memory-mapped model file that has not yet been loaded into RAM.

Inference Process Accesses Model Page Page not currently in RAM Page Fault Linux loads page Process continues

A few page faults are normal. But a large number of major page faults can hurt latency because Linux has to retrieve data from storage.

For an inference system, that could result in:

Request Major Page Fault Disk access Higher latency

This is especially interesting when discussing cold starts for large models.

Page cache and model loading

The second load can be much faster

Suppose a 20 GB model is stored on NVMe. The first time you load it:

Model File NVMe Page Cache Application

That may take time. If the model is loaded again and those pages are still cached:

Application Page Cache No need to read everything again from disk

The second load can therefore be significantly faster. This is one reason why Linux showing very little free memory on an AI server does not automatically mean there is a problem. Linux may be using memory productively as cache.

free -h

You might see:

              total   used   free   buff/cache   available
Mem:            128G    60G     5G       63G          65G

Someone new to Linux might say: “Only 5 GB is free! We are running out of memory.” But Linux may be able to reclaim a large portion of the cache. That is why available is usually much more useful than simply looking at free.

Swap and AI

Swap is particularly dangerous for AI workloads

Imagine an inference process needs 40 GB RAM but the system is under heavy memory pressure. Linux may start moving inactive pages to swap. Later, the process needs those pages again. This can be painfully slow compared with RAM.

For a latency-sensitive inference service:

Memory pressure Heavy swapping Major page faults Disk activity Inference latency increases

So when debugging a slow GenAI server, checking memory and swap activity is important:

free -h
vmstat 1

Pay particular attention to:

  • si = swap in
  • so = swap out

If these continuously increase, the system is actively swapping.

OOM and AI

The process disappeared. Did it crash, or was it killed?

GenAI applications can consume enormous amounts of memory. Imagine:

Consumer Memory
System RAM 64 GB
vLLM 30 GB
Model loading 20 GB
Tokenizer 4 GB
Page Cache 15 GB
Other processes 8 GB

Eventually the system may come under severe memory pressure. If Linux cannot reclaim enough memory, the OOM Killer may terminate a process.

You might suddenly find that the vLLM process disappeared, and initially assume the application crashed. But checking dmesg or journalctl -k might reveal:

Out of memory: Killed process 12345 (python)
GenAI workload RAM exhausted Memory reclaim fails OOM Killer Python / vLLM killed

This is a very realistic production troubleshooting scenario.

Then we have GPU memory

System RAM and GPU VRAM are different resources

System RAM

Linux Memory Manager

free -h

GPU VRAM

GPU Driver / CUDA

nvidia-smi

Why free does not show GPU memory

The free command shows system RAM, not GPU VRAM.

It gets its information from:

/proc/meminfo

So conceptually:

free /proc/meminfo Linux Memory Manager System RAM

A discrete NVIDIA GPU has its own separate memory:

CPU

CPU System RAM

GPU

GPU GPU VRAM

GPU VRAM is managed through the GPU driver/runtime, not as normal Linux system memory.

So:

free
→ System RAM

nvidia-smi
→ NVIDIA GPU VRAM

For example:

nvidia-smi

or:

nvidia-smi --query-gpu=memory.total,memory.used,memory.free --format=csv

A useful GenAI point is:

System RAM available
        ≠
GPU VRAM available

You can have plenty of free system RAM and still get a CUDA out-of-memory error if GPU VRAM is full. You could have 50 GB of CPU RAM available and 79 GB / 80 GB of GPU VRAM used, and your application may still fail with:

CUDA out of memory

because system RAM and GPU VRAM are different resources. That distinction is extremely important in GenAI troubleshooting.

Simple takeaway: free reports Linux system memory, while nvidia-smi reports NVIDIA GPU memory.
The overall picture

Even GPU work starts in Linux memory

GenAI Application Python / PyTorch / vLLM Virtual Memory Linux Memory Manager
CPU RAM Page Cache
Pinned Memory CUDA / GPU Driver GPU VRAM Model Inference
Even when the actual AI computation happens on the GPU, Linux memory management still plays a critical role in loading models, managing application memory, caching model files, handling CPU-to-GPU transfers, dealing with page faults, swap, and protecting the system from out-of-memory situations.
CPU memory to GPU memory

How a GenAI workload actually moves

When you start an AI application, the operating system and Python program are running on the CPU side first. The model file is usually stored on SSD/NVMe, so it must first be read into system memory, or at least made available through the operating system’s memory and I/O path.

A simple flow is:

Model on SSD CPU / System RAM CUDA / GPU Driver GPU VRAM GPU performs computation

Why does this happen? Because the GPU normally does not independently open your model file, start Python, manage the filesystem, and decide what to execute. The CPU runs the application, prepares the data, tells the GPU what work needs to be done, and transfers the required model weights and tensors into GPU memory.

For example:

model = load_model("model.bin")

At this point, the CPU-side application reads the model. Then:

model.to("cuda")

conceptually means: take this model from CPU-accessible memory and copy it into GPU memory. Once the model is in GPU VRAM, the GPU can perform the expensive operations such as matrix multiplication much faster.

So for GenAI inference:

User Request CPU receives request CPU tokenizes / prepares input Input transferred to GPU GPU performs model computation Result returned to CPU Application sends response

This is also why we need to monitor both:

free -h

for CPU/system RAM, and:

nvidia-smi

for GPU VRAM.

Before you go

If you remember only five things

1. Processes see virtual memory. The MMU and page tables map that onto physical RAM.

2. Pages, not bytes. Linux allocates, faults, caches, and swaps in page-sized chunks.

3. Available beats free. Page cache is often useful memory, not wasted memory.

4. Swap and OOM are last resorts. Heavy swap hurts latency. The OOM Killer is how Linux survives when reclaim fails.

5. RAM and VRAM are different. free -h and nvidia-smi answer different questions, and a GenAI failure can come from either side.