← Back to the lesson
1 / 1

Week 3

Designing an AI Inference System

A simplified platform for serving large language models.

Click or press → to reveal each idea. Use the buttons on a slide to explore.

The difference

Same building blocks. A different clock.

Traditional API Load balancer Database · cache · queue Milliseconds
LLM inference Same building blocks Expensive GPUs Several seconds

That changes scaling, load balancing, caching, scheduling, and the metrics that matter.

The API

From the outside, it looks like a normal API

POST /v1/chat/completions
model
messages
temperature
max_tokens
Generate Stream tokens back

Internally, this is not a traditional web request.

Two kinds of request

A traditional API finishes. An LLM keeps generating.

Traditional User API Server Database Return JSON
LLM inference Prompt Tokenization Prefill Token 1 → 2 → N

It generates the answer one token at a time. That single fact drives the architecture.

The lifecycle

Four stages

"What is Kubernetes?" Tokenizer · CPU [2061, 374, 58644, 30]
Input tokens Transformer model First output token + KV cache
Token 1 Token 2 depends on Token 1 Token 3 depends on 1 + 2 Token 4 depends on everything before it
Kubernetes Kubernetes is Kubernetes is a …

What it must do

Functional requirements

Prompt inText out
Chat APIOpenAI-compatible
temperaturetop_p · max_tokens · stop
Many modelsStream tokens
Client disconnects Cancel the request · stop the GPU

If the user closes the browser after ten tokens, there is no reason to generate hundreds more.

How well

The metrics that matter

User sends request First token
Token 2 Token 3 Token 4
Availability99.99%
Idle GPUsExpensive

Users will wait for the rest if the first token is fast. Large gaps between tokens feel slow. Idle GPUs make the service expensive.

Two bottlenecks

Capacity and bandwidth

GPU memory capacity Model weights KV cache Runtime state

Can it fit?

Memory bandwidth Weights GPU Next token

How fast can the data move?

These two constraints affect almost everything else.

GPU memory

It is not just the weights

GPU memory Model weights KV cache Runtime memory Temporary buffers Framework overhead

This matters as soon as many users send requests at once.

KV cache

The model’s temporary working memory

Prompt + generated tokens KV cache Next token
User 1KV cache
User 2KV cache
User 3KV cache
User 10,000KV cache

As concurrency grows, KV-cache memory becomes a capacity limit.

The shape

Gateway, router, workers

Internet API Gateway Request Router
Worker 1GPUs
Worker 2GPUs
Worker 3GPUs
Shared infrastructure
Object storageModel weights
DatabaseConversations

API Gateway

Protect the GPUs before a request enters

Client API Gateway
Auth
Validate
Rate limit
User tier
Router / Scheduler GPU workers

Rate limits

One request is not one request

Request A 100 input tokens
Request B 10,000 input tokens

Enforce requests per minute and tokens per minute. Token consumption reflects the workload better than request count alone.

When every GPU is busy

Do not treat every caller the same

GPU 1
GPU 2
GPU 3
GPU 4
EnterpriseServe first
PaidQueue
FreeThrottle / reject

A rejection can come back as HTTP 429, or HTTP 503 with Retry-After: 5, depending on why.

The gateway knows who you are. The router knows whether a GPU can take you.

Who owns what

Three questions, three components

Can this request enter?
Who
Allowed?
Tier

Attaches priority. Does not pick the GPU.

Where should this run?
Model
Worker
KV room
Queue?
Prefix
Execute the model
Prefill KV cache Batch Decode Tokens

Request Router

Stateless, so it can be copied

Client API Gateway Request Router
Router 1
Router 2
Router 3
Worker 1GPU
Worker 2GPU
Worker 3GPU

The gateway asks “is this allowed?” The router asks “where should this run?”

Model selection

Not every prompt needs the 70B

"What is Docker?" Small · 8B
Detailed root-cause analysis Large · 70B

Larger models need more GPU memory, compute, time, and cost.

Simpler requests can go to smaller models. More complex requests go to larger ones.

Worker selection

Round robin is the wrong default

Round robin Request Any server
Capacity-aware
20%KV cache
85%KV cache
Unhealthy
Best worker

Sending the next long request to the fullest worker makes the problem worse.

Two optimizations

Reuse the prefix. Drop the sick worker.

Same system prompt Router Worker A Reuse prefix cache
Health checks Worker 1 · healthy Worker 2 · healthy Worker 3 · failed Remove from pool

Admission control

Full, full, and almost full

Incoming request Is capacity available?
Yes → send to worker
No → admission control Enterprise · priority Paid · queue Free · throttle / reject
40%Good
82%Busy
95%Reject

When caches are highly utilized, queue or return an error. Do not let the system degrade silently.

Inside the worker

This is the stateful part

Inference worker Model weights KV cache Continuous batching GPU execution
GPU 0
GPU 1
GPU 2
GPU 3

A worker crash can interrupt in-flight generation. It does not erase conversation history, because that lives in the database. The KV cache is temporary.

More than one GPU

140 GB of model, 80 GB cards

Large model · 140 GB
├──Part 1 → GPU 0
├──Part 2 → GPU 1
├──Part 3 → GPU 2
└──Part 4 → GPU 3

Workers can use multiple H100 GPUs this way. Weights are downloaded when the worker starts, not on every request.

Prefill vs decode

Parallel prompt, then one token at a time

Prefill All input tokens together First token + KV cache
Decode Token 2 Token 3 Token 4
1 request

That idle time is why batching exists.

Continuous batching

The batch is not frozen

FixedA · 50B · 500C · 100
Step 3BC

A finishes

Step 4DBC

D joins

Step 6DBE

E joins

That raises GPU utilization, throughput, and users served, and it lowers cost per token.

Paged attention

Don’t reserve one giant block per request

Req A Req B Free Req C Free
A1B1A2C1B2A3

Each request takes the pages it needs. Memory use gets tighter.

Cancel and stream

Tokens go out as they are born

Worker Router Gateway Client
Client disconnect Gateway Router Stop generation Release KV cache

Otherwise the GPU keeps writing tokens nobody will read.

Two protocols

HTTPS outside. gRPC inside.

Client API Gateway Request Router Inference worker GPU

Protocol Buffers are more efficient than JSON for token IDs, health, capacity, and cancellation.

What is durable

Weights in object storage. History in a database.

Object storage Worker startup Local storage GPU
Conversation history Database PostgreSQL · DynamoDB

KV cache stays temporary

Autoscaling

Requests per second hides the real load

Request A · 10 tokens Request B · 10,000 tokens

Both count as one request

GPU % lags
40%

Healthy

75%

Watch

80%

Scale out

90%

Admission control

Scale on KV-cache utilization. 40–60% healthy, 75% watch, 80% scale out, 90% admission control. Test the thresholds. Don’t treat them as universal.

Watch these

Normal metrics, plus inference metrics

Platform
CPU
Memory
Network
Disk
GPU %
GPU mem
Inference
TTFT
KV cache
tok/s
ITL
Queue
Batch

TTFT and KV-cache utilization are the two inference indicators to emphasize.

The system

Clients to GPUs

Inference serving poster: clients, API gateway, request router, three GPU workers, and shared infrastructure.

Remember

Don’t design it like a web app

API request Tokenization Prefill KV cache Decode Continuous batching Token streaming

The gateway protects the platform. The router places the request. The workers run the model.

Open the full lesson →

Click or → to reveal