Week 3
Designing an AI Inference System
A simplified platform for serving large language models.
Click or press → to reveal each idea. Use the buttons on a slide to explore.
The difference
Same building blocks. A different clock.
Traditional API
Load balancer
Database · cache · queue
Milliseconds
LLM inference
Same building blocks
Expensive GPUs
Several seconds
That changes scaling, load balancing, caching, scheduling, and the metrics that matter.
The API
From the outside, it looks like a normal API
POST /v1/chat/completions
model
messages
temperature
max_tokens
↓
Generate
↓
Stream tokens back
Internally, this is not a traditional web request.
Two kinds of request
A traditional API finishes. An LLM keeps generating.
Traditional
User
↓
API Server
↓
Database
↓
Return JSON
LLM inference
Prompt
↓
Tokenization
↓
Prefill
↓
Token 1 → 2 → N
It generates the answer one token at a time. That single fact drives the architecture.
The lifecycle
Four stages
1 Tokenize
2 Prefill
3 Decode
4 Stream
"What is Kubernetes?"
↓
Tokenizer · CPU
↓
[2061, 374, 58644, 30]
Input tokens
↓
Transformer model
↓
First output token + KV cache
Token 1
↓
Token 2 depends on Token 1
↓
Token 3 depends on 1 + 2
↓
Token 4 depends on everything before it
Kubernetes
→
Kubernetes is
→
Kubernetes is a
→
…
What it must do
Functional requirements
Prompt in Text out
Chat API OpenAI-compatible
temperature top_p · max_tokens · stop
Many models Stream tokens
↓
Client disconnects
↓
Cancel the request · stop the GPU
If the user closes the browser after ten tokens, there is no reason to generate hundreds more.
How well
The metrics that matter
User sends request
↓ TTFT
First token
Token 2
→ 30ms
Token 3
→ 30ms
Token 4
Availability 99.99%
Idle GPUs Expensive
Users will wait for the rest if the first token is fast. Large gaps between tokens feel slow. Idle GPUs make the service expensive.
Two bottlenecks
Capacity and bandwidth
GPU memory capacity
Model weights
KV cache
Runtime state
Can it fit?
Memory bandwidth
Weights
↓
GPU
↓
Next token
How fast can the data move?
These two constraints affect almost everything else.
GPU memory
It is not just the weights
GPU memory
Model weights
KV cache
Runtime memory
Temporary buffers
Framework overhead
This matters as soon as many users send requests at once.
KV cache
The model’s temporary working memory
Prompt + generated tokens
↓
KV cache
↓
Next token
User 1 KV cache
User 2 KV cache
User 3 KV cache
User 10,000 KV cache
As concurrency grows, KV-cache memory becomes a capacity limit.
The shape
Gateway, router, workers
Internet
↓
API Gateway
↓
Request Router
Worker 1 GPUs
Worker 2 GPUs
Worker 3 GPUs
↓
Shared infrastructure
Object storage Model weights
Database Conversations
API Gateway
Protect the GPUs before a request enters
Client
↓
API Gateway
Auth
Validate
Rate limit
User tier
↓
Router / Scheduler
↓
GPU workers
Rate limits
One request is not one request
Request A
100 input tokens
Request B
10,000 input tokens
Enforce requests per minute and tokens per minute. Token consumption reflects the workload better than request count alone.
When every GPU is busy
Do not treat every caller the same
GPU 1
GPU 2
GPU 3
GPU 4
Enterprise Serve first
Paid Queue
Free Throttle / reject
A rejection can come back as HTTP 429, or HTTP 503 with Retry-After: 5, depending on why.
The gateway knows who you are. The router knows whether a GPU can take you.
Who owns what
Three questions, three components
Gateway
Router
Worker
Can this request enter?
Attaches priority. Does not pick the GPU.
Where should this run?
Model
Worker
KV room
Queue?
Prefix
Execute the model
Prefill
→
KV cache
→
Batch
→
Decode
→
Tokens
Request Router
Stateless, so it can be copied
Client
↓
API Gateway
↓
Request Router
Router 1
Router 2
Router 3
Worker 1 GPU
Worker 2 GPU
Worker 3 GPU
The gateway asks “is this allowed?” The router asks “where should this run?”
Model selection
Not every prompt needs the 70B
"What is Docker?"
↓
Small · 8B
Detailed root-cause analysis
↓
Large · 70B
Larger models need more GPU memory, compute, time, and cost.
Simpler requests can go to smaller models. More complex requests go to larger ones.
Worker selection
Round robin is the wrong default
Round robin
Request
↓
Any server
Capacity-aware
20% KV cache
85% KV cache
Unhealthy
↓
Best worker
Sending the next long request to the fullest worker makes the problem worse.
Two optimizations
Reuse the prefix. Drop the sick worker.
Same system prompt
↓
Router
↓
Worker A
↓
Reuse prefix cache
Health checks
Worker 1 · healthy
Worker 2 · healthy
Worker 3 · failed
↓
Remove from pool
Admission control
Full, full, and almost full
Incoming request
↓
Is capacity available?
Yes → send to worker
No → admission control
Enterprise · priority
Paid · queue
Free · throttle / reject
40% Good
82% Busy
95% Reject
When caches are highly utilized, queue or return an error. Do not let the system degrade silently.
Inside the worker
This is the stateful part
Inference worker
Model weights
KV cache
Continuous batching
GPU execution
A worker crash can interrupt in-flight generation. It does not erase conversation history, because that lives in the database. The KV cache is temporary.
More than one GPU
140 GB of model, 80 GB cards
Large model · 140 GB
├── Part 1 → GPU 0
├── Part 2 → GPU 1
├── Part 3 → GPU 2
└── Part 4 → GPU 3
Workers can use multiple H100 GPUs this way. Weights are downloaded when the worker starts, not on every request.
Prefill vs decode
Parallel prompt, then one token at a time
Prefill
All input tokens together
↓
First token + KV cache
Decode
Token 2
↓
Token 3
↓
Token 4
1 request
That idle time is why batching exists.
Continuous batching
The batch is not frozen
Fixed A · 50 B · 500 C · 100
That raises GPU utilization, throughput, and users served, and it lowers cost per token.
Paged attention
Don’t reserve one giant block per request
Req A
Req B
Free
Req C
Free
A1 B1 A2 C1 B2 A3
Each request takes the pages it needs. Memory use gets tighter.
Cancel and stream
Tokens go out as they are born
Worker
↓ token
Router
↓
Gateway
↓
Client
Client disconnect
↓
Gateway
↓
Router
↓
Stop generation
↓
Release KV cache
Otherwise the GPU keeps writing tokens nobody will read.
Two protocols
HTTPS outside. gRPC inside.
Client
↓ HTTPS + JSON
API Gateway
↓
Request Router
↓ gRPC + Protobuf
Inference worker
↓
GPU
Protocol Buffers are more efficient than JSON for token IDs, health, capacity, and cancellation.
What is durable
Weights in object storage. History in a database.
Object storage
↓
Worker startup
↓
Local storage
↓
GPU
Conversation history
↓
Database
PostgreSQL · DynamoDB
KV cache stays temporary
Autoscaling
Requests per second hides the real load
Request A · 10 tokens
Request B · 10,000 tokens
Both count as one request
GPU % lags
40%
Healthy
75%
Watch
80%
Scale out
90%
Admission control
Scale on KV-cache utilization. 40–60% healthy, 75% watch, 80% scale out, 90% admission control. Test the thresholds. Don’t treat them as universal.
Watch these
Normal metrics, plus inference metrics
Platform
CPU
Memory
Network
Disk
GPU %
GPU mem
Inference
TTFT
KV cache
tok/s
ITL
Queue
Batch
TTFT and KV-cache utilization are the two inference indicators to emphasize.
The system
Clients to GPUs
Remember
Don’t design it like a web app
API request
↓
Tokenization
↓
Prefill
↓
KV cache
↓
Decode
↓
Continuous batching
↓
Token streaming
The gateway protects the platform. The router places the request. The workers run the model.
Open the full lesson →