Designing an AI Inference System

A simplified AI inference platform capable of serving Large Language Models.

Open slides →

Start with how this differs from a normal API ↓

When we normally think about system design, we think about things like:

  • API servers
  • Load balancers
  • Databases
  • Caches
  • Message queues
  • Horizontal scaling

AI inference systems use many of the same building blocks.

But there is one major difference:

Instead of primarily using CPUs to process short-lived requests, an AI inference system may spend several seconds generating a response on expensive GPUs.

That changes how we think about scaling, load balancing, caching, scheduling, and even performance metrics.

In this guide, we will design a simplified AI inference platform capable of serving Large Language Models.

1. What are we trying to build?

An API that streams a generated answer

Imagine that we want to provide an API like this:

POST /v1/chat/completions

The client sends:

{
  "model": "my-large-model",
  "messages": [
    {
      "role": "user",
      "content": "Explain Kubernetes in simple terms."
    }
  ],
  "temperature": 0.7,
  "max_tokens": 500
}

Our system should generate the response and stream it back to the user.

From the outside, this looks like a normal API.

Internally, however, the request is very different from a traditional web request.

2. Traditional API vs. AI inference API

The computation does not disappear when the request ends

Consider a traditional application.

A request might look like:

User API Server Database Return JSON
A traditional request does its work, returns JSON, and the computation is done.

The application receives the request, performs some computation or database operations, and sends the response.

Once the request finishes, most of the computation associated with it disappears.

LLM inference behaves differently.

Prompt Tokenization Prefill Generate Token 1 Generate Token 2 Generate Token 3 … Generate Token N
The model generates the answer one token at a time.

The model does not retrieve a pre-existing answer.

It generates the answer one token at a time.

That single fact drives much of our architecture.

3. Understanding the request lifecycle

Four stages from prompt to streamed answer

Before designing the architecture, we need to understand what happens when someone sends a prompt.

There are four important stages.

3.1 Step 1 — Tokenization

Suppose the user writes:

What is Kubernetes?

The model does not directly understand words.

The text is first converted into numerical token IDs.

Conceptually:

"What is Kubernetes?" Tokenizer [2061, 374, 58644, 30]
Words become token IDs before the model ever runs.

The exact IDs depend on the tokenizer used by the model.

Tokenization normally happens on the CPU and is relatively inexpensive compared with running the model itself.

3.2 Step 2 — Prefill

Now the input tokens are sent to the model.

This stage is called prefill.

Input Tokens Transformer Model First output token KV Cache
Prefill produces the first output token and the KV cache.

During prefill, the model processes the prompt and calculates the information it needs to start generating the response.

Prefill is generally compute intensive because GPUs perform large matrix multiplication operations across the input.

Another very important thing is created during this stage:

KV Cache

We will use this later.

3.3 Step 3 — Decode

After prefill, the system starts generating the response.

Suppose the model wants to generate:

Kubernetes is a container orchestration platform.

It cannot normally generate the entire sentence in one operation.

Instead:

Generate: Kubernetes Generate: is Generate: a Generate: container Generate: orchestration Generate: platform
Each word of the answer is its own generation step.

Each new token depends on the tokens that came before it.

For example:

Token 1 Token 2 depends on Token 1 Token 3 depends on Token 1 + Token 2 Token 4 depends on Token 1 + Token 2 + Token 3
Autoregressive decoding: each token depends on everything before it.

This is called autoregressive decoding.

This is one of the most important concepts to understand when designing an LLM inference system.

3.4 Step 4 — Stream the response

Imagine the complete response takes seven seconds to generate.

One option would be:

0 sec — user sees nothing 7 sec — full answer
Kubernetes Kubernetes is Kubernetes is a Kubernetes is a container …
Waiting for the full answer, compared with streaming each token as it is ready.

That creates a poor user experience.

Instead, AI applications normally stream tokens as soon as they become available.

The client sees the answer being generated.

This is why applications such as AI chat interfaces appear to "type" the answer.

Streaming dramatically improves perceived latency, even though the total generation time might remain the same.

4. Functional requirements

What the service needs to do

Now that we understand inference, we can define what our service needs to do.

Our system should:

  • Accept a prompt and generate text.
  • Support an OpenAI-compatible chat-completion style API.
  • Allow inference parameters such as: temperature, top_p, max_tokens, stop sequences.
  • Support multiple models.
  • Stream generated tokens to the client.
  • Allow requests to be cancelled.
  • Stop GPU computation when the client disconnects.

Cancellation is particularly important.

If a user closes the browser after receiving ten tokens, there is normally no reason to continue generating hundreds of additional tokens.

Continuing would waste GPU resources.

5. Non-functional requirements

How well the system should work

Now we define how well the system should work.

For an inference service, I would pay attention to the following metrics.

Time to First Token — TTFT

This measures:

User sends request First token reaches user
TTFT is the wait until the first token arrives.

TTFT is extremely important.

A user may tolerate the response taking several seconds to finish if the first token appears quickly.

Inter-Token Latency

After the first token appears, we want subsequent tokens to arrive smoothly.

For example:

Token 1 Token 2 Token 3 Token 4
Smooth gaps between tokens. Large gaps make the application feel slow.

Large delays between tokens make the application feel slow.

GPU Utilization

GPUs are expensive.

If we have a fleet of GPUs sitting mostly idle, our inference service becomes unnecessarily expensive.

So we want the scheduler and batching engine to keep GPUs busy.

Availability

Our API should remain available even if:

  • a router crashes
  • an inference worker crashes
  • a GPU fails
  • a node is replaced

A possible target could be:

99.99% availability

depending on the business requirements.

Horizontal Scalability

If traffic increases:

1,000 requests 10,000 requests 100,000 requests
Growth should be met by adding inference workers.

we should be able to add more inference workers.

6. The two important bottlenecks

Memory capacity and memory bandwidth

Two things are particularly important when designing LLM inference.

GPU memory capacity. Can the model and its runtime state fit into GPU memory?

Memory bandwidth. How quickly can the GPU move the model data required for generating tokens?

These two constraints affect almost everything else in the architecture.

7. GPU memory is not just model weights

Weights are only one slice of the budget

A common beginner mistake is thinking:

GPU Memory Required = Model Weights

It is more like:

GPU Memory Model Weights KV Cache Runtime Memory Temporary Buffers Other Framework Overhead
GPU memory is the sum of these pieces, not the weights alone.

This becomes very important when many users send requests simultaneously.

8. What is the KV cache?

The model's temporary working memory

During generation, the model repeatedly needs information from the tokens it has already processed.

Recalculating everything from scratch for every generated token would be extremely inefficient.

Instead, the system stores intermediate attention information in memory.

That storage is called the:

Key-Value Cache, or KV Cache.

Think of it as the model's temporary working memory for the current request.

Conceptually:

Request A Prompt + Generated Tokens KV Cache Generate next token
The KV cache holds what this request has already processed.

Every active request consumes some KV-cache memory.

So imagine:

User 1KV Cache
User 2KV Cache
User 3KV Cache
User 4KV Cache
User 10,000KV Cache
Each active user holds a KV cache. Concurrency makes that memory a capacity limit.

As concurrency increases, KV-cache memory can become one of the most important capacity constraints in the system.

9. High-level architecture

Gateway, router, workers, shared infrastructure

Now we can design the system.

Internet API Gateway Request Router
Inference Worker 1GPUs
Inference Worker 2GPUs
Inference Worker 3GPUs
Shared Infrastructure
Object StorageModel Weights
DatabaseConversations
From the internet, through the gateway and router, onto GPU workers, with shared storage beside them.

Let's walk through each component.

10. API Gateway

Protect the expensive GPU infrastructure

The API Gateway is the entry point into our AI inference platform.

Its main goal is simple:

Protect the expensive GPU infrastructure before a request enters the inference system.

Client API Gateway
Auth
Validate
Rate Limit
User Tier
Router / Scheduler GPU Workers
The gateway checks the request before it can reach a GPU.

The API Gateway should mainly handle:

  • Authentication
  • Authorization
  • Request validation
  • Prompt/input-size limits
  • Requests-per-minute limits
  • Tokens-per-minute limits
  • Identifying the customer tier
  • Streaming connections
  • Detecting client disconnects

For example, the gateway may attach information like:

customer_id = 1234
tier        = enterprise
model       = large-model
max_tokens  = 1000

The request is then forwarded to the inference router.

Why token-based rate limiting matters

For a traditional API, we might limit:

100 requests/minute

But this is not enough for an LLM API.

Consider:

Request A 100 input tokens
Request B 10,000 input tokens
Both are one HTTP request. Request B costs far more inference.

Both are one HTTP request, but Request B consumes significantly more inference resources.

Therefore we may enforce both:

Requests Per Minute Tokens Per Minute

Token-aware rate limiting matters because token consumption better reflects AI inference workload than request count alone.

What happens if all GPUs are busy?

This is an important interview question.

Suppose our GPU cluster is almost full:

GPU 1
GPU 2
GPU 3
GPU 4
Every GPU in the cluster is busy.

At this point, we should not treat every request equally.

For example, our policy could be:

  • Priority 1 → Enterprise users
  • Priority 2 → Paid users
  • Priority 3 → Free users

If there is not enough capacity:

  • Enterprise → Queue / Serve first
  • Paid → Queue if capacity permits
  • Free → Throttle or reject temporarily

A simplified flow looks like this:

Request API Gateway Identify Tier Router / Scheduler — is GPU capacity free?
YES → Send to GPU
NO → Priority Queue Enterprise — serve first Paid — queue Free — throttle / reject
When capacity is gone, tier decides who is served, queued, or turned away.

Does this happen at the API Gateway?

Not primarily.

The API Gateway knows:

  • Who is the customer?
  • What tier are they on?
  • Are they allowed to make this request?

But the Router/Scheduler knows:

  • Which GPUs are busy?
  • How much KV cache is available?
  • How many requests are queued?
  • Which workers have capacity?
  • Which request should run next?

So the responsibility should be separated like this:

API Gateway "This is an Enterprise customer." Router / Scheduler "GPUs are busy. Enterprise gets priority." GPU Worker
The gateway names the tier. The router decides what that tier gets.

The API Gateway provides the priority information.

The Router/Scheduler makes the scheduling decision.

What happens to free users?

If the system is overloaded, the scheduler may decide not to admit free-tier traffic.

For example:

GPU Capacity 95% Full Admission Control Enterprise → Accept Paid → Queue Free → Reject
At 95% full, admission control treats the three tiers differently.

The rejection can then be returned through the API Gateway as something like:

HTTP 429 Too Many Requests

or:

HTTP 503 Service Unavailable
Retry-After: 5

depending on why the request was rejected.

Important interview point

The API Gateway identifies and authenticates the customer and attaches the service tier to the request. The inference Router/Scheduler owns GPU-aware admission control and prioritization. If GPUs are saturated, Enterprise traffic can be prioritized, paid traffic queued next, and free traffic throttled or rejected.

This keeps responsibilities clean:

API Gateway Who are you? Are you allowed? What tier are you?
Router / Scheduler Where should you run? Is there capacity? What priority should you get?
Inference Worker Run the model on the GPU
11. Request Router

Which inference worker should handle this request?

Once the API Gateway has authenticated and validated the request, the next question is:

Which inference worker should handle this request?

That is the job of the Request Router.

Client API Gateway Request Router
Worker 1GPU
Worker 2GPU
Worker 3GPU
The router chooses a worker after the gateway has let the request in.

The router should generally be stateless.

That means it does not permanently store conversation history or inference state.

Because it is stateless, we can easily run multiple router replicas:

Router 1
Router 2
Router 3
Several identical routers. If one fails, another keeps accepting requests.

If one router fails, another can continue accepting new requests.

The router stays stateless so it can be horizontally replicated.

What does the router actually do?

The router mainly makes four decisions:

  1. Which model should serve the request?
  2. Which inference worker should receive it?
  3. Does the system currently have enough capacity?
  4. Should the request be accepted, queued, or rejected?

So the API Gateway asks:

"Is this request allowed?"

The Router asks:

"Where should this request run?"

11.1 Model selection

Suppose our platform supports multiple models:

  • Small Model → 8B
  • Medium Model → 30B
  • Large Model → 70B

Not every request needs the largest model.

For example:

"What is Docker?"

while a more complex request might go to:

"Analyze this distributed-system failure and produce a detailed root-cause analysis."

"What is Docker?" Small Model
"Analyze this distributed-system failure and produce a detailed root-cause analysis." Large Model
"What is Docker?" goes to a small model. The longer analysis goes to a large model.

So the router can perform:

Incoming Request Model Selection
Simple Request Small Model
Complex Request Large Model

This matters because larger models usually require more:

  • GPU memory
  • GPU compute
  • Inference time
  • Cost

Simpler requests can be routed to smaller models and more complex requests to larger ones.

11.2 Worker selection

Once the model has been selected, the router needs to decide:

Which worker running that model should receive the request?

Suppose we have:

  • Worker 1 → GPU utilization high
  • Worker 2 → KV cache 50%
  • Worker 3 → KV cache 80%
  • Worker 4 → unhealthy

The router should not randomly send traffic.

It needs to consider worker state.

Router Check Worker State
Worker 1Busy
Worker 2Healthy
Worker 3Busy
Select Worker 2
The healthy worker with room is the one that receives the request.

Useful information could include:

  • Worker health
  • Available KV-cache capacity
  • Current queue length
  • Active requests
  • Model loaded on worker
  • Prefix-cache availability

11.3 Why not just use round robin?

For a traditional stateless web application, we might simply use:

Round Robin

For example:

  • Request 1 → Server A
  • Request 2 → Server B
  • Request 3 → Server C

But AI inference workers are not equal at every moment.

Consider:

  • Worker A — KV Cache = 20%
  • Worker B — KV Cache = 85%
  • Worker C — KV Cache = 45%

Sending the next request blindly to Worker B may make things worse.

So AI inference routing needs to be capacity aware.

Traditional load balancing Request Round Robin Any Server
AI inference routing Request Router Model? Capacity? KV Cache? Queue? Prefix Cache? Best Worker
Round robin picks any server. Inference routing picks the worker that can actually take the work.

11.4 Prefix-aware routing

This is an important AI-specific optimization.

Suppose thousands of requests use the same system prompt:

"You are an expert DevOps assistant..."

Without prefix-aware routing:

  • Request 1 → Worker A
  • Request 2 → Worker B
  • Request 3 → Worker C

Each worker may process the same prefix independently.

Instead, if Worker A already has that prefix cached:

Same System Prompt Router Worker A Reuse Prefix Cache
Similar prompts go to the worker that already cached that prefix.

The router can try to send similar requests to that worker.

This can reduce repeated prefill work and improve efficiency.

Prefix-aware load balancing lets requests sharing the same system prompt reuse cached work on the same worker.

11.5 Health checks

The router should continuously know whether inference workers are healthy.

For example:

Router
├──Health Check → Worker 1 → Healthy
├──Health Check → Worker 2 → Healthy
└──Health Check → Worker 3 → Failed
A failed worker is removed from the routing pool.

If Worker 3 repeatedly fails health checks:

Worker 3 Remove from routing pool

New requests should no longer be sent there.

Health checks can remove workers after repeated failures.

The exact health-check frequency or failure threshold is an implementation choice.

11.6 What happens if GPUs are busy?

This is where the router becomes very important.

Suppose:

  • Worker 1 → Full
  • Worker 2 → Full
  • Worker 3 → Almost Full

The router now has to perform:

Admission Control

Incoming Request Request Router Is capacity available?
YES → Send to Worker
NO → Apply Policy

If capacity is unavailable, the router may:

Queue the request

OR

Reject the request

This is also where the priority policy we discussed earlier fits.

GPUs Busy Admission Control
EnterprisePriority queue
PaidQueue
FreeThrottle / Reject

So remember:

API Gateway Router / Scheduler Inference Worker

11.7 Why KV cache matters to the router

Suppose:

  • Worker A — KV Cache = 40%
  • Worker B — KV Cache = 82%
  • Worker C — KV Cache = 95%

Sending another long-context request to Worker C might cause memory pressure.

So the router can use KV-cache utilization as one of its capacity signals.

Router Check KV Cache
40%Good
82%Busy
95%Reject
Worker A
The worker with KV cache at 40% is the one that can take another long request.

When workers' KV caches become highly utilized, the router can queue new requests or return an error instead of allowing the system to degrade silently.

11.8 What happens if a router fails?

Because the router is stateless:

Load Balancer
Router 1Failed
Router 2New request
A failed router does not stop new requests. Another replica takes them.

another router can continue handling new requests.

This is one reason we deliberately avoid storing important inference state inside the router.

11.9 Router vs inference worker

This separation is important in interviews.

What the request router decides compared with what the inference worker runs
Request RouterInference Worker
Which model?Prefill
Which worker?KV Cache
Is there capacity?Continuous Batching
Should we queue?Decode
What priority?GPU Execution
Can we reuse a prefix?Token Generation

Think of it like this:

Router "You should run on Worker 3." Worker 3 "I will execute the model." GPU

11.10 Complete router flow

A request reaches the router.

Request Request Router Which Model? Find Healthy Workers Check GPU Capacity Check KV Cache Check Prefix Locality Capacity Available?
YES → Select Worker
NO → Admission Control Enterprise — serve Paid — queue Free — throttle / reject
Model, health, capacity, KV cache, and prefix locality, then admit or hold the request.
Important interview point

The Request Router is a stateless service that decides where an inference request should execute. It selects the model and worker based on worker health, available capacity, KV-cache utilization, and potentially prefix-cache locality. If the GPU cluster is saturated, the router performs admission control and applies traffic priority policies such as prioritizing Enterprise users over paid and free users.

The most important distinction is:

  • API Gateway — Can this request enter the system?
  • Request Router — Where should this request run?
  • Inference Worker — Execute the model on the GPU

That is usually enough depth for an interview without turning the router discussion

12. Inference Workers

This is where the model actually runs

Once the Request Router selects the best destination, the request is sent to an Inference Worker.

This is where the actual LLM execution happens.

Request Router Inference Worker
Model
KV Cache
Batching
GPU Execution Generated Tokens
The worker holds the model, the cache, and the batch, then runs them on the GPU.

The Router decides:

Where should the request run?

The Inference Worker does:

Actually run the model and generate the response.

Inference workers are stateful workers where the model lives, containing continuous batching, KV-cache management, and the GPU execution engine.

What lives inside an inference worker?

A simplified worker looks like this:

Inference Worker Model Weights KV Cache Continuous Batching Engine GPU Execution Engine
GPU 0
GPU 1
GPU 2
GPU 3

The important parts are:

  1. Model weights
  2. KV cache
  3. Continuous batching
  4. GPU execution
  5. Token streaming

12.1 Model weights

Before the worker can serve requests, the model needs to be loaded into GPU memory.

For example:

Object Storage Model Files Inference Worker GPU Memory
Weights are loaded at startup, not downloaded again for every request.

The model should not be downloaded from object storage for every request.

Instead, the worker loads the model during startup or deployment and keeps it available for inference.

Model weights live in object storage and are downloaded when the cluster or worker is initialized rather than per request.

12.2 Why multiple GPUs?

Large models may not fit on a single GPU.

Suppose:

Model requires 140 GB

but each GPU only has:

80 GB VRAM

We may need multiple GPUs.

Model
GPU 0
GPU 1
GPU 2

One common technique is:

Tensor Parallelism

The model is divided across multiple GPUs, and those GPUs cooperate during each forward pass.

Conceptually:

Large Model
├──Part 1 → GPU 0
├──Part 2 → GPU 1
├──Part 3 → GPU 2
└──Part 4 → GPU 3
Tensor parallelism splits one model across GPUs that work together on each forward pass.

Workers can use multiple H100 GPUs through tensor parallelism.

12.3 Prefill happens inside the worker

Remember the inference lifecycle:

Prompt Tokenization Prefill Decode

During prefill, the model processes all input tokens.

For example:

"Explain Kubernetes in simple terms" Prefill Phase First Generated Token KV Cache

Prefill is generally more compute-intensive because many input tokens can be processed in parallel.

Prefill processes the input tokens together and produces both the first generated token and the KV cache.

12.4 KV cache

The KV cache is one of the most important things maintained by the inference worker.

Suppose the model has already processed:

"Kubernetes is a container"

When generating the next token, we do not want to recompute everything from scratch.

Instead, intermediate attention information from previous tokens is stored in the:

KV Cache

Request A KV Cache
├──Token 1 state
├──Token 2 state
├──Token 3 state
└──Token 4 state
Generate next token

Each active request has its own KV-cache state.

So with many concurrent users:

  • User 1 → KV Cache
  • User 2 → KV Cache
  • User 3 → KV Cache
  • User 4 → KV Cache

This is why the inference worker is considered stateful while requests are active.

KV cache can consume significant GPU memory during high-traffic workloads.

12.5 Why the worker is stateful

This is an important interview distinction.

The Router is stateless.

The Inference Worker is stateful.

Why?

Because while generation is happening, the worker holds:

  • KV Cache
  • Active sequences
  • Batch state
  • Model execution state

For example:

Worker 1
├──Request A KV Cache
├──Request B KV Cache
└──Request C KV Cache
Active requests live in the worker. A crash drops that temporary state.

If the worker crashes, those active requests may fail because the temporary inference state disappears.

But persistent conversation history should still live in a database, not inside the worker.

Router failure does not lose persistent data, while worker failure can interrupt inflight requests; KV cache is temporary state.

12.6 Decode phase

After prefill, the worker enters the decode phase.

This is where output tokens are generated.

Prompt Prefill Token 1 Token 2 Token 3 Token 4

The important point is:

Output tokens are generated sequentially.

The model cannot normally generate token 10 before token 9 exists.

Token 1 Token 2 depends on Token 1 Token 3 depends on Token 1 + Token 2 Token 4 depends on everything before it

This is why decoding behaves very differently from the prefill phase.

12.7 Why decode can underutilize the GPU

Imagine only one user is generating tokens.

For every token, the GPU performs another model forward pass.

But a single request may not provide enough parallel work to fully utilize a large GPU.

Conceptually:

GPU
1 req
A huge GPU, and one request that only fills a small slice of it.

A lot of the GPU's compute capability may not be efficiently used.

That leads us to one of the most important techniques in inference serving:

Batching

12.8 Traditional batching

Suppose we receive:

Request A, Request B, Request C

We could group them together:

Batch Request A Request B Request C GPU
Grouping requests lets the GPU do more work in parallel.

Processing requests together allows the GPU to perform more work in parallel.

That improves GPU utilization.

But traditional batching has a problem.

Suppose:

  • Request A → 50 output tokens
  • Request B → 500 output tokens
  • Request C → 100 output tokens

Request A finishes quickly.

Request C finishes shortly afterward.

But Request B continues for much longer.

If the batch is fixed, we may waste capacity waiting for the longest request.

12.9 Continuous batching

Modern inference systems solve this using:

Continuous Batching

Instead of keeping the batch fixed, requests can dynamically enter and leave.

Step 1ABC
Step 2ABC
Step 3BC

A finishes

Step 4DBC

D arrives

Step 5DB

C finishes

Step 6DBE

E arrives

Finished requests leave. New requests take the open slots.

The worker continuously fills available batch slots.

This keeps the GPU much busier.

Continuous batching is a tight loop running during each decode step inside the inference worker.

12.10 Why continuous batching matters

Without continuous batching:

GPU
Part of the GPU may sit idle.

Part of the GPU may sit idle.

With continuous batching:

GPU
The GPU stays on useful work.

We try to keep the GPU doing useful work.

This directly impacts:

  • GPU utilization
  • Throughput
  • Cost per token
  • Number of users served

So if an interviewer asks:

How do you improve GPU utilization in an inference system?

One strong answer is:

Use continuous batching so that new requests can join the batch when other requests finish instead of waiting for the entire batch to complete.

12.11 Paged KV cache / paged attention

Another challenge is managing KV-cache memory efficiently.

Imagine GPU memory like this:

Req A Req B Free Req C Free
Different sequence lengths leave gaps. Memory fragments.

Requests have different sequence lengths, so their KV-cache requirements are different.

Memory can become fragmented.

The KV-cache pool can be managed using paged attention.

The beginner-friendly idea is:

Instead of requiring one large continuous memory block for every request, divide KV-cache memory into smaller pages or blocks.

Conceptually:

A1B1A2C1B2A3
Each request uses the pages it needs instead of one unbroken block.

Each request can use the pages it needs.

This makes GPU memory usage more efficient.

12.12 Request cancellation

Suppose a user closes the browser while the model is still generating.

User closes browser API Gateway Router Inference Worker Request A cancelled Remove from batch Release KV Cache
A disconnect should stop generation and free the cache.

The worker should stop generating tokens.

Otherwise, the GPU might continue generating hundreds of tokens that nobody will consume.

Generation should stop when the client disconnects so GPU resources can be freed.

12.13 Streaming tokens back

The worker generates:

Token 1, Token 2, Token 3, Token 4

and streams them back through the system.

Inference Worker Router API Gateway Client

We do not want the worker to wait for the entire answer before returning anything.

Instead:

Generate Token Send Token Generate Next Token Send Token

This is what produces the familiar typewriter-like response in AI chat applications.

12.14 What happens when the worker is full?

Suppose:

  • Worker 1 KV Cache = 95%
  • Worker 2 KV Cache = 92%
  • Worker 3 KV Cache = 97%

The worker should not blindly accept more work.

Instead, it exposes its capacity information to the Router.

Inference Worker
├──Health
├──KV Cache Usage
├──Active Requests
└──Capacity
Router

Then the Router can decide:

Send request, OR queue request, OR reject request.

This is why the responsibility is split:

  • Inference Worker — Reports current capacity
  • Router — Makes admission decision

12.15 What happens if an inference worker fails?

Suppose:

Worker 2 — GPU failure Router
Worker 1
Worker 3
New requests go to the workers that are still up. Work already on Worker 2 may fail.

The Router stops sending new requests to that worker.

But requests already running on Worker 2 may fail.

Why?

Because their temporary state lived there:

  • KV Cache
  • Active sequence state
  • Current generation state

The client may receive an error and potentially retry the request.

Persistent conversation history is still safe because it lives outside the worker.

12.16 Complete inference worker flow

Let's put everything together.

Request from Router Inference Worker Model Available? Prefill Create KV Cache Add Request to Batch Continuous Batch GPU Execution Generate Token
Update KV Cache
Stream Token
Generate Next Request Finished Release KV Cache
Prefill, cache, batch, decode, stream, then release the cache when the request ends.
Important interview point

An inference worker is the stateful part of the serving system where the model is loaded and actual GPU execution happens. It manages model weights, active KV caches, continuous batching, prefill, decode, and token streaming. Multiple requests are continuously batched together to improve GPU utilization, while KV-cache memory is managed carefully because it becomes a major capacity constraint under high concurrency.

The clean separation is:

  • API Gateway — Can this request enter?
  • Request Router — Where should it execute?
  • Inference Worker — Run the model
  • GPU — Perform the computation

And inside the worker, remember this flow:

Request Prefill KV Cache Continuous Batching Decode Stream Tokens Release KV Cache

That is usually the right depth for explaining Inference Workers in a DevOps, SRE, Platform, or system-design interview.

13. Communication between router and workers

HTTPS outside, gRPC inside

Externally, clients usually communicate with the platform using:

HTTPS + JSON

This is simple, widely supported, and easy for clients to consume.

Internally, however, the communication between the Request Router and Inference Workers should be more efficient.

A common choice is:

gRPC + Protocol Buffers

The communication flow looks like this:

Client API Gateway Request Router Inference Worker GPU
Clients speak HTTPS and JSON. The router and workers speak gRPC.

Why use gRPC internally?

The Router and Inference Workers communicate frequently and may exchange information such as:

  • Token IDs
  • Request metadata
  • Model information
  • Generated tokens
  • Cancellation messages
  • Worker health and capacity

Using gRPC internally gives us a few advantages:

  • Binary serialization is generally more compact than JSON.
  • It reduces serialization/deserialization overhead.
  • It supports strongly defined request and response formats.
  • It supports streaming, which is useful because LLM responses are generated token by token.
  • It works well for service-to-service communication inside the inference cluster.

gRPC is used between the Router and workers because Protocol Buffers are more efficient than JSON for exchanging token-related data.

Streaming tokens back

The worker does not wait until the complete response is generated.

As soon as a token is generated:

Inference Worker Router API Gateway Client

Then:

Token 2, Token 3, Token 4, … continue through the same path.

So the flow becomes:

Client API Gateway Router Inference Worker Router API Gateway Client

Cancellation also travels through the same path

Suppose the user closes the browser while generation is still happening.

The cancellation should travel downstream:

Client Disconnect API Gateway Router Inference Worker Stop Generation Release KV Cache
Cancellation follows the request path so the GPU stops and the cache is freed.

This avoids wasting GPU resources on a response that nobody is waiting for.

Important interview point

Externally, I would expose HTTPS/JSON because it is easy for clients to use. Internally, between the Router and Inference Workers, I would prefer gRPC with Protocol Buffers because it is more efficient for high-frequency service-to-service communication and supports streaming naturally.

So the key distinction is:

External API HTTPS + JSON

Easy for clients

Internal Communication gRPC + Protobuf

Efficient for inference traffic

This level is probably enough — it gives the interviewer the why, without going into HTTP/2 framing, protobuf encoding, or gRPC internals unless they specifically ask.

14. Where do model weights live?

Object storage, then the GPU at startup

We should not permanently keep the master copy of every model only on individual GPU servers.

Instead:

Object Storage Model Weights Worker Startup Local Storage GPU
The master copy lives in object storage. Workers load it when they start.

Object storage might contain model files.

Workers download or prepare the models during deployment/startup rather than downloading them for every inference request.

15. Where does conversation history live?

A database, not the KV cache

Remember:

The KV cache is temporary.

It should not be our permanent conversation database.

Instead:

Conversation History Database PostgreSQL DynamoDB etc.
KV Cache

Exists primarily to make active inference efficient.

History is durable. The KV cache exists only while inference is running.

Meanwhile:

KV Cache exists primarily to make active inference efficient.

If an inference worker crashes, its KV cache can disappear.

The persistent conversation data should remain elsewhere.

16. Stateless router, stateful worker

A distinction worth remembering

This distinction is worth remembering for interviews.

Why the router is stateless and the worker is stateful
RouterInference Worker
StateStatelessStateful during inference
WhyHorizontal scaling and recovery are easier.It holds model state, KV cache, active sequences, and batch information.
If it crashesAnother router can handle new requests.Active generation on that worker may fail.

But persistent conversation history should remain safe in the storage layer.

17. How should we autoscale?

Requests per second is not enough

This is where AI infrastructure differs from many traditional services.

The obvious answer might be:

Scale based on requests per second.

But consider:

Request A 10 tokens
Request B 10,000 tokens
Both count as one request. They do not cost the same.

Both count as:

1 request

But they require very different resources.

So:

Requests Per Second alone is not enough.

18. Should we scale on GPU utilization?

KV cache utilization is the earlier signal

GPU utilization is useful for monitoring, but using it alone as the primary scaling signal is a poor choice because it can behave as a lagging indicator.

By the time GPU utilization tells us there is a problem, requests may already be waiting.

Instead, one useful signal for an inference platform is:

KV Cache Utilization

For example:

40%

Healthy

60%

Healthy

75%

Watch

80%

Scale out

90%

Admission control

KV cache utilization, from healthy through scale-out to admission control.

The exact thresholds should come from testing rather than being treated as universal numbers.

19. Full request flow

From the question to the streamed answer

Now let's put everything together.

A user asks:

"Explain Docker in simple terms."

The complete request flow becomes:

USER API Gateway Authentication · Rate Limiting · Validation Router Select Model · Select Worker · Check Capacity Inference Worker Tokenization Prefill Create/Use KV Cache Decode Generate Token Streaming USER
One question, from the gateway through decode, streamed back to the user.

Meanwhile:

Model Weights Object Storage
Conversation History Database
Metrics Monitoring System
Weights, history, and metrics live beside the request path.
20. Observability

Infrastructure metrics, plus inference metrics

An AI inference system needs the normal infrastructure metrics:

  • CPU
  • Memory
  • Network
  • Disk
  • GPU Utilization
  • GPU Memory

But we also need inference-specific metrics.

Important ones include:

  • Time to First Token
  • Tokens per Second
  • Inter-Token Latency
  • Input Tokens
  • Output Tokens
  • Request Queue Length
  • KV Cache Utilization
  • Batch Size
  • Request Cancellation Rate
  • Model Loading Time
  • GPU Memory Utilization

Among these, TTFT and KV Cache Utilization stand out as important inference-specific indicators.

21. Final architecture

The simplified system

Our simplified architecture now looks like this:

Poster of the inference serving path: clients, API gateway, request router, three GPU workers, and shared infrastructure for object storage, database, and monitoring.
Clients, the gateway, the router, three workers, and the shared layer underneath.
22. The main idea to remember

This is not a traditional web application

The biggest mistake when designing an AI inference system is treating it exactly like a traditional web application.

A traditional web application might primarily worry about:

  • Requests per second
  • CPU utilization
  • Database connections
  • API latency

An AI inference system also needs to think about:

  • GPU memory
  • Model weights
  • KV cache
  • Memory bandwidth
  • Prefill
  • Decode
  • Continuous batching
  • TTFT
  • Tokens per second
  • Request cancellation

The important mental model is:

API Request Tokenization Prefill KV Cache Decode Continuous Batching Token Streaming

Once you understand this flow, the architecture starts making much more sense.

The API gateway protects the platform.

The router decides where requests should go.

The inference workers perform the actual model execution.

The KV cache keeps active inference efficient.

The batching engine helps us use GPUs efficiently.

The storage layer keeps persistent information separate from temporary inference state.

And the autoscaler adds GPU workers before the inference layer becomes overloaded.

That is the foundation of a modern AI inference system.