What Does the “G” in ChatGPT Mean?

The G stands for Generative, not “generation.” The natural way to explain Generative is to explain how generation actually happens: one token at a time.

Let’s start with what the model is actually producing ↓

In the previous part, we explored the “Chat” in ChatGPT and looked at what actually happens when we have a conversation with a Large Language Model.

Now let’s move to the next letter:

G → Generative

The word Generative means that the model can generate new content.

When you ask ChatGPT a question, it does not simply look inside a database, find a complete answer, and return it to you. Instead, it generates the answer while you are interacting with it.

That raises an interesting question:

How does an LLM actually generate text?

At first, it may look as if ChatGPT writes an entire sentence or paragraph in one go. That is not normally what happens.

An LLM generates its response one token at a time. Understanding this one idea makes many other LLM concepts much easier to understand.

Before generation

First, what is a token?

Before understanding generation, we need to understand what the model is actually generating.

An LLM does not directly think in terms of sentences or paragraphs. It works with tokens.

A token can be an entire word, part of a word, punctuation, or another small piece of text.

For example, a sentence such as:

Explain how a rainbow is formed.

is first converted into tokens before the model processes it.

Internally, those tokens are represented using numbers called token IDs. The tokenizer converts human-readable text into these tokens before the model works with them.

For a beginner, the exact token IDs are not important. The main idea is simply:

Diagram: human text is converted into tokens, then those tokens go into the LLM.
Human text → Tokens → LLM. The model never starts from a full paragraph. It starts from tokens.
The most important idea

One token at a time

Suppose we give an LLM this input:

The capital of France is

The model needs to decide what should come next. It may predict:

Paris

Now the sequence becomes:

The capital of France is Paris

The newly generated token becomes part of the input for the next prediction. The model then asks again: What token should come next?

It predicts another token. Then another. Then another.

Diagram of autoregressive generation: predict a token, add it to the sequence, then predict again.
Predict → Add → Predict → Add → Repeat. You can think of the model as continuously completing the text that has already been generated.

This is the core idea behind autoregressive text generation.

The interesting part

How does the model know which token comes next?

The model does not usually say, “I know with 100% certainty that the next word is blue.”

Instead, it considers many possible tokens. Suppose the current text is:

The sky is

The model might consider possibilities such as:

Example next-token probabilities after the phrase The sky is
Possible next tokenExample probability
blue
65%
clear
15%
cloudy
10%
beautiful
6%
green
1%
These numbers are only an illustration. A real model may have tens or hundreds of thousands of possible tokens in its vocabulary.

For every possible token, the model first produces a score. These raw scores are commonly called logits.

You do not need to worry about the mathematics behind logits yet. For now, think of them as:

The model’s raw scores for all the possible next tokens.

A higher score generally means that the model considers that token more suitable in the current context.

From scores to a choice

From scores to probabilities

Raw scores are difficult for humans to interpret. Conceptually, those scores can be converted into probabilities.

For example, raw scores such as blue → 8.2, clear → 6.5, cloudy → 5.9 may become something easier to understand:

blue → 65%
clear → 20%
cloudy → 10%
everything else → remaining probability

Again, the numbers are only examples.

The important concept is that the model creates a probability distribution over possible next tokens. One token is then selected from those possibilities.

Suppose the selected token is blue. Our sequence becomes:

The sky is blue

Now the model repeats the entire next-token prediction process. It asks: given “The sky is blue,” what token is most likely to come next?

Maybe it generates because. Now:

The sky is blue because

Then another prediction happens. Then another. This continues until the response is finished.

A better mental model

Generation is really repeated next-token prediction

When ChatGPT writes something like:

A rainbow forms when sunlight passes through water droplets…

it may feel as if the entire sentence was produced at once. Conceptually, the process is closer to this:

A vertical chain of generated tokens: A, then rainbow, then forms, then when, then sunlight, each followed by predict next token.
Each newly generated token becomes part of the context used to predict the next one. That is the heart of generation in a Large Language Model.
From generation to inference

Does the model treat the prompt and every new token the same way?

Now we understand the basic idea: an LLM generates a response one token at a time.

But another question appears. When we send a prompt to an LLM, does it process the prompt and every generated token in exactly the same way?

Not quite. At a high level, LLM inference can be divided into two important phases: Prefill and Decode. There is also another important concept called KV Cache.

These three ideas are fundamental to understanding how modern LLM inference works. A simple way to remember them is:

Prefill = Read
Decode = Write
KV Cache = Remember
Phase 1

What is Prefill?

Suppose we ask:

Explain how a rainbow is formed in simple terms.

Before the model can generate an answer, it must first process the prompt. That initial processing stage is called Prefill.

During prefill, the model looks at the tokens in your input prompt and builds the internal information it needs to begin generating the answer.

Prefill flow: prompt, convert into tokens, process the input tokens, then predict the first output token.
Prefill happens at the beginning of generation. If your prompt contains 100 tokens, those tokens must be processed before the answer can start.

If your prompt contains 10,000 tokens, considerably more input information needs to be processed. This is why prompt length can affect how quickly an LLM begins responding.

Prefill = The model reads your prompt.
Phase 2

What is Decode?

After the prompt has been processed, the model begins generating the response. This stage is called Decode.

During decode, the model generates new tokens one at a time.

Suppose the prompt is:

Explain how a rainbow is formed.

After prefill, the first generated token might be A. Now decode continues. The model may generate rainbow, then forms, then when, then sunlight, and so on.

Prefill processes the full prompt and generates the first token. Decode then generates, adds, and repeats until the response is finished.
The model continues decoding until it reaches a stopping condition, such as a special end-of-sequence token or a configured generation limit.
Prefill processes the input. Decode generates the output.
Why two names?

Why is this distinction important?

At first, Prefill and Decode may sound like two names for almost the same thing. But they behave differently.

During Prefill, the model is processing the prompt that already exists. During Decode, the model is generating new tokens one at a time.

That difference becomes extremely important when we start thinking about performance, GPU usage, latency, memory, batching, and serving many users at the same time.

Token-by-token generation with a Gemma LLM: Step 1 Prefill processes the entire prompt and produces the first token A. Step 2 Decode repeats, appending each new token until the response is complete.
Prefill happens once: process the full prompt and generate the first token. Decode repeats: generate one token at a time, append it, and continue.

But before going deeper into those serving topics, there is one important problem we need to solve.

The expensive path

The problem: repeating the same work

Imagine that the prompt contains 10 tokens. The model processes those 10 tokens and generates Token 11. Now it needs to generate Token 12.

A very simple implementation could send all 11 existing tokens through the model again. Then the model generates Token 12. Now the sequence contains 12 tokens. To generate Token 13, the simple approach could again process tokens 1 through 12.

The sequence keeps growing. And if we keep sending the entire sequence through the model every time, the model keeps doing work for tokens it has already processed.

Without a cache, generating token 11 processes tokens 1 to 10, generating token 12 processes tokens 1 to 11 again, and generating token 13 processes tokens 1 to 12 again.
Can you see the problem? The model has already processed the earlier tokens. Repeating all that work means more GPU computation, more memory activity, and more generation time as the sequence grows.

There must be a better way. There is. This is where KV Cache comes in.

The better way

What is KV Cache?

KV Cache may initially sound like an advanced concept. But its basic purpose is very simple:

Don’t calculate the same information again if you can save it and reuse it.

Inside a Transformer, the attention mechanism calculates several internal values. Two of these are called K = Key and V = Value.

You do not need to understand the mathematics behind Keys and Values at this stage. For now, think of them as:

Useful internal information calculated for previously processed tokens.

Instead of calculating this information again and again, the model can save it. That saved information is called the Key-Value Cache, or simply KV Cache.

Without KV Cache

Process everything again, every step

Suppose the model has generated:

A rainbow forms

To generate the next token, a simple implementation might effectively work with the prompt plus A plus rainbow plus forms, process everything again, and generate the next token. Maybe it generates when.

Now the sequence becomes “A rainbow forms when.” For another token, the simple implementation again works through the whole sequence. As the sequence becomes longer, more previous information is repeatedly processed.

With KV Cache

Save the work. Reuse it. Only process what is new.

When the model processes the original prompt during Prefill, it calculates Key and Value information for the prompt tokens. Instead of discarding that information, it stores it in the KV Cache.

At the same time, the model generates its first output token. Suppose that token is A. Now we move into Decode.

For the next generation step, the model does not need to recalculate all of the previous Key and Value information. Instead, it can use saved KV Cache + new token.

Prefill processes the prompt, saves Key and Value information into the KV Cache, then decode reuses that cache plus the newest token to generate the next token.
Decode can reuse the cached Key and Value states from previous tokens instead of repeatedly recalculating those states.
Before KV Cache, the model reprocesses the entire sequence at every step. After KV Cache, prefill processes the prompt once, then each decode step only processes the new token and reuses saved Key-Value states.
Without KV Cache, the model redoes the same work for previous tokens at every step. With KV Cache, only the new token is processed each time, and past Key-Value states are reused.
A simple analogy

Don’t reread the first 100 pages every time

Imagine you are reading a 200-page book. You have already read the first 100 pages. Now you reach page 101.

Without memory, you would have to go back and reread the first 100 pages before understanding page 101. Then before reading page 102, you would again reread pages 1 through 101. That would be incredibly inefficient.

Instead, you remember useful information from what you have already read and continue from there.

KV Cache follows a similar idea. It does not store the text simply as human memory would. Instead, it stores useful internal Key and Value states calculated by the Transformer’s attention mechanism. But the high-level idea is the same:

Remember useful previous work instead of recomputing it.
The complete flow

Putting Prefill, Decode, and KV Cache together

Suppose the user asks:

Explain how a rainbow is formed.

The request begins with Prefill.

Step 1: Prefill

The model receives the entire prompt. The prompt is converted into tokens. The model processes those tokens. During this processing, useful Key and Value states are generated and stored in the KV Cache. The model also predicts the first response token. Suppose it generates A.

Step 2: Decode

Instead of processing the entire prompt again, the model can reuse the saved KV Cache. Conceptually: KV Cache + A → rainbow. The KV Cache is updated. Next: updated KV Cache + rainbow → forms. Update the cache again. Then another token. And so on.

Complete inference mental model: prompt goes through Prefill, which creates the KV Cache and the first token, then Decode reuses the cache, generates the next token, updates the cache, and repeats.
That is a much better way to think about LLM inference: read the prompt once, then write the answer one token at a time while remembering previous work.
Why it matters

Why does KV Cache matter so much?

Without KV Cache, the amount of repeated work grows as the generated sequence becomes longer. With KV Cache, prefill processes the complete prompt once, then each decode step reuses the cache and processes the new token.

Without KV Cache each iteration processes a longer sequence. With KV Cache, prefill processes the prompt once and each decode step only processes the new token. KV Cache trades memory for speed.
This greatly reduces repeated computation during decoding. But KV Cache is not free: the cache needs memory.

During GPU inference, KV Cache is typically stored in GPU memory. That means we are making a trade-off: use more memory, perform less repeated computation, generate tokens faster.

KV Cache trades memory for speed.

This trade-off becomes extremely important when an LLM is serving many users.

From one chat to thousands

One user is easy. Thousands of users are different.

Imagine one person chatting with an LLM. The model maintains KV Cache information for that active sequence.

Now imagine 10 users. 100 users. 1,000 users. 10,000 users. Each active sequence may require KV Cache memory. And the longer a sequence becomes, the more KV Cache memory may be required.

Suddenly, KV Cache is no longer just a small implementation detail. It becomes one of the major resources an LLM serving system needs to manage.

This is why modern inference frameworks such as vLLM and SGLang put so much effort into efficiently managing KV Cache memory.

We’ll encounter many of these ideas again when discussing LLM inference and serving in more depth.

Before you go

Bringing it back to the “G” in ChatGPT

We started with a simple question: What does the G in ChatGPT mean?

The G stands for Generative. And now we can give that word a much more meaningful explanation.

ChatGPT is called generative because it generates new output token by token. For every step, the model considers possible next tokens, assigns scores or probabilities to them, selects a token, adds it to the sequence, and repeats the process.

Underneath that simple-looking conversation, several important things are happening:

Next-token prediction decides what could come next.
Prefill processes the user’s input.
Decode generates the response one token at a time.
KV Cache remembers useful previous attention states so the model can avoid repeating unnecessary work.

If you remember only three things from this section, remember:

Prefill = Read
Decode = Write
KV Cache = Remember

And if you remember only one thing about the G in ChatGPT, remember this:

Generative means the model creates the response as a sequence of next-token predictions rather than retrieving a complete prewritten answer.

That is the foundation of how modern Large Language Models generate text.

Next we look at the P: why the model can make those predictions in the first place. Continue to Week 1 · Pre-trained.