What Does the “G” in ChatGPT Mean?
The G stands for Generative, not “generation.” The natural way to explain Generative is to explain how generation actually happens: one token at a time.
Let’s start with what the model is actually producing ↓
In the previous part, we explored the “Chat” in ChatGPT and looked at what actually happens when we have a conversation with a Large Language Model.
Now let’s move to the next letter:
G → Generative
The word Generative means that the model can generate new content.
When you ask ChatGPT a question, it does not simply look inside a database, find a complete answer, and return it to you. Instead, it generates the answer while you are interacting with it.
That raises an interesting question:
How does an LLM actually generate text?
At first, it may look as if ChatGPT writes an entire sentence or paragraph in one go. That is not normally what happens.
An LLM generates its response one token at a time. Understanding this one idea makes many other LLM concepts much easier to understand.
First, what is a token?
Before understanding generation, we need to understand what the model is actually generating.
An LLM does not directly think in terms of sentences or paragraphs. It works with tokens.
A token can be an entire word, part of a word, punctuation, or another small piece of text.
For example, a sentence such as:
Explain how a rainbow is formed.
is first converted into tokens before the model processes it.
Internally, those tokens are represented using numbers called token IDs. The tokenizer converts human-readable text into these tokens before the model works with them.
For a beginner, the exact token IDs are not important. The main idea is simply:
One token at a time
Suppose we give an LLM this input:
The capital of France is
The model needs to decide what should come next. It may predict:
Paris
Now the sequence becomes:
The capital of France is Paris
The newly generated token becomes part of the input for the next prediction. The model then asks again: What token should come next?
It predicts another token. Then another. Then another.
This is the core idea behind autoregressive text generation.
How does the model know which token comes next?
The model does not usually say, “I know with 100% certainty that the next word is blue.”
Instead, it considers many possible tokens. Suppose the current text is:
The sky is
The model might consider possibilities such as:
| Possible next token | Example probability |
|---|---|
| blue | 65% |
| clear | 15% |
| cloudy | 10% |
| beautiful | 6% |
| green | 1% |
For every possible token, the model first produces a score. These raw scores are commonly called logits.
You do not need to worry about the mathematics behind logits yet. For now, think of them as:
The model’s raw scores for all the possible next tokens.
A higher score generally means that the model considers that token more suitable in the current context.
From scores to probabilities
Raw scores are difficult for humans to interpret. Conceptually, those scores can be converted into probabilities.
For example, raw scores such as blue → 8.2, clear → 6.5, cloudy → 5.9 may become something easier to understand:
blue → 65%
clear → 20%
cloudy → 10%
everything else → remaining probability
Again, the numbers are only examples.
The important concept is that the model creates a probability distribution over possible next tokens. One token is then selected from those possibilities.
Suppose the selected token is blue. Our sequence becomes:
The sky is blue
Now the model repeats the entire next-token prediction process. It asks: given “The sky is blue,” what token is most likely to come next?
Maybe it generates because. Now:
The sky is blue because
Then another prediction happens. Then another. This continues until the response is finished.
Generation is really repeated next-token prediction
When ChatGPT writes something like:
A rainbow forms when sunlight passes through water droplets…
it may feel as if the entire sentence was produced at once. Conceptually, the process is closer to this:
Does the model treat the prompt and every new token the same way?
Now we understand the basic idea: an LLM generates a response one token at a time.
But another question appears. When we send a prompt to an LLM, does it process the prompt and every generated token in exactly the same way?
Not quite. At a high level, LLM inference can be divided into two important phases: Prefill and Decode. There is also another important concept called KV Cache.
These three ideas are fundamental to understanding how modern LLM inference works. A simple way to remember them is:
Prefill = Read
Decode = Write
KV Cache = Remember
What is Prefill?
Suppose we ask:
Explain how a rainbow is formed in simple terms.
Before the model can generate an answer, it must first process the prompt. That initial processing stage is called Prefill.
During prefill, the model looks at the tokens in your input prompt and builds the internal information it needs to begin generating the answer.
If your prompt contains 10,000 tokens, considerably more input information needs to be processed. This is why prompt length can affect how quickly an LLM begins responding.
Prefill = The model reads your prompt.
What is Decode?
After the prompt has been processed, the model begins generating the response. This stage is called Decode.
During decode, the model generates new tokens one at a time.
Suppose the prompt is:
Explain how a rainbow is formed.
After prefill, the first generated token might be A. Now decode continues. The model may generate rainbow, then forms, then when, then sunlight, and so on.
Prefill processes the input. Decode generates the output.
Why is this distinction important?
At first, Prefill and Decode may sound like two names for almost the same thing. But they behave differently.
During Prefill, the model is processing the prompt that already exists. During Decode, the model is generating new tokens one at a time.
That difference becomes extremely important when we start thinking about performance, GPU usage, latency, memory, batching, and serving many users at the same time.
But before going deeper into those serving topics, there is one important problem we need to solve.
The problem: repeating the same work
Imagine that the prompt contains 10 tokens. The model processes those 10 tokens and generates Token 11. Now it needs to generate Token 12.
A very simple implementation could send all 11 existing tokens through the model again. Then the model generates Token 12. Now the sequence contains 12 tokens. To generate Token 13, the simple approach could again process tokens 1 through 12.
The sequence keeps growing. And if we keep sending the entire sequence through the model every time, the model keeps doing work for tokens it has already processed.
There must be a better way. There is. This is where KV Cache comes in.
What is KV Cache?
KV Cache may initially sound like an advanced concept. But its basic purpose is very simple:
Don’t calculate the same information again if you can save it and reuse it.
Inside a Transformer, the attention mechanism calculates several internal values. Two of these are called K = Key and V = Value.
You do not need to understand the mathematics behind Keys and Values at this stage. For now, think of them as:
Useful internal information calculated for previously processed tokens.
Instead of calculating this information again and again, the model can save it. That saved information is called the Key-Value Cache, or simply KV Cache.
Process everything again, every step
Suppose the model has generated:
A rainbow forms
To generate the next token, a simple implementation might effectively work with the prompt plus A plus rainbow plus forms, process everything again, and generate the next token. Maybe it generates when.
Now the sequence becomes “A rainbow forms when.” For another token, the simple implementation again works through the whole sequence. As the sequence becomes longer, more previous information is repeatedly processed.
Save the work. Reuse it. Only process what is new.
When the model processes the original prompt during Prefill, it calculates Key and Value information for the prompt tokens. Instead of discarding that information, it stores it in the KV Cache.
At the same time, the model generates its first output token. Suppose that token is A. Now we move into Decode.
For the next generation step, the model does not need to recalculate all of the previous Key and Value information. Instead, it can use saved KV Cache + new token.
Don’t reread the first 100 pages every time
Imagine you are reading a 200-page book. You have already read the first 100 pages. Now you reach page 101.
Without memory, you would have to go back and reread the first 100 pages before understanding page 101. Then before reading page 102, you would again reread pages 1 through 101. That would be incredibly inefficient.
Instead, you remember useful information from what you have already read and continue from there.
KV Cache follows a similar idea. It does not store the text simply as human memory would. Instead, it stores useful internal Key and Value states calculated by the Transformer’s attention mechanism. But the high-level idea is the same:
Remember useful previous work instead of recomputing it.
Putting Prefill, Decode, and KV Cache together
Suppose the user asks:
Explain how a rainbow is formed.
The request begins with Prefill.
Step 1: Prefill
The model receives the entire prompt. The prompt is converted into tokens. The model processes those tokens. During this processing, useful Key and Value states are generated and stored in the KV Cache. The model also predicts the first response token. Suppose it generates A.
Step 2: Decode
Instead of processing the entire prompt again, the model can reuse the saved KV Cache. Conceptually: KV Cache + A → rainbow. The KV Cache is updated. Next: updated KV Cache + rainbow → forms. Update the cache again. Then another token. And so on.
Why does KV Cache matter so much?
Without KV Cache, the amount of repeated work grows as the generated sequence becomes longer. With KV Cache, prefill processes the complete prompt once, then each decode step reuses the cache and processes the new token.
During GPU inference, KV Cache is typically stored in GPU memory. That means we are making a trade-off: use more memory, perform less repeated computation, generate tokens faster.
KV Cache trades memory for speed.
This trade-off becomes extremely important when an LLM is serving many users.
One user is easy. Thousands of users are different.
Imagine one person chatting with an LLM. The model maintains KV Cache information for that active sequence.
Now imagine 10 users. 100 users. 1,000 users. 10,000 users. Each active sequence may require KV Cache memory. And the longer a sequence becomes, the more KV Cache memory may be required.
Suddenly, KV Cache is no longer just a small implementation detail. It becomes one of the major resources an LLM serving system needs to manage.
This is why modern inference frameworks such as vLLM and SGLang put so much effort into efficiently managing KV Cache memory.
We’ll encounter many of these ideas again when discussing LLM inference and serving in more depth.
Bringing it back to the “G” in ChatGPT
We started with a simple question: What does the G in ChatGPT mean?
The G stands for Generative. And now we can give that word a much more meaningful explanation.
ChatGPT is called generative because it generates new output token by token. For every step, the model considers possible next tokens, assigns scores or probabilities to them, selects a token, adds it to the sequence, and repeats the process.
Underneath that simple-looking conversation, several important things are happening:
Next-token prediction decides what could come next.
Prefill processes the user’s input.
Decode generates the response one token at a time.
KV Cache remembers useful previous attention states so the model can avoid repeating unnecessary work.
If you remember only three things from this section, remember:
Prefill = Read
Decode = Write
KV Cache = Remember
And if you remember only one thing about the G in ChatGPT, remember this:
Generative means the model creates the response as a sequence of next-token predictions rather than retrieving a complete prewritten answer.
That is the foundation of how modern Large Language Models generate text.
Next we look at the P: why the model can make those predictions in the first place. Continue to Week 1 · Pre-trained.