Ten Questions on ChatGPT
You’ve now explored Chat, Generative, Pre-trained, and Transformer — the four words behind ChatGPT.
This page gives you a chance to see how well you can explain those ideas in your own words.
There are no answers on this page on purpose. Try each question first, and if you are unsure, go back to the lesson, review the concept, and give it another try.
Why are we doing this?
If you can explain the name, you can explain how ChatGPT works
By now, you have learned about three of the four words in ChatGPT:
Chat is the conversation around the model.
G — Generative is how the answer is written, one token at a time.
P — Pre-trained is how the model learned before you ever typed a word.
The last letter is T — Transformer: the architecture that made this kind of model possible.
A Transformer does not read a sentence strictly one word at a time the way older RNNs and LSTMs did. It uses self-attention, which allows each token to look at other tokens and understand which ones are important for its meaning.
Think of these questions as practice for explaining these concepts naturally — especially in an interview.
Don’t worry about giving a textbook-perfect answer. Try to explain each concept as if you were explaining it to another person. If you get stuck, review the lesson and try again.
Chat · Questions 1–3
How does ChatGPT continue a conversation?
When you chat with ChatGPT, it may feel like the model remembers everything you previously said.
But the model itself is stateless.
These questions will help you understand how chat history, context windows, and long-term memory work together to create the experience of an ongoing conversation.
- You tell ChatGPT your name. A few messages later you ask, “What is my name?” and it answers correctly. Does the model actually remember you? If not, what is the application doing?
- What is a context window, and why is sending every earlier message — even when it still fits — often a bad idea?
- A user says they prefer Python examples, closes the chat, and comes back a week later. How can the system still know that preference, even though the model itself is stateless?
Generative · Questions 4–6
How does ChatGPT create an answer?
The G in GPT means Generative.
When you ask ChatGPT a question, the model is not simply finding a completed paragraph somewhere and returning it to you.
Instead, it generates the response step by step — predicting one token, then another, then another.
Concepts such as Prefill, Decode, and KV Cache help us understand what is happening behind the scenes while that response is being generated.
- When ChatGPT writes a paragraph, is it producing the whole answer at once? What is the model actually doing, token by token?
- Prefill and Decode are two different phases of generation. What does each one do, and why does that distinction matter when the prompt is very long?
- What problem does KV Cache solve, and why does it become a major concern when thousands of people are using the model at the same time?
Pre-trained · Questions 7–8
Where does the model’s knowledge come from?
The P in GPT means Pre-trained.
Long before you start a conversation with the model, it has already gone through a large training process.
These questions will help you think about the difference between training and inference, and what we really mean when we say that an LLM has “knowledge.”
- If an LLM was trained on a huge amount of internet text, does it look up pages from that data when you ask a question? Where does the model’s knowledge actually live?
- Why isn’t a bigger model automatically a better model? And why isn’t a pretrained base model the same thing as ChatGPT?
Transformer · Questions 9–10
What is happening underneath GPT?
The T in GPT means Transformer.
Before Transformers, architectures such as RNNs and LSTMs were commonly used for sequence-based data.
Transformers introduced self-attention, which made it much easier for models to understand relationships between words or tokens, even when those words appear far apart.
For example, context helps the model understand that “Apple” in one sentence may refer to a company, while “apple” in another sentence may refer to a fruit.
There is also one important idea to keep in mind:
A Transformer can look at the relationships between tokens in the available context, while ChatGPT can still generate its response one token at a time.
Those ideas do not contradict each other — they describe different parts of the process.
- RNNs and LSTMs already worked with sequences. What problem did they have, and what did the Transformer change?
- People say a Transformer looks at the whole sentence at once, but ChatGPT still generates one token at a time. How can both of those statements be true? Use this pair of sentences in your answer: “Apple released a new laptop.” and “I ate an apple.”
How to Use This Page
Ten questions. Four words. One goal: understand what is really happening.
You don’t need to memorize definitions.
Instead, try to explain each answer naturally, using your own words and examples.
If you can comfortably talk through these ten questions, you are already building a strong understanding of what the name ChatGPT actually represents:
Chat — how we interact with it.
Generative — how it creates a response.
Pre-trained — how it learned before we started chatting.
Transformer — the architecture underneath it.
The goal is not to give the perfect answer. The goal is to understand the concept well enough that you can explain it in your own words.