What Does the “P” in ChatGPT Mean?
The P stands for Pre-trained, not “pre-training.” To explain what Pre-trained means, we need to understand the pre-training process that happened before you ever started chatting.
Let’s start with where the model’s knowledge comes from ↓
In the previous part, we explored the “G” in ChatGPT:
G → Generative
We learned that an LLM generates a response one token at a time by repeatedly predicting what token should come next.
Now let’s move to the next letter:
P → Pre-trained
This word tells us something very important about how ChatGPT became capable of generating those responses in the first place.
Before ChatGPT could answer questions, write code, explain Linux, summarize an article, or talk about history, it first had to learn language and knowledge from an enormous amount of data.
That learning happens before we start chatting with the model.
That is why it is called:
Pre-trained
You can think of it very simply:
Pre-trained = the model learned before you started talking to it.
But what exactly does the model learn? And where does all that knowledge come from?
To understand that, we first need to talk about something that is sometimes overshadowed by model architecture:
Data
Architecture is important, but data is what gets learned
When people talk about Large Language Models, the conversation often focuses on things such as Transformers, attention, Mixture of Experts, number of parameters, GPUs, and model architecture.
All of those things matter. Architecture gives the model the ability to learn.
But the actual information the model learns comes from data.
A sophisticated model architecture trained on poor-quality or limited data will still produce poor results.
This is why training data has become one of the most important parts of building a modern Large Language Model.
What happens during pre-training?
Pre-training is the first major stage of training a Large Language Model.
During this stage, the model is exposed to enormous amounts of text. That data may include web pages, Wikipedia, books, code, news articles, discussion forums, and large public datasets.
The scale can reach billions or trillions of words or tokens.
The purpose is not to manually teach the model every possible question and answer. Instead, the model learns patterns from the data.
Remember what we discussed in the Generative section? The model learns to predict:
What token is likely to come next?
Pre-training performs this kind of learning at enormous scale.
What does a real pre-training mix look like?
We just listed sources such as web pages, Wikipedia, books, and code. That can still feel abstract until you see an actual recipe.
The original LLaMA paper from Meta (2023) is one of the clearer public examples. They published the data mixture used to train the model on about 1.4 trillion tokens.
Read the table as three different questions.
Disk size is how much raw data that source occupies. CommonCrawl is huge: about 3.3 TB. Wikipedia is tiny next to that: about 83 GB.
Sampling proportion is how often the trainers actually pick from that source while the model is learning. CommonCrawl plus C4 together make up 82% of the mix. GitHub, Wikipedia, and books are each 4.5%. Scientific papers from arXiv are 2.5%. Stack Exchange questions and answers are 2%.
Epochs tell us how many times the model sees that subset. For most of LLaMA’s data, each token was used about once. Wikipedia and books were the exception: they were seen a little more than twice.
That is a deliberate quality choice. Wikipedia is much smaller than a web crawl, but it is treated as higher-value, so the trainers sample it more often per byte. The huge crawl still dominates the mix, but each webpage is not endlessly repeated.
The sources themselves match the story of pre-training: lots of cleaned web text, plus code, encyclopedias, books, scientific papers, and high-quality Q&A. LLaMA also filtered along the way — for example keeping English CommonCrawl pages that looked more like Wikipedia references, using public GitHub projects under Apache, BSD, and MIT licenses, covering Wikipedia in 20 languages, and sorting Stack Exchange answers by score.
Pre-training data is mixed on purpose. It is not “dump the entire internet in equally.”
This is not ChatGPT’s private recipe. It is one published mix that makes the idea concrete: the model learns from many kinds of text, in carefully chosen proportions.
What is robots.txt?
Look back at that LLaMA table. The largest slice is CommonCrawl — public web pages. That raises a practical question:
How do crawlers know which parts of a website they are allowed to visit?
Many sites answer that with a simple text file called robots.txt.
It usually lives at the root of a website, for example https://www.google.com/robots.txt.
Inside the file, rules such as User-agent, Allow, and Disallow give instructions to crawlers. Google’s robots.txt, for instance, tells crawlers not to access certain URLs such as /search, while still allowing some specific pages. It can also point to the website’s sitemap.
Think of robots.txt as a sign for web crawlers:
“Please enter here, but don’t go there.”
That matters for pre-training because a huge amount of LLM training text comes from crawled web pages. robots.txt is one of the conventions websites use to say which URLs they want crawlers to skip.
There is an important catch. robots.txt is only a set of instructions. Well-behaved crawlers try to follow it. It is not a security mechanism, and it should not be used to protect private or sensitive information. If a page is truly private, it needs real access control — not a polite note in a text file.
Common Crawl
Common Crawl is a nonprofit project that regularly crawls billions of web pages and stores the collected data in a large, publicly available archive. You can think of it as a huge snapshot of the public web that researchers, developers, and AI companies can use instead of crawling the entire internet themselves. Its datasets contain web pages, links, metadata, and extracted text, and new crawls are released regularly—for example, the current archive is listed as CC-MAIN-2026-34. Common Crawl is especially important in AI because large collections of web text like these can be used as one of the data sources for training and researching large language models.
Raw Common Crawl data is huge and contains a lot of things we may not want when training an LLM, such as HTML, duplicate pages, navigation text, advertisements, spam, and low-quality content. Because of this, researchers usually clean, filter, and deduplicate Common Crawl before using it for model training. One good example is FineWeb, a publicly available dataset on Hugging Face containing trillions of tokens of cleaned English text collected from Common Crawl. Other popular cleaned Common Crawl datasets include C4 and Falcon RefinedWeb. So instead of downloading and processing raw Common Crawl data yourself, you can often start with one of these already-cleaned datasets.
Wikipedia dumps
Wikipedia also makes its content available as periodic data dumps. Instead of visiting and downloading millions of Wikipedia pages one by one, developers and researchers can download large datasets containing Wikipedia articles, page metadata, links, and even revision history. Wikimedia regularly creates these dumps and makes them publicly available at dumps.wikimedia.org. You can think of a Wikipedia dump as a snapshot of Wikipedia taken at a particular point in time. These datasets are useful for research, search engines, data analysis, and AI projects that need access to a large collection of human-written text.
GitHub archives
GitHub also has large public archives of open-source software and activity. One example is GH Archive, which continuously records public activity happening on GitHub, such as commits, pull requests, issues, forks, and repository stars. This data is collected into hourly archives and can be downloaded or analyzed through Google BigQuery. For the actual source code, Google also provides a large GitHub public dataset in BigQuery containing millions of open-source repositories and billions of files. GitHub also runs the GitHub Archive Program, where public repositories are preserved through organizations such as Software Heritage and the Internet Archive. You can think of these archives as large snapshots and historical records of the open-source world, which can be useful for research, software analysis, and building or training AI systems.
arXiv
arXiv is another large source of publicly available text, especially scientific research papers. It hosts papers from areas such as computer science, mathematics, physics, statistics, electrical engineering, and many other research fields. Instead of downloading papers one by one, researchers can access a machine-readable arXiv dataset containing information such as paper titles, authors, categories, abstracts, and full-text PDFs. The dataset is also available through Kaggle and Google Cloud Storage and is updated regularly. You can think of arXiv as a huge digital library of scientific and technical knowledge, which makes it a valuable source for research, data analysis, and training language models on academic and technical content.
The model is not storing a giant list of answers
Suppose the model sees sentences such as:
Paris is the capital of France.
France is a country in Europe.
The Eiffel Tower is located in Paris.
Paris is known for art, culture, and architecture.
It may encounter information about Paris thousands or millions of times across different documents.
Over time, the model begins learning relationships between concepts such as Paris and France, Paris and city, France and Europe, Paris and the Eiffel Tower.
This is one reason an LLM can respond to questions phrased in ways that did not appear exactly in its training data.
Pre-training is next-token prediction at massive scale
This is where the G and P parts of GPT connect beautifully.
In the previous section, we learned that during generation the model predicts one token at a time. During pre-training, the model learns how to make those predictions.
Imagine training text such as:
The capital of France is Paris.
The model could receive “The capital of France is” and try to predict “Paris.” If its prediction is wrong, the training process adjusts the model. Then the model tries again across another example. And another. And another.
Not millions of times. Potentially trillions of tokens worth of training.
What does the model learn during pre-training?
Pre-training teaches the model much more than individual facts. Through huge amounts of text, the model begins learning several things at the same time.
Vocabulary
The model learns how words and concepts are used. For example, Kubernetes is commonly related to containers, pods, clusters, deployments, nodes, services, and orchestration.
Similarly, Python could refer to a programming language or a snake depending on the surrounding context.
Grammar
The model learns patterns of language. It begins learning that “I am going to the store” is a normal English sentence, while unusual arrangements of the same words are much less likely.
Nobody needs to manually write every grammatical rule into the model. Many language patterns emerge from seeing enormous amounts of text.
General knowledge
The model encounters information about countries, cities, science, technology, history, programming, literature, mathematics, and many other subjects. This exposure helps build broad general knowledge.
Relationships between concepts
The model also begins learning how concepts are connected. For example: Docker → containers, Kubernetes → orchestration, AWS → cloud computing, Linux → operating system, Paris → France.
These relationships help the model produce useful responses later.
Language patterns
Suppose the model has seen many documents that begin with something like “To install this package…” It learns that instructions may follow in a certain style.
Similarly, after “Once upon a time…” a story is likely to follow. After “The root cause of the outage was…” a technical explanation may follow.
The model learns these patterns through training data rather than through manually written rules.
Does someone label all this data by hand?
No. And this is one of the major differences between traditional machine learning and Large Language Model pre-training.
Traditional supervised machine learning often relies heavily on labeled examples. Humans may need to label thousands or millions of examples.
| Label | |
|---|---|
| Win $1 million today! | Spam |
| Meeting at 3 PM | Not Spam |
Large Language Model pre-training works differently. Instead of manually labeling every sentence, enormous quantities of naturally occurring text can be used to teach the model language patterns.
This is one reason LLM training can scale to massive datasets. The shift is away from manually labeling millions of examples toward collecting huge amounts of text, training general-purpose models, and then applying smaller amounts of fine-tuning later.
Why is it called “pre”-training?
The word pre is important. Pre-training happens before the model becomes the assistant we eventually interact with.
A model that has completed pre-training has learned a huge amount about language, patterns, concepts, facts, code, and relationships between ideas.
But that does not automatically mean it will behave like ChatGPT.
A raw pretrained model may be very good at continuing text, but being a useful conversational assistant requires additional training. This brings us to an important distinction.
Pre-training is only one stage
Modern language models are generally not trained in one single step. Training is often divided into three broad stages:
Stage 1 → Pre-training
Stage 2 → Mid-training
Stage 3 → Post-training
For understanding the P in GPT, pre-training is the most important one. But understanding the other stages helps explain why a raw pretrained model is different from something like ChatGPT.
Pre-training
This is where the model learns broad language patterns and general knowledge from enormous datasets. Think of it as:
Building the model’s foundation.
During this stage, the model learns things such as vocabulary, grammar, general knowledge, basic reasoning patterns, relationships between concepts, and patterns found across different types of text.
At the end of pre-training, we have something much more capable than an untrained neural network. But we do not necessarily have a polished chatbot yet.
Mid-training
After broad pre-training, some models go through another stage using more carefully selected or specialized data.
This can improve specific capabilities such as coding, mathematics, long-document understanding, multilingual performance, and other targeted skills.
For example, training with high-quality programming data can significantly strengthen coding capabilities. Training with mathematical problems can improve mathematical reasoning. Training with long documents can help the model become better at working with larger contexts.
Pre-training = broad education
Mid-training = skill development
Post-training
Now imagine we have a model that knows a lot. That still does not guarantee that it will respond like a helpful assistant.
It needs to learn things such as: when a user asks a question, answer it clearly; follow the user’s instructions; format responses appropriately; avoid harmful responses; behave helpfully during a conversation.
This happens during post-training.
That stage often includes techniques such as Supervised Fine-Tuning (SFT) for instruction following and Reinforcement Learning from Human Feedback (RLHF) for shaping helpful behavior.
Pre-training teaches knowledge.
Mid-training strengthens skills.
Post-training shapes behavior.
An easy school analogy
Imagine a student. During many years at school, the student reads textbooks, stories, science books, history, mathematics, newspapers, and programming material. This gives the student broad knowledge. Think of that as pre-training.
Later, the student chooses computer science and spends extra time learning programming, algorithms, operating systems, and databases. Think of that as mid-training.
Finally, suppose that person joins a customer-support organization and receives training on how to answer customers, how to communicate clearly, what they should and should not say, and how to respond appropriately. Think of that as post-training.
Why does the amount of training data matter?
If data is so important, a natural question is: why don’t we simply build bigger and bigger models?
Because model size is only one part of the equation. A very large model without enough training data may not use its capacity efficiently.
This became especially clear through research commonly associated with the Chinchilla scaling laws.
The key idea was that, for a given compute budget, model size and the amount of training data need to be balanced. A Chinchilla-era rule of thumb is around:
20 training tokens per model parameter
| Parameters | Approx. training tokens |
|---|---|
| 1 billion | ~20 billion tokens |
| 7 billion | ~140 billion tokens |
| 70 billion | ~1.4 trillion tokens |
The exact recipes used by modern models can differ, but the bigger lesson is more important:
Making the model bigger is not enough. You also need enough good training data.
A smaller model can sometimes beat a bigger model
This was an important insight from Chinchilla.
GPT-3 had around 175 billion parameters and was trained on roughly 300 billion tokens. Chinchilla used around 70 billion parameters but trained on approximately 1.4 trillion tokens.
Despite having far fewer parameters, Chinchilla achieved stronger performance under its comparison setting.
But more data is not automatically better
Suppose you have two datasets.
That means modern data pipelines perform a great deal of work before training ever begins. They may clean text, remove duplicates, filter poor-quality sources, balance different types of content, remove unwanted material, and decide how much data should come from each domain.
The LLaMA mix we saw earlier is a good example. Wikipedia is tiny on disk compared with CommonCrawl, but it is sampled more often per byte. The trainers were not treating every token as equally valuable.
This is one reason training data preparation is such an important part of modern AI development.
Why do AI companies talk so little about their data?
Something interesting happens when companies release information about Large Language Models. They may tell us quite a lot about number of layers, parameter count, attention architecture, tokenizer, Mixture-of-Experts design, and training infrastructure.
But detailed information about the training dataset is often much harder to find.
Why? Because the exact data mixture and preparation pipeline can be a competitive advantage.
Two companies could use very similar Transformer architectures but end up with models having very different capabilities because they used different data sources, filtering techniques, data mixtures, deduplication strategies, quality controls, and training stages.
The reasons for keeping data pipelines private often fall into three broad areas:
Competitive dynamics
Legal constraints
Trade secrets
The original LLaMA paper is a useful exception: it published the sampling mix in Table 1. Many later frontier models do not.
So when we talk about the intelligence of an LLM, architecture is only part of the story. The training data is another enormous part of it.
Open-Source vs Open-Weight Models
You will often hear AI models described as open source, but not every downloadable model is truly open source. An open-source AI model gives researchers much more than just the finished model. Ideally, it provides the model weights, training code, architecture, training recipes, and information about the data used to train it, so people can study how the model was created, modify it, and reproduce parts of the process. A good example is Ai2’s OLMo family, which releases its model weights, training code, training data, evaluation tools, checkpoints, and other details about the training process.
An open-weight model, on the other hand, mainly gives you access to the model’s trained weights. This means you can download the model, run it on your own machine or servers, and often fine-tune it for your own use. However, you may not receive the complete training dataset or everything required to recreate the model from scratch. Models from families such as Meta Llama are commonly described as open-weight models because their weights can be downloaded, but they use their own licenses and do not provide everything required to reproduce the original training process.
Pre-training does not mean the model “remembers the internet”
This distinction is especially important for beginners.
When we say an LLM was trained on large amounts of text, we should not imagine that it contains a normal database with every webpage stored inside.
It is not working like: question → search training database → find matching paragraph → return paragraph.
Instead, during training, the model’s internal parameters are adjusted based on patterns found across the training data. After training, those learned parameters are what the model uses to predict tokens.
That is very different from a traditional search engine.
Connecting “G” and “P”
G → Generative. The model generates the response token by token.
P → Pre-trained. The model can make those predictions because it was trained beforehand on enormous amounts of data.
Think about the phrase “The capital of France is…” Why can the model predict “Paris”? Because during pre-training, relationships between concepts such as France, Paris, capital, and Europe were learned from enormous amounts of language data.
Without training, there would be no useful next-token predictions.
But where is the knowledge stored?
This leads to another important idea. After pre-training, what exactly contains everything the model learned?
The answer is: parameters.
A Large Language Model may contain billions of parameters. During training, these parameters are repeatedly adjusted.
This distinction is worth remembering because it connects pre-training directly to the generation process we covered in the previous section.
Training vs inference
We can now clearly separate two concepts.
Training: the model is learning. Its parameters are being updated. This requires enormous amounts of data, computation, GPUs, time, and infrastructure.
Inference: the model is using what it already learned. Its parameters generally remain fixed while it processes your prompt and generates tokens.
Bringing it back to the “P” in ChatGPT
We started with a simple question: What does the P in ChatGPT mean?
The answer is: P → Pre-trained.
Before the model interacts with you, it has already gone through a massive learning process using enormous amounts of text and other training data.
During pre-training, it learns vocabulary, grammar, language patterns, relationships between concepts, general knowledge, and patterns useful for reasoning.
Later training stages can strengthen particular capabilities and shape the model into something that behaves more like a helpful assistant. But the foundation comes from pre-training.
If you remember only one thing from this section, remember:
Pre-trained means the model learned from enormous amounts of data before you ever started chatting with it.
And now the name begins to make even more sense:
Chat → We interact with it conversationally.
G — Generative → It generates new responses one token at a time.
P — Pre-trained → It learned language, knowledge, and patterns before we started using it.
There is now only one major part of the name left:
T → Transformer
And that explains the architecture that makes much of this possible. When you are ready, ten questions cover Chat, Generative, Pre-trained, and Transformer together.