What Does the “P” in ChatGPT Mean?

The P stands for Pre-trained, not “pre-training.” To explain what Pre-trained means, we need to understand the pre-training process that happened before you ever started chatting.

Let’s start with where the model’s knowledge comes from ↓

In the previous part, we explored the “G” in ChatGPT:

G → Generative

We learned that an LLM generates a response one token at a time by repeatedly predicting what token should come next.

Now let’s move to the next letter:

P → Pre-trained

This word tells us something very important about how ChatGPT became capable of generating those responses in the first place.

Before ChatGPT could answer questions, write code, explain Linux, summarize an article, or talk about history, it first had to learn language and knowledge from an enormous amount of data.

That learning happens before we start chatting with the model.

That is why it is called:

Pre-trained

You can think of it very simply:

Pre-trained = the model learned before you started talking to it.

But what exactly does the model learn? And where does all that knowledge come from?

To understand that, we first need to talk about something that is sometimes overshadowed by model architecture:

Data
What actually gets learned

Architecture is important, but data is what gets learned

When people talk about Large Language Models, the conversation often focuses on things such as Transformers, attention, Mixture of Experts, number of parameters, GPUs, and model architecture.

All of those things matter. Architecture gives the model the ability to learn.

But the actual information the model learns comes from data.

A sophisticated model architecture trained on poor-quality or limited data will still produce poor results.

Architecture determines how the model can learn. Data determines what the model learns.
Architecture determines how the model can learn. Data determines what the model learns.

This is why training data has become one of the most important parts of building a modern Large Language Model.

The first major stage

What happens during pre-training?

Pre-training is the first major stage of training a Large Language Model.

During this stage, the model is exposed to enormous amounts of text. That data may include web pages, Wikipedia, books, code, news articles, discussion forums, and large public datasets.

The scale can reach billions or trillions of words or tokens.

The purpose is not to manually teach the model every possible question and answer. Instead, the model learns patterns from the data.

Remember what we discussed in the Generative section? The model learns to predict:

What token is likely to come next?

Pre-training performs this kind of learning at enormous scale.

A published example

What does a real pre-training mix look like?

We just listed sources such as web pages, Wikipedia, books, and code. That can still feel abstract until you see an actual recipe.

The original LLaMA paper from Meta (2023) is one of the clearer public examples. They published the data mixture used to train the model on about 1.4 trillion tokens.

Table 1 from the LLaMA paper: pre-training data mix when training on 1.4 trillion tokens. CommonCrawl 67%, C4 15%, GitHub 4.5%, Wikipedia 4.5%, Books 4.5%, ArXiv 2.5%, StackExchange 2.0%, with sampling proportion, epochs, and disk size for each source.
Table 1 from LLaMA: Open and Efficient Foundation Language Models (Touvron et al., 2023). Sampling proportion is how often a source is drawn during training — not simply how large the files are on disk.

Read the table as three different questions.

Disk size is how much raw data that source occupies. CommonCrawl is huge: about 3.3 TB. Wikipedia is tiny next to that: about 83 GB.

Sampling proportion is how often the trainers actually pick from that source while the model is learning. CommonCrawl plus C4 together make up 82% of the mix. GitHub, Wikipedia, and books are each 4.5%. Scientific papers from arXiv are 2.5%. Stack Exchange questions and answers are 2%.

Epochs tell us how many times the model sees that subset. For most of LLaMA’s data, each token was used about once. Wikipedia and books were the exception: they were seen a little more than twice.

That is a deliberate quality choice. Wikipedia is much smaller than a web crawl, but it is treated as higher-value, so the trainers sample it more often per byte. The huge crawl still dominates the mix, but each webpage is not endlessly repeated.

The sources themselves match the story of pre-training: lots of cleaned web text, plus code, encyclopedias, books, scientific papers, and high-quality Q&A. LLaMA also filtered along the way — for example keeping English CommonCrawl pages that looked more like Wikipedia references, using public GitHub projects under Apache, BSD, and MIT licenses, covering Wikipedia in 20 languages, and sorting Stack Exchange answers by score.

Pre-training data is mixed on purpose. It is not “dump the entire internet in equally.”

This is not ChatGPT’s private recipe. It is one published mix that makes the idea concrete: the model learns from many kinds of text, in carefully chosen proportions.

How the web is collected

What is robots.txt?

Look back at that LLaMA table. The largest slice is CommonCrawl — public web pages. That raises a practical question:

How do crawlers know which parts of a website they are allowed to visit?

Many sites answer that with a simple text file called robots.txt.

It usually lives at the root of a website, for example https://www.google.com/robots.txt.

Inside the file, rules such as User-agent, Allow, and Disallow give instructions to crawlers. Google’s robots.txt, for instance, tells crawlers not to access certain URLs such as /search, while still allowing some specific pages. It can also point to the website’s sitemap.

Example robots.txt file: User-agent star, Disallow /search, Allow /search/about, and a Sitemap URL.
A simplified robots.txt. User-agent says which crawler the rules apply to. Disallow and Allow say where it should and should not go. Sitemap points to a map of the site.

Think of robots.txt as a sign for web crawlers:

“Please enter here, but don’t go there.”

That matters for pre-training because a huge amount of LLM training text comes from crawled web pages. robots.txt is one of the conventions websites use to say which URLs they want crawlers to skip.

There is an important catch. robots.txt is only a set of instructions. Well-behaved crawlers try to follow it. It is not a security mechanism, and it should not be used to protect private or sensitive information. If a page is truly private, it needs real access control — not a polite note in a text file.

A public web archive

Common Crawl

Common Crawl is a nonprofit project that regularly crawls billions of web pages and stores the collected data in a large, publicly available archive. You can think of it as a huge snapshot of the public web that researchers, developers, and AI companies can use instead of crawling the entire internet themselves. Its datasets contain web pages, links, metadata, and extracted text, and new crawls are released regularly—for example, the current archive is listed as CC-MAIN-2026-34. Common Crawl is especially important in AI because large collections of web text like these can be used as one of the data sources for training and researching large language models.

Web crawler architecture: the World Wide Web feeds web pages into a multi-threaded downloader. A scheduler sends URLs to the downloader, discovered URLs go back into a queue, and text and metadata are written to storage.
A crawler takes URLs from a scheduler and queue, downloads pages from the web, discovers more URLs, and stores text and metadata. Common Crawl runs this kind of process at internet scale and publishes the result.

Raw Common Crawl data is huge and contains a lot of things we may not want when training an LLM, such as HTML, duplicate pages, navigation text, advertisements, spam, and low-quality content. Because of this, researchers usually clean, filter, and deduplicate Common Crawl before using it for model training. One good example is FineWeb, a publicly available dataset on Hugging Face containing trillions of tokens of cleaned English text collected from Common Crawl. Other popular cleaned Common Crawl datasets include C4 and Falcon RefinedWeb. So instead of downloading and processing raw Common Crawl data yourself, you can often start with one of these already-cleaned datasets.

Another public snapshot

Wikipedia dumps

Wikipedia also makes its content available as periodic data dumps. Instead of visiting and downloading millions of Wikipedia pages one by one, developers and researchers can download large datasets containing Wikipedia articles, page metadata, links, and even revision history. Wikimedia regularly creates these dumps and makes them publicly available at dumps.wikimedia.org. You can think of a Wikipedia dump as a snapshot of Wikipedia taken at a particular point in time. These datasets are useful for research, search engines, data analysis, and AI projects that need access to a large collection of human-written text.

Open-source archives

GitHub archives

GitHub also has large public archives of open-source software and activity. One example is GH Archive, which continuously records public activity happening on GitHub, such as commits, pull requests, issues, forks, and repository stars. This data is collected into hourly archives and can be downloaded or analyzed through Google BigQuery. For the actual source code, Google also provides a large GitHub public dataset in BigQuery containing millions of open-source repositories and billions of files. GitHub also runs the GitHub Archive Program, where public repositories are preserved through organizations such as Software Heritage and the Internet Archive. You can think of these archives as large snapshots and historical records of the open-source world, which can be useful for research, software analysis, and building or training AI systems.

Scientific papers

arXiv

arXiv is another large source of publicly available text, especially scientific research papers. It hosts papers from areas such as computer science, mathematics, physics, statistics, electrical engineering, and many other research fields. Instead of downloading papers one by one, researchers can access a machine-readable arXiv dataset containing information such as paper titles, authors, categories, abstracts, and full-text PDFs. The dataset is also available through Kaggle and Google Cloud Storage and is updated regularly. You can think of arXiv as a huge digital library of scientific and technical knowledge, which makes it a valuable source for research, data analysis, and training language models on academic and technical content.

A simple example

The model is not storing a giant list of answers

Suppose the model sees sentences such as:

Paris is the capital of France.
France is a country in Europe.
The Eiffel Tower is located in Paris.
Paris is known for art, culture, and architecture.

It may encounter information about Paris thousands or millions of times across different documents.

Over time, the model begins learning relationships between concepts such as Paris and France, Paris and city, France and Europe, Paris and the Eiffel Tower.

The model sees many sentences about Paris and learns relationships such as Paris to France, Paris to city, France to Europe, and Paris to the Eiffel Tower.
It is not simply storing one giant list of questions and answers. It is learning statistical patterns and relationships between tokens and concepts.

This is one reason an LLM can respond to questions phrased in ways that did not appear exactly in its training data.

Where G and P meet

Pre-training is next-token prediction at massive scale

This is where the G and P parts of GPT connect beautifully.

In the previous section, we learned that during generation the model predicts one token at a time. During pre-training, the model learns how to make those predictions.

Imagine training text such as:

The capital of France is Paris.

The model could receive “The capital of France is” and try to predict “Paris.” If its prediction is wrong, the training process adjusts the model. Then the model tries again across another example. And another. And another.

Not millions of times. Potentially trillions of tokens worth of training.

During pre-training the model receives a prefix such as The capital of France is and tries to predict Paris. Wrong predictions cause the model to be adjusted, across potentially trillions of tokens.
Pre-training teaches the model how to predict. Generation uses what the model learned to produce a response.
More than facts

What does the model learn during pre-training?

Pre-training teaches the model much more than individual facts. Through huge amounts of text, the model begins learning several things at the same time.

Vocabulary

The model learns how words and concepts are used. For example, Kubernetes is commonly related to containers, pods, clusters, deployments, nodes, services, and orchestration.

Similarly, Python could refer to a programming language or a snake depending on the surrounding context.

Grammar

The model learns patterns of language. It begins learning that “I am going to the store” is a normal English sentence, while unusual arrangements of the same words are much less likely.

Nobody needs to manually write every grammatical rule into the model. Many language patterns emerge from seeing enormous amounts of text.

General knowledge

The model encounters information about countries, cities, science, technology, history, programming, literature, mathematics, and many other subjects. This exposure helps build broad general knowledge.

Relationships between concepts

The model also begins learning how concepts are connected. For example: Docker → containers, Kubernetes → orchestration, AWS → cloud computing, Linux → operating system, Paris → France.

These relationships help the model produce useful responses later.

Language patterns

Suppose the model has seen many documents that begin with something like “To install this package…” It learns that instructions may follow in a certain style.

Similarly, after “Once upon a time…” a story is likely to follow. After “The root cause of the outage was…” a technical explanation may follow.

The model learns these patterns through training data rather than through manually written rules.

A major difference

Does someone label all this data by hand?

No. And this is one of the major differences between traditional machine learning and Large Language Model pre-training.

Traditional supervised machine learning often relies heavily on labeled examples. Humans may need to label thousands or millions of examples.

Traditional supervised learning uses human-labeled examples such as spam versus not spam
Email Label
Win $1 million today! Spam
Meeting at 3 PM Not Spam

Large Language Model pre-training works differently. Instead of manually labeling every sentence, enormous quantities of naturally occurring text can be used to teach the model language patterns.

This is one reason LLM training can scale to massive datasets. The shift is away from manually labeling millions of examples toward collecting huge amounts of text, training general-purpose models, and then applying smaller amounts of fine-tuning later.

The word “pre” matters

Why is it called “pre”-training?

The word pre is important. Pre-training happens before the model becomes the assistant we eventually interact with.

A model that has completed pre-training has learned a huge amount about language, patterns, concepts, facts, code, and relationships between ideas.

But that does not automatically mean it will behave like ChatGPT.

A raw pretrained model may be very good at continuing text, but being a useful conversational assistant requires additional training. This brings us to an important distinction.

Not one single step

Pre-training is only one stage

Modern language models are generally not trained in one single step. Training is often divided into three broad stages:

Stage 1 → Pre-training
Stage 2 → Mid-training
Stage 3 → Post-training

For understanding the P in GPT, pre-training is the most important one. But understanding the other stages helps explain why a raw pretrained model is different from something like ChatGPT.

Three training stages: pre-training builds the foundation, mid-training develops skills such as coding and math, and post-training shapes helpful assistant behavior.
Pre-training = broad education. Mid-training = skill development. Post-training = how to behave.
Stage 1

Pre-training

This is where the model learns broad language patterns and general knowledge from enormous datasets. Think of it as:

Building the model’s foundation.

During this stage, the model learns things such as vocabulary, grammar, general knowledge, basic reasoning patterns, relationships between concepts, and patterns found across different types of text.

At the end of pre-training, we have something much more capable than an untrained neural network. But we do not necessarily have a polished chatbot yet.

Stage 2

Mid-training

After broad pre-training, some models go through another stage using more carefully selected or specialized data.

This can improve specific capabilities such as coding, mathematics, long-document understanding, multilingual performance, and other targeted skills.

For example, training with high-quality programming data can significantly strengthen coding capabilities. Training with mathematical problems can improve mathematical reasoning. Training with long documents can help the model become better at working with larger contexts.

Pre-training = broad education
Mid-training = skill development
Stage 3

Post-training

Now imagine we have a model that knows a lot. That still does not guarantee that it will respond like a helpful assistant.

It needs to learn things such as: when a user asks a question, answer it clearly; follow the user’s instructions; format responses appropriately; avoid harmful responses; behave helpfully during a conversation.

This happens during post-training.

That stage often includes techniques such as Supervised Fine-Tuning (SFT) for instruction following and Reinforcement Learning from Human Feedback (RLHF) for shaping helpful behavior.

Pre-training teaches knowledge.
Mid-training strengthens skills.
Post-training shapes behavior.
A mental model

An easy school analogy

Imagine a student. During many years at school, the student reads textbooks, stories, science books, history, mathematics, newspapers, and programming material. This gives the student broad knowledge. Think of that as pre-training.

Later, the student chooses computer science and spends extra time learning programming, algorithms, operating systems, and databases. Think of that as mid-training.

Finally, suppose that person joins a customer-support organization and receives training on how to answer customers, how to communicate clearly, what they should and should not say, and how to respond appropriately. Think of that as post-training.

School years of broad reading map to pre-training, extra computer science study maps to mid-training, and customer-support training maps to post-training.
The analogy is not technically exact, but it gives us a useful mental model: learn broadly, improve specific skills, then learn how to behave.
Why data volume matters

Why does the amount of training data matter?

If data is so important, a natural question is: why don’t we simply build bigger and bigger models?

Because model size is only one part of the equation. A very large model without enough training data may not use its capacity efficiently.

This became especially clear through research commonly associated with the Chinchilla scaling laws.

The key idea was that, for a given compute budget, model size and the amount of training data need to be balanced. A Chinchilla-era rule of thumb is around:

20 training tokens per model parameter
Chinchilla-era guideline of about 20 training tokens per model parameter
Parameters Approx. training tokens
1 billion ~20 billion tokens
7 billion ~140 billion tokens
70 billion ~1.4 trillion tokens

The exact recipes used by modern models can differ, but the bigger lesson is more important:

Making the model bigger is not enough. You also need enough good training data.
An important insight

A smaller model can sometimes beat a bigger model

This was an important insight from Chinchilla.

GPT-3 had around 175 billion parameters and was trained on roughly 300 billion tokens. Chinchilla used around 70 billion parameters but trained on approximately 1.4 trillion tokens.

Despite having far fewer parameters, Chinchilla achieved stronger performance under its comparison setting.

GPT-3 had about 175 billion parameters and 300 billion training tokens. Chinchilla had about 70 billion parameters and 1.4 trillion training tokens, and performed more strongly in its comparison setting.
Model size alone does not determine intelligence or performance. How much data the model sees — and how effectively that data is prepared — matters enormously.
Quantity is not the whole story

But more data is not automatically better

Suppose you have two datasets.

Dataset A has 10 trillion tokens of duplicated pages, spam, and low-quality text. Dataset B has 5 trillion tokens of carefully filtered technical material, discussions, and code.
A larger dataset does not automatically mean a better model. Quality matters as much as quantity.

That means modern data pipelines perform a great deal of work before training ever begins. They may clean text, remove duplicates, filter poor-quality sources, balance different types of content, remove unwanted material, and decide how much data should come from each domain.

The LLaMA mix we saw earlier is a good example. Wikipedia is tiny on disk compared with CommonCrawl, but it is sampled more often per byte. The trainers were not treating every token as equally valuable.

This is one reason training data preparation is such an important part of modern AI development.

A competitive advantage

Why do AI companies talk so little about their data?

Something interesting happens when companies release information about Large Language Models. They may tell us quite a lot about number of layers, parameter count, attention architecture, tokenizer, Mixture-of-Experts design, and training infrastructure.

But detailed information about the training dataset is often much harder to find.

Why? Because the exact data mixture and preparation pipeline can be a competitive advantage.

Two companies could use very similar Transformer architectures but end up with models having very different capabilities because they used different data sources, filtering techniques, data mixtures, deduplication strategies, quality controls, and training stages.

The reasons for keeping data pipelines private often fall into three broad areas:

Competitive dynamics
Legal constraints
Trade secrets

The original LLaMA paper is a useful exception: it published the sampling mix in Table 1. Many later frontier models do not.

So when we talk about the intelligence of an LLM, architecture is only part of the story. The training data is another enormous part of it.

What “open” actually means

Open-Source vs Open-Weight Models

You will often hear AI models described as open source, but not every downloadable model is truly open source. An open-source AI model gives researchers much more than just the finished model. Ideally, it provides the model weights, training code, architecture, training recipes, and information about the data used to train it, so people can study how the model was created, modify it, and reproduce parts of the process. A good example is Ai2’s OLMo family, which releases its model weights, training code, training data, evaluation tools, checkpoints, and other details about the training process.

An open-weight model, on the other hand, mainly gives you access to the model’s trained weights. This means you can download the model, run it on your own machine or servers, and often fine-tune it for your own use. However, you may not receive the complete training dataset or everything required to recreate the model from scratch. Models from families such as Meta Llama are commonly described as open-weight models because their weights can be downloaded, but they use their own licenses and do not provide everything required to reproduce the original training process.

Putting the letters together

Connecting “G” and “P”

G → Generative. The model generates the response token by token.

P → Pre-trained. The model can make those predictions because it was trained beforehand on enormous amounts of data.

Think about the phrase “The capital of France is…” Why can the model predict “Paris”? Because during pre-training, relationships between concepts such as France, Paris, capital, and Europe were learned from enormous amounts of language data.

G means Generative: the model generates a response token by token. P means Pre-trained: it can make those predictions because it was trained beforehand on enormous amounts of data.
Pre-training creates the knowledge and language patterns. Generation uses those learned patterns to create the response.

Without training, there would be no useful next-token predictions.

After training

But where is the knowledge stored?

This leads to another important idea. After pre-training, what exactly contains everything the model learned?

The answer is: parameters.

A Large Language Model may contain billions of parameters. During training, these parameters are repeatedly adjusted.

Training data leads to predictions, prediction errors are measured, parameters are adjusted, and the loop repeats across enormous amounts of data.
Training modifies the parameters. Inference uses those parameters. The knowledge is not a stored webpage. It lives in the parameters.

This distinction is worth remembering because it connects pre-training directly to the generation process we covered in the previous section.

Two different jobs

Training vs inference

We can now clearly separate two concepts.

Training: the model is learning. Its parameters are being updated. This requires enormous amounts of data, computation, GPUs, time, and infrastructure.

Inference: the model is using what it already learned. Its parameters generally remain fixed while it processes your prompt and generates tokens.

Training means the model is learning and updating parameters. Inference means the model is using what it already learned, with parameters generally remaining fixed.
Training = Learn. Inference = Use what was learned. When you ask ChatGPT to explain Kubernetes, that is inference. The enormous pre-training process happened earlier. That’s the P in GPT.
Before you go

Bringing it back to the “P” in ChatGPT

We started with a simple question: What does the P in ChatGPT mean?

The answer is: P → Pre-trained.

Before the model interacts with you, it has already gone through a massive learning process using enormous amounts of text and other training data.

During pre-training, it learns vocabulary, grammar, language patterns, relationships between concepts, general knowledge, and patterns useful for reasoning.

Later training stages can strengthen particular capabilities and shape the model into something that behaves more like a helpful assistant. But the foundation comes from pre-training.

If you remember only one thing from this section, remember:

Pre-trained means the model learned from enormous amounts of data before you ever started chatting with it.

And now the name begins to make even more sense:

Chat → We interact with it conversationally.
G — Generative → It generates new responses one token at a time.
P — Pre-trained → It learned language, knowledge, and patterns before we started using it.

There is now only one major part of the name left:

T → Transformer

And that explains the architecture that makes much of this possible. When you are ready, ten questions cover Chat, Generative, Pre-trained, and Transformer together.