llm ai/books/designingllmapplication ai
Tldr
This chapter argues that an LLM’s capabilities, limitations, privacy risks, biases, and downstream performance are deeply shaped by what data it was trained on and, just as importantly, how that data was collected, cleaned, filtered, deduplicated, mixed, and ordered.
1. Why pre-training data matters
The quality and composition of an LLM’s pre-training data strongly influence its downstream behavior. Understanding how the model was trained helps explain its strengths, weaknesses, biases, and failure modes.
An LLM can be thought of as having four major ingredients:
- Pre-training data — What information does it learn from?
- Vocabulary/tokenizer — How is text broken into tokens?
- Learning objective — What is the model trained to predict or accomplish?
- Architecture — How is the model internally structured?

2. Base models vs. specialized models
The language models trained using the process described in this chapter and the next are called base models. Lately, model providers have been augmenting the base model by fine-tuning it on much smaller datasets to steer them toward being more aligned with human needs and preferences.
3. Why LLMs need enormous amounts of data
Modern LLMs are highly data-hungry. Creating enough human-labeled examples is impractical, so pre-training primarily uses self-supervised learning(01 - Introduction > Supervised Learning), where the training data itself provides the learning signal.
LLMs are expected to learn many abilities from this data:
- factual knowledge
- language understanding
- reasoning
- mathematics
- coding
- creativity
- handling ambiguity
A major question is whether text alone contains enough information to learn all these abilities.
4. The grounding problem
Text describes the world but often leaves information unstated. Humans use:
- common sense
- world knowledge
- context
- emotional understanding
- real-world experience
to fill in these gaps.
This is called grounding: connecting linguistic information to the real world.
Multimodal models that combine text, images, video, and speech may help with grounding, but it remains an open research question whether massive amounts of text alone can provide sufficient grounding.
5. Reasoning and ambiguity
LLMs can learn reasoning from examples such as:
- mathematical proofs
- explanations
- puzzle solutions
- step-by-step reasoning
But naturally occurring text contains relatively little explicit reasoning. Techniques such as Chain of Thought (CoT) and process supervision can help compensate.
LLMs also struggle with **ambiguity**, because human communication frequently leaves things unsaid. Simply increasing model and dataset size may not completely solve this problem.
6. Are we running out of training data?
Not yet. There are still enormous amounts of publicly available information that have not been incorporated into most training datasets, including:
- court judgments
- parliamentary proceedings
- government documents
- SEC filings
- other specialized sources
However, at sufficiently large scales, naturally occurring human-written data may eventually become insufficient.
This has led to growing interest in synthetic data, where LLMs generate training examples for other models. Synthetic data can increase available training material, but excessive use can cause models to drift away from the true distribution of human-generated data.
7. Quality matters more than simply increasing quantity
High-quality data can make models more sample-efficient, meaning they can learn more from fewer examples.
Therefore, training data is carefully:
- cleaned
- filtered
- deduplicated
- quality-ranked
- balanced across domains
- checked for privacy issues
- checked for benchmark contamination
The chapter emphasizes that data preprocessing may be one of the most important parts of LLM development.
8. Multiple epochs
An epoch means the model has seen the entire training dataset once.
Info
In machine learning, an epoch is one complete pass through the entire training dataset.
For example, if you have 10,000 training examples:
- 1 epoch = the model sees all 10,000 examples once.
- 5 epochs = the model sees the dataset 5 times.
- During each epoch, the model updates its parameters based on the training examples.
Simple analogy
Imagine studying 100 flashcards:
- Studying all 100 once = 1 epoch
- Going through all 100 five times = 5 epochs
Important: More epochs don’t always mean a better model. Too many can cause overfitting, where the model learns the training data too closely and performs worse on new data.
LLMs can generally reuse high-quality data for several epochs—research suggests roughly 4–5 epochs can sometimes be useful without significant performance degradation.
However:
- repeated exposure eventually provides diminishing returns
- larger models can overfit more easily
- high-quality data is limited
Therefore, training cannot simply rely on repeatedly showing the same small collection of excellent data.
9. Major pre-training datasets
Some important general-purpose datasets include:
| Dataset | Main source | Key idea |
|---|---|---|
| C4 | Common Crawl | Large cleaned web dataset |
| The Pile | Web + books + academic + code + forums, etc. | High diversity |
| WebText/OpenWebText | Reddit-linked webpages | Attempts to select higher-quality web content |
| Wikipedia | Wikipedia | Structured factual knowledge |
| BooksCorpus | Books | Long-form literary text |
| FineWeb | Common Crawl | Very large, aggressively cleaned dataset |
| RedPajama | Multiple public sources | Open reproduction-oriented dataset |
| RefinedWeb | Common Crawl | Large cleaned web corpus |
| ROOTS | Multiple sources | Dataset used for BLOOM |
A key lesson is that many LLMs use overlapping sources, particularly web data.
10. Training data is disappearing
Access to high-quality web data is becoming harder because of:
- copyright restrictions
- paywalls
- terms of service
- robots.txt restrictions
- websites charging for AI data access
Some historically important datasets are no longer publicly available in their original form.
This creates a growing tension between AI development, copyright, data ownership, and access to information.
11. Synthetic data
LLMs can generate training data for other LLMs.
Examples include:
- Microsoft Phi: used large amounts of synthetic data.
- Cosmopedia: Hugging Face’s synthetic dataset using sources such as educational materials and web data.
Synthetic data can provide:
- more training examples
- targeted educational content
- reasoning examples
- greater diversity
But it can also introduce factual errors, reasoning errors, and distributional drift.
12. Data preprocessing
The typical preprocessing pipeline involves:

Raw data → extraction/cleaning → language filtering → spam/toxicity filtering → quality filtering → deduplication → PII removal → decontamination → data mixing → ordering
A. Boilerplate removal
Web pages contain lots of useless material:
- navigation menus
- advertisements
- headers/footers
- repeated website elements
- HTML artifacts
Tools such as jusText, Dragnet, Trafilatura, html2text, and Newspaper can help extract meaningful text.
Interestingly, researchers have also experimented with keeping some HTML structure because tags can contain useful information.
B. Language identification
Datasets may try to remove unwanted languages using tools such as: langdetect, langid, FastText, pycld2
But language detection is imperfect, especially with code-switching.
For example, an “English-only” dataset can still contain substantial amounts of other languages.
This helps explain why some supposedly monolingual LLMs can unexpectedly perform well in other languages.
C. Spam and low-quality filtering
Common filtering techniques remove:
- SEO spam
- keyword stuffing
- repetitive text
- very short documents
- meaningless boilerplate
- pornography
- abusive/toxic material
- malformed documents
Simple heuristics are efficient but can introduce false positives and bias.
13. Quality filtering
Several approaches are used to identify high-quality documents.
K-L divergence
Documents whose token distributions are very different from a reference high-quality dataset can be removed.
Classifiers
A classifier can be trained using:
- high-quality sources → positive examples
- random web data → negative examples
The classifier then identifies high-quality documents.

Perplexity
A language model can assign a perplexity score to text.
- Lower perplexity → text is more predictable to the reference model.
- Higher perplexity → text is less similar to the reference style.
However, low perplexity does not necessarily mean high quality. Repetitive or simplistic text can also have low perplexity.
14. Deduplication
Web datasets contain enormous amounts of duplicated content.
Three types are identified:
- Exact duplicates — identical text.
- Approximate duplicates — nearly identical text.
- Semantic duplicates — same meaning expressed differently.
Deduplication is important because it:
- reduces dataset size
- prevents train/test overlap
- reduces memorization
- reduces overfitting
- makes evaluation more reliable
Techniques such as MinHash are commonly used.
An important result is that sequence-level deduplication can dramatically reduce a model’s tendency to reproduce training data verbatim.
15. Memorization and privacy
LLMs can memorize parts of their training data.
This creates two important attack types:
Membership inference
An attacker tries to determine whether a particular piece of text appeared in the training data.
Training-data extraction
An attacker tries to make the model reproduce memorized information.
This is particularly concerning when the training data contains personally identifiable information (PII).

Examples of PII include:
- names
- addresses
- phone numbers
- emails
- government IDs
- credit card information
- medical information
- precise location information
16. PII remediation
Possible approaches include:
- replacing PII with special tokens
- replacing it with realistic fake information
- shuffling entities
- removing documents containing excessive PII
However, removing PII is difficult because privacy is context-dependent.

Something publicly available isn’t necessarily appropriate to expose through an LLM. A piece of information may have been shared in a limited context but become much more discoverable when included in a massive training dataset.

Differential privacy
Another approach is differential privacy, which adds mathematical privacy guarantees through techniques such as DP-SGD.
The tradeoff is that differential privacy can:
- slow training
- reduce model performance
- affect some groups disproportionately
Other approaches include model unlearning, adversarial training, and specialized decoding techniques.
17. Training-data poisoning
Because much LLM training data comes from the web, attackers may deliberately insert content into it.
This is called data poisoning.
Even a relatively small amount of poisoned data can potentially influence model behavior or make other information easier to extract.
This also creates an incentive for LLM SEO: organizations may attempt to create web content that is more likely to survive dataset filtering and influence future models.
18. Benchmark contamination
A model’s evaluation can become misleading if benchmark questions were already present in its training data.
Two forms are:
- Input + label contamination: both questions and answers appear in training.
- Input contamination: only the questions appear.
Common methods such as n-gram matching can detect obvious overlap, but contamination can also occur through:
- paraphrasing
- translation
- modified versions of benchmark questions
Therefore, benchmark scores may sometimes overestimate a model’s true generalization ability.
19. Data mixtures
LLM datasets combine different domains in specific proportions.
For example, the chapter gives a Llama 3 mixture approximately consisting of:
- 50% general knowledge
- 25% mathematics/reasoning
- 17% code
- 8% non-English data
The exact mixture matters because different domains contribute different capabilities.
Interestingly, code in the training data can improve non-coding tasks as well.
20. Dataset ordering
The order in which examples are presented can affect learning.
This is known as curriculum learning.
Possible strategies include:
- starting with shorter sequences
- gradually increasing sequence length
- introducing common words before rare words
However, sophisticated curriculum strategies are not yet standard across most LLM training.
21. Pre-training data affects downstream performance
One of the chapter’s most important conclusions is:
What appears frequently in the training data tends to be easier for the model.
Models perform better on tasks that resemble patterns frequently encountered during pre-training.
For example, they may perform better on:
- common arithmetic operations than unusual ones
- normal alphabetical sorting than reverse sorting
- common facts than obscure facts
- high-frequency phrases than unusual phrases
So, training-data frequency can predict downstream performance.
22. Bias and fairness
Pre-training data doesn’t just teach language—it also teaches the model something about the world and society.
Internet data contains: stereotypes, discrimination, hate, cultural biases, unequal representation
Models can also amplify these biases rather than merely reproduce them.
Bias can enter at multiple stages:
- which sources are selected
- which sources are excluded
- how data is filtered
- which languages are represented
- which communities are represented
- which words are considered “toxic”
For example, simplistic keyword filtering can accidentally remove legitimate content from minority communities or dialects.
Therefore, data cleaning itself is not neutral. Every filtering and selection decision embeds assumptions and values.
Related
- 1. An Introduction to LLM: Language Modeling (Pretraining) — where this chapter’s data becomes a model
- 1. Introduction to LLMs — the big picture this chapter sits inside