llm ai/books/designingllmapplication ai

1. Why pre-training data matters

The quality and composition of an LLM’s pre-training data strongly influence its downstream behavior. Understanding how the model was trained helps explain its strengths, weaknesses, biases, and failure modes.

An LLM can be thought of as having four major ingredients:

  1. Pre-training data — What information does it learn from?
  2. Vocabulary/tokenizer — How is text broken into tokens?
  3. Learning objective — What is the model trained to predict or accomplish?
  4. Architecture — How is the model internally structured?

2. Base models vs. specialized models

The language models trained using the process described in this chapter and the next are called base models. Lately, model providers have been augmenting the base model by fine-tuning it on much smaller datasets to steer them toward being more aligned with human needs and preferences.

LLM Post Training

3. Why LLMs need enormous amounts of data

Modern LLMs are highly data-hungry. Creating enough human-labeled examples is impractical, so pre-training primarily uses self-supervised learning(01 - Introduction > Supervised Learning), where the training data itself provides the learning signal.

LLMs are expected to learn many abilities from this data:

  • factual knowledge
  • language understanding
  • reasoning
  • mathematics
  • coding
  • creativity
  • handling ambiguity
    A major question is whether text alone contains enough information to learn all these abilities.

4. The grounding problem

Text describes the world but often leaves information unstated. Humans use:

  • common sense
  • world knowledge
  • context
  • emotional understanding
  • real-world experience

to fill in these gaps.

This is called grounding: connecting linguistic information to the real world.

Multimodal models that combine text, images, video, and speech may help with grounding, but it remains an open research question whether massive amounts of text alone can provide sufficient grounding.

5. Reasoning and ambiguity

LLMs can learn reasoning from examples such as:

  • mathematical proofs
  • explanations
  • puzzle solutions
  • step-by-step reasoning

But naturally occurring text contains relatively little explicit reasoning. Techniques such as Chain of Thought (CoT) and process supervision can help compensate.

LLMs also struggle with **ambiguity**, because human communication frequently leaves things unsaid. Simply increasing model and dataset size may not completely solve this problem.

6. Are we running out of training data?

Not yet. There are still enormous amounts of publicly available information that have not been incorporated into most training datasets, including:

  • court judgments
  • parliamentary proceedings
  • government documents
  • SEC filings
  • other specialized sources

However, at sufficiently large scales, naturally occurring human-written data may eventually become insufficient.

This has led to growing interest in synthetic data, where LLMs generate training examples for other models. Synthetic data can increase available training material, but excessive use can cause models to drift away from the true distribution of human-generated data.

7. Quality matters more than simply increasing quantity

High-quality data can make models more sample-efficient, meaning they can learn more from fewer examples.

Therefore, training data is carefully:

  • cleaned
  • filtered
  • deduplicated
  • quality-ranked
  • balanced across domains
  • checked for privacy issues
  • checked for benchmark contamination

The chapter emphasizes that data preprocessing may be one of the most important parts of LLM development.

8. Multiple epochs

An epoch means the model has seen the entire training dataset once.

Info

In machine learning, an epoch is one complete pass through the entire training dataset.

For example, if you have 10,000 training examples:

  • 1 epoch = the model sees all 10,000 examples once.
  • 5 epochs = the model sees the dataset 5 times.
  • During each epoch, the model updates its parameters based on the training examples.

Simple analogy

Imagine studying 100 flashcards:

  • Studying all 100 once = 1 epoch
  • Going through all 100 five times = 5 epochs

Important: More epochs don’t always mean a better model. Too many can cause overfitting, where the model learns the training data too closely and performs worse on new data.

LLMs can generally reuse high-quality data for several epochs—research suggests roughly 4–5 epochs can sometimes be useful without significant performance degradation.

However:

  • repeated exposure eventually provides diminishing returns
  • larger models can overfit more easily
  • high-quality data is limited

Therefore, training cannot simply rely on repeatedly showing the same small collection of excellent data.

9. Major pre-training datasets

Some important general-purpose datasets include:

DatasetMain sourceKey idea
C4Common CrawlLarge cleaned web dataset
The PileWeb + books + academic + code + forums, etc.High diversity
WebText/OpenWebTextReddit-linked webpagesAttempts to select higher-quality web content
WikipediaWikipediaStructured factual knowledge
BooksCorpusBooksLong-form literary text
FineWebCommon CrawlVery large, aggressively cleaned dataset
RedPajamaMultiple public sourcesOpen reproduction-oriented dataset
RefinedWebCommon CrawlLarge cleaned web corpus
ROOTSMultiple sourcesDataset used for BLOOM

A key lesson is that many LLMs use overlapping sources, particularly web data.

10. Training data is disappearing

Access to high-quality web data is becoming harder because of:

  • copyright restrictions
  • paywalls
  • terms of service
  • robots.txt restrictions
  • websites charging for AI data access

Some historically important datasets are no longer publicly available in their original form.
This creates a growing tension between AI development, copyright, data ownership, and access to information.

11. Synthetic data

LLMs can generate training data for other LLMs.

Examples include:

  • Microsoft Phi: used large amounts of synthetic data.
  • Cosmopedia: Hugging Face’s synthetic dataset using sources such as educational materials and web data.

Synthetic data can provide:

  • more training examples
  • targeted educational content
  • reasoning examples
  • greater diversity

But it can also introduce factual errors, reasoning errors, and distributional drift.

12. Data preprocessing

The typical preprocessing pipeline involves:

Raw data → extraction/cleaning → language filtering → spam/toxicity filtering → quality filtering → deduplication → PII removal → decontamination → data mixing → ordering

A. Boilerplate removal

Web pages contain lots of useless material:

  • navigation menus
  • advertisements
  • headers/footers
  • repeated website elements
  • HTML artifacts

Tools such as jusText, Dragnet, Trafilatura, html2text, and Newspaper can help extract meaningful text.

Interestingly, researchers have also experimented with keeping some HTML structure because tags can contain useful information.

B. Language identification

Datasets may try to remove unwanted languages using tools such as: langdetect, langid, FastText, pycld2

But language detection is imperfect, especially with code-switching.
For example, an “English-only” dataset can still contain substantial amounts of other languages.
This helps explain why some supposedly monolingual LLMs can unexpectedly perform well in other languages.

C. Spam and low-quality filtering

Common filtering techniques remove:

  • SEO spam
  • keyword stuffing
  • repetitive text
  • very short documents
  • meaningless boilerplate
  • pornography
  • abusive/toxic material
  • malformed documents

Simple heuristics are efficient but can introduce false positives and bias.

13. Quality filtering

Several approaches are used to identify high-quality documents.

K-L divergence

Documents whose token distributions are very different from a reference high-quality dataset can be removed.

Classifiers

A classifier can be trained using:

  • high-quality sources → positive examples
  • random web data → negative examples

The classifier then identifies high-quality documents.

Perplexity

A language model can assign a perplexity score to text.

  • Lower perplexity → text is more predictable to the reference model.
  • Higher perplexity → text is less similar to the reference style.

However, low perplexity does not necessarily mean high quality. Repetitive or simplistic text can also have low perplexity.

14. Deduplication

Web datasets contain enormous amounts of duplicated content.

Three types are identified:

  1. Exact duplicates — identical text.
  2. Approximate duplicates — nearly identical text.
  3. Semantic duplicates — same meaning expressed differently.

Deduplication is important because it:

  • reduces dataset size
  • prevents train/test overlap
  • reduces memorization
  • reduces overfitting
  • makes evaluation more reliable

Techniques such as MinHash are commonly used.

An important result is that sequence-level deduplication can dramatically reduce a model’s tendency to reproduce training data verbatim.

15. Memorization and privacy

LLMs can memorize parts of their training data.
This creates two important attack types:

Membership inference

An attacker tries to determine whether a particular piece of text appeared in the training data.

Training-data extraction

An attacker tries to make the model reproduce memorized information.
This is particularly concerning when the training data contains personally identifiable information (PII).

Examples of PII include:

  • names
  • addresses
  • phone numbers
  • emails
  • government IDs
  • credit card information
  • medical information
  • precise location information

16. PII remediation

Possible approaches include:

  • replacing PII with special tokens
  • replacing it with realistic fake information
  • shuffling entities
  • removing documents containing excessive PII

However, removing PII is difficult because privacy is context-dependent.

Something publicly available isn’t necessarily appropriate to expose through an LLM. A piece of information may have been shared in a limited context but become much more discoverable when included in a massive training dataset.

Differential privacy

Another approach is differential privacy, which adds mathematical privacy guarantees through techniques such as DP-SGD.

The tradeoff is that differential privacy can:

  • slow training
  • reduce model performance
  • affect some groups disproportionately

Other approaches include model unlearning, adversarial training, and specialized decoding techniques.

17. Training-data poisoning

Because much LLM training data comes from the web, attackers may deliberately insert content into it.

This is called data poisoning.

Even a relatively small amount of poisoned data can potentially influence model behavior or make other information easier to extract.

This also creates an incentive for LLM SEO: organizations may attempt to create web content that is more likely to survive dataset filtering and influence future models.

18. Benchmark contamination

A model’s evaluation can become misleading if benchmark questions were already present in its training data.

Two forms are:

  • Input + label contamination: both questions and answers appear in training.
  • Input contamination: only the questions appear.

Common methods such as n-gram matching can detect obvious overlap, but contamination can also occur through:

  • paraphrasing
  • translation
  • modified versions of benchmark questions

Therefore, benchmark scores may sometimes overestimate a model’s true generalization ability.

19. Data mixtures

LLM datasets combine different domains in specific proportions.

For example, the chapter gives a Llama 3 mixture approximately consisting of:

  • 50% general knowledge
  • 25% mathematics/reasoning
  • 17% code
  • 8% non-English data

The exact mixture matters because different domains contribute different capabilities.
Interestingly, code in the training data can improve non-coding tasks as well.

20. Dataset ordering

The order in which examples are presented can affect learning.

This is known as curriculum learning.

Possible strategies include:

  • starting with shorter sequences
  • gradually increasing sequence length
  • introducing common words before rare words

However, sophisticated curriculum strategies are not yet standard across most LLM training.

21. Pre-training data affects downstream performance

One of the chapter’s most important conclusions is:

What appears frequently in the training data tends to be easier for the model.

Models perform better on tasks that resemble patterns frequently encountered during pre-training.

For example, they may perform better on:

  • common arithmetic operations than unusual ones
  • normal alphabetical sorting than reverse sorting
  • common facts than obscure facts
  • high-frequency phrases than unusual phrases

So, training-data frequency can predict downstream performance.

22. Bias and fairness

Pre-training data doesn’t just teach language—it also teaches the model something about the world and society.

Internet data contains: stereotypes, discrimination, hate, cultural biases, unequal representation

Models can also amplify these biases rather than merely reproduce them.

Bias can enter at multiple stages:

  • which sources are selected
  • which sources are excluded
  • how data is filtered
  • which languages are represented
  • which communities are represented
  • which words are considered “toxic”

For example, simplistic keyword filtering can accidentally remove legitimate content from minority communities or dialects.

Therefore, data cleaning itself is not neutral. Every filtering and selection decision embeds assumptions and values.