llm ai/books/designingllmapplication ai

2. Tokens and Embeddings

Core Idea

Subword tokenization is the compromise LLMs use: common words stay whole, unfamiliar words split into pieces the model already knows, so the vocabulary stays finite and sequences stay short.

  • Whole words break on new words (out-of-vocabulary); characters never break but make sequences far too long. Subwords avoid both.
  • Pipeline: raw text β†’ normalization β†’ pre-tokenization β†’ tokenization β†’ postprocessing β†’ token IDs β†’ embeddings.
  • BPE merges the most frequent adjacent pairs until the target size is reached; byte-level BPE (GPT family) starts from 256 bytes, so nothing needs <UNK>.
  • Token count drives KV cache memory at inference time.

Vocablulary and Tokens

  • A model’s vocabulary is the collection of units it can recognize.
  • Modern language models generally use tokens, rather than whole dictionary words.
  • Tokens can be:
    • Individual characters
    • Whole words
    • Subwords (parts of words)
    • Special symbols or sequences
  • Each token has a numerical token ID/index, which is mapped to an embedding representing its meaning and syntactic properties; these embeddings are processed by learned weights.
  • Tokens can be case-sensitive and may include the preceding space as part of the token.

Why not use words or characters?

Using only complete words creates an **out-of-vocabulary (OOV)** problem because language constantly produces new words, word forms, combinations, and domain-specific terminology.

Using individual characters avoids OOV problems but produces far more tokens. This makes sequences longer, reduces how much information fits into a fixed context length, and can increase training and inference costs.

Subword tokenization provides a compromise: unfamiliar words can be broken into smaller pieces while common words or word parts can remain as efficient tokens. It is the predominant approach described in the chapter.

Vocabulary size

A larger vocabulary generally means fewer tokens per piece of text, improving compression and allowing a model to process more text for the same compute. However, very large vocabularies contain more rare tokens, which may have poorly learned representations.

The chapter notes that optimal vocabulary size tends to increase with model size and compute.

What a tokenizer does

A tokenizer has two main roles:

  1. During tokenizer training, it processes text to help construct the vocabulary.
  2. During model training/inference, it converts raw text into tokens and token IDs that can be passed to the model.

Tokenization Pipeline

The tokenization pipeline Β· Hugging Face

Raw text β†’ normalization β†’ pre-tokenization β†’ tokenization β†’ postprocessing β†’ token IDs β†’ embeddings β†’ model

Normalization

Different types of normalization applied include:

  • Converting text to lowercase (if you are using an uncased model)
    • Not common, there are risk of loss some information.
  • Stripping off accents from characters, like from the word PeΓ±a
  • Unicode normalization

Pre-Tokenization

These are some optional steps before the tokenization.

A common step is to first perform word tokenization and then feed the output of it to the subword tokenization algorithm. This step is called pre-tokenization.

Tokenization

After the optional pre-tokenization step, the actual tokenization step is performed. Some of the important algorithms in this space are byte pair encoding (BPE), byte-level BPE, WordPiece, and Unigram LM.

The tokenizer comprises a set of rules that is learned during a pre-training phase over a pre-training dataset.

Example: BPE Training and Inference

Training stage:

  • The training text is first normalized and pre-tokenized.
  • The tokenizer identifies unique characters and creates the initial vocabulary.
  • It counts how frequently consecutive token pairs occur.
  • The most frequent pair is merged into a new token.
  • This process repeats, adding new subword tokens, until the desired vocabulary size is reached.
  • Example: a + p β†’ ap, because ap occurs most frequently; then a + t β†’ at.

Inference stage:

  • New input is first normalized and pre-tokenized.
  • The text is split into individual characters.
  • The learned merge rules are applied in order.
  • The resulting subword tokens are the final tokens passed to the language model.

Subword tokenization in practice

For example, FLAN-T5 does not necessarily have a token for a number such as 937; it can split it into "9" and "37". Likewise, a misspelled word can be broken into multiple subwords.

This means subword tokenization handles OOV words, but it can be brittle with typos. Larger models can nevertheless be more robust because their training data contains many naturally occurring misspellings.

Major tokenization algorithms

The chapter discusses several approaches:

  • BPE (Byte Pair Encoding): repeatedly merges the most frequent adjacent token pairs until the desired vocabulary size is reached.
  • Byte-level BPE: begins with 256 byte tokens, allowing arbitrary Unicode text to be represented without <UNK> tokens. GPT-family models use this approach according to the chapter.
  • WordPiece: similar to BPE but chooses merges using a likelihood-based score rather than raw frequency.
  • Unigram LM: listed as another important tokenization approach, though the chapter’s detailed discussion focuses mainly on BPE and WordPiece.

Special and unusual tokens

Tokenizers can contain special tokens such as:

  • <PAD> β€” padding
  • <EOS> β€” end of sequence
  • <UNK> β€” out-of-vocabulary term
  • Tool-call and tool-result markers

Glitch/Undertrained tokens: tokens that exist in a vocabulary but had little or no useful representation learned during model pre-training.

Tokenization-free models

The chapter briefly examines models that avoid conventional tokenization:

  • CANINE β€” uses Unicode codepoints and hashed embeddings.
  • ByT5 β€” operates on bytes.
  • Charformer β€” operates on bytes and learns latent subword representations.

Evaluating tokenizers

Two important metrics are introduced:

  • Fertility: average number of tokens needed per word. Higher fertility means poorer compression.
  • Parity: compares the number of tokens required to represent equivalent data in two languages.

A major issue is that tokenizers trained primarily on English can tokenize other languages inefficiently, sometimes requiring substantially more tokens for equivalent content. Domain-specific text can have the same problem when specialized terms are split into many pieces.

Domain-specific vocabularies

For specialized domains such as healthcare or scientific literature, new domain-specific tokens can be added to an existing tokenizer. However, adding the tokens alone is not enough: their embedding vectors initially contain no learned information. The model must undergo fine-tuning or continued pre-training so those new tokens acquire useful representations.