ai ai/agenticai llm ai/books/designingllmapplication

Why fine-tuning is needed

Few-shot prompting can fail when a task is complex or requires extensive domain knowledge. Continuously adding instructions makes prompts too long and difficult for the model to follow. Fine-tuning solves this by updating the model’s weights using input-output examples, allowing it to learn a particular task or behavior.

Fine-tuning is particularly useful for:

  • Learning a specific input → output mapping
  • Adapting to a new textual/domain style
  • Developing more complex capabilities or behaviors

It is not the best choice for adding new or frequently changing factual knowledge; techniques such as 12. Retrieval-Augmented Generation (RAG) are better suited for that. Fine-tuning can also cause regression in the model’s existing capabilities, so the resulting model should be tested carefully.

Example: Political promise detector

The chapter demonstrates fine-tuning a Llama 2 7B model to determine whether a statement is a political promise.

A promise must be:

  • Tangible
  • Specific
  • Something the government has the agency to accomplish

For example, promising to build 10,000 km of subway lines qualifies, while predictions, vague statements, or things outside government control do not.

Important fine-tuning techniques and hyperparameters

The chapter covers several groups of training parameters:

Optimization

  • AdamW and Adafactor are common optimizers.
  • 8-bit AdamW can dramatically reduce memory requirements.
  • Paged optimizers can move memory between GPU and CPU when GPU memory is insufficient.
  • Learning rate and weight decay strongly affect training.

Learning-rate scheduling

  • Constant
  • Constant with warmup
  • Cosine
  • Cosine with restarts
  • Linear

Warmup is particularly useful with AdamW, and the chapter notes that cosine annealing can outperform linear decay.

Memory optimization

  • Gradient checkpointing: saves memory by recomputing activations.
  • Gradient accumulation: simulates larger batch sizes by accumulating gradients across multiple batches.
  • Quantization: reduces memory usage while maintaining reasonable performance.

Regularization

  • Label smoothing helps prevent overfitting and improves calibration.
  • Noise embeddings / NEFTune reduce overfitting to the wording and formatting of a small dataset.

Batch size
Larger batches are faster but require more memory and can sometimes contribute to overfitting. The chapter recommends balancing batch size, learning rate, memory, and convergence behavior.

Parameter-Efficient Fine-Tuning (PEFT)

Instead of updating every model parameter, PEFT updates only a small subset, greatly reducing computational and memory requirements while retaining much of the performance of full fine-tuning.

The example uses LoRA (Low-Rank Adaptation) with:

  • r = 64
  • lora_alpha = 8
  • lora_dropout = 0.1

The example also uses 4-bit quantization (NF4) through bitsandbytes to further reduce memory usage.

Complete fine-tuning setup

The example combines:

  • Llama 2 7B
  • PEFT/LoRA
  • 4-bit quantization
  • paged AdamW
  • cosine learning-rate scheduling
  • gradient accumulation
  • gradient checkpointing
  • BF16
  • label smoothing
  • NEFTune
  • batch size of 8
  • 3 training epochs

The key lesson is that hyperparameters interact in complex ways, so experimentation is necessary. However, the chapter emphasizes that improving the quality of training data is generally more valuable than endlessly optimizing the final few percentage points of model performance.

Fine-tuning datasets

A good fine-tuning dataset generally contains:

  1. Instruction — explains the task and desired output.
  2. Input — the text/data the model must process.
  3. Output — the correct response.

Datasets can be:

  • Single-task — focused on one particular task.
  • Multi-task — used for broader instruction tuning.

Instruction tuning helps bridge the gap between the objectives used during LLM pre-training and the way humans actually use LLMs. It also makes model behavior more controllable and helps it learn desired output formats.

Creating instruction-tuning datasets

The chapter describes three main approaches:

  • Use existing public instruction-tuning datasets.
  • Convert traditional datasets into instruction format.
  • Create high-quality seed examples manually and use LLMs to generate additional examples.

Popular datasets discussed include FLAN, OIG, P3, Natural Instructions, Self-Instruct, Evol-Instruct, Alpaca, Guanaco, Vicuna, and OpenAssistant.

Using LLMs to generate training data

LLMs can generate synthetic instruction datasets from a small set of high-quality examples. Two important approaches are:

  • Input-first: generate an input, then generate its label/output.
  • Output-first: generate the desired output/label first, then generate an input that matches it.

Combining both approaches can help avoid label imbalance.

However, synthetic generation can produce irrelevant or low-quality examples. The dataset therefore needs quality filtering and careful control of its distribution.

The chapter highlights Evol-Instruct, which improves seed instructions through:

  1. Instruction evolution
  2. Response generation
  3. Candidate filtering
    Evolution can increase constraints, reasoning requirements, specificity, or input complexity, or broaden topic coverage.