ai llm agents ai/books/designingllmapplication

Different types of LLMs

The chapter distinguishes several model variants:

  • Base models: Original pretrained models without additional task-specific tuning.
  • Instruction-tuned models: Trained to follow natural-language instructions more effectively.
  • Chat models: Instruction-tuned models optimized for multi-turn conversations.
  • Long-context models: Designed to process much larger amounts of text.
  • Domain/task-adapted models: Fine-tuned for specific applications such as financial analysis or named-entity recognition.

Instruction tuning generally improves usefulness, but it can also cause regressions in some capabilities, such as reasoning. Alignment tuning can also make a model reflect the provider’s values rather than those of your organization.

FLAN Dataset

The FLAN Dataset is a collection of datasets used to fine-tune language models to follow natural-language instructions. “FLAN” stands for Fine-tuned Language Net.

Instead of training a model only to predict the next word, FLAN-style training gives the model examples like:

Instruction: Translate “Good morning” into French.
Answer: Bonjour.

Open-source vs. proprietary LLMs

A truly open LLM should ideally provide:

  1. Model Weight
  2. Model code
  3. Training data
    However, most so-called open models do not publish their complete training datasets. Copyright, licensing, and competitive concerns are major reasons.

Model Weight

An open-weight model is an AI model where the trained model parameters (weights) are publicly available, so people can download and run the model themselves.

Open-weight ≠ fully open-source

Open-weightFully open-source
Model weights availableYesYes
Can download modelUsually yesYes
Architecture availableUsuallyYes
Training code availableNot necessarilyUsually
Training dataset availableNot necessarilyIdeally
Can inspect everythingNoMuch more
License restrictionsMay existDepends on license

License Types

Common license types include:

  • Noncommercial: Research/personal use only.
  • Copyleft: Commercial use is allowed, but derivatives must use the same license.
  • Permissive: Allows commercial use and proprietary modifications.
  • Restricted licenses: Allow usage but prohibit particular applications.

For commercial applications, the chapter recommends carefully checking the license and generally preferring permissive licenses such as Apache 2.0 or MIT when appropriate.

How to choose an LLM

The choice should depend on your specific application, not simply which model is currently at the top of a leaderboard.

Important criteria include:

  • Cost
  • Speed / time per output token (TPOT)
  • Task performance
  • Task type
  • Required capabilities, such as reasoning or planning
  • Licensing
  • Available ML/MLOps expertise
  • Safety, security, and privacy

The chapter emphasizes that open-source models provide flexibility and transparency, but self-hosting can require substantial computing resources and engineering effort.

Evaluating LLMs

LLM evaluation is difficult because existing benchmarks can be incomplete, easily gamed, contaminated, or misleading.

Examples of evaluation frameworks include:

  • LM Evaluation Harness: Tests hundreds of different tasks.
  • MMLU: Tests knowledge across many subjects.
  • ARC: Tests science reasoning.
  • HellaSwag: Tests commonsense reasoning.
  • TruthfulQA: Tests truthfulness.
  • Winogrande: Tests commonsense reasoning.
  • GSM8K: Tests mathematical reasoning.
  • HELM: Evaluates many dimensions including accuracy, robustness, fairness, toxicity, efficiency, and summarization.

Key lesson: Don’t blindly trust benchmark leaderboards. Models can be optimized specifically for benchmarks, and different evaluation methods can produce very different results. The chapter recommends creating internal benchmarks based on your actual use case.

Human evaluation

The Elo rating system can be used to compare models through human preferences, as in Chatbot Arena.

However, human evaluation also has biases. People may:

  • Prefer longer answers.
  • Be influenced by confident or authoritative writing.
  • Choose ties too often.
  • Be influenced by answer order.

Therefore, human ratings also need to be interpreted carefully.

Loading and running LLMs

Running LLMs efficiently often requires GPUs and careful inference optimization.

Memory requirements depend on numerical precision and:

  • Model size
  • Numerical precision
  • Whether you’re doing inference or training

For example, a 7B model needs roughly 7 GB GPU RAM in 8-bit mode and 14 GB in BF16 for inference; full fine-tuning requires considerably more memory.

Tools such as Hugging Face Accelerate can move parts of a model between GPU, CPU, and disk when the model doesn’t fit entirely in GPU memory. Ollama provides another convenient way to run models locally.

Decoding strategies

Decoding determines how an LLM chooses its next token.

MethodMain ideaMain issue
GreedyAlways choose highest-probability tokenCan become repetitive
Beam searchKeep several high-probability sequencesOften still sounds unnatural
Top-kSample from the k most probable tokensFixed k can include unsuitable tokens
Top-pDynamically sample from tokens making up probability pMore flexible and widely used

Top-p, also called nucleus sampling, is described as the most popular sampling strategy among these approaches.

Reliability and nondeterminism

LLM outputs can vary even when the same prompt is used because generation is probabilistic.

Lowering temperature makes outputs more predictable but can reduce creativity. In production, the chapter recommends generating multiple answers and using techniques such as majority voting or self-consistency when reliability matters.

Structured outputs

LLMs can be instructed to produce structured formats such as JSON.

Methods include:

  • JSON schemas
  • Jsonformer
  • LMQL
  • Guidance
  • Regular expressions
  • Context-free grammars (CFGs)

These techniques are useful when LLM output must be consumed by another software system.

Interpretability and debugging

LLM interpretability is still developing. One practical approach is to examine how outputs change when inputs are slightly modified and to inspect intermediate model behavior.

LIT-NLP can help with:

  • Attention visualization
  • Salience maps
  • Embedding visualization
  • Counterfactual analysis

The chapter also introduces mechanistic interpretability, where researchers study neurons and combinations of neuron activations called features to better understand model behavior.