ai llm agents ai/books/designingllmapplication
Tldr
Choosing an LLM is not simply about finding the biggest or highest-ranked model.
You should select a model based on your task, required capabilities, cost, speed, licensing, infrastructure, safety, and evaluation results. Then test it using your own representative data, carefully evaluate reliability, choose an appropriate decoding strategy, and use interpretability/debugging tools to understand failures.
Different types of LLMs
The chapter distinguishes several model variants:
- Base models: Original pretrained models without additional task-specific tuning.
- Instruction-tuned models: Trained to follow natural-language instructions more effectively.
- Chat models: Instruction-tuned models optimized for multi-turn conversations.
- Long-context models: Designed to process much larger amounts of text.
- Domain/task-adapted models: Fine-tuned for specific applications such as financial analysis or named-entity recognition.
Instruction tuning generally improves usefulness, but it can also cause regressions in some capabilities, such as reasoning. Alignment tuning can also make a model reflect the provider’s values rather than those of your organization.
FLAN Dataset
The FLAN Dataset is a collection of datasets used to fine-tune language models to follow natural-language instructions. “FLAN” stands for Fine-tuned Language Net.
Instead of training a model only to predict the next word, FLAN-style training gives the model examples like:
Instruction: Translate “Good morning” into French.
Answer: Bonjour.

Open-source vs. proprietary LLMs
A truly open LLM should ideally provide:
- Model Weight
- Model code
- Training data
However, most so-called open models do not publish their complete training datasets. Copyright, licensing, and competitive concerns are major reasons.
Model Weight
An open-weight model is an AI model where the trained model parameters (weights) are publicly available, so people can download and run the model themselves.
Open-weight ≠fully open-source
| Open-weight | Fully open-source | |
|---|---|---|
| Model weights available | Yes | Yes |
| Can download model | Usually yes | Yes |
| Architecture available | Usually | Yes |
| Training code available | Not necessarily | Usually |
| Training dataset available | Not necessarily | Ideally |
| Can inspect everything | No | Much more |
| License restrictions | May exist | Depends on license |
License Types
Common license types include:
- Noncommercial: Research/personal use only.
- Copyleft: Commercial use is allowed, but derivatives must use the same license.
- Permissive: Allows commercial use and proprietary modifications.
- Restricted licenses: Allow usage but prohibit particular applications.
For commercial applications, the chapter recommends carefully checking the license and generally preferring permissive licenses such as Apache 2.0 or MIT when appropriate.
How to choose an LLM
The choice should depend on your specific application, not simply which model is currently at the top of a leaderboard.
Important criteria include:
- Cost
- Speed / time per output token (TPOT)
- Task performance
- Task type
- Required capabilities, such as reasoning or planning
- Licensing
- Available ML/MLOps expertise
- Safety, security, and privacy
The chapter emphasizes that open-source models provide flexibility and transparency, but self-hosting can require substantial computing resources and engineering effort.
Evaluating LLMs
LLM evaluation is difficult because existing benchmarks can be incomplete, easily gamed, contaminated, or misleading.
Examples of evaluation frameworks include:
- LM Evaluation Harness: Tests hundreds of different tasks.
- MMLU: Tests knowledge across many subjects.
- ARC: Tests science reasoning.
- HellaSwag: Tests commonsense reasoning.
- TruthfulQA: Tests truthfulness.
- Winogrande: Tests commonsense reasoning.
- GSM8K: Tests mathematical reasoning.
- HELM: Evaluates many dimensions including accuracy, robustness, fairness, toxicity, efficiency, and summarization.
Key lesson: Don’t blindly trust benchmark leaderboards. Models can be optimized specifically for benchmarks, and different evaluation methods can produce very different results. The chapter recommends creating internal benchmarks based on your actual use case.
Human evaluation
The Elo rating system can be used to compare models through human preferences, as in Chatbot Arena.
However, human evaluation also has biases. People may:
- Prefer longer answers.
- Be influenced by confident or authoritative writing.
- Choose ties too often.
- Be influenced by answer order.
Therefore, human ratings also need to be interpreted carefully.
Loading and running LLMs
Running LLMs efficiently often requires GPUs and careful inference optimization.
Memory requirements depend on numerical precision and:
- Model size
- Numerical precision
- Whether you’re doing inference or training
For example, a 7B model needs roughly 7 GB GPU RAM in 8-bit mode and 14 GB in BF16 for inference; full fine-tuning requires considerably more memory.
Tools such as Hugging Face Accelerate can move parts of a model between GPU, CPU, and disk when the model doesn’t fit entirely in GPU memory. Ollama provides another convenient way to run models locally.
Decoding strategies
Decoding determines how an LLM chooses its next token.
| Method | Main idea | Main issue |
|---|---|---|
| Greedy | Always choose highest-probability token | Can become repetitive |
| Beam search | Keep several high-probability sequences | Often still sounds unnatural |
| Top-k | Sample from the k most probable tokens | Fixed k can include unsuitable tokens |
| Top-p | Dynamically sample from tokens making up probability p | More flexible and widely used |
Top-p, also called nucleus sampling, is described as the most popular sampling strategy among these approaches.
Reliability and nondeterminism
LLM outputs can vary even when the same prompt is used because generation is probabilistic.
Lowering temperature makes outputs more predictable but can reduce creativity. In production, the chapter recommends generating multiple answers and using techniques such as majority voting or self-consistency when reliability matters.
Structured outputs
LLMs can be instructed to produce structured formats such as JSON.
Methods include:
- JSON schemas
- Jsonformer
- LMQL
- Guidance
- Regular expressions
- Context-free grammars (CFGs)
These techniques are useful when LLM output must be consumed by another software system.
Interpretability and debugging
LLM interpretability is still developing. One practical approach is to examine how outputs change when inputs are slightly modified and to inspect intermediate model behavior.
LIT-NLP can help with:
- Attention visualization
- Salience maps
- Embedding visualization
- Counterfactual analysis
The chapter also introduces mechanistic interpretability, where researchers study neurons and combinations of neuron activations called features to better understand model behavior.