llm ai/books/designingllmapplication ai

What is an LLM?

  • A language model is a model trained on large amounts of text to learn patterns of human language, including grammar and meaning.
  • Its basic training objective is next-token prediction: given a sequence of tokens, predict what token is likely to come next.
  • The model actually produces a probability distribution over the vocabulary, rather than simply choosing one word.
  • Repeated next-token prediction training allows the model to learn surprisingly complex capabilities.
  • Modern LLMs are primarily based on neural networks, especially the Transformer Architecture
  • The same approach can model sequences other than natural language, such as programming code, chess moves, DNA sequences, and airline schedules.

What makes a language model “large”?

There is no universally accepted definition of “large.” The book uses more than 1 billion parameters as its working definition.

The chapter discusses scaling laws:

  • Increasing model size, training data, and compute generally improves performance.
  • Early scaling research emphasized increasing model size.
  • Later research showed that training-data size needs to grow alongside model size for compute-optimal training.
  • Smaller models can still benefit from additional high-quality training data, especially when speed, energy efficiency, or cost matters.

It also introduces emergent capabilities: abilities that appear to become significantly stronger as models grow, such as arithmetic and logical reasoning. However, the chapter notes that researchers disagree about whether these capabilities truly “emerge” suddenly or whether the apparent jumps are partly caused by the evaluation methods used.

Brief history of LLMs

The development of LLMs follows the broader history of Natural Language Processing (NLP):

1950s → Machine translation and symbolic/rule-based NLP
1960s → ELIZA demonstrates rule-based conversational interaction
1990s–2000s → Statistical machine learning becomes increasingly important
2010s → Deep learning replaces much manual feature engineering
2017 → Transformer architecture introduced
Modern era → Pre-training, fine-tuning, increasingly large LLMs, and open-source development.

The GPT progression illustrates this evolution:

  • GPT-1: large-scale unsupervised pre-training followed by supervised fine-tuning.
  • GPT-2: large web-data training and stronger zero-shot capabilities.
  • GPT-3: dramatically larger and popularized few-shot/in-context learning.
  • GPT-4: emphasized alignment training and controllability.
  • o1: focused more strongly on reasoning and inference-time computation.

A major trend is consolidation: more NLP tasks are increasingly handled end-to-end by a single general model rather than separate task-specific systems.

Impact and enterprise use

The chapter argues that LLM adoption is not merely another technology hype cycle. Its analysis of large U.S. public companies found 2,195 companies discussing or adopting LLMs, compared with far fewer discussing Web3 or crypto.

Major enterprise applications include:

  1. Employee productivity — coding assistants, marketing content, knowledge-base Q&A.
  2. Report generation — summarizing reports, research, meetings, paperwork, and contracts.
  3. Chatbots — customer support and interfaces to company information.
  4. Information extraction — entity extraction, relation extraction, sentiment analysis, NER.
  5. Translation and transformation — translating languages or changing writing style.
  6. Workflow automation — LLM-based agents that can search, retrieve information, run code, and interact with other systems.

However, adoption also has employment implications. Companies may use AI to reduce costs and workforce requirements even when the AI system performs worse than humans. The chapter warns that replacing humans purely for cost savings can negatively affect customer satisfaction.

Prompting

Prompting is the process of interacting with an LLM by providing it with input/context.

The key idea is:

LLMs are next-token predictors, so the text provided in the context strongly influences the output.

Good prompts should generally be:

  • Explicit
  • Detailed
  • Structured
  • Unambiguous
  • Clear about the desired output format.

Important prompting techniques:

TechniqueMain idea
Zero-shotGive an instruction without examples
Few-shotGive examples to demonstrate the desired behavior
Chain-of-thoughtEncourage step-by-step reasoning
Prompt chainingBreak a complex task into multiple prompts
Adversarial promptingConstruct prompts that attempt to bypass model behavior/alignment
The chapter emphasizes that prompt engineering is not magic. Good prompting can significantly improve performance, but it does not unlock unlimited hidden capabilities.

It also introduces **prompt drift**: a prompt that works well with one model or model version may perform differently after the underlying model changes. Therefore, prompts should be treated as something that may need to be tested and version-controlled.

LLM APIs

LLMs can be accessed programmatically through APIs.

The chapter introduces common API concepts such as:

  • System message — establishes high-level behavior.
  • User message — contains the user’s request.
  • Assistant message — model output.
  • Tool message — interaction with external tools.
  • Temperature — higher values generally produce more diverse/creative outputs; lower values produce more predictable outputs.
  • top_p — another decoding-control parameter.
  • n — number of completions generated.
  • stop / max tokens — control output length.
  • presence/frequency penalties — help reduce repetition.
  • logit_bias — increases or decreases the probability of particular tokens.
  • logprobs — provides probability information about generated tokens.