ai llm machinelearning

Reinforcement Learning from Human Feedback
Illustrating Reinforcement Learning from Human Feedback (RLHF)

1. Why RLHF? The Fundamental Problem

Pretrained Language Models (LLMs) are trained on massive internet text using a simple objective: predict the next token (cross-entropy loss).

While this creates models with immense knowledge, a raw base model doesn’t inherently know how to be a helpful assistant. If you ask a base model:

“The president of the United States in 2006 was”

It might continue with internet metadata or list of random trivia instead of a concise, helpful answer:

”…George W. Bush, the governor of Florida in 2006 was Jeb Bush, and John McCain was an Arizona senator… September 1 – U.S. President Bush signs an executive order…”

Traditional automatic evaluation metrics like BLEU or ROUGE compare generated tokens to reference texts using rigid n-gram matching rules. They cannot measure subjective attributes like creativity, truthfulness, empathy, or safety. Writing a mathematical loss function directly for human values is intractable.

The Core Idea of RLHF

Can we optimize complex AI behaviors using simple human preference signals as a guide?
RLHF solves hard-to-specify alignment problems by using human feedback to train a reward function, and then optimizing the language model against that reward.

2. Core Concepts & Analogies for Beginners

To understand how RLHF works without getting lost in math, consider these foundational analogies:

Analogy 1: The F1 Race Car (Elicitation Theory)

  • The Base Model = F1 Engine & Chassis: Base models contain almost all the raw intelligence, knowledge, and capabilities learned during pretraining on billions of pages of text.
  • Post-Training / RLHF = Aerodynamics & Race Tuning: F1 teams spend the season refining aerodynamics and suspension around a static engine to dramatically boost lap times. Similarly, RLHF doesn’t teach the model new world knowledge from scratch—it elicits and reshapes the latent knowledge inside the base model into an engaging, structured, and helpful format.
       +---------------------------------------------+
       |           Base Model (Pretrained)           |
       |  Raw Knowledge, Unformatted Next-Token Pred. |
       +---------------------------------------------+
                              |
                              v  (Post-Training / RLHF)
       +---------------------------------------------+
       |            Instruct / Chat Model            |
       |    Helpful, Harmless, High-Empathy Style    |
       +---------------------------------------------+

Analogy 2: Teaching vs. Grading (Contrastive Loss)

  • Supervised Learning (SFT) is like showing a student exact textbook answers: “Predict the exact next word.”
  • RLHF is like grading essay submissions: The model generates multiple candidate answers, and a reviewer indicates which answer is better. This contrastive learning teaches the model what makes a response superior and what bad patterns to avoid.

Analogy 3: Thermostat & CartPole (Classic RL vs. RLHF)

  • In classic RL (like a thermostat adjusting temperature or CartPole balancing a pole), the environment provides a fixed numerical reward (e.g., for balancing, for falling).
  • In LLM RLHF, the “environment” does not provide a native score. We must build a neural network (Reward Model) to act as the judge.

3. The Canonical 3-Step RLHF Pipeline

Popularized by OpenAI’s InstructGPT (2022), the canonical RLHF workflow consists of three main sequential stages:

Canonical RLHF 3-Step Recipe
Figure 1: The canonical 3-step RLHF recipe (InstructGPT / Anthropic / DeepMind).

Step 1: Pretraining & Instruction Fine-Tuning (SFT / IFT)

Before RLHF begins, we need a base model that can understand prompt-response structures.

  1. Pretraining: Train a base transformer model on large web text via next-token prediction.
  2. Supervised Fine-Tuning (SFT): Fine-tune the base model on ~10k–100k curated question-answering examples to teach it the assistant persona.

SFT Pretraining
Figure 2: Step 1 - Pretraining a base model and fine-tuning on instruction prompts.

Chat Templates & Token Formatting

To format multi-turn conversations into flat token sequences that the model understands, frameworks use Chat Templates (e.g., ChatML format):

<|im_start|>system
You are a helpful and harmless assistant.<|im_end|>
<|im_start|>user
How can I improve my sleep quality?<|im_end|>
<|im_start|>assistant

Role of SFT

SFT teaches the model the shape and structure of how to answer (formatting), while RLHF will refine the quality, style, and safety of the answers.

Step 2: Preference Data Collection & Reward Model (RM) Training

The goal of Step 2 is to create a Reward Model (RM)—a neural network that takes a prompt and completion , and outputs a scalar score representing human preferability.

Reward Model Training
Figure 3: Step 2 - Gathering human comparisons and training a scalar Reward Model.

Why Ranking instead of Direct Scalar Ratings?

If human annotators give absolute scores (e.g., to ) to text completions, scores become noisy and uncalibrated because different humans have different baseline expectations.

Instead, annotators evaluate model completions in head-to-head pairwise matchups (Option A vs. Option B):

  1. Sample a prompt from a dataset.
  2. Generate two (or more) model completions: (preferred / winning) and (dispreferred / losing).
  3. Train the Reward Model using a Bradley-Terry preference loss:

where is the sigmoid function. This maximizes the reward gap between the winning and losing completions.

Step 3: Fine-Tuning with Reinforcement Learning (PPO)

With an instruction-tuned model (Policy ) and a trained Reward Model (), we use Reinforcement Learning to optimize the model parameters.

RLHF PPO Training
Figure 4: Step 3 - Fine-tuning the language model policy with PPO and a KL divergence penalty.

Formulating LLM Generation as an RL Problem:

  • Policy (): The language model being fine-tuned.
  • State / Observation (): The input prompt sequence.
  • Action (): The sequence of output tokens generated by the model.
  • Action Space: The entire vocabulary of the tokenizer (~50,000+ tokens).
  • Reward (): Combined scalar reward from the Reward Model minus a KL divergence penalty.

The KL Divergence Penalty ()

If we only optimize for maximum reward from the RM, the policy will exploit flaws in the Reward Model—generating repetitive, sycophantic, or gibberish text that gets a high score (a phenomenon called Reward Hacking or Over-Optimization).

To prevent this, a penalty term measures how far the active policy has drifted from the initial reference policy using Kullback-Leibler (KL) Divergence:

Where:

  • is the scalar score from the Reward Model.
  • is a hyperparameter controlling constraint strength.
  • penalizes large token probability shifts from the baseline model.

Reward Hacking

Without the KL penalty, the model learns to “trick” the reward model by spamming keywords or ballooning output lengths (Length Bias).

4. Key Differences: Standard RL vs. Language Model RLHF

AspectStandard RL (e.g., Robotics / CartPole)RLHF (Language Models)
Initial PolicyLearned from scratch (random initialization)Fine-tuned from a pretrained LLM prior
Reward SignalEnvironmental rule function Learned Reward Model from human preference data
State TransitionsMulti-step environment physics Single-turn prompt sampled from dataset (no environment transition)
ActionSingle discrete or continuous action Complete text sequence
Reward GranularityStep-by-step per actionResponse-level (bandit-style, scalar per completion)
Discounting (balances short vs long-term reward) (no discounting over generated tokens)

5. Evolution of Post-Training Recipes

Post-training has evolved significantly beyond the initial 3-step InstructGPT workflow.

RLHF Timeline
Figure 5: Historical timeline of preference learning and RLHF development (2008–2026).

1. Direct Alignment (DPO - Direct Preference Optimization)

Introduced in 2023, DPO mathematically re-formulates the RLHF objective to optimize the policy directly on pairwise preference data without needing an explicit intermediate Reward Model or PPO loop:

2. Multi-Stage Open Recipes (e.g., Tülu 3)

Modern state-of-the-art open models like Tülu 3 use multi-stage pipelines incorporating synthetic data, preference tuning, and targeted RL:

Tülu 3 Recipe
Figure 6: The multi-stage post-training pipeline of Tülu 3 (SFT -> On-Policy DPO -> RLVR).

  1. SFT on ~1M examples: Massive synthetic & human instruction mixture.
  2. On-Policy Preference Tuning: DPO/PPO on ~1M preference pairs to boost chat quality.
  3. RLVR (Reinforcement Learning with Verifiable Rewards): Targeted RL on math/code prompts with verifiable ground-truth rewards.

3. Reasoning & RLVR Era (e.g., DeepSeek R1 & OpenAI o1)

With models like DeepSeek R1, post-training incorporates verifiable rewards (unit tests, math checkers) and cold-start reasoning chains. Compute is shifted heavily into post-training RL, enabling models to generate long internal Chains of Thought (CoT) before answering.


Reinforcement Learning from Human Feedback explained with math derivations and the PyTorch code.