Reinforcement Learning from Human Feedback
Illustrating Reinforcement Learning from Human Feedback (RLHF)
Abstract
Reinforcement Learning from Human Feedback (RLHF) is a post-training technique used to align Large Language Models (LLMs) with complex, hard-to-specify human values (such as helpfulness, honesty, harmlessness, warmth, and formatting).
Instead of relying purely on cross-entropy next-token prediction loss, RLHF uses human preferences to train a Reward Model, which then guides an RL algorithm (typically PPO) to update the language model while penalizing excessive deviation from its baseline using KL Divergence.
1. Why RLHF? The Fundamental Problem
Pretrained Language Models (LLMs) are trained on massive internet text using a simple objective: predict the next token (cross-entropy loss).
While this creates models with immense knowledge, a raw base model doesn’t inherently know how to be a helpful assistant. If you ask a base model:
“The president of the United States in 2006 was”
It might continue with internet metadata or list of random trivia instead of a concise, helpful answer:
”…George W. Bush, the governor of Florida in 2006 was Jeb Bush, and John McCain was an Arizona senator… September 1 – U.S. President Bush signs an executive order…”
Traditional automatic evaluation metrics like BLEU or ROUGE compare generated tokens to reference texts using rigid n-gram matching rules. They cannot measure subjective attributes like creativity, truthfulness, empathy, or safety. Writing a mathematical loss function directly for human values is intractable.
The Core Idea of RLHF
Can we optimize complex AI behaviors using simple human preference signals as a guide?
RLHF solves hard-to-specify alignment problems by using human feedback to train a reward function, and then optimizing the language model against that reward.
2. Core Concepts & Analogies for Beginners
To understand how RLHF works without getting lost in math, consider these foundational analogies:
Analogy 1: The F1 Race Car (Elicitation Theory)
- The Base Model = F1 Engine & Chassis: Base models contain almost all the raw intelligence, knowledge, and capabilities learned during pretraining on billions of pages of text.
- Post-Training / RLHF = Aerodynamics & Race Tuning: F1 teams spend the season refining aerodynamics and suspension around a static engine to dramatically boost lap times. Similarly, RLHF doesn’t teach the model new world knowledge from scratch—it elicits and reshapes the latent knowledge inside the base model into an engaging, structured, and helpful format.
+---------------------------------------------+
| Base Model (Pretrained) |
| Raw Knowledge, Unformatted Next-Token Pred. |
+---------------------------------------------+
|
v (Post-Training / RLHF)
+---------------------------------------------+
| Instruct / Chat Model |
| Helpful, Harmless, High-Empathy Style |
+---------------------------------------------+
Analogy 2: Teaching vs. Grading (Contrastive Loss)
- Supervised Learning (SFT) is like showing a student exact textbook answers: “Predict the exact next word.”
- RLHF is like grading essay submissions: The model generates multiple candidate answers, and a reviewer indicates which answer is better. This contrastive learning teaches the model what makes a response superior and what bad patterns to avoid.
Analogy 3: Thermostat & CartPole (Classic RL vs. RLHF)
- In classic RL (like a thermostat adjusting temperature or CartPole balancing a pole), the environment provides a fixed numerical reward (e.g., for balancing, for falling).
- In LLM RLHF, the “environment” does not provide a native score. We must build a neural network (Reward Model) to act as the judge.
3. The Canonical 3-Step RLHF Pipeline
Popularized by OpenAI’s InstructGPT (2022), the canonical RLHF workflow consists of three main sequential stages:

Figure 1: The canonical 3-step RLHF recipe (InstructGPT / Anthropic / DeepMind).
Step 1: Pretraining & Instruction Fine-Tuning (SFT / IFT)
Before RLHF begins, we need a base model that can understand prompt-response structures.
- Pretraining: Train a base transformer model on large web text via next-token prediction.
- Supervised Fine-Tuning (SFT): Fine-tune the base model on ~10k–100k curated question-answering examples to teach it the assistant persona.

Figure 2: Step 1 - Pretraining a base model and fine-tuning on instruction prompts.
Chat Templates & Token Formatting
To format multi-turn conversations into flat token sequences that the model understands, frameworks use Chat Templates (e.g., ChatML format):
<|im_start|>system
You are a helpful and harmless assistant.<|im_end|>
<|im_start|>user
How can I improve my sleep quality?<|im_end|>
<|im_start|>assistantRole of SFT
SFT teaches the model the shape and structure of how to answer (formatting), while RLHF will refine the quality, style, and safety of the answers.
Step 2: Preference Data Collection & Reward Model (RM) Training
The goal of Step 2 is to create a Reward Model (RM)—a neural network that takes a prompt and completion , and outputs a scalar score representing human preferability.

Figure 3: Step 2 - Gathering human comparisons and training a scalar Reward Model.
Why Ranking instead of Direct Scalar Ratings?
If human annotators give absolute scores (e.g., to ) to text completions, scores become noisy and uncalibrated because different humans have different baseline expectations.
Instead, annotators evaluate model completions in head-to-head pairwise matchups (Option A vs. Option B):
- Sample a prompt from a dataset.
- Generate two (or more) model completions: (preferred / winning) and (dispreferred / losing).
- Train the Reward Model using a Bradley-Terry preference loss:
where is the sigmoid function. This maximizes the reward gap between the winning and losing completions.
Step 3: Fine-Tuning with Reinforcement Learning (PPO)
With an instruction-tuned model (Policy ) and a trained Reward Model (), we use Reinforcement Learning to optimize the model parameters.

Figure 4: Step 3 - Fine-tuning the language model policy with PPO and a KL divergence penalty.
Formulating LLM Generation as an RL Problem:
- Policy (): The language model being fine-tuned.
- State / Observation (): The input prompt sequence.
- Action (): The sequence of output tokens generated by the model.
- Action Space: The entire vocabulary of the tokenizer (~50,000+ tokens).
- Reward (): Combined scalar reward from the Reward Model minus a KL divergence penalty.
The KL Divergence Penalty ()
If we only optimize for maximum reward from the RM, the policy will exploit flaws in the Reward Model—generating repetitive, sycophantic, or gibberish text that gets a high score (a phenomenon called Reward Hacking or Over-Optimization).
To prevent this, a penalty term measures how far the active policy has drifted from the initial reference policy using Kullback-Leibler (KL) Divergence:
Where:
- is the scalar score from the Reward Model.
- is a hyperparameter controlling constraint strength.
- penalizes large token probability shifts from the baseline model.
Reward Hacking
Without the KL penalty, the model learns to “trick” the reward model by spamming keywords or ballooning output lengths (Length Bias).
4. Key Differences: Standard RL vs. Language Model RLHF
| Aspect | Standard RL (e.g., Robotics / CartPole) | RLHF (Language Models) |
|---|---|---|
| Initial Policy | Learned from scratch (random initialization) | Fine-tuned from a pretrained LLM prior |
| Reward Signal | Environmental rule function | Learned Reward Model from human preference data |
| State Transitions | Multi-step environment physics | Single-turn prompt sampled from dataset (no environment transition) |
| Action | Single discrete or continuous action | Complete text sequence |
| Reward Granularity | Step-by-step per action | Response-level (bandit-style, scalar per completion) |
| Discounting | (balances short vs long-term reward) | (no discounting over generated tokens) |
5. Evolution of Post-Training Recipes
Post-training has evolved significantly beyond the initial 3-step InstructGPT workflow.

Figure 5: Historical timeline of preference learning and RLHF development (2008–2026).
1. Direct Alignment (DPO - Direct Preference Optimization)
Introduced in 2023, DPO mathematically re-formulates the RLHF objective to optimize the policy directly on pairwise preference data without needing an explicit intermediate Reward Model or PPO loop:
2. Multi-Stage Open Recipes (e.g., Tülu 3)
Modern state-of-the-art open models like Tülu 3 use multi-stage pipelines incorporating synthetic data, preference tuning, and targeted RL:

Figure 6: The multi-stage post-training pipeline of Tülu 3 (SFT -> On-Policy DPO -> RLVR).
- SFT on ~1M examples: Massive synthetic & human instruction mixture.
- On-Policy Preference Tuning: DPO/PPO on ~1M preference pairs to boost chat quality.
- RLVR (Reinforcement Learning with Verifiable Rewards): Targeted RL on math/code prompts with verifiable ground-truth rewards.
3. Reasoning & RLVR Era (e.g., DeepSeek R1 & OpenAI o1)
With models like DeepSeek R1, post-training incorporates verifiable rewards (unit tests, math checkers) and cold-start reasoning chains. Compute is shifted heavily into post-training RL, enabling models to generate long internal Chains of Thought (CoT) before answering.
Reinforcement Learning from Human Feedback explained with math derivations and the PyTorch code.