The Bitter Lesson Summary
Don't be distracted by human knowledge, as AI has been historically. Instead focus on methods for creating knowledge that scale with computation, like search and learning.
— Richard S. Sutton (@RichardSSutton)
Rich Sutton’s “The Bitter Lesson” argues that, across the histroy of AI, general methods that can exploit increasing amounts of computation tend to beat systems built around lots of hand-crafted human knowledge.
Building Chess AI
-
One approach is to encode thousands of expert rules: “knights are good here,” “protect this pawn structure,” and so on.
-
Another approach is to give the system a general mechanism for searching many possible moves and enough computing power to search very deeply.
Sutton says AI history repeatedly shows that the second kind of approach eventually wins, even when the hand-designed system initially looks more intelligent or elegant.
The lesson is “bitter” because researchers naturally want their own understanding of a problem to matter. It is satisfying to put human insights directly into a machine. Sutton argues that this often helps in the short term, but eventually those hand-crafted assumptions can become a ceiling: the system becomes tied to what humans already know and has difficulty benefiting from much more computation. In his summary, breakthroughs tend to arrive from approaches that instead scale search and learning.

Instead of directly transferring human knowledge into AI, develop general methods that allow AI to learn, search, and discover useful knowledge on its own.
Richard Sutton
He is a computer scientist and one of the foundational researchers in reinforcement learning, where an AI learns through interaction, feedback, rewards, and experience rather than being explicitly told every rule. He is a professor of computing science at the University of Alberta. Apps at UAlberta
He is also well known for co-authoring the textbook Reinforcement Learning: An Introduction with Andrew Barto. Sutton and Barto received the 2024 ACM A.M. Turing Award for their foundational contributions to reinforcement learning.
Rich Sutton Interview Summary
Sutton argues that intelligence should primarily come from continual interaction with the world—acting, observing consequences, receiving rewards, and learning from experience(10 - Reinforcement Learning > Agent and Enviroment). Dwarkesh Patel pushes back that today’s LLMs may already provide a powerful foundation that could later be combined with reinforcement learning and continual experience.
1. Sutton thinks current LLMs imitate people more than they understand the world
Sutton draws a sharp distinction between predicting human-generated text and predicting what actually happens in the environment. In his view, an LLM mostly learns from records of what humans said or did rather than from experiencing the consequences of its own actions.
“Reinforcement learning is about understanding your world. They’re not about figuring out what to do.”
2. Bitter Lesson
LLMs appear to fit The Bitter Lesson because they use enormous amounts of computation and relatively general learning methods. But Sutton sees a tension: their training data contains an enormous amount of human-produced knowledge.
He therefore wonders whether future systems that learn directly from experience will eventually replace systems trained primarily on human-generated data.
His alternative is very simple:
“The scalable method is you learn from experience. You try things, you see what works.”
So this extends the idea you identified earlier:
- Instead of transferring human knowledge → build AI that discovers knowledge.
In this interview, Sutton takes it further:
- Human data → AI imitates knowledge
versus - Experience + actions + consequences + learning → AI discovers knowledge
3. Sutton thinks animals provide a better model of intelligence than language does
Patel argues that humans clearly learn partly through imitation and cultural transmission. Children watch other people, copy language and behaviors, and inherit knowledge accumulated over generations.
Sutton emphasizes something more basic: before sophisticated cultural learning exists, animals already learn by trying things and experiencing consequences.
He puts it simply: “The child tries things and sees what happens.”
4. Sutton’s alternative is the “era of experience”
Sutton describes intelligence as a continuous stream:
sensation → action → consequence/reward → learning → new action → …
The agent doesn’t have a separate period where it finishes training and is then deployed forever. Its entire life is learning.
“This is what reinforcement learning paradigm is, learning from experience.”
5. Reward isn’t the only source of information
Patel raises an important objection: a simple reward signal seems far too low-bandwidth to teach an intelligent agent everything about the world.
Sutton agrees. The reward tells the agent how well things are going, but the agent learns much more from all its sensory experience.
He describes four important components of an intelligent agent:
- Policy: what should I do in this situation?
- Value function: how well am I doing / how promising is this state?
- Perception or state representation: what situation am I currently in?
- Transition/world model: if I do something, what will happen next?
The world model is learned from the rich stream of observations, not merely from reward. Pasted text
This is an important clarification of Sutton’s position: RL does not mean that the AI learns everything from a single scalar reward. It observes large amounts of information about the world; reward helps determine which outcomes matter.
6. Why continual learning matters: the world is too big to pretrain on everything
Sutton calls attention to what he describes as the big world hypothesis.
You cannot teach an AI everything it will ever need before deploying it because its particular environment will contain countless details—specific coworkers, customers, preferences, organizations, situations, and new events—that could not have been anticipated during training
His point is essentially:
- Pretraining: learn about the average/general world.
- Continual learning: learn this particular world while living in it.
This is one of Sutton’s strongest criticisms of relying solely on pretrained models.
Whole Interview Idea
Sutton’s view can be condensed to:
Today’s LLMs primarily learn by absorbing what humans already know. Sutton thinks the deeper path to intelligence is to build agents with goals that continually act in the world, predict consequences, observe what actually happens, and update themselves from that experience.
Patel’s main counterargument is:
LLMs may not be the final form of intelligence, but their enormous pretrained knowledge could be the starting point on top of which this experiential, continual-learning system is built.
My view is that intelligence develops through two complementary forms of learning.
Humans learn partly by interacting directly with the world, but they also inherit knowledge accumulated by previous generations. For example, a tribe living in a forest may know which plants are poisonous or which snakes are dangerous because that knowledge has been passed down over generations. Without this inherited knowledge, learning only through trial and error could be costly or even dangerous.
In AI, pre-trained language models play a similar role to inherited human knowledge. They begin with knowledge collected from large amounts of human-generated data. However, pretraining alone is not enough. AI should also be able to interact with the world, observe the consequences of its actions, adapt to new situations, and create new knowledge from experience.
Therefore, LLMs and reinforcement learning should complement each other. LLMs provide accumulated knowledge from the past, while RL and continual learning allow AI to test, refine, and extend that knowledge through real-world experience.
In simple terms: LLM = inherited knowledge + RL = learning from experience —> More complete intelligence = inherited knowledge + continual experience