AI Age · AI-05
Across seventy years of AI research, general methods that leverage more computation have consistently beaten methods built on human-crafted domain knowledge — a pattern researchers keep re-learning the hard way.
The observation, drawn from the history of AI research, that approaches relying on scaling general-purpose learning and search methods with more computation have repeatedly and decisively outperformed approaches that build in human domain knowledge and hand-crafted heuristics — and that researchers, generation after generation, have resisted this conclusion because it devalues the specific expertise and cleverness they'd invested in building.
Coined and argued in Rich Sutton's widely-cited 2019 essay 'The Bitter Lesson,' drawing on historical examples across chess, Go, speech recognition, and computer vision spanning several decades of AI research.
The Mechanism
Two curves that keep re-crossing across decades of AI subfields
The hand-crafted approach often leads early — it encodes real, useful expertise from the start. But it plateaus once that expertise is exhausted, while the general, compute-scaled method — initially unimpressive — keeps climbing as more computation becomes available, and eventually and repeatedly overtakes it. Sutton's essay documents this exact crossing pattern recurring across chess, Go, speech, and vision.
01 · IT'S AN EMPIRICAL PATTERN, NOT A PRIORI THEORY
Sutton's argument rests on historical track record, not abstract principle
The essay's force comes from repeatedly pointing to concrete history: chess engines built on human grandmaster heuristics were eventually overtaken by search-and-learning systems (culminating in Deep Blue and later AlphaZero-style approaches); hand-built speech recognition features gave way to statistical and later deep-learning approaches trained on scale. Each time, researchers initially resisted the general method because it seemed to 'waste' available domain expertise.
02 · THE RESISTANCE ITSELF IS PART OF THE PATTERN, AND IS PSYCHOLOGICALLY EXPLAINED
Researchers have real incentives to prefer the approach that uses their expertise
Sutton explicitly notes that researchers tend to prefer methods that leverage their own specific knowledge and cleverness — an understandable but systematically misleading bias, since it makes the eventually-superior general method look naive or wasteful in the short run, right up until enough compute becomes available for it to overtake the hand-crafted approach decisively.
03 · IT DOESN'T CLAIM DOMAIN KNOWLEDGE IS NEVER USEFUL — ONLY THAT IT DOESN'T SCALE THE SAME WAY
A frequently missed nuance
The lesson isn't that human expertise is worthless — it's that baking expertise in as a fixed structural assumption tends to eventually become a ceiling, whereas methods designed to keep improving as more computation becomes available don't hit that same ceiling. Domain knowledge remains useful for bootstrapping, interpretability, and efficiency at smaller scales; it's specifically the long-run scaling comparison where the pattern holds most consistently.
Where It Fails / Inversion
Where it fails / inversion
The lesson has real limits: in domains where additional compute genuinely isn't available or isn't the binding constraint (small, well-understood physical systems; problems with hard external limits on the achievable data), hand-crafted domain expertise doesn't get overtaken because the scaling that would overtake it never arrives — treating 'just add compute and generality' as a universal strategy ignores that Sutton's own examples are specifically domains where compute kept increasing sharply over sustained periods.
How To Use It
Worked example · deciding whether to hand-tune a system or invest in a more general, scalable approach
A team building an internal AI tool for document classification faces a real choice: hand-craft rules and features based on deep knowledge of the specific document types (fast to build, works well immediately, plateaus as document variety grows), or invest in a more general learned approach that will underperform the hand-crafted version initially but improve as more training data and compute become available. The Bitter Lesson's practical guidance: if you expect meaningfully more data and compute to become available over the tool's lifetime, bias toward the general, scalable approach even though it looks worse today — the crossing point, per the historical pattern, tends to arrive faster than intuition expects.
How to use it
Before investing heavily in a hand-crafted, expertise-based solution to a problem that's likely to have more data and compute available in the future, ask honestly whether you're choosing it because it's actually the better long-run bet, or because it lets you use expertise you already have. If more scale is coming, bias toward the general method even when it currently looks less impressive — the historical pattern says it usually wins the race you haven't finished yet.
See Also