Game Theory · GT-09

Prisoner's Dilemma

Game Theory

Two rational actors, doing the individually smart thing, land on the one outcome worse for both.

Each party does better by betraying the other regardless of what the other does — yet mutual betrayal leaves both worse off than mutual cooperation would have. The canonical demonstration that individually rational choices can produce a collectively irrational result.

Formalized in 1950 by Merrill Flood and Melvin Dresher at RAND Corporation as an experiment in strategic behavior; the prison-sentence framing and the name were added shortly after by mathematician Albert Tucker, who used it to explain the result to a lay audience.

The Mechanism

Two coffee chains, deciding whether to discount — click a cell

Chain B chooses Hold Price Discount Hold Price Discount $8M, $8M R,R — Reward for mutual cooperation best joint outcome — unreachable alone $2M, $11M S,T — sucker / temptation $11M, $2M T,S — temptation / sucker $4M, $4M P,P — punishment for mutual defection ✗ THE ONLY NASH EQUILIBRIUM T (temptation) > R (reward) > P (punishment) > S (sucker) — the payoff ranking that defines every Prisoner's Dilemma

Discounting dominates for both chains, regardless of what the other does — so both discount, both land in the $4M/$4M cell, and both would have made more ($8M each) by silently holding price together. Neither can get there alone: whoever holds price while the other discounts gets the worst outcome of all ($2M, "the sucker").

01 · DOMINANT STRATEGY

Defect wins the argument every time you run it

Ask "if the other chain holds price, am I better off discounting?" Yes: $11M beats $8M. Ask "if the other chain discounts, am I better off discounting too?" Also yes: $4M beats $2M. Discounting is a dominant strategy — it wins regardless of the other player's move, so a rational actor never needs to guess what the opponent will do.

02 · INDIVIDUALLY RATIONAL, COLLECTIVELY WORSE

Both players reasoning correctly still lose together

Both chains follow the same airtight logic, both land on Discount, and both end up at $4M — worse than the $8M they'd both have gotten by holding price. Nobody made a mistake. That's what makes this different from ordinary bad decision-making: it's the *correct* individual reasoning that produces the bad joint outcome.

03 · ONE SHOT vs. REPEATED

The trap is specific to a single, isolated encounter

Everything above assumes this happens once, with no future and no reputation at stake. Play it repeatedly with the same counterpart, and the calculation changes completely — future retaliation becomes possible, and strategies like Tit-for-Tat can sustain cooperation indefinitely (see Repeated Games & the Folk Theorem).

Where It Fails / Inversion

The trap: importing one-shot logic into a repeated relationship

The classic result is a description of a single, anonymous, one-time interaction with no memory and no future — and it's routinely misapplied to situations that don't meet that description. Real business relationships, marriages, and long-standing partnerships are repeated games with reputational consequences, which is exactly why cooperation survives far more often in practice than the one-shot matrix predicts. Robert Axelrod's 1980 computer tournaments showed simple, forgiving reciprocal strategies (Tit-for-Tat) consistently outperforming pure defection once the game repeats.

The second, subtler trap: treating a genuinely repeated or reputation-bearing interaction as if it were one-shot, and defecting pre-emptively "to be safe" — which can manufacture the very breakdown in trust the theory describes, where none was structurally necessary.

How To Use It

Worked example · the airline fare war pattern

Two airlines serving the same route both do best if both hold fares high; each individually does even better by undercutting the other while fares stay high elsewhere; and if both undercut, both lose margin industry-wide (the $4M/$4M cell). This exact structure has played out repeatedly in aviation, and it explains why price wars are hard to prevent through appeals to "rational restraint" alone — restraint is not the dominant strategy in a single round.

What actually stabilizes prices in real markets that keep repeating this game isn't a handshake agreement (illegal collusion aside) — it's the credible expectation of retaliation in future rounds, transparent pricing that makes defection instantly visible, and a long enough time horizon that the future cost of a price war outweighs this quarter's temptation.

How to use it

Before assuming a partner, competitor, or colleague will "do the smart, cooperative thing," check whether you're actually in a repeated game with visible history and future consequences, or a genuine one-shot interaction. If it's one-shot and anonymous, expect the dominant strategy to win — and if you want cooperation instead, your real lever is changing the game's structure (make it repeated, make defection visible, add a future) rather than appealing to goodwill.

See Also

Nash Equilibrium → Tit-for-Tat Strategy → Repeated Games & the Folk Theorem → Tragedy of the Commons →