About the Prisoner's Dilemma
The prisoner's dilemma is the most famous puzzle in game theory. Two players each choose to cooperate or defect. Each individually scores higher by defecting, yet if both defect they both end up worse than if both had cooperated. This simulation plays the iterated, spatial version on a population so you can watch trust evolve.
Strategies compete on a grid: always-cooperate, always-defect, tit-for-tat, grudger and random. Each cell plays its neighbours over several rounds, then copies the most successful nearby strategy. Cooperation can survive — even thrive — when it clusters, because cooperators reap the high mutual-reward payoff among themselves.
Adjust the payoffs T, R, P and S, the number of rounds, and the mutation rate, and the balance of power between exploiters and trusters shifts before your eyes.
Frequently asked questions
What is the prisoner's dilemma?
It is a game where two players each choose to cooperate or defect. Each individually gains by defecting, but if both defect they both do worse than if both had cooperated. This tension between individual and collective interest is the dilemma.
Why does mutual defection happen if cooperation is better?
Defection is the dominant strategy in a one-shot game: whatever the other player does, you score higher by defecting. Rational self-interest drives both players to defect, producing an outcome worse for both — a Nash equilibrium that is not Pareto optimal.
What do the payoffs T, R, P and S mean?
T is the Temptation to defect against a cooperator, R is the Reward for mutual cooperation, P is the Punishment for mutual defection, and S is the Sucker payoff for cooperating against a defector. The dilemma requires T > R > P > S, and for the iterated game 2R > T + S.
What is tit-for-tat?
Tit-for-tat starts by cooperating, then copies whatever its opponent did on the previous move. It is nice, retaliatory, forgiving and clear — the qualities that won Robert Axelrod's iterated prisoner's dilemma tournaments in 1980.
How does the spatial version work?
Each cell on the grid plays a repeated game against its eight neighbours and accumulates a score. Then every cell copies the strategy of whichever neighbour (including itself) scored highest. This imitation rule lets cooperator clusters grow or shrink over generations.
What is a grudger strategy?
A grudger (also called grim trigger) cooperates until its opponent defects even once, then defects forever afterwards. It is unforgiving but very stable against exploitation.
Why do cooperators survive in the spatial game but not the well-mixed one?
In a well-mixed population defectors always win, but on a grid cooperators form clusters where they mostly meet other cooperators and earn the high mutual-reward payoff. The cluster edge can resist invasion, so cooperation persists despite being globally dominated.
What does the mutation slider do?
Mutation randomly reassigns a small fraction of cells to a random strategy each generation. A little mutation keeps diversity alive and can seed new cooperator clusters; too much mutation overwhelms selection and homogenises the population into noise.
Who invented the prisoner's dilemma?
The game was framed by Merrill Flood and Melvin Dresher at RAND in 1950, and given its memorable prisoner story by Albert Tucker. Robert Axelrod's 1980s computer tournaments popularised the iterated version and the success of tit-for-tat.
What real-world situations does it model?
Arms races, price wars, climate-treaty free-riding, doping in sport, overfishing and the evolution of biological cooperation all share the dilemma's structure: short-term self-interest undermines a mutually better outcome unless repetition, reputation or kinship reward cooperation.