Players, strategies and the pay-off matrix
A game has players, the strategies (choices) each player can make, and pay-offs (what each gets).
For two players with a few choices, we use a pay-off matrix. Player A chooses a row, player B a column. Each cell shows (A's pay-off, B's pay-off).
| B: Silent | B: Confess | |
|---|---|---|
| A: Silent | (3, 3) | (0, 5) |
| A: Confess | (5, 0) | (1, 1) |
In a zero-sum game, what one wins the other loses, so we write only A's pay-off. A game tree is used instead when players move one after another (a sequential game).
Dominance: removing bad choices
A strategy is dominated if another strategy gives a pay-off at least as good against every choice of the opponent, and better against at least one. A sensible player never uses it, so we cross it out. This makes the matrix smaller.
In the table above, compare A's rows: against Silent, Confess gives 5 > 3; against Confess, it gives 1 > 0. So Confess dominates Silent for A. It is a dominant strategy: best whatever B does. By symmetry the same is true for B.
For the column player in a zero-sum game, remember that B wants A's number to be small: a column is dominated if its entries are all bigger.
Nash equilibrium and the prisoner's dilemma
A Nash equilibrium is a pair of strategies where each player is already giving their best reply to the other. Nobody can do better by changing alone.
To find it: for each column, mark A's best row; for each row, mark B's best column. A cell with both marks is an equilibrium.
In the prisoner's dilemma, both confessing (1, 1) is the only equilibrium, although both staying silent (3, 3) is better for both. Each player's private interest leads to a worse shared result. This explains price wars, arms races and over-fishing.
Some games have more than one equilibrium. In a coordination game (both pick tea or both pick coffee) both matching cells are stable.
When the same game is repeated many times, players can reward cooperation and punish cheating ("tit for tat"), so cooperation can last. A credible threat or commitment is one a player would really carry out.
Zero-sum games: play-safe strategies and saddle points
In a zero-sum game a cautious player assumes the worst.
- Row player (A): find the minimum of each row, then choose the row with the largest minimum (maximin).
- Column player (B): find the maximum of each column, then choose the column with the smallest maximum (minimax).
These are the play-safe strategies. If maximin = minimax, the game has a saddle point and a stable solution: neither player gains by changing. That common number is the value of the game.
Example: A's pay-offs [[3, 1], [4, 2]]. Row minimums 1, 2 → maximin 2. Column maximums 4, 2 → minimax 2. Saddle point at (A2, B2), value 2.
Mixed strategies by graph
If maximin < minimax, there is no stable pure solution. A player should then mix: play each row with a fixed probability, chosen at random, so the opponent cannot guess.
Let A play row 1 with probability p and row 2 with 1 − p. Against each of B's columns, A's expected pay-off is a straight line in p. B will pick the column that is worse for A, so A gets the lower of the lines. A chooses p at the highest point of this lower edge, usually where the lines cross.
Example [[4, 1], [2, 3]]: against B1, E = 4p + 2(1 − p) = 2 + 2p. Against B2, E = p + 3(1 − p) = 3 − 2p. Set equal: 2 + 2p = 3 − 2p → p = 1/4. Value = 2.5.
For larger games (e.g. 3×3), the same idea is written as a linear programming problem: maximise the value V subject to each column giving A at least V, and solved by the simplex method.
Try it: the 3D and at home
In the 3D: on the last step pick each game and predict the equilibrium before the rings appear. In the mixed game, slide p from 0 to 1 and find where the red dot is highest.
At home (with a friend): play the ultimatum game with 10 sweets. One person offers a split; the other accepts (both keep their shares) or rejects (nobody gets any). Theory says any offer above 0 should be accepted, but real people often reject unfair offers. This shows that fairness and reputation matter, not only pay-offs.
Exam focus
Expect to: write a pay-off matrix from a story; reduce it by dominance; find play-safe strategies and test for a saddle point; find a Nash equilibrium; find the optimal mixed strategy and value by graph or by solving two equations; explain the prisoner's dilemma in a real context.
Key formulas and definitions
- Maximin (row) = largest of the row minimums
- Minimax (column) = smallest of the column maximums
- Saddle point ⇔ maximin = minimax = value of game
- E(column j) = p × a(1,j) + (1 − p) × a(2,j)
- Optimal p: E(column 1) = E(column 2)
- Nash equilibrium: each strategy is a best reply to the other
Worked examples
1. In the matrix [[3, 1], [4, 2]] (A's pay-offs, zero-sum), find the play-safe strategies and test for a saddle point.
Row minimums: 1, 2 → A plays row 2 (maximin 2). Column maximums: 4, 2 → B plays column 2 (minimax 2). Equal, so the saddle point is (row 2, column 2) and the value is 2.
2. Reduce by dominance: A's pay-offs [[2, 5, 4], [1, 3, 2], [3, 6, 1]].
Row 2 (1, 3, 2) is dominated by row 1 (2, 5, 4): remove it. Now [[2, 5, 4], [3, 6, 1]]. For B (wants small), column 2 (5, 6) is worse than column 1 (2, 3) in both rows: remove it. Left: [[2, 4], [3, 1]].
3. Find the Nash equilibria of the coordination game: (Tea, Tea) = (4, 4), (Tea, Coffee) = (0, 0), (Coffee, Tea) = (0, 0), (Coffee, Coffee) = (3, 3).
If B picks Tea, A's best is Tea; if B picks Coffee, A's best is Coffee. Same for B. Both (Tea, Tea) and (Coffee, Coffee) have both best replies, so there are two equilibria.
4. Solve the zero-sum game [[4, 1], [2, 3]] for A.
Maximin = 2, minimax = 3: no saddle point. Let A play row 1 with probability p. E(B1) = 2 + 2p, E(B2) = 3 − 2p. Equal when p = 1/4. A plays row 1 with probability 1/4 and row 2 with 3/4; value = 2 + 2(1/4) = 2.5.
5. For the same game, find B's optimal mix.
Let B play column 1 with probability q. A's row 1 gives 4q + 1(1 − q) = 1 + 3q; row 2 gives 2q + 3(1 − q) = 3 − q. Equal when 4q = 2, q = 1/2. B plays each column half the time; value = 2.5, the same as for A.
6. Two firms choose High or Low price. (High, High) = (6, 6), (High, Low) = (2, 8), (Low, High) = (8, 2), (Low, Low) = (3, 3). What happens and why?
Low dominates High for each firm (8 > 6 and 3 > 2), so both choose Low: equilibrium (3, 3). Both would earn 6 with High prices, so this is a prisoner's dilemma; repeated play or agreements (collusion) may keep prices high.
Common mistakes
- Using the column player's maximin. In a zero-sum matrix of A's pay-offs, B wants SMALL numbers, so B uses minimax.
- Thinking a Nash equilibrium is always the best outcome for everyone. The prisoner's dilemma shows it may not be.
- Taking the HIGHER of the expected-pay-off lines. The opponent picks the reply that is worse for you, so use the lower edge.
- Removing a row that is better in some columns and worse in others. Dominance needs it to be at least as good in EVERY column.