Computer vision
Computer vision helps a computer understand pictures and video. A picture is a grid of pixel numbers (full details in the lesson on computer vision). The usual chain is:
- Pixels: numbers for brightness or colour.
- Edges and features: places where the numbers jump, such as the outline of a tree.
- Shapes and objects: a model matches patterns and decides "tree" or "house".
- Labels: the answer with a confidence score, such as "tree 97%".
Main jobs: classification (what is it?), detection (where is it?) and segmentation (which pixels?). Uses: face unlock, number-plate reading, medical scans, crop checks, self-driving cars. Modern systems learn the patterns from many labelled photos using neural networks.
Language processing (NLP)
Natural language processing (NLP) lets a computer work with human language. People say the same thing in many ways, so this is hard. The steps are:
- Tokenisation: cut the sentence into pieces called tokens. "I like cold tea" becomes 4 tokens.
- Tagging (parts of speech): mark each word as noun, verb, adjective and so on.
- Parsing: find how the words connect (who does what).
- Meaning: find what the sentence says, its feeling (sentiment, here positive) or its answer.
Uses: translation, search, voice assistants, spelling and grammar help, spam filters, chatbots. Hard parts: one word can have many meanings ("bank"), sarcasm, and mixing languages like Hinglish. Modern systems learn language patterns from huge amounts of text.
Reasoning
Reasoning means drawing a new fact that must follow from known facts and rules. This is deduction.
Rule: IF it rains THEN the road is wet. Fact: it rains. Conclusion: the road is wet. This is valid: it cannot fail.
Now the other way. Rule: IF it rains THEN the road is wet. Fact: the road is wet. Can we say it rained? No. A water tanker could have wet the road. This mistake is called affirming the consequent. A reasoning system must tell valid steps from risky ones.
Other ideas: planning (choose a list of steps to reach a goal), reasoning under uncertainty (using probability, as in Bayes) and common-sense reasoning, which is still hard for computers. Uses: rule-based helpers, schedule makers, puzzle solvers, checking that a design follows rules.
Game playing AI
Games are clear tests for AI because the rules are fixed and the winner is known. A program draws a game tree: each ball is a position, each arrow is a move.
Stick game: take 1 or 2 sticks; whoever takes the last stick wins. Look at the positions from the bottom:
- 0 sticks left on your turn: the other player took the last one, so you have lost.
- A position is winning if you can move to a losing position for the other player. A position is losing if every move leads to a winning position for the other player.
This bottom-up scoring is minimax: you pick the move that is best for you, assuming the other player also picks the move that is best for them. With 1 stick you win, with 2 you win (take 2), with 3 you lose (leave 1 or 2), with 4 you win (leave 3). Every multiple of 3 is a losing position, so the winner always leaves a multiple of 3.
Big games like chess have too many moves to draw the whole tree. So programs look only a few moves ahead, use a score for the board (a heuristic), and skip hopeless branches (pruning). Modern programs also learn from millions of practice games. Deep Blue beat the chess champion Garry Kasparov in 1997, and AlphaGo beat the Go champion Lee Sedol in 2016.
Try it
Play the stick game with a friend using 7 coins. Take turns taking 1 or 2 coins; whoever takes the last coin wins. Can you always leave your friend a multiple of 3? Then play against the computer in step 6 of the 3D. Next, test the rain rule: write one more rule of your own, such as "IF I eat sweets THEN I feel happy", and check which two conclusions are valid and which are risky.
Key formulas and definitions
- Vision chain: pixels → edges → shapes → labels
- NLP chain: tokens → tags → parse → meaning
- Valid: IF P THEN Q; P is true → Q is true
- Not safe: IF P THEN Q; Q is true → P may be false
- Minimax: win if some move reaches a losing position for the opponent
- Stick game (take 1 or 2): leave a multiple of 3
Worked examples
1. List the four NLP steps for "I like cold tea" and say what each gives.
Tokenise: I | like | cold | tea (4 tokens). Tag: pronoun, verb, adjective, noun. Parse: "I" does "like" to "cold tea". Meaning: the speaker enjoys tea, a positive feeling.
2. Rule: IF the battery is flat THEN the torch is off. Fact: the torch is off. Is it certain that the battery is flat?
No. The bulb could be broken or the switch off. This is a risky step. Only "battery flat, so torch off" is valid.
3. Stick game (take 1 or 2, last stick wins). Is 5 sticks winning or losing for the player to move?
Take 2 to leave 3, a multiple of 3. The other player is now in a losing position. So 5 is winning.
4. Stick game with 9 sticks, you move first. Who wins with best play?
9 is a multiple of 3, so it is a losing position for the mover. The second player wins by always leaving a multiple of 3 (if you take 1, they take 2; if you take 2, they take 1).
5. A vision system outputs "ball 91%". What do the three words tell you?
The label (ball), the confidence (91%) and that this is classification. It is 91% sure, not 100% sure.
Common mistakes
- Thinking the computer sees a photo like we do. It sees numbers (pixels).
- Thinking NLP means the computer truly understands feelings. It finds patterns in words and context.
- Using the reverse of a rule as if it were true. "If rain then wet" does not mean "if wet then rain".
- Thinking game AI checks every possible game. For big games it looks a few moves ahead and uses a score.