What is computer vision?
Computer vision (CV) is a branch of artificial intelligence. It helps a computer get meaning from pictures and videos.
Our eyes catch light and the brain understands it. In CV, a camera catches light and a program tries to understand it. The program has learned from many example images.
- Input: an image or a video (a video is many images per second).
- Process: turn pixels into numbers and look for patterns.
- Output: an answer, like a label ("cat"), a box, or a decision ("unlock").
Images are made of pixels
A pixel (picture element) is the smallest dot of a digital image. Pixels are placed in rows and columns, like squares on graph paper.
Resolution
Resolution = number of pixels across Γ number of pixels down. A 1920 Γ 1080 image has 1920 Γ 1080 = 2,073,600 pixels (about 2 megapixels). More pixels show more detail.
Pixel values
- Grayscale: one number per pixel, from 0 (black) to 255 (white). That is 256 levels, which fit in 1 byte (8 bits).
- Colour (RGB): three numbers per pixel, one each for Red, Green and Blue, each 0β255. (255, 0, 0) is pure red, (255, 255, 0) is yellow, (0, 0, 0) is black, (255, 255, 255) is white.
- So a colour image is three grids stacked together, called channels.
Features: what the computer looks for
A feature is a piece of the image that helps tell things apart. Common features:
- Edges: where pixel values change suddenly (dark next to light).
- Corners: where two edges meet. Corners are very good features because they are easy to find again in another photo.
- Colour and texture: a ripe mango is yellow; grass has a rough green texture.
- Shapes: circles, lines, faces.
A small grid of numbers called a filter (or kernel) slides over the image. At each place it multiplies and adds pixel values. The result is large where the feature is present. This sliding idea is called convolution. A Convolutional Neural Network (CNN) learns its own filters from thousands of labelled images: early layers find edges, later layers find eyes, wheels or leaves.
CV tasks and applications
Main tasks
- Image classification: one label for the whole image ("dog").
- Classification + localisation: the label plus one box showing where the object is.
- Object detection: boxes and labels for many objects (3 cars, 2 people).
- Instance segmentation: marks the exact pixels of each object.
Applications
- Face unlock and photo tagging
- Self-driving cars and lane-keeping
- Medical imaging: spotting signs of disease in X-rays and eye scans (a doctor confirms)
- Agriculture: finding leaf disease or counting fruit from drone photos
- Retail: self-checkout, stock counting, barcode and QR reading
- Traffic: number-plate reading, counting vehicles
- Google Lensβtype apps: translate signboards, identify plants
Limits and ethics
CV can be fooled by bad light, blur or unusual angles. It can be unfair if training images miss some groups of people. Cameras also raise privacy questions, so rules and consent matter.
Try it at home
Open any photo on a phone and zoom in as far as you can. You will see small squares: pixels. Now use the brightness tool and watch every square get lighter together. That is exactly what the 3D scene shows in its last step.
Key formulas and definitions
- Pixel: smallest dot of a digital image
- Resolution = width (pixels) Γ height (pixels)
- Grayscale pixel: 0 (black) to 255 (white), 1 byte
- RGB pixel: (R, G, B), each 0β255; 3 channels
- Feature: edge, corner, colour, texture or shape that helps recognise an object
- Tasks: classification (what), detection (where), segmentation (which pixels)
Worked examples
1. An image is 640 Γ 480 pixels. How many pixels does it have?
640 Γ 480 = 307,200 pixels (about 0.3 megapixels).
2. A grayscale image is 100 Γ 100. Each pixel uses 1 byte. How much memory does it need? What if it were RGB?
Grey: 100 Γ 100 Γ 1 = 10,000 bytes. RGB uses 3 bytes per pixel: 100 Γ 100 Γ 3 = 30,000 bytes.
3. A pixel has value 200. You add brightness of 80. What is the new value?
200 + 80 = 280, but the maximum is 255, so the pixel becomes 255 (pure white).
4. A parking camera must say how many cars are present and where each one is. Which CV task is this?
Object detection: it gives a box and a label for every car, so it can count them too.
Common mistakes
- Thinking the computer sees a picture like we do. It only gets a grid of numbers.
- Saying grey pixels go from 0 to 256. There are 256 levels, but the range is 0 to 255.
- Mixing up classification and detection. Classification gives a label only; detection also says where.
- Thinking more pixels always means a better result. Very large images are slow; CV often shrinks them first.