πŸ“˜ CodingMarble Learn

Computer Vision: How Machines See

Computer vision (CV) is the part of AI that lets computers understand images and videos. A digital image is a grid of pixels; each pixel is a number (0–255 for grey) or three numbers (R, G, B) for colour. Resolution is width Γ— height in pixels. CV finds features such as edges, corners and colours, then does tasks like classification (what), object detection (where) and segmentation (which pixels). It is used in face unlock, self-driving cars, medical scans, farming, shops and traffic cameras.

🎬 Step-by-step story

  1. A computer does not see a picture. It sees a grid of tiny squares called pixels. This heart has 8 Γ— 8 = 64 pixels.
  2. Each pixel is only a number. In a grey image, 0 is black and 255 is white. A colour pixel has three numbers: red, green and blue.
  3. Features are useful patterns, like edges and corners. A small filter slides over the image and lights up where the numbers jump.
  4. CV does three main jobs. Classification says what is in the image. Detection draws a box around it. Segmentation marks its exact pixels.
  5. We use CV every day: face unlock, self-driving cars, medical scans, farms and shops.
  6. Try it: add brightness. The same number is added to every pixel. Too much and the shape disappears.

Tip: drag the 3D scene to turn it. Use two fingers to zoom.

πŸ€” Common doubts, cleared

If the computer only has numbers, how does it know it is a heart?

It learns from many labelled examples which number patterns go with which object. The pattern of dark pixels in a heart shape matches what it learned.

Why does a colour image need three numbers per pixel?

Screens make every colour by mixing red, green and blue light. One number per colour gives three numbers.

What exactly is an edge for a computer?

A place where neighbouring pixel numbers jump a lot, like 40 next to 255. The filter lights up there.

Why can't the computer see the shape in a very bright photo?

When every pixel is pushed up to 255, the dark and light parts become equal. With no difference, there are no edges to find.

Is a picture really made of squares?

Yes. Zoom out and the squares blend together, which is why we see a smooth picture.

What is computer vision?

Computer vision (CV) is a branch of artificial intelligence. It helps a computer get meaning from pictures and videos.

Our eyes catch light and the brain understands it. In CV, a camera catches light and a program tries to understand it. The program has learned from many example images.

Images are made of pixels

A pixel (picture element) is the smallest dot of a digital image. Pixels are placed in rows and columns, like squares on graph paper.

Resolution

Resolution = number of pixels across Γ— number of pixels down. A 1920 Γ— 1080 image has 1920 Γ— 1080 = 2,073,600 pixels (about 2 megapixels). More pixels show more detail.

Pixel values

Features: what the computer looks for

A feature is a piece of the image that helps tell things apart. Common features:

A small grid of numbers called a filter (or kernel) slides over the image. At each place it multiplies and adds pixel values. The result is large where the feature is present. This sliding idea is called convolution. A Convolutional Neural Network (CNN) learns its own filters from thousands of labelled images: early layers find edges, later layers find eyes, wheels or leaves.

CV tasks and applications

Main tasks

Applications

Limits and ethics

CV can be fooled by bad light, blur or unusual angles. It can be unfair if training images miss some groups of people. Cameras also raise privacy questions, so rules and consent matter.

Try it at home

Open any photo on a phone and zoom in as far as you can. You will see small squares: pixels. Now use the brightness tool and watch every square get lighter together. That is exactly what the 3D scene shows in its last step.

Key formulas and definitions

Worked examples

1. An image is 640 Γ— 480 pixels. How many pixels does it have?

640 Γ— 480 = 307,200 pixels (about 0.3 megapixels).

2. A grayscale image is 100 Γ— 100. Each pixel uses 1 byte. How much memory does it need? What if it were RGB?

Grey: 100 Γ— 100 Γ— 1 = 10,000 bytes. RGB uses 3 bytes per pixel: 100 Γ— 100 Γ— 3 = 30,000 bytes.

3. A pixel has value 200. You add brightness of 80. What is the new value?

200 + 80 = 280, but the maximum is 255, so the pixel becomes 255 (pure white).

4. A parking camera must say how many cars are present and where each one is. Which CV task is this?

Object detection: it gives a box and a label for every car, so it can count them too.

Common mistakes

Practice quiz

1. The smallest unit of a digital image is a:
2. In a grayscale image, the value 0 means:
3. How many numbers describe one RGB pixel?
4. Drawing boxes around every person in a photo is:
5. An edge in an image is a place where:

Practice: answer these yourself

Type or choose your answer, then press Check. Use a hint if you are stuck; the full solution appears after you answer.

Frequently asked questions

What is computer vision in simple words?

It is AI that lets computers understand photos and videos, for example recognising faces, reading text or spotting objects.

What is the difference between image processing and computer vision?

Image processing changes an image (brighter, sharper). Computer vision tries to understand what is in it.

Why is the pixel range 0 to 255?

Each grey value is stored in 8 bits (1 byte), which can hold 2⁸ = 256 different values: 0 to 255.

Where this is taught

CBSE (India)Class 10Part B: Computer Vision
CBSE (India)Class 12Making Machines See

Related lessons

All Computer Science lessons