By the end of this lesson you will be able to
- say in plain words what computer vision is and what a robot uses it for
- describe the chain from camera to numbers to picture processing to decisions to action
- give five reasons why seeing is hard for a robot, even though it feels easy to you
- name Bramble's cameras and say what each one gives
- run a vision workbench, change its code and run it again
- explain why a safe robot checks what it sees and stops when it is unsure
A new job for Bramble
It is a quiet morning at the Ashdown Robotics Workshop. Bramble, the workshop's humanoid robot, stands by the long bench with its arms at its sides. Eleanor Price, one of the technicians, has spent months learning how Bramble's motors, joints and balance work. Today Henry Wren, the senior engineer, has a new task for her.
“Put a blue cup on the far bench,” Henry says, “and ask Bramble to fetch it.” Eleanor does. Bramble turns its head slowly, takes a few careful steps, and stops. It does not reach for the cup.
“It walks perfectly well,” says Eleanor. “Why doesn't it go and get the cup?”
“Because it doesn't know where the cup is,” Henry replies. “It knows how to move. It doesn't yet know how to look. That's what you and I are going to teach it, one small step at a time.”
Margaret Ellis, the workshop manager, looks up from her clipboard. “And while you're at it, teach it to notice people. I want it to stop long before it gets near anyone.”
What is computer vision?
Computer vision means getting useful information out of pictures with a computer program. The pictures come from a camera. The information might be “there is a blue cup”, “the cup is 30 cm to the left and 190 cm ahead”, “there is a crate in the way”, or “a person is standing in front of me”.
A robot uses vision to answer questions about the world around it so that it can act. It is not enough to take a picture: the robot must work something out from the picture, and then do something sensible with the answer.
From camera to action
Every vision system in this course follows the same chain of five links:
- Camera. Light from the scene passes through a lens and lands on a sensor, a small chip covered in a grid of tiny light-measuring points (a phone camera has millions of them).
- Numbers. Each point measures how much light reached it and turns that into a number. The camera hands the computer a grid of numbers, not a picture. You will see those numbers for yourself in the next lesson.
- Picture processing. A program works through the numbers: it tidies them, looks for colours, edges and shapes, and measures what it finds. Most of this course is about this link.
- Decisions. From the measurements, the program decides what is there and what to do: “that blob is the blue cup, it is a little to the left, so turn left a little.”
- Action. The robot moves: it turns, steps, reaches or stops. Then it takes a new picture and the chain starts again.
That last point matters. A robot does not look once and then walk blindly. It looks, decides and acts over and over, several times a second, so that it can correct small mistakes as it goes. You will build a loop like this yourself in the last module.
You can see a red cup, a blue box, a green ball and a steel spanner on a wooden bench. Bramble cannot “see” any of that yet. All it has is 3072 small groups of numbers. Turning those numbers back into “red cup, blue box, green ball, spanner” is the job of the programs you will write.
How the workbench boxes work
The box above is a vision workbench. It holds a short program written in JavaScript, a common programming language. Press Run and the program runs in your browser; anything it shows or prints appears underneath.
- You can change the code and press Run again. Nothing you do can break Bramble or the page, so experiment freely.
- If you make a mistake, the box tells you what went wrong in plain words and which line it happened on.
- Lines starting with
//are notes for you. The computer ignores them.
Try it now. In the box above, change "workbench" to "workbench-dim" and press Run. You will see the same bench with the lights turned down. Change it back when you have had a look.
This lesson only asks you to press Run and change a word or two. Lesson 1.2 shows how pictures are stored, and lesson 1.3 teaches, gently and from the very beginning, all the programming you need for the rest of the course.
Why seeing is hard for a robot
You can glance at a table and spot a cup instantly, without effort. It is tempting to think that must be easy for a computer too. It is not. Your brain has had years of practice and uses a large part of its power for seeing. A robot starts with nothing but numbers. Here are five things that make vision hard:
- Lighting. The same cup gives very different numbers in bright light, in dim light, in orange evening light or in shadow. A rule such as “the cup's pixels are about this bright” stops working the moment someone switches off a lamp.
- Clutter. Real benches are covered in things. The program must pick out the one object it wants from everything else, including the pattern of the wood itself.
- Look-alikes. A blue box and a blue cup can have exactly the same colour. A red sticker can look like a small red cup. The program needs more than one clue to tell them apart.
- Things partly hidden. When one object stands in front of another, the one behind shows only part of itself, and it looks smaller and a different shape.
- Movement. People walk past, balls roll and the robot itself moves, so the picture keeps changing. The program has to keep up and cope with blur.
Each of these problems has its own lesson or two later in the course. The good news is that you do not need to solve them all at once: a robot that works well in a tidy, well-lit workshop is a fine start, and you will make it tougher step by step.
Bramble's cameras
Bramble carries several cameras, each with its own job. This is Bramble's camera data sheet. You will come back to it often, so do not worry about remembering it now.
| Camera | What it gives | Details |
|---|---|---|
| Head camera | colour pictures, 64 × 48 pixels | looks straight down at the workbench, a parts tray or the floor, giving “top-down” pictures |
| Front camera | colour pictures, 64 × 48 pixels | on the chest, 100 cm above the floor, looking straight ahead; its focal length is 50 pixels and the centre of its picture is at x = 32, y = 24 |
| Stereo cameras | two grey pictures, 64 × 48 pixels | on the chest, 20 cm apart (10 cm either side of the front camera), at the same height and looking the same way |
| Depth camera | a grid of distances in centimetres, 64 × 48 | low on the chest, 60 cm above the floor, looking straight ahead; reads from 20 cm to 800 cm; 0 means “no reading” |
| Lens | slight barrel bending | straight lines near the edges of the picture bow outwards a little; Module 6 shows how to correct this |
Some of these words (focal length, stereo, barrel bending) will be explained properly when you need them. For now, notice that Bramble has two kinds of eyes: cameras that measure light, giving colour or grey pictures, and a depth camera that measures distance.
A few other facts about Bramble will matter later. It is 1.6 m tall and about 44 cm wide. Its arm can reach a cup up to 65 cm from the middle of its body. And one rule is fixed: it must not move towards a person who is closer than 130 cm in front of it.
Real robot cameras give pictures with hundreds of thousands or millions of pixels. Bramble's pictures in this course are only 64 × 48 so that you can print the numbers, look at them and check your working by eye. The methods you learn are exactly the ones used on bigger pictures; they just take the computer longer.
Here is what the front camera sees when Bramble stands in the workshop facing a bench across the room:
The blue cup Eleanor put out is the small blue patch on the bench, a little left of the middle. It is nearly 2 metres away, so it covers only a handful of pixels. Finding small, distant things like this reliably is one of the main goals of the course.
And here is the depth camera looking down a corridor with two crates in it. Instead of colours, every pixel holds a distance. Near things are drawn bright and far things dark:
The low crate is 220 cm ahead, the taller one 400 cm, and the end wall 700 cm. A robot that wants to walk down the corridor can read these numbers directly to find out what is in its way. You will do exactly that in Module 7.
What you will learn, module by module
| Module | What it covers |
|---|---|
| 1 Getting Started | how pictures are stored as numbers, and the JavaScript you need |
| 2 Pixels and Colour | brightness, the red, green and blue parts of a colour, and picking out one colour |
| 3 Cleaning Up Pictures | black-and-white masks, removing grain and specks, filters, and coping with poor lighting |
| 4 Edges and Shapes | finding edges and corners, grouping pixels into objects, and measuring them |
| 5 Finding Things | searching for a pattern, coloured markers and printed tags, following a line, sorting parts on a tray |
| 6 The Camera and Geometry | how a camera makes a picture, turning pixels into directions and distances, lens bending, and depth from two cameras |
| 7 Depth and the 3-D World | depth pictures, finding the floor and obstacles, and positions relative to the robot |
| 8 Motion | spotting movement, following a moving object, and predicting where it will go |
| 9 Recognising Objects | describing shapes with numbers, telling objects apart, look-alikes, and staying safe when unsure |
| 10 Putting It Together | steering by sight, testing vision properly, and a final project: Bramble finds and fetches the blue cup |
Along the way you will help the rest of the workshop: Oliver Grant brings trays of parts to sort, Alice Bennett wants the stores counted, and Thomas Hale, the test engineer, makes sure that everything you build works every time, not just once.
Vision can be wrong: a word on safety
No vision system is perfect. Pictures can be too dark, a lens can be dirty, a look-alike can fool the program, and a person can step out from behind a crate. A robot that trusts every answer blindly will sooner or later do something it should not.
Safe robots are built on a simple idea: check, and stop when unsure. They look again before acting, they test whether an answer makes sense (is that “cup” the right size for its distance?), and when they cannot be sure, they stop or ask a person rather than guess. Margaret's rule about people is a good example: whatever the camera thinks it sees, Bramble keeps its distance from anyone in front of it.
You will meet this idea throughout the course, and Module 9 is given over to it. Good vision is not only about finding things; it is also about knowing when you have not found them.
Quick check
Choose an answer to see whether you are right and why.
What does Bramble's camera actually hand to the computer?
The camera only measures light. It gives a grid of numbers, and working out which objects are there, and where, is the job of the vision program.
Which is the right order for the chain from camera to action?
Light reaches the camera, which turns it into numbers; a program processes the numbers, decides what is there and what to do, and then the robot acts. Then it all repeats with a new picture.
Which of Bramble's cameras measures distances rather than light?
The depth camera gives a grid of distances in centimetres, with 0 meaning no reading. The head and front cameras give colour pictures. (The stereo cameras measure light too, but two pictures together can be used to work out distance, as Module 6 shows.)
Bramble's program is not sure whether a blob in the picture is the blue cup or the blue box. What should a safe robot do?
When vision is unsure, a safe robot does not guess. It checks more clues, looks again from another place, or stops and asks. Module 9 shows how to measure how sure a program is.
Summary
- Computer vision means getting useful information out of pictures with a program, so that a robot can act on it.
- The chain is camera → numbers → picture processing → decisions → action, repeated many times a second.
- Seeing is hard for a robot because of changing light, clutter, look-alikes, things partly hidden, and movement.
- Bramble has a head camera looking down, a front camera looking ahead, two stereo cameras, and a depth camera that measures distances in centimetres. Its pictures are 64 × 48 pixels.
- A vision workbench runs a short program when you press Run; you can change the code and run it as often as you like.
- Vision can be wrong, so safe robots check what they see and stop when they are unsure.
Found a mistake on this page, or something unclear? Report a problem and mention “Robot Vision Lesson 1.1”.
