Humanoid Robot Vision › Module 1: Getting Started › Lesson 1.1

How Bramble Sees

Free previewLesson 1.1 · about 25 minutes

This lesson has no exercises. It sets the scene for the rest of the course.

Get the full course

10 modules of hands-on robot vision — a vision workbench where you write JavaScript that works on Bramble's camera pictures, with checked exercises in every lesson, in your browser, nothing to install. Lessons 1.1 and 1.2 are free; the rest of the course is Rs. 300 + 18% GST.

Enroll Now — Rs. 354.00

By the end of this lesson you will be able to

  • say in plain words what computer vision is and what a robot uses it for
  • describe the chain from camera to numbers to picture processing to decisions to action
  • give five reasons why seeing is hard for a robot, even though it feels easy to you
  • name Bramble's cameras and say what each one gives
  • run a vision workbench, change its code and run it again
  • explain why a safe robot checks what it sees and stops when it is unsure

A new job for Bramble

It is a quiet morning at the Ashdown Robotics Workshop. Bramble, the workshop's humanoid robot, stands by the long bench with its arms at its sides. Eleanor Price, one of the technicians, has spent months learning how Bramble's motors, joints and balance work. Today Henry Wren, the senior engineer, has a new task for her.

“Put a blue cup on the far bench,” Henry says, “and ask Bramble to fetch it.” Eleanor does. Bramble turns its head slowly, takes a few careful steps, and stops. It does not reach for the cup.

“It walks perfectly well,” says Eleanor. “Why doesn't it go and get the cup?”

“Because it doesn't know where the cup is,” Henry replies. “It knows how to move. It doesn't yet know how to look. That's what you and I are going to teach it, one small step at a time.”

Margaret Ellis, the workshop manager, looks up from her clipboard. “And while you're at it, teach it to notice people. I want it to stop long before it gets near anyone.”

What is computer vision?

Computer vision means getting useful information out of pictures with a computer program. The pictures come from a camera. The information might be “there is a blue cup”, “the cup is 30 cm to the left and 190 cm ahead”, “there is a crate in the way”, or “a person is standing in front of me”.

A robot uses vision to answer questions about the world around it so that it can act. It is not enough to take a picture: the robot must work something out from the picture, and then do something sensible with the answer.

From camera to action

Every vision system in this course follows the same chain of five links:

  1. Camera. Light from the scene passes through a lens and lands on a sensor, a small chip covered in a grid of tiny light-measuring points (a phone camera has millions of them).
  2. Numbers. Each point measures how much light reached it and turns that into a number. The camera hands the computer a grid of numbers, not a picture. You will see those numbers for yourself in the next lesson.
  3. Picture processing. A program works through the numbers: it tidies them, looks for colours, edges and shapes, and measures what it finds. Most of this course is about this link.
  4. Decisions. From the measurements, the program decides what is there and what to do: “that blob is the blue cup, it is a little to the left, so turn left a little.”
  5. Action. The robot moves: it turns, steps, reaches or stops. Then it takes a new picture and the chain starts again.

That last point matters. A robot does not look once and then walk blindly. It looks, decides and acts over and over, several times a second, so that it can correct small mistakes as it goes. You will build a loop like this yourself in the last module.

You can see a red cup, a blue box, a green ball and a steel spanner on a wooden bench. Bramble cannot “see” any of that yet. All it has is 3072 small groups of numbers. Turning those numbers back into “red cup, blue box, green ball, spanner” is the job of the programs you will write.

How the workbench boxes work

The box above is a vision workbench. It holds a short program written in JavaScript, a common programming language. Press Run and the program runs in your browser; anything it shows or prints appears underneath.

Try it now. In the box above, change "workbench" to "workbench-dim" and press Run. You will see the same bench with the lights turned down. Change it back when you have had a look.

You do not need to know JavaScript yet

This lesson only asks you to press Run and change a word or two. Lesson 1.2 shows how pictures are stored, and lesson 1.3 teaches, gently and from the very beginning, all the programming you need for the rest of the course.

Why seeing is hard for a robot

You can glance at a table and spot a cup instantly, without effort. It is tempting to think that must be easy for a computer too. It is not. Your brain has had years of practice and uses a large part of its power for seeing. A robot starts with nothing but numbers. Here are five things that make vision hard:

Each of these problems has its own lesson or two later in the course. The good news is that you do not need to solve them all at once: a robot that works well in a tidy, well-lit workshop is a fine start, and you will make it tougher step by step.

Bramble's cameras

Bramble carries several cameras, each with its own job. This is Bramble's camera data sheet. You will come back to it often, so do not worry about remembering it now.

CameraWhat it givesDetails
Head cameracolour pictures, 64 × 48 pixelslooks straight down at the workbench, a parts tray or the floor, giving “top-down” pictures
Front cameracolour pictures, 64 × 48 pixelson the chest, 100 cm above the floor, looking straight ahead; its focal length is 50 pixels and the centre of its picture is at x = 32, y = 24
Stereo camerastwo grey pictures, 64 × 48 pixelson the chest, 20 cm apart (10 cm either side of the front camera), at the same height and looking the same way
Depth cameraa grid of distances in centimetres, 64 × 48low on the chest, 60 cm above the floor, looking straight ahead; reads from 20 cm to 800 cm; 0 means “no reading”
Lensslight barrel bendingstraight lines near the edges of the picture bow outwards a little; Module 6 shows how to correct this

Some of these words (focal length, stereo, barrel bending) will be explained properly when you need them. For now, notice that Bramble has two kinds of eyes: cameras that measure light, giving colour or grey pictures, and a depth camera that measures distance.

A few other facts about Bramble will matter later. It is 1.6 m tall and about 44 cm wide. Its arm can reach a cup up to 65 cm from the middle of its body. And one rule is fixed: it must not move towards a person who is closer than 130 cm in front of it.

Why such small pictures?

Real robot cameras give pictures with hundreds of thousands or millions of pixels. Bramble's pictures in this course are only 64 × 48 so that you can print the numbers, look at them and check your working by eye. The methods you learn are exactly the ones used on bigger pictures; they just take the computer longer.

Here is what the front camera sees when Bramble stands in the workshop facing a bench across the room:

The blue cup Eleanor put out is the small blue patch on the bench, a little left of the middle. It is nearly 2 metres away, so it covers only a handful of pixels. Finding small, distant things like this reliably is one of the main goals of the course.

And here is the depth camera looking down a corridor with two crates in it. Instead of colours, every pixel holds a distance. Near things are drawn bright and far things dark:

The low crate is 220 cm ahead, the taller one 400 cm, and the end wall 700 cm. A robot that wants to walk down the corridor can read these numbers directly to find out what is in its way. You will do exactly that in Module 7.

What you will learn, module by module

ModuleWhat it covers
1 Getting Startedhow pictures are stored as numbers, and the JavaScript you need
2 Pixels and Colourbrightness, the red, green and blue parts of a colour, and picking out one colour
3 Cleaning Up Picturesblack-and-white masks, removing grain and specks, filters, and coping with poor lighting
4 Edges and Shapesfinding edges and corners, grouping pixels into objects, and measuring them
5 Finding Thingssearching for a pattern, coloured markers and printed tags, following a line, sorting parts on a tray
6 The Camera and Geometryhow a camera makes a picture, turning pixels into directions and distances, lens bending, and depth from two cameras
7 Depth and the 3-D Worlddepth pictures, finding the floor and obstacles, and positions relative to the robot
8 Motionspotting movement, following a moving object, and predicting where it will go
9 Recognising Objectsdescribing shapes with numbers, telling objects apart, look-alikes, and staying safe when unsure
10 Putting It Togethersteering by sight, testing vision properly, and a final project: Bramble finds and fetches the blue cup

Along the way you will help the rest of the workshop: Oliver Grant brings trays of parts to sort, Alice Bennett wants the stores counted, and Thomas Hale, the test engineer, makes sure that everything you build works every time, not just once.

Vision can be wrong: a word on safety

No vision system is perfect. Pictures can be too dark, a lens can be dirty, a look-alike can fool the program, and a person can step out from behind a crate. A robot that trusts every answer blindly will sooner or later do something it should not.

Safe robots are built on a simple idea: check, and stop when unsure. They look again before acting, they test whether an answer makes sense (is that “cup” the right size for its distance?), and when they cannot be sure, they stop or ask a person rather than guess. Margaret's rule about people is a good example: whatever the camera thinks it sees, Bramble keeps its distance from anyone in front of it.

You will meet this idea throughout the course, and Module 9 is given over to it. Good vision is not only about finding things; it is also about knowing when you have not found them.

Quick check

Choose an answer to see whether you are right and why.

What does Bramble's camera actually hand to the computer?

Which is the right order for the chain from camera to action?

Which of Bramble's cameras measures distances rather than light?

Bramble's program is not sure whether a blob in the picture is the blue cup or the blue box. What should a safe robot do?

Summary

  • Computer vision means getting useful information out of pictures with a program, so that a robot can act on it.
  • The chain is camera → numbers → picture processing → decisions → action, repeated many times a second.
  • Seeing is hard for a robot because of changing light, clutter, look-alikes, things partly hidden, and movement.
  • Bramble has a head camera looking down, a front camera looking ahead, two stereo cameras, and a depth camera that measures distances in centimetres. Its pictures are 64 × 48 pixels.
  • A vision workbench runs a short program when you press Run; you can change the code and run it as often as you like.
  • Vision can be wrong, so safe robots check what they see and stop when they are unsure.

Found a mistake on this page, or something unclear? Report a problem and mention “Robot Vision Lesson 1.1”.