By the end of this lesson you will be able to
- say in plain words what statistics is and what it is for
- explain why a single story (an anecdote) is weaker evidence than data collected from many cases
- name the kinds of question this course will teach you to answer
- see why statistics helps you judge claims in the news and answers given by AI assistants
- run a live workbench and know how the checked exercises in later lessons work
A question at the farm shop
It is Emily Hartwell's first morning at Brookfield Growers, a co-operative of fruit and vegetable farms with a busy farm shop. Before she has even taken her coat off, two people have told her something with great confidence.
Ruth Oakley, who manages the farm shop, is worried. “A customer told me yesterday that nobody would recommend us any more since we changed the opening hours. She was very cross. Should we change them back?”
Ten minutes later Samuel Pike, who runs the orchard, leans in at the door. “Everyone knows the north block grows the biggest apples. Put the north block apples at the front of the display.”
Daniel Mercer, who keeps the co-operative's records, smiles at Emily. “Two claims before nine o'clock. Welcome to the job. One cross customer is a story, and ‘everyone knows’ is a feeling. Neither is evidence yet. Shall we look at what the records actually say?”
That, in a sentence, is what this course is about.
What statistics is
Statistics is the craft of learning from data. It has three parts, and you will practise all three:
- Collecting data in a fair way, so that it tells you about the thing you care about. Who was asked? How was it measured? Was anything left out?
- Describing data: turning a table of hundreds of numbers into a few clear summaries and charts. What is typical? How much do the values vary? Are there any odd ones?
- Drawing conclusions from data, while being honest about how sure you can be. A difference seen in a sample might be real, or it might be the luck of which cases happened to be measured. Statistics gives you tools to tell these apart.
The word also has a second, everyday meaning: a single number such as “65 per cent of customers would recommend the shop” is often called a statistic. In this course you will learn both to produce such numbers and to question them.
Everything in this course starts from school arithmetic: adding, dividing, fractions and percentages. The workbench does the long calculations for you. What matters most is thinking clearly about what a number means, and that is a skill anyone can learn.
Why it matters
Statistics is not only for scientists. It helps with three kinds of situation that almost everyone meets every week.
Everyday decisions
Should you take a coat if the forecast says there is a 30 per cent chance of rain? Is a bus that is “usually on time” reliable enough for a job interview? Is a tomato variety that gave a bigger crop in one garden really better, or did that garden just have a sunny year? Each of these is a question about data and chance, and each has a sensible answer once you know how to think about it.
Claims in the news
News reports are full of numbers: “sales up 40 per cent”, “a new study links coffee to longer life”, “crime has doubled”. Some of these are careful and fair. Others compare the wrong things, pick a flattering starting point, rest on a handful of cases, or confuse two things happening together with one causing the other. By the end of the course you will know the questions to ask: out of how many? compared with what? could it be chance? who was measured?
Answers from AI assistants
More and more people ask AI assistants questions, and the answers often contain statistics: an average, a percentage, a claim that one thing is “significantly” better than another. These answers can sound authoritative while being wrong: a number may be made up, worked out incorrectly, or true of a different group from the one you asked about. An assistant is a useful helper, but it is not a source of truth. Knowing how an average, a percentage or a margin of error is calculated lets you check such an answer yourself, and the workbench in this course lets you do the recalculation in a few lines. Lesson 10.1 practises exactly this.
Data versus anecdote
An anecdote is a single story: one customer's complaint, a neighbour whose car never broke down, a relative who smoked and lived to ninety. Anecdotes are vivid and easy to remember, which is exactly why they mislead. The cross customer Ruth met is real, but she tells us about one person. The customers who were happy with the new hours had no reason to come and say so.
Data are observations collected in a planned way from many cases, so that each case counts once and nobody is left out just because they were quiet. (The word data is plural; one observation is a datum. Many people now treat it as singular, and both are accepted.)
Suppose Ruth asks four friends who shop there, and three say they would recommend the shop. That is 3 out of 4, or 3 ÷ 4 = 0.75, which is 75 per cent. It sounds like a clear answer, but with only four people, one different reply would change it to 50 or 100 per cent. Her friends may also be kinder than the average customer.
Now suppose the shop asked 200 customers, and 130 said yes. That is 130 ÷ 200 = 0.65, or 65 per cent. One different reply would change it only to 64.5 or 65.5 per cent. More cases, collected fairly, give a steadier answer.
In fact the farm shop did run a survey of 200 customers after the opening hours changed. Here is your first live workbench. Press Run and see what those customers said. (Every workbench on these pages runs inside your browser. You will learn exactly what each line means in the next few lessons; for now, just read the results.)
The table counts the answers: 130 of the 200 customers (65 per cent) said they would recommend the shop, and 70 (35 per cent) said they would not. So the cross customer is not alone, since about a third of customers would not recommend it, but “nobody would recommend us” is plainly false. Ruth now has something far more useful than a story: a number, out of a known total, that she can compare with next year's survey.
Common mistake: letting the loudest story decide
People who are very pleased or very annoyed are the ones who speak up, write reviews and send letters. If you judge by what you hear, you hear mostly from the extremes. Before acting on a claim, ask: how many cases is this based on, and how were they chosen? This does not mean ignoring stories. A complaint can point you to a real problem worth measuring. It means checking the story against data before deciding.
Checking a claim with data
Now Samuel's claim that the north block grows the biggest apples. The co-operative weighed 120 apples from the orchard's four blocks. The mean (the ordinary average: add up the weights and divide by how many there are) can be worked out for each block separately. Press Run:
Rounded to the nearest gram, the mean weights are: east 149 g, north 143 g, south 147 g and west 147 g. Far from growing the biggest apples, the north block has the lowest average of the four in this sample.
But notice how close the numbers are: the gap between the highest and the lowest is only about 6 grams, and a single apple can weigh 50 grams more than its neighbour. Is a gap that small a real difference between the blocks, or just the luck of which apples happened to be weighed? That is a genuine statistical question, and answering it properly needs ideas you will meet later in the course (spread in Module 3, samples in Module 7, tests in Module 8). For now, the honest summary is: the data give no support to Samuel's claim, and the four blocks look fairly similar.
Workbenches are for playing with. In the box above, change mean to max and press Run again to see the heaviest apple from each block. You will find that the east block holds one giant apple of 238 grams, more than 40 grams heavier than any other, and on its own it lifts that block's average by about 3½ grams: without it, east would fall from first place to third. Lesson 2.3 looks at what to do about values like this. Reset puts the original commands back. You cannot break anything.
The big questions this course answers
Over ten modules, Emily will learn to answer the co-operative's questions, and you will learn alongside her. The questions fall into a few big groups:
| Question | For example | Where |
|---|---|---|
| What does the data look like? | What kinds of data does the survey hold, and how are the answers spread? | Module 1 |
| What is typical? | What does a typical apple weigh? Should Ruth report the mean or the median number of visits? | Module 2 |
| How much do things vary? | Which sack-filling machine is more consistent? | Module 3 |
| How can it be shown honestly? | Which chart suits the data, and how can a chart mislead? | Module 4 |
| How likely is it? | If a plant tests positive for blight, how likely is it to really have blight? | Modules 5 and 6 |
| What can a sample tell us about everyone? | What can 200 customers tell us about all the shop's customers, and how far off might they be? | Module 7 |
| Is a difference real, or just chance? | Does the new bean seed really give a bigger crop than the old one? | Module 8 |
| Are two things related? | Do sunnier tomato plants give more fruit, and does that mean sunshine causes it? | Module 9 |
| How do we report it well? | How to spot traps in the news and in AI answers, and write an honest report | Module 10 |
The course ends with a final project: a full investigation for Brookfield Growers, from question to written report.
How the workbenches and exercises work
Every lesson is built around live statistics workbenches like the two above. Each one understands a small, friendly language made for this course: one short command per line, such as mean(apples.weight). There is nothing to install, and every result appears in your browser straight away.
- Run (or Ctrl + Enter) works through the commands from the top and shows the result of each one underneath, including tables and charts.
- Reset puts back the commands the lesson started with.
- If a line has a mistake, the workbench stops there and explains in plain words what it did not understand, often suggesting the fix. The lines above still show their results.
- Commands that involve chance, such as tossing a simulated coin, give the same results every time you run the same commands, so what you see matches what the lesson describes.
From Lesson 1.2 onwards, each lesson ends with five exercises. Each exercise has its own workbench and a question from the co-operative, such as “show the average weight of the Russet apples” or “keep the number of late deliveries under the name late_count”. You type your commands and press Check my answer. The checker runs your commands and tells you whether the result is right. If it is not, it tells you what is still missing and may give a hint about a likely slip. There is usually more than one correct way to get an answer, and the checker accepts any of them. If you are stuck, Show solution reveals one correct answer, often with a note on what it means.
After the exercises comes a short quick check of four questions about the ideas in the lesson, and then a summary. This lesson has only the quick check.
Meet the people you will work with
Emily and Daniel will be with you throughout the course. Emily is new, curious and not afraid to ask “but what does that number actually mean?”. Daniel has kept the co-operative's records for years and has learnt, sometimes the hard way, which numbers to trust. Along the way you will also meet:
- Ruth Oakley, who manages the farm shop and its customer survey;
- Samuel Pike, who runs the orchard and has strong opinions about apples;
- Agnes Fairweather, who is running a trial of a new bean seed;
- Walter Bright, the delivery manager, and his drivers Albert, Beatrice, Cecil and Dora;
- and the two packing sheds, Hill and Vale, whose damaged boxes hold one of the course's best surprises.
Their records are the course's datasets: apples, rainfall, deliveries, a customer survey, the seed trial, sacks of potatoes, eggs, tomatoes, a blight test, packing boxes, and hedges and birds. Each one was built to show a real statistical idea clearly.
“So,” says Emily, looking at the two results on the screen, “the opening hours are not a disaster, and the north block apples are nothing special.”
“That is what these numbers say,” Daniel replies. “Next we will learn to read the records properly, row by row. Then you will be able to check claims like these yourself.”
Quick check
Four questions to confirm the main ideas. Choose an answer to see the explanation.
A friend says a restaurant is terrible because they once waited an hour for their food. What is the best way to describe this evidence?
One experience is an anecdote. It may be true and even worth looking into, but it does not tell you how long most customers wait. That would need many visits, recorded fairly.
In the farm shop survey, 130 of 200 customers said they would recommend the shop. Why is this better evidence than asking four friends?
With 200 customers, one different answer moves the percentage by only half a percentage point; with four friends it moves it by 25. Surveys can still be biased, depending on who takes part (Lesson 7.1), and 65 per cent is far from everyone.
The apple data showed mean weights of about 149, 143, 147 and 147 grams for the four blocks. What is the most honest conclusion at this stage?
The north block has the lowest mean here, so the claim gets no support. But a few grams' difference could come from which apples happened to be weighed. Deciding whether it is real needs the tools of Modules 7 and 8.
An AI assistant tells you that “the average apple in the dataset weighs 160 grams”. What is the most sensible response?
An assistant's answer is a claim like any other. When the data are available, recalculating takes a moment and settles the matter. For these apples, the mean is about 147 grams, not 160.
Summary
- Statistics is learning from data: collecting it fairly, describing it clearly, and drawing conclusions while being honest about uncertainty.
- It helps with everyday decisions and with judging claims in the news and in answers from AI assistants.
- An anecdote is one story; data come from many cases collected in a planned way. Always ask how many cases a claim rests on and how they were chosen.
- The survey showed 65 per cent of 200 customers would recommend the shop, and the apple data gave no support to the claim about the north block.
- Each lesson has live workbenches (Run, Reset) and, from Lesson 1.2, five checked exercises and a quick check.
Found a mistake on this page, or something unclear? Report a problem and mention “Statistics Lesson 1.1”.
