Hero background

Clean Data, Clear Conclusions

Maths • 60 • 25 students • Created with AI following Aligned with New Zealand Curriculum

Download now

Free PDF · we'll email you a copy

Maths
60
25 students
17 August 2026

Teaching Instructions

Create a detailed Year 9 New Zealand Mathematics and Statistics lesson plan on Cleaning and Organising Data Sets. Align to Te Mātaiaho Phase 4 Statistics, especially developing knowledge from data, visualisation of data, and interpreting data (NZ-TMA-MATHEMATIC-Y9-10-statistics-099-DOC141, NZ-TMA-MATHEMATIC-Y9-10-statistics-113-DOC141, NZ-TMA-MATHEMATIC-Y9-10-statistics-127-DOC141; official source: https://newzealandcurriculum.tahurangi.education.govt.nz/new-zealand-curriculum-online/nzc---mathematics-and-statistics-phase-4-years-9-10/5637291579.p). Include: clear learning intentions and success criteria; prior knowledge; vocabulary; resources; an engaging starter; explicit teacher modelling using a deliberately messy data set; guided practice with teacher questioning and error analysis; differentiated independent tasks; extension and support; assessment questions with answers/marking guidance; formative assessment checkpoints; misconceptions; plenary/exit ticket; culturally responsive NZ context; and approximate timings for a 60-minute lesson. Focus on identifying variables, correcting inconsistent entries, handling missing/invalid data, sorting and tabulating data, and explaining how cleaning choices affect conclusions. Use accessible but appropriately challenging Year 9 language.

Overview

Students investigate how poor-quality data can distort a statistical conclusion. Using a deliberately messy dataset about Year 9 students’ ways of travelling to school, they identify variables, correct inconsistent entries, manage missing or invalid values, sort and tabulate data, and explain how cleaning decisions affect interpretation.

Learning intentions

Students will:

  • identify variables and distinguish valid, missing and invalid data;
  • apply consistent rules to clean and organise a dataset;
  • sort data and construct a frequency table;
  • explain how data-cleaning choices can change a conclusion.

Success criteria

  • I can name each variable and describe its type.
  • I can correct, remove or flag data entries using a stated rule.
  • I can organise cleaned data into a frequency table.
  • I can justify how a cleaning decision affects what the data suggests.

Prior knowledge

Students should be able to read simple tables, compare numbers and categories, calculate a frequency, and describe a data display using words such as “most”, “least” and “difference”.

Vocabulary

Variable, categorical, numerical, dataset, entry, frequency, category, consistent, missing, invalid, outlier, cleaning rule, conclusion, bias.

Curriculum links

  • Te Mātaiaho Mathematics and Statistics, Phase 4 Statistics: developing knowledge from data.
  • Te Mātaiaho Mathematics and Statistics, Phase 4 Statistics: visualising and representing data.
  • Te Mātaiaho Mathematics and Statistics, Phase 4 Statistics: interpreting data and communicating findings.
  • Mathematical and statistical thinking, reasoning, communicating and using digital tools purposefully.

Lesson structure (60 minutes)

  1. 0–6 min · Engaging starter. Open with the provocative before-and-after data hook and display two possible claims: “Most students walk to school” and “Most students use active transport.” Students silently decide which claim is safer, then discuss what information they would need before trusting either claim.

  2. 6–14 min · Model the problem. Show this messy dataset on the messy dataset modelling slide:

Mode: Walk, walk, Wlk, Bus, bus, Car, bike, Bicycle, BIK,?, Train, car, Bus, Bike, walk, 12, car, Bus, absent, Walk

Explain that each entry represents one student’s reported travel mode. Model identifying the variable (travel mode), recognising categories, and creating cleaning rules: standardise capitalisation and equivalent labels; treat “?” and “absent” as missing; flag 12 as invalid because it is not a travel mode; do not silently guess missing values. Students annotate the entries on the data-cleaning investigation sheet.

  1. 14–25 min · Guided practice and error analysis. Revisit the dataset using the cleaning rules and questioning slides. Students suggest corrections and defend them while the teacher asks: “Is ‘bike’ the same category as ‘Bicycle’?”, “What evidence allows us to change ‘BIK’?”, “Should a missing response be counted as ‘none’?”, and “What could go wrong if we simply delete unusual entries?” Present three deliberately flawed approaches: counting Bus and bus separately, changing ? to Car, and counting 12 as a category. Pairs identify the error, correct it and explain its possible effect on the conclusion. Emphasise that cleaning rules must be transparent and applied consistently.

  2. 25–39 min · Collaborative organisation. Pairs complete the first section of the data-cleaning investigation sheet: record variables, classify entries as valid/missing/invalid, apply agreed cleaning rules, and sort the valid responses into a frequency table. The teacher circulates and checks that students do not count missing or invalid entries as valid categories. Pause at minute 33 for a checkpoint: students hold up fingers for the number of valid responses and explain how they found it.

  3. 39–53 min · Differentiated independent task. Students complete the table, answer interpretation questions and write a short conclusion on the data-cleaning investigation sheet. Support students receive a reduced dataset, a category bank (Walk, Bus, Car, Bike, Train) and sentence starters: “The most common valid response is…”, “I did not count… because…”. Core students compare the cleaned table with a table that counts inconsistent labels separately. Challenge students create a second defensible cleaning rule for BIK or missing responses, recalculate the relevant frequencies, and explain whether the overall conclusion changes. Students may use a spreadsheet if available, but must still state their rules.

  4. 53–60 min · Plenary and exit ticket. Use the conclusion and exit-ticket slides to compare findings. Students complete the final three questions on the worksheet, then share one rule and its consequence. Collect responses as the exit ticket.

Resources

  • data-cleaning lesson deck covering the hook, dataset, modelling, questions, instructions and plenary
  • the data-cleaning investigation sheet
  • Projector or interactive display
  • Calculators or spreadsheet software
  • Highlighters in two or three colours
  • Mini-whiteboards or scrap paper
  • Pens and rulers

Assessment

  • During modelling, ask students to identify the variable and justify whether an entry is valid, missing or invalid.
  • At the guided-practice checkpoint, accept: 16 valid responses, 3 missing/invalid responses (?, absent, 12), with equivalent labels standardised. Check that students explain rather than guess.
  • Exit-ticket answers and marking guidance:
  • “Name the variable”: travel mode (1 mark).
  • “Give one cleaning rule”: for example, combine walk and Walk, or treat ? as missing (1 mark).
  • “Why can cleaning affect the conclusion?” Expected answer: combining or excluding entries changes frequencies and therefore may change which category appears most common (2 marks: change identified and linked to conclusion).

Differentiation

  • Support with a worked example, colour-coding, category bank, reduced dataset and sentence starters; read instructions aloud and pair students strategically.
  • Provide extension through alternative defensible rules and comparison of resulting conclusions, rather than simply adding more entries.
  • For EAL learners, use icons for travel modes, explicitly model vocabulary and allow oral rehearsal before written explanations.
  • For students requiring additional support, provide a partially completed frequency table and check one row at a time; avoid assuming that an unusual value is automatically an error.

Misconceptions to address

  • “Different capitalisation means different categories”: explain that a category label should be standardised.
  • “Missing means zero or none”: missing means the response is unknown.
  • “Every unusual entry must be deleted”: investigate whether it is valid before removing it.
  • “Cleaning makes data objective”: cleaning involves choices, so rules and their effects must be reported.

Create Your Own AI Lesson Plan

Join thousands of teachers using Kuraplan AI to create personalized lesson plans that align with Aligned with New Zealand Curriculum in minutes, not hours.

AI-powered lesson creation
Curriculum-aligned content
Ready in minutes

Created with Kuraplan AI

Generated using openai/gpt-5.6-luna

🌟 Trusted by 1000+ Schools

Join educators across New Zealand