Skip to content
Marlbridge

Study Guides

Cambridge O-Level Statistics: Data and Its Collection (4040)

Sampling, survey design and data classification -- the opening topic of Cambridge O Level Statistics (4040), a twelve-topic syllabus closely mirroring sibling IGCSE Statistics 0479.

Subject
Statistics
Level
O LEVELS
Topic
Topic 1 – Data and Its Collection
Updated

Aligned to Cambridge O Level Statistics (4040), 2025-2027. Official specification .

Found an error? Report a correction.

This guide covers Topic 1 Data and Its Collection, the first of twelve topic areas in Cambridge O Level Statistics (4040), for examination 2025-2027. The syllabus closely mirrors sibling Cambridge IGCSE Statistics (0479), sharing the same overall topic sequence with slightly different topic-name wording.

Where this fits in 4040

Topic 1 sets up how statistical data is gathered and classified before the course progresses through Summary representation of data, formation of frequency distributions, statistical measures, and eventually probability and time series. Candidates take two compulsory components, both of which draw on the sampling and classification skills developed here.

Syllabus coverage

CAMBRIDGE O LEVEL STATISTICS (4040) — TOPIC 1 DATA AND ITS COLLECTION

Topic 1 covers how statistical data is collected, including sampling methods and their appropriate use, survey design considerations, and the classification of data into different types – the groundwork needed before data can be summarised, represented or analysed in the topics that follow.

How to approach it

Because sampling questions typically present a real-world scenario and ask which method is most appropriate, build a habit of justifying method choice with reference to the scenario’s specific constraints (cost, time, access to a sampling frame) rather than reciting general advantages and disadvantages. Practise working through the mechanics of selecting a sample step by step, since exam questions often require showing the working, not just stating a final sample. Because correctly classifying data (discrete versus continuous, qualitative versus quantitative) determines which statistical techniques are appropriate later in the course, treat this classification skill as something to get automatic early, rather than revisiting it only when it resurfaces in later topics.

Official syllabus

Cambridge O Level Statistics (4040) syllabus for examination 2025, 2026 and 2027 — cambridgeinternational.org.

Classifying data

Qualitative data describes a category and cannot be measured — hair colour, town of birth. Quantitative data is numerical, and splits into:

  • Discrete — from counting, taking only certain values. Number of siblings, goals scored.
  • Continuous — from measuring, taking any value in a range. Height, mass, time.

Continuous data is always recorded to an accuracy, so class boundaries must be handled correctly. A time recorded as 12.4 seconds to 1 decimal place lies in 12.35 <= t < 12.45.

Variables may also be described as bivariate where two are recorded per item, which is what makes scatter diagrams and correlation possible.

Primary and secondary sources

Primary data is collected by the investigator for the current purpose: relevant, with known reliability, but slow and costly. Secondary data already exists in published statistics, records or databases: quick and cheap, but possibly outdated, incomplete, or gathered under different definitions.

Definitional differences matter more than students expect — two sources reporting “unemployment” may count different people.

Population, census and sample

A census covers every member of the population: complete and accurate, but expensive, slow, and impossible where testing destroys the item.

A sample covers part of it. The sampling frame is the list from which it is drawn, and gaps in that frame are a common source of bias.

Method Procedure Limitation
Simple random Random numbers; all equally likely Requires a complete frame
Systematic Every nth item from a random start Bias if the list is periodic
Stratified Proportional numbers from each group Strata must be identifiable
Quota Set numbers per category, interviewer chooses Not random; interviewer bias
Cluster Whole groups selected at random Less precise if clusters differ
Opportunity Whoever is available Unrepresentative

Collecting data and avoiding bias

Questionnaires should use clear language, avoid leading and ambiguous questions, offer response options that are exhaustive and non-overlapping, and place sensitive items last. A pilot survey tests the instrument before full deployment.

Bias arises from an incomplete sampling frame, non-response, self-selection, leading questions, and interviewer influence. Increasing sample size reduces sampling error but does not remove bias.

Worked example. A researcher wants to survey the reading habits of adults in a town, so she stands outside a bookshop on a Saturday morning and asks every 5th person who passes – systematic sampling applied to a self-selected location. This sample is unlikely to be representative for two separate reasons: it only includes people near a bookshop, who are already more likely to be interested in reading than the general adult population, so the results will overstate reading habits; and it is conducted only on a Saturday morning, which excludes adults who work or are otherwise unavailable at that time, so the sample also fails to represent all adults’ schedules and habits. The fix is not a bigger sample but a different sampling frame – surveying at several different locations and at different times and days, so the frame is no longer skewed toward people who were already likely to be readers.

Data may be recorded using a tally chart or a data collection sheet, then organised into a grouped frequency table with equal class widths where possible.

Worked example

A factory has 240 workers in production, 90 in sales and 30 in admin. Take a stratified sample of 36.

Total = 360,  sampling fraction = 36 / 360 = 1/10

Production: 240 / 10 = 24
Sales:       90 / 10 =  9
Admin:       30 / 10 =  3
                       ---
                        36

Each stratum is then sampled randomly within itself — a stratified sample is still random, just random within groups.

Common mistakes

Calling shoe size continuous. Writing overlapping classes such as 10–20 and 20–30. Taking equal numbers from each stratum instead of proportional numbers. Forgetting that stratified sampling still requires randomness within each group. Claiming a larger sample eliminates bias. Omitting the random start in systematic sampling.

Quick revision checklist

  • Classify data as qualitative or quantitative, discrete or continuous, primary or secondary.
  • State class boundaries for continuous data recorded to a given accuracy.
  • Compare census and sample, and describe each sampling method with its limitation.
  • Calculate a stratified sample so the parts total correctly.
  • Identify sources of bias and explain why sample size does not remove them.
  • Design an unbiased questionnaire and explain the role of a pilot survey.

Related resources

Related articles

Working through Statistics? Tutoring covers the same material with a teacher.

Find Learning Support