Practice Questions
Data and Its Collection: Practice Questions
Original exam-style practice questions with full worked answers on sampling methods, data types, bias and questionnaire design.
- Subject
- Statistics
- Level
- IGCSE
- Topic
- Topic 1 – Data and Its Collection
- Author
- Marlbridge Academic Team
- Updated
Aligned to Cambridge IGCSE Statistics (0479), 2027. Official specification .
These are original questions written for Marlbridge, in the style and at the standard of the examination. They are not reproduced past-paper questions — examination boards hold copyright in their own papers. Use these alongside the official past papers available free from your board.
Related: Data and Its Collection revision notes
Section A
1. Distinguish between qualitative and quantitative data, and between discrete and continuous data, giving an example of each. [4]
2. Distinguish between a population, a census and a sample. [3]
Section B
3. Explain two advantages and two disadvantages of taking a sample rather than a census. [4]
4. Explain the following sampling methods and give one advantage and one disadvantage of each: simple random, systematic, stratified, quota. [12]
5. A school of 900 students has 400 in Key Stage 3, 300 in Key Stage 4 and 200 in Key Stage 5. A stratified sample of 90 is required.
(a) Calculate how many students should be taken from each stage. [3] (b) Explain how the individuals within each stratum should be selected. [2]
6. A questionnaire asks: “Don’t you agree that the school canteen has improved this year?”
(a) Identify two faults with this question. [2] (b) Rewrite it appropriately. [2] (c) Explain two other principles of good questionnaire design. [4]
Section C
7. A student’s height is recorded as 152 cm to the nearest cm. State the class boundaries within which the true height must lie. [2]
8. Before running a survey of 500 households, a researcher tests the questionnaire on 20 households first.
(a) State the name for this preliminary test. [1] (b) Explain one reason it is carried out. [2]
9. A researcher argues that increasing a biased sample from 50 to 500 will remove the bias. Explain why this is incorrect, and state two other sources of bias besides a poorly worded question. [4]
Answers
1. Qualitative data is non-numerical, describing a quality, e.g. eye colour [1]; quantitative data is numerical, e.g. height [1]. Discrete data can only take particular separate values, usually from counting, e.g. number of siblings [1]; continuous data can take any value in a range, usually from measuring, e.g. mass [1].
2. A population is the entire group being studied [1]. A census collects data from every member of that population [1]. A sample is a selected subset of the population from which conclusions about the whole are drawn [1].
3. Advantages: a sample is much cheaper and quicker to collect and process [1]; it is often the only option, since a census is impractical for a large population or destructive testing [1]. Disadvantages: a sample may be unrepresentative, so conclusions may not generalise [1]; there is always sampling error, so results are estimates with a margin of uncertainty rather than exact values [1].
4. Simple random — every member of the population has an equal chance of selection, using random numbers [1]. Advantage: free from bias [1]. Disadvantage: requires a complete list (sampling frame) of the population, and may by chance miss a subgroup [1]. Systematic — select every nth member from a list after a random start [1]. Advantage: simple and quick [1]. Disadvantage: if the list contains a periodic pattern matching the interval, the sample will be biased [1]. Stratified — the population is divided into strata and a sample is taken from each in proportion to its size [1]. Advantage: most representative, since every subgroup is included proportionally [1]. Disadvantage: the strata must be known in advance and the method is more time-consuming [1]. Quota — the interviewer is told how many people of each type to obtain but chooses them freely [1]. Advantage: quick and needs no sampling frame [1]. Disadvantage: not random — interviewer choice introduces bias, since they approach people who look approachable [1].
5. (a) Sampling fraction = 90 ÷ 900 = 0.1 [1]. KS3: 40; KS4: 30 [1]; KS5: 20 [1]. (b) Within each stratum, students should be selected by simple random sampling, for example by numbering them and using random numbers [1] [1].
6. (a) It is a leading question — “Don’t you agree” pressures the respondent towards agreement [1]; it has no time frame or defined scale, and “improved” is vague — improved in what respect? [1] (b) For example: “How would you rate the school canteen this year? Very good / Good / Satisfactory / Poor / Very poor” [1] [1]. (c) Any two, 2 marks each: avoid overlapping or incomplete response categories — options such as 0–10, 10–20 leave the respondent unsure where 10 belongs [1] [1]. Avoid personal, sensitive or embarrassing questions, or place them at the end and make responses anonymous, since they otherwise reduce the response rate and truthfulness [1] [1]. Keep the questionnaire short and use simple, unambiguous language with no jargon or double negatives, so every respondent interprets each question the same way [1] [1].
7. A height of 152 cm to the nearest cm lies in the interval 151.5 ≤ h < 152.5 [2] (1 mark for each correct bound, using ≤ on the lower bound and < on the upper so adjacent classes never overlap).
8. (a) A pilot survey [1]. (b) Any one: it tests the questionnaire on a small group before the full survey runs, exposing ambiguous wording, questions that are misunderstood, or response categories that do not cover every answer, so these can be fixed before the real data collection begins [1] [1].
9. Increasing the sample size reduces sampling error, the random variation between samples, but does not remove bias, which is a systematic fault in the method itself — a biased method stays biased at any size, however large [2]. Any two other sources, 1 mark each: an incomplete sampling frame (the list used to select the sample does not cover the whole population); non-response (those who reply may differ systematically from those who do not); self-selection (only strongly motivated individuals choose to take part); the interviewer effect (respondents answer differently depending on who is asking) [1] [1].
Where marks are usually lost
- Calling shoe size continuous data.
- Confusing stratified with quota sampling.
- Rounding stratified sample sizes so they no longer total the required sample.
- Rewriting a leading question but leaving the response options unbalanced.
Related resources
-
Study Guides
Cambridge IGCSE Statistics: Data and Its Collection (0479)
Sampling methods, survey design and classifying data -- the opening topic of Cambridge IGCSE Statistics (0479), a twelve-topic syllabus assessed across two compulsory papers.
Statistics · Cambridge · IGCSE
-
Study Guides
Cambridge O-Level Statistics: Data and Its Collection (4040)
Sampling, survey design and data classification -- the opening topic of Cambridge O Level Statistics (4040), a twelve-topic syllabus closely mirroring sibling IGCSE Statistics 0479.
Statistics · Cambridge · O LEVELS
-
Practice Questions
Data and Its Collection (O Level 4040): Practice Questions
Original exam-style practice questions with full worked answers on data types, sampling and census methods for Cambridge O Level Statistics 4040.
Statistics · Cambridge · O LEVELS
Related articles
-
curriculum guides
Choosing subjects at IGCSE and A Level
How subject choices at 14 and 16 affect university options later, and how to keep pathways open without overloading a timetable.
28 July 2026
-
study skills
How to revise for a science examination
Most science revision fails because it rereads notes instead of retrieving them. A practical method for revising physics, chemistry and biology in the weeks before a paper.
14 July 2026
Working through Statistics? Tutoring covers the same material with a teacher.
Find Learning Support