Skip to content
Marlbridge

Practice Questions

Data and Its Collection: Practice Questions

Original exam-style practice questions with full worked answers on sampling methods, data types, bias and questionnaire design.

Subject
Statistics
Level
IGCSE
Topic
Topic 1 – Data and Its Collection
Updated

Aligned to Cambridge IGCSE Statistics (0479), 2027. Official specification .

Found an error? Report a correction.

These are original questions written for Marlbridge, in the style and at the standard of the examination. They are not reproduced past-paper questions — examination boards hold copyright in their own papers. Use these alongside the official past papers available free from your board.

Related: Data and Its Collection revision notes


Section A

1. Distinguish between qualitative and quantitative data, and between discrete and continuous data, giving an example of each. [4]

2. Distinguish between a population, a census and a sample. [3]

Section B

3. Explain two advantages and two disadvantages of taking a sample rather than a census. [4]

4. Explain the following sampling methods and give one advantage and one disadvantage of each: simple random, systematic, stratified, quota. [12]

5. A school of 900 students has 400 in Key Stage 3, 300 in Key Stage 4 and 200 in Key Stage 5. A stratified sample of 90 is required.

(a) Calculate how many students should be taken from each stage. [3] (b) Explain how the individuals within each stratum should be selected. [2]

6. A questionnaire asks: “Don’t you agree that the school canteen has improved this year?”

(a) Identify two faults with this question. [2] (b) Rewrite it appropriately. [2] (c) Explain two other principles of good questionnaire design. [4]

Section C

7. A student’s height is recorded as 152 cm to the nearest cm. State the class boundaries within which the true height must lie. [2]

8. Before running a survey of 500 households, a researcher tests the questionnaire on 20 households first.

(a) State the name for this preliminary test. [1] (b) Explain one reason it is carried out. [2]

9. A researcher argues that increasing a biased sample from 50 to 500 will remove the bias. Explain why this is incorrect, and state two other sources of bias besides a poorly worded question. [4]


Answers

1. Qualitative data is non-numerical, describing a quality, e.g. eye colour [1]; quantitative data is numerical, e.g. height [1]. Discrete data can only take particular separate values, usually from counting, e.g. number of siblings [1]; continuous data can take any value in a range, usually from measuring, e.g. mass [1].

2. A population is the entire group being studied [1]. A census collects data from every member of that population [1]. A sample is a selected subset of the population from which conclusions about the whole are drawn [1].

3. Advantages: a sample is much cheaper and quicker to collect and process [1]; it is often the only option, since a census is impractical for a large population or destructive testing [1]. Disadvantages: a sample may be unrepresentative, so conclusions may not generalise [1]; there is always sampling error, so results are estimates with a margin of uncertainty rather than exact values [1].

4. Simple random — every member of the population has an equal chance of selection, using random numbers [1]. Advantage: free from bias [1]. Disadvantage: requires a complete list (sampling frame) of the population, and may by chance miss a subgroup [1]. Systematic — select every nth member from a list after a random start [1]. Advantage: simple and quick [1]. Disadvantage: if the list contains a periodic pattern matching the interval, the sample will be biased [1]. Stratified — the population is divided into strata and a sample is taken from each in proportion to its size [1]. Advantage: most representative, since every subgroup is included proportionally [1]. Disadvantage: the strata must be known in advance and the method is more time-consuming [1]. Quota — the interviewer is told how many people of each type to obtain but chooses them freely [1]. Advantage: quick and needs no sampling frame [1]. Disadvantage: not random — interviewer choice introduces bias, since they approach people who look approachable [1].

5. (a) Sampling fraction = 90 ÷ 900 = 0.1 [1]. KS3: 40; KS4: 30 [1]; KS5: 20 [1]. (b) Within each stratum, students should be selected by simple random sampling, for example by numbering them and using random numbers [1] [1].

6. (a) It is a leading question — “Don’t you agree” pressures the respondent towards agreement [1]; it has no time frame or defined scale, and “improved” is vague — improved in what respect? [1] (b) For example: “How would you rate the school canteen this year? Very good / Good / Satisfactory / Poor / Very poor” [1] [1]. (c) Any two, 2 marks each: avoid overlapping or incomplete response categories — options such as 0–10, 10–20 leave the respondent unsure where 10 belongs [1] [1]. Avoid personal, sensitive or embarrassing questions, or place them at the end and make responses anonymous, since they otherwise reduce the response rate and truthfulness [1] [1]. Keep the questionnaire short and use simple, unambiguous language with no jargon or double negatives, so every respondent interprets each question the same way [1] [1].

7. A height of 152 cm to the nearest cm lies in the interval 151.5 ≤ h < 152.5 [2] (1 mark for each correct bound, using ≤ on the lower bound and < on the upper so adjacent classes never overlap).

8. (a) A pilot survey [1]. (b) Any one: it tests the questionnaire on a small group before the full survey runs, exposing ambiguous wording, questions that are misunderstood, or response categories that do not cover every answer, so these can be fixed before the real data collection begins [1] [1].

9. Increasing the sample size reduces sampling error, the random variation between samples, but does not remove bias, which is a systematic fault in the method itself — a biased method stays biased at any size, however large [2]. Any two other sources, 1 mark each: an incomplete sampling frame (the list used to select the sample does not cover the whole population); non-response (those who reply may differ systematically from those who do not); self-selection (only strongly motivated individuals choose to take part); the interviewer effect (respondents answer differently depending on who is asking) [1] [1].


Where marks are usually lost

  • Calling shoe size continuous data.
  • Confusing stratified with quota sampling.
  • Rounding stratified sample sizes so they no longer total the required sample.
  • Rewriting a leading question but leaving the response options unbalanced.

Related resources

Related articles

Working through Statistics? Tutoring covers the same material with a teacher.

Find Learning Support