Skip to content
Marlbridge

Practice Questions

IB DP Mathematics: Applications and Interpretation -- Statistics and Probability Practice Questions

Original practice questions with full worked answers covering descriptive statistics, probability, distributions and inferential statistics, for IB Diploma Programme Mathematics: Applications and Interpretation.

Level
IB
Topic
Statistics and probability
Updated

Aligned to International Baccalaureate IB Diploma Programme Mathematics: Applications and Interpretation (DP Mathematics: Applications and Interpretation), First assessment 2021. Official specification .

Found an error? Report a correction.

These are original questions written for Marlbridge, in the style and at the standard of the examination. They are not reproduced past-paper questions – the IB holds copyright in its own papers. Use these alongside the official past papers available through your school or the IB store.

Related: Statistics and Probability revision notes and the IB DP Mathematics: Applications and Interpretation syllabus guide.

Section A

1. A dataset is strongly skewed with several extreme outliers. State, with a reason, which measure of central tendency is more appropriate: the mean or the median. [2]

2. Two events A and B are independent, with P(A) = 0.4 and P(B) = 0.5. Find P(A and B). [2]

3. State one difference between a discrete distribution and a continuous distribution. [1]

Section B

4. A bag contains 5 red balls and 3 blue balls. Two balls are drawn without replacement.

(a) Draw or describe a tree diagram for this situation, showing the four possible outcomes. [2] (b) Calculate the probability that both balls drawn are red. [2] (c) Calculate the probability that the two balls drawn are different colours. [3]

5. A dataset of 8 exam scores is: 62, 65, 68, 70, 71, 73, 74, 95.

(a) Calculate the mean of the dataset. [1] (b) Calculate the median of the dataset. [1] (c) Explain, with reference to your answers, why the median might be reported as a more representative “typical” score than the mean for this dataset. [2]

6. A study finds a strong positive correlation between the number of hours a group of students spends on a revision app and their exam scores.

(a) State what this correlation shows about the relationship between the two variables. [1] (b) Explain why this correlation, by itself, does not establish that using the app caused the higher scores. [2]

Section C

7. A factory machine produces components with a known 3% defect rate, and each component is independent of the others. A sample of 20 components is taken.

(a) State which discrete distribution is appropriate to model the number of defective components in the sample, and identify its two parameters for this scenario. [2] (b) Explain what information would be needed to instead model the height of components using a normal distribution. [2] (c) A quality manager claims “since the defect rate is only 3%, we would never expect more than 2 defective components in a sample of 20.” Evaluate this claim, referring to the nature of a binomial distribution. [3]

Worked answers

1. The median, because the mean is sensitive to outliers and extreme values, which would pull it away from where most of the data actually sits; the median is unaffected by extreme values in the same way, making it more representative of a typical value in a skewed dataset. [2]

2. Since A and B are independent, P(A and B) = P(A) x P(B) = 0.4 x 0.5 = 0.20. [2]

3. A discrete distribution models outcomes that can be counted (e.g. a whole number of successes); a continuous distribution models outcomes that can take any value within a range (e.g. a measurement). [1]

4. (a) First draw: P(Red) = 5/8, P(Blue) = 3/8. Second draw (without replacement) depends on the first: if Red first, P(Red second) = 4/7, P(Blue second) = 3/7; if Blue first, P(Red second) = 5/7, P(Blue second) = 2/7 – giving four branch outcomes: RR, RB, BR, BB. [2] (b) P(RR) = (5/8) x (4/7) = 20/56 = 5/14. [2] (c) P(different colours) = P(RB) + P(BR) = (5/8)(3/7) + (3/8)(5/7) = 15/56 + 15/56 = 30/56 = 15/28. [3] (1 mark for identifying the two relevant branches, 2 for correct calculation)

5. (a) Mean = (62+65+68+70+71+73+74+95)/8 = 578/8 = 72.25. [1] (b) Ordering the data (already ordered), the median is the average of the 4th and 5th values: (70+71)/2 = 70.5. [1] (c) The score of 95 is a clear outlier relative to the rest of the dataset (which clusters between 62 and 74), and this single value pulls the mean up to 72.25, above all but one of the actual scores; the median (70.5) sits much closer to where most of the data actually lies, making it more representative of a “typical” score in this dataset. [2]

6. (a) It shows that, within this dataset, students who spent more hours on the revision app tended to also have higher exam scores. [1] (b) A correlation only shows that two variables tend to change together; it does not rule out other explanations, such as a third factor (e.g. general study motivation) causing both increased app use and higher scores, or the direction of causation being reversed (e.g. already-strong students choosing to use the app more) – so causation cannot be concluded from correlation alone. [2]

7. (a) The binomial distribution is appropriate, since there is a fixed number of independent trials (n = 20) each with the same probability of an outcome (defect). Parameters: n = 20, p = 0.03. [2] (b) A normal distribution would need the mean and standard deviation of the component heights, since a normal distribution is defined by these two parameters and describes continuous data clustering symmetrically around the mean. [2] (c) This claim is not well supported: a binomial distribution with n = 20 and p = 0.03 assigns some (small but non-zero) probability to every outcome from 0 to 20 defective components, including 3 or more – it does not rule out higher counts, it merely makes them progressively less likely. The manager’s claim conflates a low probability with an impossibility; a fuller answer would calculate or reference the actual probability of 3 or more defects (a small positive value) rather than assuming it is zero. [3]

Why this set stays interpretation-heavy

Nearly every worked answer ends with a sentence interpreting what the result means in context (5c, 6b, 7c in particular), because the revision notes identify this as the “other half” of a complete answer in this strand – assessment objectives in this course reward communication and interpretation alongside pure calculation, so a set of practice questions that stopped at the numerical answer would misrepresent what full marks actually requires.

Official syllabus

International Baccalaureate Organization, Diploma Programme Subject Brief – Mathematics: Applications and Interpretation, first assessment 2021, (c) 2019 – the same source cited by the Statistics and Probability revision notes and the IB DP Mathematics: Applications and Interpretation syllabus guide.

Related resources

Related articles

Working through Mathematics: Applications and Interpretation? Tutoring covers the same material with a teacher.

Find Learning Support