Skip to content
Marlbridge

Study Guides

Pearson Edexcel International GCSE Mathematics A 4MA1: Statistics and probability – Study Guide

Study guide for Edexcel IGCSE Maths 4MA1 topic 6: charts, histograms, cumulative frequency, averages, quartiles, Venn diagrams and tree diagrams.

Subject
Mathematics
Level
IGCSE
Topic
Statistics and probability
Updated

Aligned to Pearson Edexcel IGCSE Mathematics (4MA1), Specification Issue 2, November 2017. Official specification .

Syllabus page (what it covers and how it is assessed): Pearson Edexcel IGCSE Mathematics.

Syllabus points this page covers

4MA1

  • 6 Statistics and probability (whole topic)
  • 6.1 Graphical representation of data
  • 6.2 Statistical measures
  • 6.3 Probability

Found an error? Report a correction.

Need help with this topic? Request a free trial class for IGCSE Mathematics (4MA1).

This study guide teaches topic 6, Statistics and probability, of the Pearson Edexcel International GCSE Mathematics A (4MA1) specification, Issue 2 (November 2017), first assessed in June 2018 with papers available in January and June. It covers section 6.1 (Graphical representation of data), 6.2 (Statistical measures) and 6.3 (Probability). Each section has Foundation content, which both tiers need, and extra content marked Higher tier only.

Each tier sits two 2-hour papers of 100 marks (1F and 2F, or 1H and 2H), each worth 50% of the qualification. A calculator may be used on every paper. No statistics or probability formula is printed on the formulae sheet, so learn the ones below.

See also the revision notes, the practice questions, the Edexcel IGCSE Mathematics hub, the printable 4MA1 checklist and the free 10-minute diagnostics.

What this topic covers

Section What you must be able to do Tier
6.1 A–C Present data as pictograms, bar charts, pie charts and two-way tables; tabulate data; interpret diagrams Both tiers
6.1 A–C Construct and interpret histograms (continuous data, unequal class intervals); construct and use cumulative frequency diagrams Higher tier only
6.2 A–D Understand averages; mean, median, mode and range of discrete data; estimate the mean of grouped data; modal class Both tiers
6.2 A–D Median from a cumulative frequency diagram; measures of spread; interquartile range from a list and from a cumulative frequency diagram Higher tier only
6.3 A–J Probability language and scale; theoretical and experimental probability; Venn diagrams; sample spaces; complements; addition rule for mutually exclusive events; expected frequency Both tiers
6.3 A–D Tree diagrams; independent events; simple conditional probability (without replacement); applying probability to problems Higher tier only

6.1 Graphical representation of data

Charts and tables (both tiers)

  • Pictogram: symbols stand for a stated number of items. Always give a key, and use part-symbols for part-amounts.
  • Bar chart: equal-width bars with gaps, for categories or discrete data. The height is the frequency.
  • Pie chart: each sector angle = (frequency ÷ total) × 360°.
  • Two-way table: shows two categories at once, with row and column totals. Fill it in by using the totals.

Worked example 1 (pie chart). 72 students choose a sport: football 26, tennis 14, swimming 20, other 12.

360° ÷ 72 = 5° per student
Football 26 × 5 = 130°, Tennis 14 × 5 = 70°,
Swimming 20 × 5 = 100°, Other 12 × 5 = 60°
Check: 130 + 70 + 100 + 60 = 360°

Worked example 2 (two-way table). 90 students are in Years 10 and 11. 48 are in Year 10. 50 have a school lunch and the rest bring a packed lunch. 17 Year 11 students bring a packed lunch. Year 11 has 90 − 48 = 42 students, so 42 − 17 = 25 have school lunch. Year 10 then has 50 − 25 = 25 school lunches and 48 − 25 = 23 packed lunches.

To interpret a diagram, read values carefully from the scale, compare categories, and answer in the context of the question.

Histograms (Higher tier only)

A histogram is for continuous data. The area of each bar represents the frequency, so with unequal class widths the vertical axis is frequency density:

frequency density = frequency ÷ class width
frequency = frequency density × class width

Worked example 3. Journey times t (minutes) for 40 people:

Time t 0 < t ≤ 10 10 < t ≤ 20 20 < t ≤ 30 30 < t ≤ 50
Frequency 6 14 12 8
Class width 10 10 10 20
Frequency density 0.6 1.4 1.2 0.4

Draw bars with no gaps, from 0 to 10, 10 to 20 and so on, with heights 0.6, 1.4, 1.2 and 0.4. The last bar is twice as wide and so only 0.4 high.

Cumulative frequency diagrams (Higher tier only)

Cumulative frequency is the running total. Plot each running total at the upper class boundary, start at the lower boundary of the first class with cumulative frequency 0, and join the points with a smooth curve or straight lines.

6.2 Statistical measures

Averages and range (both tiers)

  • Mean = total of the values ÷ number of values. For a frequency table, mean = Σfx ÷ Σf.
  • Median: the middle value when the data are in order; for n values it is the ((n + 1)/2)th value.
  • Mode: the most common value.
  • Range = largest − smallest.

The mean uses every value, but an extreme value distorts it; the median is then a better average. The mode suits non-numerical data.

Worked example 4. Goals scored in 20 matches:

Goals x 0 1 2 3 4
Frequency f 4 7 5 3 1
Σfx = 0 + 7 + 10 + 9 + 4 = 30      mean = 30 ÷ 20 = 1.5 goals
Median = the 10.5th value: the 10th and 11th values are both 1 (running totals 4, 11), so median = 1
Mode = 1 (highest frequency)       Range = 4 − 0 = 4

Grouped data (both tiers)

You do not know the exact values in a group, so use the midpoint of each class. The result is an estimate of the mean. The modal class is the class with the highest frequency.

Worked example 5. Using the journey times in worked example 3:

Midpoints:   5, 15, 25, 40
f × midpoint: 30, 210, 300, 320     Σ = 860
Estimated mean = 860 ÷ 40 = 21.5 minutes
Modal class: 10 < t ≤ 20

Spread and quartiles (Higher tier only)

A measure of spread shows how varied the data are. The range uses only two values; the interquartile range (IQR) = upper quartile − lower quartile covers the middle half of the data and ignores extreme values.

For a list of n values in order, the lower quartile is the ((n + 1)/4)th value and the upper quartile is the (3(n + 1)/4)th value.

Worked example 6. 3, 5, 6, 8, 9, 11, 12, 14, 15, 18, 21 (n = 11).

Lower quartile: (11 + 1)/4 = 3rd value = 6
Median:         (11 + 1)/2 = 6th value = 11
Upper quartile: 3 × (11 + 1)/4 = 9th value = 15
IQR = 15 − 6 = 9

From a cumulative frequency diagram with total n, read across at n/2 for the median, n/4 for the lower quartile and 3n/4 for the upper quartile, then down to the horizontal axis.

Worked example 7. Heights h (cm) of 80 plants:

Height h ≤ 10 ≤ 20 ≤ 30 ≤ 40 ≤ 50 ≤ 60
Cumulative frequency 4 20 40 60 72 80

Plot (0, 0), (10, 4), (20, 20), (30, 40), (40, 60), (50, 72), (60, 80). Read across at 40 for the median: 30 cm. Read at 20 and 60 for the quartiles: 20 cm and 40 cm, so the IQR is 20 cm. To find how many plants are taller than a value, read the cumulative frequency at that value and subtract from 80.

6.3 Probability

Language, scale and theory (both tiers)

An experiment has outcomes; an event is one or more outcomes. Outcomes are equally likely in a fair, random selection. Probabilities lie on a scale from 0 to 1: P(impossibility) = 0 and P(certainty) = 1. Write them as fractions, decimals or percentages; a ratio such as 1 : 4 is not a probability and usually loses the mark.

For equally likely outcomes, P(event) = number of favourable outcomes ÷ total number of outcomes.

  • Complement: P(A′) = 1 − P(A).
  • Mutually exclusive events cannot happen together: P(A or B) = P(A) + P(B).
  • Expected frequency = probability × number of trials.

Sample spaces (both tiers)

A sample space lists every possible outcome. For two coins it is (H, H), (H, T), (T, H), (T, T). For two successive events, list systematically or use a table.

Worked example 8. Spinner X shows 1, 2, 3 or 4; spinner Y shows 1, 2 or 3. Both are fair. There are 4 × 3 = 12 equally likely pairs. The pairs that total 5 are (2, 3), (3, 2) and (4, 1), so P(total = 5) = 3/12 = 1/4.

Experimental probability (both tiers)

Relative frequency = number of times the event happens ÷ number of trials. It estimates the probability, and the estimate improves with more trials.

Worked example 9. A drawing pin lands point up 132 times in 400 throws. Relative frequency = 132 ÷ 400 = 0.33. Expected number of “point up” results in 600 throws = 0.33 × 600 = 198.

Venn diagrams (both tiers)

Put the overlap in first, then the “only” regions, then the region outside both circles.

Worked example 10. In a group of 50 people, 17 are in A only, 8 are in both A and B, 12 are in B only and 13 are in neither.

P(A) = (17 + 8)/50 = 25/50 = 1/2
P(A or B or both) = (17 + 8 + 12)/50 = 37/50
P(B but not A) = 12/50 = 6/25

Tree diagrams and independent events (Higher tier only)

Two events are independent if one does not change the probability of the other. Then P(A and B) = P(A) × P(B). On a tree diagram, multiply along the branches and add the end results of different routes.

Worked example 11. The probability that a bus is late on any day is 0.2, independently of other days. For two days:

P(late both days)    = 0.2 × 0.2 = 0.04
P(late on exactly one) = 0.2 × 0.8 + 0.8 × 0.2 = 0.32
P(late at least once)  = 1 − 0.8 × 0.8 = 0.36

Conditional probability without replacement (Higher tier only)

If an item is not replaced, the second-stage probabilities change: both the number of that item and the total go down by one.

Worked example 12. A bag has 5 red and 3 blue counters. Two are taken without replacement.

P(both red)  = 5/8 × 4/7 = 20/56
P(both blue) = 3/8 × 2/7 = 6/56
P(same colour) = 26/56 = 13/28

Common errors

  • Drawing a histogram with frequency, not frequency density, on the vertical axis when class widths differ.
  • Plotting cumulative frequency at class midpoints instead of upper boundaries.
  • Dividing Σfx by the number of classes instead of Σf.
  • Giving the median as the middle frequency in a table, rather than the middle value.
  • Adding probabilities for events that are not mutually exclusive, or multiplying for events that are not independent.
  • Forgetting to reduce both numerator and denominator on the second branch without replacement.
  • Using “at least one” by listing some routes and missing others; 1 − P(none) is quicker and safer.

Next steps

Recap with the revision notes, then try the practice questions.

Official syllabus

Pearson Edexcel International GCSE in Mathematics (Specification A) (4MA1), Specification Issue 2, November 2017, Pearson Education Limited. Topic 6, Statistics and probability: 6.1 Graphical representation of data, 6.2 Statistical measures and 6.3 Probability, with Foundation and Higher tier content.

Get free revision emails (optional)

Occasional emails with practice questions, worked explanations and links to free resources for the qualification and subjects you choose. No spam, and you can unsubscribe from any email. The free tools on this site never need an email.

Subjects (optional, up to 6)

Choose a qualification to see its subjects.

Related resources

Related articles

Studying this with a teacher

Working through Mathematics IGCSE?

This page is free and stays free. If you would rather be taught it, Marlbridge runs Mathematics classes one-to-one and in small groups of up to 15, online in your own time zone. The first trial class is free. WhatsApp replies within an hour (8am–11pm Pakistan time, every day); email the same day.