Skip to content
Marlbridge

Revision Notes

AQA GCSE Mathematics 8300: Statistics – Revision Notes

Condensed AQA GCSE Maths 8300 Statistics revision notes: averages, charts, histograms, box plots and scatter graphs, with a self-test and answers.

Subject
Mathematics
Level
GCSE
Topic
Statistics
Updated

Aligned to AQA GCSE Mathematics (8300), For first teaching 2015. Official specification .

Syllabus page (what it covers and how it is assessed): AQA GCSE Mathematics.

Syllabus points this page covers

8300

  • 6 Statistics (whole topic)

Found an error? Report a correction.

Need help with this topic? Request a free trial class for GCSE Mathematics (8300).

These revision notes cover Topic 6, Statistics (S1 to S6), of the AQA GCSE Mathematics (8300) specification, for teaching from September 2015 with exams from May/June 2017 (version 1.0). Most of the topic is for Foundation and Higher tier. Histograms and cumulative frequency graphs (S3), and box plots, quartiles and interquartile range (S4) are Higher tier only. Statistics can appear on the non-calculator Paper 1 or the calculator Papers 2 and 3, in the May/June and November series.

For full explanations and worked examples, read the Statistics study guide. When you are ready, try the Statistics practice questions. The AQA GCSE Mathematics hub, the printable checklist and the free diagnostics help you plan what to revise next. The Probability revision notes cover the rest of this area.

Key words

Term Meaning
Population The whole group you want to find out about
Sample The part of the population you collect data from
Random sample Every member has an equal chance of being chosen
Biased sample Some members are more likely to be chosen, so results are skewed
Primary data Data you collect yourself
Secondary data Data someone else collected
Discrete data Separate values, usually counts
Continuous data Measured values that can take any value in a range
Outlier A value far from the rest of the data
Bivariate data Pairs of values, plotted on a scatter graph

Sampling (S1, S5)

  • A sample is quicker and cheaper than the whole population, but only gives an estimate.
  • Bias comes from where, when and who you ask: one place, one time of day, or a group with a shared interest.
  • A bigger random sample gives a more reliable estimate.
  • To scale up: proportion in the sample × population size.
  • To describe a population (S5), use the sample’s averages and spread as estimates, and say they are estimates.

Small reminder: 9 out of a random sample of 60 bulbs fail a test. For a batch of 3000 bulbs, the estimate is 9/60 × 3000 = 450 failures. A second sample could give a different figure, so say “about 450”.

Charts and diagrams (S2)

Data Suitable diagram
Categories Bar chart, pictogram, pie chart
Ungrouped discrete numbers Vertical line chart, bar chart
Values over time Line graph (time series)
Grouped continuous data Histogram (Higher tier only)
Running totals, to find median and quartiles Cumulative frequency graph (Higher tier only)
Two variables together Scatter graph

Pie chart angle = (frequency ÷ total) × 360°. Check that the angles add up to 360°.

For a time series, describe the overall trend and any repeating seasonal pattern.

Small reminder: quarterly visitors to a park of 2, 5, 9, 4 (thousands) in one year and 3, 6, 11, 5 in the next show a rising trend, with the same peak in the third quarter each year.

Formulas and rules

What How
Mean Σx ÷ n
Mean from a frequency table Σ(x × f) ÷ Σf
Estimated mean (grouped) Σ(midpoint × f) ÷ Σf
Median position (n + 1)/2
Range Largest − smallest
Frequency density Frequency ÷ class width (Higher tier only)
Frequency from a histogram Frequency density × class width (Higher tier only)
Quartile positions (list of n values) LQ at (n + 1)/4, UQ at 3(n + 1)/4 (Higher tier only)
Interquartile range UQ − LQ (Higher tier only)

Method in steps: estimated mean from grouped data (S4)

  1. Find each class midpoint: (lower + upper) ÷ 2.
  2. Multiply each midpoint by its frequency.
  3. Add these products.
  4. Divide by the total frequency, not the number of classes.
  5. Check that your answer lies inside the range of the classes.

Small reminder: the midpoint of 20 < t ≤ 30 is 25.

Method in steps: median from a frequency table (S4)

  1. Add the frequencies to get n.
  2. Find the position (n + 1)/2.
  3. Keep a running total of the frequencies until you pass that position.
  4. The value (or class, for grouped data) where you pass it holds the median.

Small reminder: frequencies 3, 6, 5, 2 for the values 1, 2, 3, 4 give n = 16. The position is 8.5. Running totals 3, 9, … so the 8th and 9th values are both 2, and the median is 2. The mean is (3 + 12 + 15 + 8) ÷ 16 = 38 ÷ 16 = 2.375.

Method in steps: histogram (S3) – Higher tier only

  1. Work out class width for each class.
  2. Frequency density = frequency ÷ class width.
  3. Plot frequency density on the vertical axis, with no gaps between bars.
  4. To read a frequency back, find the bar’s area: density × width.
  5. For part of a class, take the matching fraction of that bar’s frequency.

Small reminder: the class 20 < t ≤ 40 has frequency 30, so its frequency density is 30 ÷ 20 = 1.5. To estimate how many values lie between 20 and 25, take 5/20 of the class: 5/20 × 30 = 7.5, so about 8. This assumes the values are spread evenly across the class.

Method in steps: cumulative frequency (S3) – Higher tier only

  1. Add up the frequencies as you go.
  2. Plot each running total at the upper class boundary; start at 0 at the first lower boundary.
  3. Join with a smooth curve or straight lines.
  4. Median: read across at n/2. LQ at n/4. UQ at 3n/4.
  5. To find “how many more than x”, read the cumulative frequency at x and subtract from n.

Box plots (S4) – Higher tier only

  • Five values: minimum, LQ, median, UQ, maximum.
  • Draw above a scale; the box runs from LQ to UQ with a line at the median.
  • The IQR is not affected by extreme values, so it is a better measure of spread than the range when there are outliers.
  • Reading a box plot: the median is the line inside the box, not the middle of the whole diagram. About a quarter of the data lies in each section.

Scatter graphs (S6)

  • Correlation: positive, negative or none; strong or weak.
  • Line of best fit: straight, by eye, through the middle of the points; ignore outliers.
  • Interpolation (inside the data) is fairly reliable. Extrapolation (outside the data) may not be.
  • Correlation does not mean one variable causes the other. A third factor, such as the season, may link them.
  • On Paper 1 you estimate from your own line by reading the graph. Show the lines you drew to read the value, so a method mark can be given if the reading is slightly off.

Must-know distinctions

This Is not the same as
Discrete (counted) Continuous (measured)
Primary (you collected it) Secondary (someone else did)
Bar chart: gaps, height = frequency Histogram: no gaps, area = frequency
Range (all data; hit by outliers) IQR (middle half; not hit by outliers)
Interpolation (inside the data) Extrapolation (outside the data)
Correlation (values move together) Causation (one makes the other change)

Comparing two distributions

Make one comparison of an average and one of spread, each with a sentence in context. For example: “On average Group B took longer, as its median is higher. Group A’s times were more consistent, as its IQR is smaller.”

Quick self-test

  1. Find the mean, median, mode and range of 4, 7, 7, 9, 13.
  2. A pie chart shows 45 people. What angle represents one person?
  3. Write the midpoint of the class 20 < t ≤ 30.
  4. (Higher tier only) A class 10 < h ≤ 25 has frequency 30. Find its frequency density.
  5. (Higher tier only) A histogram bar has frequency density 0.8 and width 15. Find the frequency.
  6. (Higher tier only) Find the quartiles and IQR of 2, 4, 5, 8, 9, 11, 14.
  7. As outdoor temperature rises, heating costs fall. Name the type of correlation.
  8. A random sample of 50 trees from a wood of 1500 has 8 diseased trees. Estimate the number of diseased trees in the wood.
  9. The mean of six numbers is 12. A seventh number, 19, is added. Find the new mean.
  10. Is “the number of cars passing a school in an hour” discrete or continuous?

Answers

  1. Mean = 40 ÷ 5 = 8; median 7; mode 7; range 13 − 4 = 9
  2. 360 ÷ 45 = 8°
  3. (20 + 30) ÷ 2 = 25
  4. 30 ÷ 15 = 2
  5. 0.8 × 15 = 12
  6. LQ (2nd value) = 4, median (4th) = 8, UQ (6th) = 11, IQR = 7
  7. Negative correlation
  8. 8/50 × 1500 = 240
  9. (6 × 12 + 19) ÷ 7 = 91 ÷ 7 = 13
  10. Discrete (it is a count)

Where marks are usually lost

  • Dividing by the number of classes instead of the total frequency for an estimated mean.
  • Using class upper bounds instead of midpoints.
  • Plotting cumulative frequency at midpoints.
  • Labelling the vertical axis of an unequal-width histogram as “frequency”.
  • Comparing only averages, or writing comparisons with no reference to the context.
  • Pie chart angles that do not add up to 360°, with no check.
  • Using a line of best fit to predict far outside the data without comment.
  • Claiming causation from correlation.
  • Forgetting that a grouped mean is an estimate, when asked to explain why.

Official syllabus

AQA GCSE Mathematics (8300) specification, for teaching from September 2015, exams from May/June 2017, version 1.0, published by AQA – section 3.6 Statistics (S1 to S6).

Get free revision emails (optional)

Occasional emails with practice questions, worked explanations and links to free resources for the qualification and subjects you choose. No spam, and you can unsubscribe from any email. The free tools on this site never need an email.

Subjects (optional, up to 6)

Choose a qualification to see its subjects.

Related resources

Related articles

Studying this with a teacher

Working through Mathematics GCSE?

This page is free and stays free. If you would rather be taught it, Marlbridge runs Mathematics classes one-to-one and in small groups of up to 15, online in your own time zone. The first trial class is free. WhatsApp replies within an hour (8am–11pm Pakistan time, every day); email the same day.