16 August 2026
GCSE Maths Statistics Revision Guide
A complete GCSE statistics revision guide covering averages from tables, cumulative frequency, histograms, scatter graphs, sampling and probability.
Statistics and probability make up a substantial and predictable chunk of every GCSE maths paper, yet many students lose easy marks here through careless reading of tables and graphs rather than a genuine lack of understanding. This guide works through the main statistics topics in the order they tend to build on one another: averages, cumulative frequency and box plots, histograms, scatter graphs, sampling, and finally probability trees and Venn diagrams.
If you're working through this guide as part of a tight revision schedule, our emergency one-week revision plan shows where statistics fits alongside everything else you need to cover.
Averages from frequency tables
The three averages — mean, median and mode — are familiar from a raw list of numbers, but GCSE questions usually present data grouped into a frequency table, and this is where marks are lost.
Mean from an ungrouped frequency table
When data is given as a value and a frequency (for example, "3 goals scored by 5 matches"), the mean is:
mean = (sum of value × frequency) / (sum of frequency)
Add a column to your table for value × frequency, total that column, and divide by the total frequency. The most common error is dividing by the number of rows in the table instead of the total frequency — always check you're dividing by the number of data items, not the number of categories.
Mean from a grouped frequency table (estimated mean)
When data is grouped into class intervals (for example, "10 ≤ x < 20"), you cannot find the exact mean because you don't know the individual values. Instead, use the midpoint of each class as a representative value:
estimated mean = (sum of midpoint × frequency) / (sum of frequency)
Because this uses midpoints rather than actual data, the answer is always described as an "estimate" — write that word in your answer, as it is sometimes explicitly credited.
Median and mode from grouped data
The modal class is simply the class interval with the highest frequency — you don't need a formula, just careful reading. For the median from grouped data at GCSE, you typically identify the class interval containing the middle value using cumulative frequency, without needing the full interpolation formula used at A-Level (though Higher tier papers sometimes expect an estimate using interpolation — check your specification).
Cumulative frequency graphs and box plots
Cumulative frequency graphs plot the running total of frequency against the upper class boundary of each group. Common mistakes include plotting against the midpoint instead of the upper boundary, and forgetting to include a point at the very start of the scale where cumulative frequency is zero.
Once plotted, you can read off:
- The median (value at half the total frequency)
- The lower quartile (value at one quarter of the total frequency)
- The upper quartile (value at three quarters of the total frequency)
- The interquartile range (upper quartile minus lower quartile)
A box plot then summarises this five-number summary visually: minimum, lower quartile, median, upper quartile and maximum. Common exam tasks include:
- Drawing a box plot from a cumulative frequency graph or a given five-number summary
- Comparing two box plots, where you must comment on both a measure of average (usually the median) and a measure of spread (usually the interquartile range), and relate both back to the context of the question
- Identifying outliers using the standard rule involving 1.5 × interquartile range, where this is specified
The comparison question is worth practising specifically, because a huge number of marks are lost simply by comparing only the median and forgetting the spread, or by making a comparison without referring to what the data actually represents (for example, "class B's revision times were more spread out than class A's" rather than just "the IQR is bigger").
Histograms with unequal class widths
Histograms differ from bar charts because the area of each bar, not the height, represents frequency. This matters specifically when class widths are unequal, which is exactly when GCSE questions test histograms.
The key relationship is:
frequency density = frequency ÷ class width
To draw a histogram, calculate frequency density for each class and plot that on the vertical axis against the class boundaries on the horizontal axis. To read a histogram (the more common exam task), you often need to do the reverse: use the height (frequency density) and the width of a bar to recover the frequency, or use a known frequency for one bar to find the scale for the others.
A frequent trap is questions where you are given the frequency for one bar and asked to find missing frequencies for others using the same frequency-density scale, or asked to find a frequency represented by only part of a bar (for example, "how many people took between 25 and 30 minutes" when the bar covers 20 to 30 minutes). In that case, assume the frequency is spread evenly across the bar and calculate the density × the sub-width you need.
Scatter graphs and correlation
Scatter graph questions test three separate skills, and papers often examine all three in the same question:
- Describing correlation: positive, negative or no correlation, and commenting on strength (strong or weak).
- Drawing a line of best fit: by eye, roughly through the middle of the data with a similar number of points on each side.
- Interpolation and extrapolation: using the line of best fit to estimate a value within the range of the data (interpolation, generally reliable) versus outside the range of the data (extrapolation, which should be flagged as less reliable).
Always comment on reliability when a question asks you to estimate a value using extrapolation — this is a common "explain" mark that is easy to secure once you know to look for it. Also remember that correlation describes a relationship between two variables, not necessarily a causal one; if a question asks you to comment on this, say that correlation does not imply causation and suggest another explanation if one is obvious from the context.
Sampling
Sampling questions are conceptual rather than calculation-heavy, and are often under-revised as a result. Know the difference between:
- Population: the entire group you are interested in.
- Sample: a smaller group selected from the population, used because surveying the whole population is impractical.
- Random sampling: every member of the population has an equal chance of selection, often carried out by numbering the population and using random numbers.
- Stratified sampling: the population is divided into groups (strata), and a sample is taken from each group in proportion to its size within the population.
For stratified sampling calculations, the method is:
number sampled from a group = (size of group ÷ size of population) × total sample size
Always double-check your groups sum back to the correct total sample size, and consider whether rounding is required — most mark schemes expect sensible rounding to whole numbers of people, with a comment if this causes the total to be slightly off.
Also be ready to comment on sampling method problems — for example, why sampling only people leaving a supermarket on a weekday morning might not represent the whole population, since it will oversample certain groups (such as those not in typical weekday employment) and undersample others.
Probability trees and Venn diagrams
Probability trees
Tree diagrams are used for combined events, particularly where events happen in sequence. Key rules:
- Multiply along branches to find the probability of a particular sequence of outcomes.
- Add the probabilities of separate branches (separate routes through the tree) that all satisfy the event you're interested in.
- Probabilities on branches from the same point must sum to 1.
- For "without replacement" problems, the probabilities on the second set of branches change depending on the first outcome — this is the most common source of errors, so always recalculate the denominator for the second event.
Venn diagrams
Venn diagrams organise outcomes into sets, usually two or three overlapping circles inside a rectangle representing the whole sample space. Standard notation to know:
- A ∪ B: A or B (union)
- A ∩ B: A and B (intersection)
- A': not A (complement)
- P(A|B): probability of A given B (conditional probability, more common at Higher tier)
When completing a Venn diagram from worded information, always start from the intersection (the overlap) and work outward, since the overlap is usually the piece of information you're given directly or can deduce first, and every other region depends on it.
Bringing it together: exam strategy for statistics
Statistics questions are often multi-part, moving from a straightforward calculation to an interpretation or comparison in the final part. Don't rush the final "comment on" or "compare" parts — they are usually worth as many marks as the calculation before them, and are marked on the quality and specificity of your written explanation, not just correct arithmetic.
If you want to practise these skills under timed conditions, use full papers from past papers and pay particular attention to the mark scheme wording for comparison and explanation questions, since it shows you exactly the level of detail examiners expect. Structured resources can also help if you want topic-by-topic practice before attempting full papers, and once you've completed some timed practice, the grade boundary predictor will give you a sense of where your marks would currently place you.
If statistics consistently costs you marks despite practice, a tutor can quickly identify whether the issue is calculation, graph-reading, or written explanation — A-Level Maths Tutoring works with students on exactly this kind of targeted diagnosis.
Takeaway
GCSE statistics rewards careful reading as much as calculation: dividing by the correct total in a frequency table, reading cumulative frequency graphs from the upper class boundary, using frequency density rather than raw height for histograms, and commenting on both average and spread when comparing distributions. Learn the mechanical methods first, then practise the written "explain" and "compare" questions specifically, since these are where most marks are lost even by otherwise strong students.