Small Group Tutorials

Here to help students catch up, keep up, and move ahead. Book a consultation here.

Secondary Mathematics: Averages, Spread and Data Interpretation

Secondary Mathematics · Worked Repair Guide 07

A data set can be summarised correctly and still be misunderstood. The mean may be mathematically correct while hiding a large spread. Two groups may have the same average but very different distributions. A percentage can look impressive until the denominator is revealed. A graph can exaggerate a small difference through its scale.

This guide develops one central habit: never interpret a summary without asking what information it compresses and what information it leaves out. The goal is not merely to calculate mean, median, mode and range, but to decide what they do and do not justify.

All data sets and contexts below are invented teaching examples. They are not reports about actual students, schools or businesses.

1. A data set has values, units and a question

Consider the numbers 4, 7, 7, 9, 13. Without context, they are simply values. If they represent minutes, marks or distances, the interpretation changes even though the arithmetic is identical.

Before calculating a summary, name the variable and its unit. Then ask what the problem wants to compare: typical value, total, most common value, central position, spread or change.

Entry check: for 4, 7, 7, 9, 13, the mean is 8, the median is 7, the mode is 7 and the range is 9. These four statistics are not competing answers. They describe different features.

2. The mean redistributes the total equally

The arithmetic mean equals the total of the observations divided by their number. For 4, 7, 7, 9, 13, the total is 40 and there are five observations, giving mean 8.

A useful interpretation is balance or equal redistribution. If the total 40 were shared equally across five positions, each would receive 8.

This interpretation makes reverse problems easy. If six values have mean 12, their total is 72. If five of the values sum to 61, the missing value is 11.

Worked example: four test scores have mean 68. A fifth score of 88 is added. The original total is 272; the new total is 360; the new mean is 72.

3. The median depends on order, not magnitude alone

The median is the middle value when the data are ordered. For an odd number of observations, it is the central item. For an even number, it is the mean of the two central items.

For 2, 4, 9, 15, the median is (4+9)/2 = 6.5. The value 6.5 need not appear in the original data.

Because the median depends on position, a very large extreme value may affect it less than it affects the mean. Compare 2, 4, 5, 6, 8 with 2, 4, 5, 6, 80. Both medians are 5, while the means are 5 and 19.4.

This does not make the median universally superior. It means the choice of summary depends on the distribution and the question.

4. The mode identifies frequency, not centrality

The mode is the most frequently occurring value or category. A data set can have one mode, more than one mode or no mode if no value occurs more frequently than the others.

For 2, 3, 3, 5, 5, 8, the data are bimodal: both 3 and 5 occur twice. It is inaccurate to average them and announce a mode of 4.

The mode can also summarise categorical data where numerical means are meaningless. If shoe colours are black, white, black, blue, black, the modal colour is black.

Use mode when frequency itself matters. Do not use it merely because it is easy to identify.

5. Range describes one simple form of spread

Range = maximum − minimum. For 4, 7, 7, 9, 13, the range is 9.

Range is sensitive to extreme values because it depends only on the endpoints. The data 4, 7, 7, 9, 13 and 4, 7, 7, 9, 103 differ drastically in range even though four values are identical.

Range alone also ignores the internal arrangement. The sets 0, 5, 5, 5, 10 and 0, 0, 5, 10, 10 both have range 10 but distribute values differently.

A spread measure should therefore be interpreted with a location measure and, where possible, with the distribution itself.

6. Two groups can have the same mean and different stories

Group A: 48, 49, 50, 51, 52. Group B: 10, 30, 50, 70, 90. Both means are 50. Both medians are 50. Their spreads are very different.

Group A is tightly clustered; Group B is widely dispersed. Saying only “both groups averaged 50” removes the difference that may matter most.

When comparing groups, use at least one measure of centre and one measure of spread when the question concerns consistency or variability.

Worked comparison: Class X has mean 72 and range 8; Class Y has mean 74 and range 30. Class Y has the higher mean, but the simple range suggests substantially more variation. That observation alone does not establish why the variation exists.

7. Frequency tables compress repeated values

Suppose a value x occurs with frequency f. The total contribution is fx. Therefore the mean from a frequency table is Σfx / Σf.

Consider values 1, 2, 3, 4 with frequencies 2, 3, 4, 1. Total frequency = 10. Total of observations = 1×2 + 2×3 + 3×4 + 4×1 = 24. Mean = 2.4.

The frequency column is not another data value. Multiplying value by frequency reconstructs the contribution that would appear if all observations were written separately.

For the median, use cumulative frequency to locate the middle position rather than averaging the displayed row labels.

8. Grouped data gives estimates when exact observations are hidden

If data are grouped into intervals such as 0≤x<10, 10≤x<20 and 20≤x<30, the exact individual values are no longer available. A common mean estimate uses class midpoints.

If frequencies are 3, 5 and 2, use midpoints 5, 15 and 25. Estimated total = 3×5 + 5×15 + 2×25 = 140. Total frequency = 10. Estimated mean = 14.

This is an estimate because the observations in each class are not known to equal the midpoint. A precise calculator display does not turn estimated inputs into exact data.

Keep the word “estimate” in the final interpretation when the method is based on grouped intervals.

9. Weighted means arise when groups have different sizes

If one class of 20 students has mean 70 and another class of 30 students has mean 80, the combined mean is not (70+80)/2 = 75.

The first total is 20×70 = 1400. The second total is 30×80 = 2400. Combined total = 3800 over 50 students, giving mean 76.

The larger class contributes more weight because it contains more observations.

This is the same mathematical issue encountered when combining percentages from unequal group sizes. Convert summaries back to totals before recombining.

10. A mean can be reconstructed after a value changes

Suppose eight observations have mean 15, so their total is 120. If one recorded value 12 should have been 20, the corrected total is 128 and the corrected mean is 16.

Do not recalculate all eight observations if only one value changed and the original total can be recovered.

Removal example: six values have mean 14, total 84. Removing a value 9 leaves total 75 across five observations, so the new mean is 15.

These reverse-total techniques are useful because the mean is fundamentally a total divided by a count.

11. Truncated axes can exaggerate visual differences

Suppose two values are 96 and 100. On a vertical axis running from 0 to 120, the difference looks small. On an axis running from 95 to 101, the same four-unit difference occupies most of the graph height.

Neither scale is automatically dishonest. A restricted range can make small differences easier to inspect. The problem appears when the viewer mistakes visual height for relative magnitude without reading the scale.

Always inspect the axis origin, tick intervals and units before comparing bars or line segments.

A bar chart is especially sensitive to a truncated vertical baseline because bar length is visually interpreted from the baseline. A line graph may legitimately focus on a narrow range, but the scale must remain explicit.

12. Percentages can conceal different absolute counts

Ten per cent of 50 is 5. Ten per cent of 5000 is 500. The same percentage can represent very different counts.

Conversely, the same count can represent different percentages. Twenty cases out of 100 is 20%; twenty out of 1000 is 2%.

When interpreting a reported rate, ask for the denominator. The percentage alone does not determine the absolute scale.

The Ratio, Percentage and the Correct Base guide develops this denominator discipline in detail.

13. Association does not automatically establish cause

If two measured quantities tend to rise together, the data may show association. That alone does not establish that changing one quantity causes the other to change.

There may be a third variable, reverse influence, selection effects or coincidence. A scatter plot is evidence about the pattern of paired observations, not a complete causal explanation.

For a school mathematics task, describe only what the graph supports. “There is a positive association” is different from “X causes Y”.

This distinction protects mathematical communication from making claims stronger than the data.

14. Interpolation and extrapolation carry different risks

Interpolation estimates between observed values. Extrapolation predicts beyond the observed range.

If observations cover x from 2 to 10, estimating at x=6 is interpolation. Estimating at x=20 is extrapolation.

Extrapolation requires the additional assumption that the observed relationship continues into an unobserved region. The calculation can be algebraically correct while the real-world model becomes poor.

State when a result is an estimate and whether it lies inside or outside the range used to build the model.

15. Averages should not be interpreted without the distribution

Consider two invented groups of weekly practice minutes.

Group P: 58, 59, 60, 61, 62. Group Q: 20, 40, 60, 80, 100. Both means and medians are 60. Their consistency differs substantially.

If the question is “Which group has the higher typical value?”, the summaries tie. If the question is “Which group is more consistent?”, spread becomes decisive.

A good comparison sentence names both features: “The groups have the same mean and median, but Group P is much less spread out.”

Do not invent explanations such as motivation or teaching quality from the numerical summaries alone.

16. A capstone comparison with unequal groups

An invented programme has Group A with 40 participants and mean score 72, and Group B with 60 participants and mean score 78. Find the combined mean. Then one participant with score 98 is removed from Group B; find the new combined mean.

Original totals: A = 40×72 = 2880; B = 60×78 = 4680. Combined total = 7560 across 100 participants. Combined mean = 75.6.

After removing 98 from Group B, the combined total becomes 7462 across 99 participants. The new combined mean is 7462/99 ≈ 75.37.

Notice that removing a value above the old combined mean lowers the combined mean. That directional prediction is a useful check before calculating.

It would be wrong to subtract 98 from the mean or to average the two group means equally. The mean must be reconstructed through totals and counts.

17. Independent practice

  1. Find the mean, median, mode and range of 3, 5, 5, 8, 14.
  2. Six values have mean 11. Find their total.
  3. Five values have mean 18. Four sum to 71. Find the missing value.
  4. Find the median of 2, 6, 9, 13.
  5. State the mode or modes of 1, 2, 2, 4, 4, 7.
  6. Compare the ranges of 5, 6, 7, 8, 9 and 1, 5, 7, 9, 13.
  7. Values 1,2,3 have frequencies 2,5,3. Find the mean.
  8. Grouped classes 0–10, 10–20, 20–30 have frequencies 4,3,3. Estimate the mean using midpoints.
  9. A group of 25 has mean 64 and a group of 15 has mean 76. Find the combined mean.
  10. Seven values have mean 20. One value 14 is corrected to 21. Find the new mean.
  11. A group has mean 30 across eight values. Removing one value 44 leaves seven values. Find the new mean.
  12. Explain why a chart starting its vertical axis at 95 can make 96 and 100 look more different.
  13. Find the count represented by 12% of 250.
  14. Twenty-four successes out of 60 trials represent what percentage?
  15. Two groups have means 50 and 50, but ranges 4 and 40. What can you conclude about their spread?
  16. A line of best fit is built from x-values 2 to 12. Is a prediction at x=8 interpolation or extrapolation?
  17. Is a prediction at x=20 interpolation or extrapolation?
  18. Group A has 30 observations with mean 80; Group B has 70 observations with mean 65. Find the combined mean.

18. Worked answers

1. Mean 7; median 5; mode 5; range 11. The total is 35 across five values.

2. 66. Total = mean × count.

3. 19. Required total is 90; subtract 71.

4. 7.5. Average the two central ordered values 6 and 9.

5. 2 and 4. Both have the highest frequency.

6. 4 and 12. Range uses maximum minus minimum.

7. 2.1. Total contribution is 2+10+9=21 over ten observations.

8. 14. Estimated total is 4×5 + 3×15 + 3×25 = 140 over ten observations.

9. 68.5. Totals are 1600 and 1140, giving 2740/40.

10. 21. Original total 140; correction adds 7, giving 147/7.

11. 28. Original total 240; remove 44 to get 196 across seven values.

12. The truncated scale expands the four-unit difference visually. The numerical difference remains four.

13. 30. Compute 0.12×250.

14. 40%. Use 24/60×100%.

15. The second group has a much larger simple range. The equal means do not imply equal consistency.

16. Interpolation. Eight lies inside the observed x-range.

17. Extrapolation. Twenty lies outside the observed range.

18. 69.5. Combined total = 30×80 + 70×65 = 6950 across 100 observations.

19. Diagnose data mistakes by the compressed information they misuse

Common errors include averaging averages without weights, reporting a grouped-data estimate as exact, confusing percentage points with relative percentage change, interpreting correlation as causation, or comparing charts without reading their scales.

A strong correction names the missing information: “Group sizes were unequal”, “Exact values were unavailable inside the intervals”, or “The y-axis began at 95”. That description tells the learner what to inspect next time.

For practice, change the data while preserving the structural issue. A repaired weighted-mean skill should survive different group sizes and different means.

20. Continue through the BTT Mathematics library

Return to the BTT Mathematics Hub. Use Ratio, Percentage and the Correct Base for denominator discipline, Graphs, Tables and Relationships for graphical representation, and Probability, Sample Spaces and Independence for uncertainty rather than observed data summaries.

Use the BTT Mathematical Lab when the same interpretation failure recurs across several data questions.

21. Sources and scope

All data sets, charts described in words and comparison scenarios are original teaching examples. The calculations shown provide the support for the numerical conclusions; no empirical claim about a real population is intended.

For the current Singapore Secondary curriculum doorway, see MOE: Curriculum for secondary schools. Match statistical notation and extension work to the learner’s actual subject level and school programme.