KNOWLEDGE WAREHOUSE · OBJECT 12
Statistics and Data
Statistics is the Mathematics of learning from data under uncertainty. It combines representation, summary, variability, sampling, modelling and inference so evidence can support—but not exceed—the conclusions justified by the data.
A calculation can be correct while the statistical conclusion is wrong.
Core ideas
- Data types and data collection.
- Tables, charts, histograms, box plots and scatter plots.
- Mean, median, mode, range, quartiles and standard deviation.
- Distribution shape, centre, spread and outliers.
- Correlation and regression.
- Sampling, probability models, confidence and hypothesis testing at advanced stages.
Prerequisites and representations
Prerequisites: number sense, percentage, ratio, algebra, graph reading and probability. Representations include raw data, frequency tables, plots, summary statistics, probability distributions and model output.
Failure signatures
- Chooses mean automatically even when skew/outliers make median more informative.
- Reads a graph without checking axis scale or sample size.
- Confuses correlation with causation.
- Computes standard deviation but cannot interpret what spread means.
- Performs a hypothesis test mechanically without stating the model, hypotheses or conclusion in context.
Diagnostic probes
- Construct two datasets with the same mean but very different spread.
- When would median be a better summary than mean?
- What information is lost when raw data are reduced to one average?
- Does a strong correlation prove one variable causes the other? Give a counterexample.
- Explain a p-value or hypothesis-test conclusion in the language of the original context rather than formula notation.
Repair and transfer
Repair by reconnecting calculation to the question being asked. Require the learner to describe distribution shape, compare representations and justify the chosen summary/model before calculating. Transfer is verified when the learner can critique unfamiliar data displays, identify limitations and communicate a bounded conclusion.
Stage progression
Primary introduces tables, graphs and simple averages. Secondary develops statistical diagrams and summary measures. H1 Mathematics gives substantial weight to Probability and Statistics; H2 extends probability distributions, sampling, hypothesis testing, correlation and regression. Current SEAB H1 specifically frames Mathematics and Statistics as preparation for informed decision-making and business/social-science study.
Downstream dependencies
Data science, research, economics, social science, medicine, engineering, AI evaluation and evidence-based decision-making all depend on statistical reasoning.
TECHNOLOGY: spreadsheets, graphing calculators, statistical software and AI can accelerate analysis; model choice, assumptions, interpretation and causal restraint must remain humanly accountable.
PHASE 4 · STATISTICS & DATA READER GUIDE
Quick Read: what is statistics really doing?
Statistics turns data into evidence. The learner must decide what the data represent, how to summarise variation and what conclusions the evidence does—and does not—justify.
A calculator can produce means, standard deviations, regressions and test statistics quickly. The mathematical work is not finished there. The student still has to know which summary is appropriate, what assumptions are being made and how to translate the result back into careful language.
One-sentence answer: statistics becomes secure when the learner can move from data to representation to summary to interpretation without claiming more than the data support.
Data first: what was measured?
- Variable: what quantity or category was observed?
- Population: which larger group is of interest?
- Sample: which observations were actually collected?
- Measurement process: how were the values obtained?
- Context: what conditions might affect interpretation?
A numerical summary is only as meaningful as the data-generating process beneath it. Poor sampling or ambiguous measurement cannot be repaired by more sophisticated calculation.
Centre and spread answer different questions
| Summary | Question it answers | Typical risk |
|---|---|---|
| Mean | Where is the numerical balance point? | Can be pulled by extreme values. |
| Median | What is the middle ordered value? | Can hide how widely values vary. |
| Range / IQR | How spread out are the observations? | A single spread measure cannot describe the whole distribution. |
| Standard deviation | How dispersed are values around the mean? | Meaning is lost if treated as calculator output only. |
Two datasets can have the same mean and very different variability. Strong interpretation therefore considers location and spread together.
A graph is an argument about the data
Bar charts, histograms, box plots, scatterplots and cumulative displays reveal different structures. Choosing a graph is therefore part of statistical reasoning.
- Use a representation that matches the variable type.
- Check axes, scales and bin choices.
- Ask what patterns become visible and what information is hidden.
- Do not infer causation from a visual association alone.
A misleading scale can exaggerate a small difference; a wide histogram bin can hide structure. Representation choices affect what the reader sees.
Correlation is not causation
A statistical association tells us that variables move together in the observed data. It does not by itself establish that changing one causes the other. Confounding variables, selection effects or reverse direction can create or distort the pattern.
A regression line can describe association. Causal explanation requires stronger design and evidence.
Three statistics students who need different repair
- Student A calculates the mean correctly but cannot decide whether it is representative. Repair interpretation and distribution reading.
- Student B produces a regression output but states that one variable causes the other. Repair evidence language and the distinction between association and causation.
- Student C understands the concept but misreads calculator output. The issue may be tool control, labels or notation rather than statistical meaning.
Again, the same low mark can hide different mechanisms. Written interpretation and the choice of representation are as diagnostic as arithmetic.
From description to inference
Descriptive statistics summarise observed data. Inferential statistics use sample evidence to reason about a wider population or model under stated assumptions.
- Define the question.
- Understand the sample and population.
- Choose an appropriate model or procedure.
- Compute the statistic or interval.
- Interpret the result in context.
- State uncertainty and limitations.
The important development is epistemic: the learner must separate what was directly observed from what is inferred.
Statistics across the school journey
| Stage | Statistical demand | Key transition |
|---|---|---|
| Primary | Read and construct simple tables and graphs; compare data. | Move from individual values to patterns in a set. |
| Secondary | Measures of centre/spread, data representations and interpretation. | Compare distributions, not only single summaries. |
| JC Mathematics | Probability distributions, sampling, correlation/regression and inference. | Use sample evidence to make cautious model-based conclusions. |
Calculator output must be translated
- Know which variable each output refers to.
- Check whether the selected model matches the question.
- Distinguish sample statistics from population parameters.
- Read signs and magnitude in context.
- State conclusions using the language justified by the procedure.
- Question implausible output instead of treating the calculator as authority.
The calculator should reduce computational burden while leaving model choice and interpretation with the learner.
What parents can notice
- Can the learner explain what the data measure?
- Can they say why one summary is more useful than another?
- Can they describe spread as well as average?
- Do they distinguish correlation from causation?
- Can they interpret calculator output in ordinary language?
- Can they identify limitations in sampling or measurement?
A useful question is: “What does the data actually allow you to claim?” That keeps statistical reasoning disciplined.
Frequently asked questions
Why can two datasets with the same mean feel different?
Because centre does not describe spread or shape. Variability and distribution structure matter.
Is a strong correlation proof of cause?
No. Correlation measures association. Causal claims need stronger evidence about design, confounding and direction.
Why do statistics questions require so much writing?
Because the conclusion is part of the Mathematics. A numerical result without interpretation may not answer the statistical question.
How do we know statistical understanding has transferred?
The learner can inspect an unfamiliar dataset, choose a justified representation or summary, use appropriate technology and state a conclusion with correct uncertainty.
The larger idea: statistics is disciplined compression of evidence
Data are detailed and noisy. Statistics compress them into summaries, models and conclusions. Every compression loses some information, so the learner must know what was preserved and what was hidden.
The mature student therefore treats a statistical result as an evidence statement with conditions, not as a number detached from how the data were produced.
Statistics is strongest when the calculation becomes a careful statement about evidence rather than a machine-produced answer.
JC / H2 routes from Statistics and Data
IB and International routes from Statistics and Data
World Mathematics route: return to the World Mathematics Atlas to see how this knowledge object appears across school levels, examinations, competitions, international curricula and university Mathematics.
Beginner statistics route: from averages to regression
Statistics is not a staircase of formulas. Each stage asks a different question about evidence. Start at the first stage the learner cannot explain or use on changed data, then move forward only when the earlier interpretation is stable.
- Ask what was measured. Identify variables, units, cases, population, sample and how the data were produced. Begin with statistical questions, data collection and variability when this foundation is weak.
- Organise and display the data. Choose a table, plot or diagram because it reveals the feature under investigation—not because the chart type was named in advance.
- Describe centre. Mean, median and mode answer different questions. A typical value is useful only when the learner understands what information it compresses.
- Describe spread and distribution. Range, interquartile range, variability, outliers and shape explain what a centre alone hides. Use Distributions, Center, Spread and Typical Values for the conceptual bridge.
- Connect probability to uncertainty. Probability models describe possible behaviour; observed data provide evidence. Do not collapse those two roles into one calculation.
- Move from description to inference carefully. Ask what can legitimately be concluded beyond the observed data and what sampling, design or assumptions limit the claim.
- Study association without jumping to cause. Scatterplots and correlation describe relationships; they do not by themselves establish a causal mechanism.
- Use regression as a model, not a magic line. Interpret slope and intercept in context, inspect residual behaviour, distinguish interpolation from extrapolation, and state the model’s limits. Continue to H2 Mathematics: Correlation and Linear Regression when the learner is ready for the formal route.
International learners can continue through IB Mathematics AA and AI: Statistics and Probability or AP Statistics. The canonical knowledge object remains this page: school, examination and international pages are stage-specific routes through the same statistical ideas.
Owner decision: “Statistics for Beginners: From Averages to Regression” is resolved by deepening this Statistics and Data knowledge object. A second general statistics overview would duplicate the same mathematical job.
One worked dataset: from averages to regression
A beginner statistics route becomes clearer when one small dataset is carried through several stages instead of replacing the data every time the technique changes. The following numbers are a synthetic teaching example, not research evidence.
| Student | Study hours in one week, x | Quiz score, y |
|---|---|---|
| A | 1 | 55 |
| B | 2 | 61 |
| C | 3 | 58 |
| D | 4 | 70 |
| E | 5 | 73 |
Step 1: ask what the variables mean
The variable x records study hours for one week. The variable y records one quiz score. Five students were observed. Before calculating anything, notice the limits: this is a tiny sample, the study hours may not have been measured perfectly, and many other factors could affect quiz performance.
That first step matters because statistics is not only arithmetic performed on columns of numbers. It is reasoning about what those numbers represent and what claims the data can support.
Step 2: describe the quiz scores with centre
The mean quiz score is (55 + 61 + 58 + 70 + 73) ÷ 5 = 63.4. Ordered as 55, 58, 61, 70, 73, the median is 61. The mean is the arithmetic balance point; the median is the middle ordered observation.
Step 3: add spread
The range is 73 − 55 = 18. Reporting only “the average score was 63.4” hides the fact that the observed scores span eighteen marks. Centre describes location; spread describes variability.
Step 4: study the relationship between two variables
Plot study hours on the horizontal axis and quiz score on the vertical axis. The points show a generally upward pattern. The Pearson correlation for these five points is approximately r = 0.916, describing a strong positive linear association in this synthetic dataset. It does not establish that extra study hours caused the higher scores.
Step 5: fit a regression model
The least-squares regression line is approximately ŷ = 49.9 + 4.5x. The slope says that within this fitted model, one additional study hour is associated with about 4.5 additional quiz-score points on average. It does not say that forcing any student to study one more hour will cause exactly 4.5 extra marks.
The intercept, 49.9, is the fitted prediction at zero study hours. Whether it has useful real-world meaning depends on the context and whether zero is a sensible point to interpret.
Step 6: distinguish interpolation from extrapolation
The observed study-hour values run from 1 to 5. At 3.5 hours, the model predicts ŷ = 49.9 + 4.5(3.5) = 65.65; this is interpolation. At 10 hours, the same line gives 94.9, but that value lies far outside the observed range. The linear pattern may not continue, so that extrapolation is much less defensible.
What this example teaches
- Measurement comes first: know what each variable represents.
- Centre compresses; mean and median do not describe the whole distribution.
- Spread restores information that centre hides.
- A scatterplot changes the question from one-variable description to association.
- Correlation describes association, not cause.
- Regression creates a fitted model whose slope and intercept need contextual interpretation.
- Interpolation is usually safer than extrapolation.
- Every statistical summary has limits and assumptions.
Fresh transfer check
Change the context and numbers—temperature and electricity use, plant height and growth, or travel distance and journey time. Ask the learner to identify variables and units, calculate a justified centre, describe spread, choose a useful graph, describe association without claiming cause, interpret a regression slope, distinguish interpolation from extrapolation and state at least one limitation.
The release test is not “Can the student press the regression buttons?” It is whether the learner can move from data → representation → summary → variability → association → model → bounded conclusion on a changed dataset while keeping causal claims within the evidence.
Same Average, Different Story: Spread, Outliers and Aggregated Data
An average can be correct and still tell the wrong story.
That does not make averages useless. It means a summary statistic answers one question while leaving others open. A mean can describe centre while hiding spread. A median can resist outliers while ignoring how far extreme values lie from the middle. An overall average can hide subgroup differences. A single number compresses a dataset; compression always discards information.
The statistical habit is therefore:
Do not ask only “What is the average?” Ask “What information does this average preserve, and what information did it remove?”
Dataset Pair 1: same mean, completely different spread
Consider:
Dataset A: (4,4,4,4,14)
Dataset B: (6,6,6,6,6)
Both have mean (6).
For A:
[ rac{4+4+4+4+14}{5}=rac{30}{5}=6. ]
For B:
[ rac{6+6+6+6+6}{5}=6. ]
But the distributions are very different. Dataset B has no spread at all. Dataset A contains four values below the mean and one value far above it.
If a learner reports only “both groups average 6”, an important structural difference disappears.
Range exposes one layer of the difference
Dataset A has range:
[ 14-4=10. ]
Dataset B has range:
[ 6-6=0. ]
The mean is identical. The range is not.
Standard deviation asks a different question
Spread measures such as standard deviation use all observations rather than only the extreme pair. A full calculation may belong to the relevant course, but the interpretation comes first:
How tightly are the values clustered around the centre?
A high mean with wide spread may describe a very different situation from the same mean with tight consistency.
Dataset Pair 2: same mean, different shape
Consider:
Dataset C: (2,4,6,8,10)
Dataset D: (6,6,6,6,6)
Both have mean (6). Dataset C is evenly spread around the centre. Dataset D is concentrated entirely at the centre.
The average alone cannot distinguish a smooth spread from perfect concentration.
Outliers can pull the mean
Delivery times in minutes are:
[ 12, 13, 13, 14, 48. ]
The mean is:
[ rac{12+13+13+14+48}{5}=rac{100}{5}=20. ]
The median is (13).
If the question is “What delivery time is typical on an ordinary day?”, (20) minutes may feel unrepresentative because one extreme delay pulls the mean upward.
If the question is “What average time burden should be planned across all deliveries including disruptions?”, the mean may still be relevant.
No statistic is best without a reader job.
Median is robust, not magically superior
Median is less sensitive to extreme values. That is useful in skewed distributions. But median can also hide magnitude.
Compare:
E: (1,2,3,4,100)
F: (1,2,3,4,10,000)
Both have median (3). The extreme values are radically different.
If the tail matters, median alone is inadequate.
Outliers should be investigated before they are deleted
An unusual value can come from:
- measurement error;
- data-entry error;
- a genuinely rare event;
- a different subgroup;
- a changed process;
- a unit mismatch;
- an important failure mode.
Deleting it automatically can remove the most informative observation in the dataset.
The outlier question is not “Does this look ugly?”
The better questions are:
- Is the value plausible?
- Was it measured under the same conditions?
- Does it belong to the target population?
- Would removing it change the conclusion?
- Can its origin be verified?
Same mean, different risk
Suppose two processes both average (100) units per hour.
Process G stays between (98) and (102) most of the time.
Process H alternates between (60) and (140).
The same mean can correspond to very different operational risk. Spread is not decorative information.
Aggregated averages can be wrong when sample sizes differ
Suppose Class A has (10) students with an average score of (80). Class B has (30) students with an average score of (60).
A careless calculation averages the two averages:
[ rac{80+60}{2}=70. ]
But the classes have different sizes.
Class A contributes:
[ 10(80)=800 ]
total marks.
Class B contributes:
[ 30(60)=1800. ]
Combined total:
[ 2600. ]
Combined number of students:
[ 40. ]
So the true combined mean is:
[ rac{2600}{40}=65. ]
The average of averages is valid only under appropriate weighting conditions.
Weighting is part of the model
Whenever subgroup summaries are combined, ask what each subgroup represents.
- Equal group weight?
- Equal individual weight?
- Equal time weight?
- Equal financial exposure?
- Equal number of observations?
The weights encode what the combined statistic means.
Aggregated data can hide subgroup structure
Imagine a school reports one overall average improvement after an intervention. That number may combine groups with different starting points, class sizes, attendance patterns or subject levels.
The overall average may be accurate while hiding that one group improved strongly and another did not.
Before interpreting an aggregate, inspect the important subgroups when the reader job depends on them.
Do not invent subgroup stories
The opposite error is also possible. If only an overall average is available, do not speculate about hidden subgroups without evidence. Say what the aggregate can establish and what cannot be recovered from it.
Mean, median and mode answer different questions
| Statistic | Useful when | Can hide |
|---|---|---|
| Mean | All values should contribute proportionally to centre | Skew, outliers, multimodality |
| Median | Order and central position matter, especially with skew | Magnitude of tails |
| Mode | Most common category or value matters | Overall numerical balance |
Range, IQR and standard deviation answer different spread questions
Range uses the two extremes. Interquartile range describes the middle half. Standard deviation measures dispersion around the mean.
Choose based on distribution and purpose rather than habit.
Box plots compress but preserve more than one number
A box plot can show median, quartiles, spread and potential outliers. It still does not reveal every data point, but it preserves more distribution structure than a single average.
Compression can be layered.
Histograms show shape that averages cannot
Two datasets can share mean and standard deviation yet differ in modality or local structure. A histogram can reveal clusters, gaps and skewness.
Always inspect a representation appropriate to the data before trusting one summary number.
Same mean, different sample size
Two groups may both have mean (70). One contains (5) observations; the other (5,000).
The equal means do not carry equal evidence about population behaviour. Sample size affects uncertainty.
Do not read a summary statistic without its data context.
Same percentage, different denominators
A success rate of (80%) can mean (4/5) or (800/1000). The percentages are identical. The evidential stability is not.
Report denominators when they matter.
Percentage change can exaggerate small bases
An increase from (1) case to (2) cases is a (100%) increase. The absolute increase is (1).
An increase from (1000) to (1100) is a (10%) increase. The absolute increase is (100).
Relative and absolute change answer different questions.
Averages of rates need care
If a journey covers equal distances at different speeds, the average speed is not generally the arithmetic mean of the speeds.
For (60) km at (30) km/h and (60) km at (60) km/h:
Time:
[ 2 ext{ h}+1 ext{ h}=3 ext{ h}. ]
Total distance:
[ 120 ext{ km}. ]
Average speed:
[ 40 ext{ km/h}. ]
But the arithmetic mean of (30) and (60) is (45).
The correct average depends on how the quantity accumulates.
Average of percentages can also mislead
If two tests have different total marks, averaging percentage scores may or may not match the desired combined score. Define the weighting.
Original dataset lab: two classes
Class P: (52,58,60,60,70)
Class Q: (40,60,60,60,80)
Both have mean (60).
Questions:
- Which class has greater spread?
- Which class has the same median?
- What does the mean fail to show?
- Which representation would make the difference most visible?
Both medians are (60). Class Q has a larger range. A dot plot or box plot would reveal the spread difference more clearly than the mean.
Original dataset lab: outlier decision
Daily repair times are:
[ 18, 19, 20, 21, 22, 95. ]
Before deleting (95), ask:
- Was there a recorded major breakdown?
- Was the unit entered correctly?
- Was the repair process the same?
- Does the analysis seek ordinary performance or total operational burden?
The outlier may be an error or the most important failure event.
Original aggregated-data lab
Branch R processes (20) cases with mean waiting time (4) min. Branch S processes (80) cases with mean waiting time (7) min.
Combined mean:
[ rac{20(4)+80(7)}{100} =rac{80+560}{100} =6.4 ext{ min}. ]
The unweighted average of branch means would be (5.5) min, which answers a different question: the average of the two branch-level averages with equal branch weight.
Question first, statistic second
Ask what the user of the data needs to know.
- Typical individual experience?
- Total resource burden?
- Consistency?
- Worst-case risk?
- Central performance?
- Subgroup difference?
- Change over time?
The appropriate summary follows from the question.
Same average can hide a trend
Suppose weekly values are (2,4,6,8,10). Mean (=6).
If the same values occur in reverse order (10,8,6,4,2), the mean is still (6). The time trend is opposite.
When order matters, a mean alone can erase direction.
Same average can hide clusters
Dataset:
[ 1,1,1,11,11,11. ]
Mean (=6). Yet no observation is near (6).
A central average can fall in a gap between clusters.
Do not call the mean “typical” automatically
“Typical” can mean most common, central, expected, ordinary, representative or middle. Those meanings are not identical.
Choose language that matches the statistic.
Correlation and averages solve different problems
A pair of variables can have similar averages while their relationship differs. When the question concerns association, a univariate average is not enough.
Use scatter plots and regression tools when the reader job concerns how variables move together.
Aggregation can reverse apparent comparisons
In advanced statistics, combining subgroups can sometimes produce a different apparent relationship from the relationships inside the groups. This is one reason subgroup structure and weighting matter.
The learner does not need to memorise a paradox name to adopt the habit: inspect how the aggregate was constructed.
Data provenance matters
Before interpreting a summary, ask:
- Who was measured?
- How were observations selected?
- What was excluded?
- What unit was used?
- Was the variable measured consistently?
- Were groups combined?
- Were missing values handled?
A correct mean of biased data remains a biased summary.
A five-part summary statement
A strong statistical conclusion can include:
- centre;
- spread;
- shape;
- important unusual values;
- contextual limitation.
Example:
“The median delivery time was (13) minutes, with most observations clustered between (12) and (14) minutes, but one (48)-minute delay created strong right-skew and raised the mean to (20) minutes.”
This is far more informative than “average delivery time was 20 minutes.”
Fresh transfer: choose the summary
For each dataset, choose at least two useful summaries and explain why.
- (5,5,5,5,25)
- (1,3,5,7,9)
- (10,10,10,10,10)
- (2,2,2,8,8,8)
- (45,46,47,48,95)
Do not calculate mechanically. Begin by describing the shape or special feature.
Fresh transfer: combine weighted means
Group A has (12) learners with mean (72). Group B has (28) learners with mean (64).
Combined mean:
[ rac{12(72)+28(64)}{40} =rac{864+1792}{40} =rac{2656}{40} =66.4. ]
Explain why ((72+64)/2=68) is not the individual-weighted combined mean.
Fresh transfer: same centre, changed spread
Create two five-value datasets with mean (10):
- one with zero spread;
- one with range at least (20).
The construction itself shows that centre does not determine spread.
Statistical checking loop
Question → raw data → representation → centre → spread → unusual values → subgroup/weighting check → conclusion in context.
If the conclusion changes substantially when one observation is removed, say so. If the combined mean depends on unequal group sizes, weight it. If the time order matters, preserve it. If the distribution is clustered, do not let one central number hide that structure.
The durable target
A statistic is a compression of data, not the data itself.
The mature learner can calculate the average, inspect what it erased, and choose a summary that preserves the feature the reader actually needs.
Distribution Audit: What the Average Cannot Tell You by Itself
The next step after calculating centre and spread is to ask whether the summary survives a change of representation. A statistic should not become more convincing merely because the raw data have disappeared.
1. Reconstruct a plausible data story
Suppose two classes both have a mean score of 70. Before comparing them, request at least one measure of spread and one distribution display. A class clustered between 67 and 73 tells a different story from a class split between scores near 50 and scores near 90. The mean alone cannot tell which structure produced 70.
2. Same median, different tail
Compare:
A: 10, 11, 12, 13, 14
B: 1, 11, 12, 13, 100
Both medians are 12. Dataset B contains much greater tail risk. If the reader cares about unusual delays, losses, errors or waiting times, the median is not enough.
3. Same range, different internal spread
Compare:
C: 0, 49, 50, 51, 100
D: 0, 0, 50, 100, 100
Both have range 100, but the internal arrangement is different. Range is a boundary statistic. It does not describe how values are distributed between the extremes.
4. Same mean and range, different shape
It is possible to build datasets with the same mean and range but different clustering. This is why a distribution cannot generally be reconstructed from two summary values.
For an unfamiliar dataset, ask for a dot plot, histogram, box plot or raw table when the decision depends on distribution shape.
5. Aggregation can hide unequal exposure
Suppose one machine operates for 2 hours at an error rate of 1% and another operates for 18 hours at an error rate of 4%. Averaging the rates as (1% + 4%)/2 = 2.5% gives equal weight to unequal operating periods.
If the reader needs the error rate per unit of operating exposure, the denominator must reflect the actual exposure. Weighting is not cosmetic; it defines the population to which the average refers.
6. The denominator audit
Before accepting a percentage or rate, write the denominator in words.
- errors per what?
- successes among whom?
- cases over what period?
- cost per transaction or per customer?
- average mark per student or per class?
Many misleading summaries are denominator failures disguised as arithmetic.
7. Missing data can change the average
If observations are missing systematically, the calculated mean may be accurate for the recorded cases and unrepresentative of the intended population.
For example, if only students who completed every optional quiz are included in an “average improvement” calculation, the summary may describe completers rather than all enrolled students. Mathematics cannot repair a population definition that was never checked.
8. Averages across time require a time model
If values are recorded hourly, daily and monthly, the weighting may differ. A simple average of monthly percentages can misrepresent a year when months contain different numbers of observations.
Ask whether the intended summary weights time intervals, events, people or total exposure equally.
9. Averages across ratios require special care
Ratios often need reconstruction from their numerators and denominators before aggregation.
Team A solves 8 of 10 cases. Team B solves 40 of 100 cases. Their success rates are 80% and 40%.
Unweighted average rate:
(80% + 40%)/2 = 60%.
Combined individual-case success rate:
(8 + 40)/(10 + 100) = 48/110 ≈ 43.6%.
Both calculations answer legitimate but different questions. One treats teams equally; the other treats cases equally.
10. Outlier influence test
Calculate the summary with and without a suspected outlier. Do not automatically delete the value. Compare how much the conclusion changes.
If one observation changes the mean from 18 to 31 while the median barely moves, report that sensitivity. The reader should know that the conclusion depends strongly on one case.
11. Robustness language
Useful reporting phrases include:
- The mean is strongly influenced by one extreme value.
- The median remains stable when the extreme observation is removed.
- The aggregate hides substantial subgroup differences.
- The weighted mean differs from the unweighted mean because group sizes are unequal.
- The centre is similar, but the spread is materially different.
- The time trend is lost when observations are reduced to one average.
12. Do not turn illustrative data into research
The datasets on this page are worked mathematical examples. They illustrate how summaries behave. They are not evidence about real schools, learners, companies or populations. A statistical example teaches a method; an empirical claim requires real data, a sampling process and appropriate source information.
13. Fresh dataset: same average, different reliability
Dataset E: 9, 10, 10, 10, 11.
Dataset F: 0, 5, 10, 15, 20.
Both have mean 10. Ask:
- Which process is more consistent?
- Which has greater range?
- Which mean is more representative of most observations?
- What additional statistic would you report?
14. Fresh weighted-average task
Course A has 8 students with mean 82. Course B has 32 students with mean 68.
Combined mean:
[8(82) + 32(68)] / 40 = (656 + 2176)/40 = 70.8.
The unweighted mean of 82 and 68 is 75. That number gives each course equal weight rather than each student equal weight.
15. Fresh outlier task
Response times are 3, 3, 4, 4, 5, 26 seconds.
Calculate mean and median. Then write two sentences: one describing typical performance and one describing tail risk. A strong answer does not force one statistic to perform both jobs.
16. Fresh time-order task
Series A: 2, 4, 6, 8, 10.
Series B: 10, 8, 6, 4, 2.
Both have the same mean, median and range. Their trends are opposite. If time matters, order is part of the data.
17. Fresh cluster task
Values: 2, 2, 3, 3, 17, 17, 18, 18.
The mean lies between two clusters where no observation may occur. Describe the distribution before using the average as a representative value.
18. Summary-statistic decision tree
- Is the question about centre? Consider mean or median.
- Is the distribution skewed or outlier-heavy? Compare mean and median.
- Does consistency matter? Add a spread measure.
- Does subgroup size differ? Use appropriate weighting.
- Does order matter? Preserve the sequence or time plot.
- Are there clusters or gaps? Inspect a distribution display.
- Does the denominator differ across rates? Reconstruct counts or exposure.
- Would removing one observation materially change the conclusion? Report sensitivity.
19. Final interpretation test
Before writing “the average shows…”, replace the word average with the exact statistic: mean, median, weighted mean, rate or proportion. Then ask whether the sentence still says only what that statistic supports.
Precision in statistical language is part of precision in statistical reasoning.
20. Transfer beyond school exercises
The same discipline applies to waiting times, investment returns, test scores, household spending, production rates, rainfall, medical measurements, website performance and any other setting where one summary number can hide a distribution.
The durable skill is to calculate the summary, inspect what it compresses, and recover the structure needed for the decision.
Final Transfer Clinic: Read the Distribution Before Naming the Average
For the final check, remove the chapter label. Each problem below requires the learner to decide what summary is appropriate before calculating.
Case A: household spending
Monthly values are 410, 420, 425, 430 and 980. The mean is pulled upward by one unusually large month. Report both mean and median, identify the extreme value and explain which summary better answers “ordinary month” versus “total budget exposure”.
Case B: production consistency
Line 1 produces 98, 99, 100, 101, 102 units. Line 2 produces 80, 90, 100, 110, 120 units. Both means are 100. The comparison is therefore about spread and operational consistency, not centre.
Case C: unequal groups
Site A has 4 observations with mean 90; Site B has 36 observations with mean 60. The individual-weighted combined mean is (4×90 + 36×60)/40 = 63. Do not average the site means unless the intended unit of analysis is the site itself.
Case D: time trend
Values 5, 6, 7, 8, 9 and values 9, 8, 7, 6, 5 have identical centre and spread but opposite direction over time. When sequence matters, preserve the sequence.
Case E: clustered data
Values 1, 1, 2, 2, 18, 18, 19, 19 have mean 10. No observation is near 10. A distribution plot is more informative than calling 10 “typical”.
Case F: rate denominator
A 90% success rate from 9/10 and a 90% success rate from 900/1000 share the same percentage but not the same sample size. Report the denominator when evidential stability matters.
Case G: outlier sensitivity
Calculate the mean with and without the most extreme value, but do not remove it automatically. State how much the conclusion depends on that observation and investigate whether it is error, rare event or a different process.
Case H: conclusion writing
Write one sentence containing centre, one containing spread, one identifying an unusual feature, and one naming a limitation. This prevents one average from carrying the entire story.
Final rule
Never let a correct statistic become a substitute for reading the data. Centre, spread, shape, order, subgroup size and denominator all belong to the model of what the numbers mean.
Release Check: Centre, Spread and Aggregation in One Decision
A final statistical answer should survive one last question: if the raw data were hidden and only my sentence remained, would the reader still understand the important structure?
For a skewed dataset, that may require median and an outlier note. For two groups with the same mean, it may require a spread comparison. For combined groups, it may require the weighting rule. For time data, it may require the trend. For rates, it may require the denominator.
Use this final sequence:
- State the reader question.
- Name the statistic exactly.
- Add the spread measure that matters.
- Identify outliers or clusters that alter interpretation.
- Check subgroup sizes before aggregation.
- Check whether order or time was lost.
- Write the conclusion with the limitation still visible.
Example: “The two groups have the same mean of 10, but Group B is much more variable; therefore equal average performance does not imply equal consistency.”
Example: “The weighted combined mean is 66.4 because the two groups contain different numbers of observations; the unweighted mean of the subgroup means answers a different question.”
Example: “The median is 13 minutes, but one 48-minute delay creates a long right tail and raises the mean to 20 minutes.”
These sentences do more than report calculations. They preserve the distribution feature that matters to the decision.
The release standard is simple: do not let one correct average erase the Mathematics still present in the data.
Distribution Decision Lab: Same Average, Different Structure
The average is often the first number readers ask for because it is compact. The danger begins when compact becomes complete. A mean, median or percentage can summarise one feature while hiding the distribution, subgroup structure, time order, sample size or uncertainty that actually determines the decision.
The exercises below use original synthetic data. They are teaching examples, not learner outcomes or empirical research.
Lab 1: same mean, different consistency
Set A: 48, 49, 50, 51, 52.
Set B: 10, 30, 50, 70, 90.
Both have mean 50 and median 50.
Yet Set A is tightly concentrated while Set B is widely dispersed. If the question is “Which process is more predictable?”, the mean and median do not answer it. Spread becomes the relevant evidence.
The range is 4 for A and 80 for B. This is already enough to show why centre and consistency are different questions.
Lab 2: same mean, same range, different internal distribution
Set C: 0, 40, 50, 60, 100.
Set D: 0, 0, 50, 100, 100.
Both have mean 50, median 50 and range 100.
The common summary measures still fail to make the datasets equivalent. C places three values around the centre. D concentrates four observations at the extremes. A dot plot or histogram immediately exposes the difference.
This is a useful statistical lesson: even centre and a simple spread measure together may not preserve distribution shape.
Lab 3: the median can hide tail magnitude
Set E: 7, 8, 9, 10, 30.
Set F: 7, 8, 9, 10, 3000.
Both have median 9. If the reader cares about a typical middle observation, the median is useful. If the extreme event has operational or financial consequences, the identical median conceals the most important difference.
A robust statistic is robust precisely because it is insensitive to certain values. That strength becomes a limitation when those values matter to the decision.
Lab 4: the mean can lie where almost nobody is
Set G: 2, 2, 2, 18, 18, 18.
The mean is 10. No observation equals 10, and the dataset has two clear clusters.
Calling 10 a “typical value” would be misleading. The mean balances the dataset arithmetically but does not describe the modal structure.
When a distribution is multimodal, the shape may be more informative than one centre.
Lab 5: same average, opposite trend
Week H: 2, 4, 6, 8, 10.
Week I: 10, 8, 6, 4, 2.
Both have mean 6 and exactly the same set of values. The time order is reversed.
If the question is “What was the average?”, the datasets are equivalent. If the question is “Is performance improving or deteriorating?”, they tell opposite stories.
Aggregation can destroy sequence information. Time-series data should not be reduced to an average when direction matters.
Lab 6: same average, different failure frequency
Suppose two synthetic systems each average 100 units.
System J: 98, 99, 100, 101, 102.
System K: 50, 100, 100, 100, 150.
The mean is 100 for both.
If failure occurs below 80 or above 120, J never fails while K fails twice in five observations. The average is identical; threshold behaviour is not.
When a decision depends on a limit, count threshold crossings instead of relying only on centre.
Lab 7: same percentage, different denominator
Two groups both report 80% success.
- Group L: 4 successes out of 5.
- Group M: 800 successes out of 1000.
The percentages match. The amount of evidence does not.
The smaller group is much more sensitive to one additional observation. If L records one more failure, the rate falls to 4/6, about 66.7%. If M records one more failure, the percentage barely moves.
Always recover the denominator when interpreting a percentage.
Lab 8: weighted averages and unequal group sizes
Group N contains 8 observations with mean 90. Group O contains 32 observations with mean 60.
An unweighted mean of the group means is:
(90+60)/2 = 75.
But the individual-level combined mean is:
[8(90)+32(60)]/40 = (720+1920)/40 = 66.
Both 75 and 66 can be valid summaries of different questions. Seventy-five is the mean of the two group-level means when each group gets equal weight. Sixty-six is the mean across all forty individuals.
Weighting is not an arithmetic detail. It defines the object being averaged.
Lab 9: average of rates can require harmonic structure
A vehicle travels 60 km at 30 km/h and another 60 km at 60 km/h.
The arithmetic mean of the speeds is 45 km/h.
Actual average speed is total distance divided by total time:
120 km / (2 h + 1 h) = 40 km/h.
The equal-distance structure gives more time weight to the slower speed. Averaging rates correctly requires knowing what is being held equal.
Lab 10: average cost can depend on quantity
Suppose 10 items cost $4 each and 90 items cost $8 each.
The average of the two prices is $6 if each price category is given equal category weight.
The average paid per item across all 100 items is:
[10(4)+90(8)]/100 = 760/100 = $7.60.
The correct average follows the reader job: average category price or average item cost.
Lab 11: percentage change needs a base
A rise from 2 to 4 is a 100% increase. A rise from 200 to 240 is a 20% increase.
The first relative change is larger; the second absolute change is larger.
Statements about “largest increase” are incomplete until the comparison scale is defined.
Lab 12: outlier as error versus event
Repair times in minutes:
18, 19, 20, 21, 22, 95.
The value 95 is unusual. Before excluding it, ask whether it is:
- a transcription error;
- a different unit;
- a genuinely severe failure;
- a case from another process;
- a measurement taken under different conditions.
If 95 represents the exact failure mode the organisation wants to control, removing it because it distorts the average would erase the operational problem.
Lab 13: outlier influence audit
For the repair-time set above, compare the mean with and without 95.
With 95:
(18+19+20+21+22+95)/6 = 195/6 = 32.5.
Without 95:
(18+19+20+21+22)/5 = 100/5 = 20.
The difference is large. That does not automatically tell us which mean to report. It tells us that interpretation depends strongly on how the unusual observation is classified.
Lab 14: median and IQR as a different summary
For a skewed dataset, median and interquartile range can describe the middle structure with less influence from extremes.
That does not make them universally superior. If total burden or tail risk matters, the extremes remain relevant.
A statistical summary should match the decision, not a habit about which statistic is “better”.
Lab 15: same mean, different interquartile structure
Set P: 0, 49, 50, 51, 100.
Set Q: 0, 0, 50, 100, 100.
Both have mean 50 and median 50. Their quartile structure differs. A box plot would reveal that the middle halves are arranged differently.
This teaches a useful hierarchy: as the decision becomes more sensitive to distribution shape, move from one number to richer summaries and visualisations.
Lab 16: sample size and apparent stability
A mean of 70 based on 4 observations and a mean of 70 based on 400 observations have the same centre but different inferential strength about a broader population, assuming comparable sampling conditions.
Sample size does not guarantee representativeness. A large biased sample can still mislead. But a tiny sample should not be spoken about with the same certainty as a much larger appropriate sample.
Lab 17: aggregation can hide subgroup reversal
Consider a synthetic comparison of two methods across two difficulty bands.
| Easy band | Hard band | |
|---|---|---|
| Method R | 90/100 = 90% | 19/20 = 95% |
| Method S | 18/20 = 90% | 94/100 = 94% |
Within the easy band the methods tie; within the hard band R is slightly higher. Yet the overall rates are:
R: 109/120 ≈ 90.8%.
S: 112/120 ≈ 93.3%.
The overall comparison is influenced by how cases are distributed across difficulty bands.
The point is not to memorise a named paradox. It is to inspect subgroup composition before treating an aggregate as a complete comparison.
Lab 18: subgroup adjustment requires a reason
Do not split data into arbitrary subgroups until a preferred conclusion appears. The subgroup variable should have a defensible relationship to the process or reader question.
Statistical resolution can clarify a mechanism; it can also become cherry-picking if groups are invented after seeing the result without justification.
Lab 19: pooled data can be legitimate
Aggregation is not inherently wrong. If the reader job is total system burden, pooled data may be exactly the correct level.
The statistical question is not “aggregate or disaggregate?” in the abstract. It is “Which level of aggregation preserves the information needed for this decision?”
Lab 20: a scatter plot can reveal what averages cannot
Two datasets may have the same mean x-value and mean y-value while showing very different relationships between x and y.
If the question concerns association, correlation or prediction, marginal averages are not enough. Plot the paired data.
The H2 route Correlation and Linear Regression owns the detailed regression layer.
Lab 21: regression is not a replacement for distribution reading
A fitted line summarises a relationship. It can hide clusters, non-linearity, influential points and changing variance.
Before interpreting slope or correlation, inspect the scatter plot and ask whether a linear model is a reasonable summary.
Lab 22: one influential point can rotate a line
Suppose most synthetic points lie close to a mild positive trend, but one point sits far to the right and high above the rest. That point can strongly influence the regression slope.
The correct response is not automatic deletion. Investigate the point, compare fits with and without it, and report the sensitivity if it matters.
Lab 23: correlation does not establish mechanism
A high correlation measures strength of linear association in the observed data. It does not by itself prove that changing one variable causes the other to change.
Confounding variables, reverse direction or common causes may be possible. Statistical description should not silently become causal explanation.
Lab 24: same mean before and after, different volatility
Before: 49, 50, 50, 50, 51.
After: 20, 40, 50, 60, 80.
The mean remains 50. A before-and-after report saying “average unchanged” is true but incomplete. Variability has increased dramatically.
When a change intervention affects consistency rather than centre, spread is the signal.
Lab 25: same spread, different centre
Dataset U: 8, 9, 10, 11, 12.
Dataset V: 48, 49, 50, 51, 52.
The distributions have the same pattern of deviations around their means but different centres.
Centre and spread are independent dimensions. Reporting one cannot stand in for the other.
Lab 26: same mean and standard deviation can still hide shape
Different datasets can share multiple summary statistics while differing in local arrangement, clustering or modality. This is why visual inspection remains important even after numerical summaries are calculated.
Statistics is not a contest to compress the data into the fewest possible numbers. It is a discipline for compressing without discarding the information required by the question.
Lab 27: repeated measures versus independent observations
Ten measurements from one individual are not automatically equivalent to one measurement from each of ten independent individuals.
The sample size may be ten in both tables, but the dependence structure differs.
Before applying an inferential method, ask what one observation represents and whether observations can reasonably be treated as independent for the intended model.
Lab 28: missing data can move an average
If extreme values are more likely to be missing, the observed mean may be systematically different from the full-data mean.
A complete-looking table can therefore be biased by what was not observed.
Missingness is not always random noise. Its mechanism matters.
Lab 29: rounding can create artificial ties
Values recorded to the nearest whole unit may appear identical even when the underlying measurements differ.
If a decision depends on small differences, check measurement precision before interpreting a cluster as exact equality.
Lab 30: category averages can hide composition change
Suppose an overall average rises from one month to the next. The increase may result from improvement within each subgroup, or simply from a larger share of observations coming from a subgroup that already had a higher average.
To understand the mechanism, compare like with like before claiming within-group improvement.
Lab 31: denominator drift changes percentages
A percentage can move because the numerator changes, the denominator changes, or both.
When reporting a rate over time, keep the denominator visible. A 20% rate from 20/100 and 20% from 200/1000 describe different volumes even though the proportions match.
Lab 32: base-rate awareness
An event can have high conditional accuracy and still produce many false positives when the underlying event is rare.
At more advanced levels, probability and statistics meet here: rates should be interpreted in the context of base frequency.
The relevant probability owner is H2 Mathematics Probability.
Lab 33: data collection quality precedes summary quality
A perfectly calculated mean cannot repair a sample that excludes the relevant population, mixes units, uses inconsistent definitions or records measurements under changing conditions.
Before asking for centre and spread, ask what generated the data.
Lab 34: “typical” needs a definition
Typical can mean:
- most common;
- middle observation;
- arithmetic balance point;
- expected long-run value;
- ordinary non-outlier case.
These are not identical. Statistical language should say which one is intended.
Lab 35: decision table for choosing a summary
| Reader job | Useful first evidence | What else to inspect |
|---|---|---|
| Central performance | Mean or median | Spread and skew |
| Consistency | IQR, SD, range | Shape and outliers |
| Tail risk | Extreme quantiles / max | Frequency and cause of extremes |
| Group comparison | Group centres | Group sizes, spread, composition |
| Trend | Time plot | Seasonality, order, changing variance |
| Association | Scatter plot, correlation | Non-linearity, influential points, causal limits |
Lab 36: one-number reporting rule
If only one statistic can be shown, add a sentence naming what it does not establish.
Example: “Median waiting time was 13 minutes; this does not show the size or frequency of long delays.”
This preserves intellectual honesty even when space is limited.
Lab 37: two-number reporting rule
When possible, pair a measure of centre with a measure of spread. “Mean 68, standard deviation 4” tells a different story from “mean 68, standard deviation 20”.
Centre without spread often creates false confidence about consistency.
Lab 38: visualisation rule
Choose a visualisation that preserves the relevant structure. Use a histogram for shape, a box plot for median and quartiles, a scatter plot for paired association, a time plot for sequence and change.
A chart is not decoration. It is another statistical representation with its own compression choices.
Lab 39: synthetic transfer task
Two classes each have mean score 70.
Class A: 68, 69, 70, 71, 72.
Class B: 20, 60, 70, 90, 110.
- State one way in which the classes are statistically similar.
- State two ways in which they differ.
- Choose a representation that would make the difference visible.
- Explain why “the classes performed the same on average” is mathematically true but potentially misleading.
Lab 40: aggregation transfer task
Site X has 50 observations with mean 40. Site Y has 150 observations with mean 70.
- Find the combined individual-level mean.
- Find the unweighted mean of the two site means.
- Explain what each number represents.
Combined mean:
[50(40)+150(70)]/200 = (2000+10500)/200 = 62.5.
Unweighted site-level mean:
(40+70)/2 = 55.
Neither is “wrong” without a reader job. They answer different weighting questions.
Lab 41: outlier transfer task
Measurements are:
30, 31, 32, 32, 33, 90.
Do not begin by deleting 90. First write three plausible explanations for it, one piece of evidence that would support each explanation, and how the reporting would change if the observation were retained or excluded.
This turns outlier handling from cosmetic cleaning into evidence-based classification.
Lab 42: trend transfer task
Month A values: 10, 20, 30, 40, 50.
Month B values: 50, 40, 30, 20, 10.
Both have mean 30. Explain why a time-series reader may regard them as opposite.
The mean compresses away order.
Lab 43: final statistical release check
Before publishing or accepting a statistical summary, ask:
- What exactly was measured?
- What does one observation represent?
- How many observations are there?
- Is centre being confused with consistency?
- Could an outlier be meaningful?
- Are group sizes unequal?
- Has weighting been defined?
- Could aggregation hide subgroup structure?
- Does time order matter?
- Would a visual representation reveal something the summary hides?
- Is correlation being overread causally?
- What claim is not supported by the available data?
Closing principle
Averages are powerful because they compress. Statistical maturity begins when the reader remembers what the compression removed.
Read centre, spread, shape, sample size, subgroup structure and time order according to the decision. The goal is not to distrust averages. It is to use them with enough surrounding evidence that the story they tell remains faithful to the data.
Final Interpretation Clinic: What the Summary Number Cannot Decide for You
Clinic 1: centre is not fairness
Two groups can have the same mean while one contains a narrow band of outcomes and the other contains extremes. Whether that distribution is fair, equitable or acceptable is not determined by the mean itself. Statistics can describe the pattern; a separate decision rule determines what pattern matters.
Clinic 2: a larger average is not automatically better
If the quantity is waiting time, defect count or error rate, a smaller mean may be preferable. If the quantity is output, score or capacity, a larger mean may be preferable. Statistical direction only gains meaning after the measured variable and decision objective are defined.
Clinic 3: lower spread is not automatically better
Consistency can be desirable, but a tightly clustered poor outcome is not automatically better than a wider distribution with a substantially stronger centre. Read centre and spread together.
Clinic 4: a stable average can hide structural change
Suppose a process average remains 50 for three months. In Month 1 values cluster near 50. In Month 2 a second cluster appears near 80 while another moves near 20. In Month 3 the extremes grow further. The average appears stable while the process becomes less homogeneous.
A time plot or distribution plot can reveal structural drift that a monthly mean hides.
Clinic 5: subgroup size can change the overall story even if subgroup means do not
Suppose Group A always averages 80 and Group B always averages 60. If the population mix shifts from mostly A to mostly B, the overall mean can fall even though neither subgroup changed internally.
This is composition change, not within-group decline.
Clinic 6: a changing overall average may therefore need decomposition
When the aggregate changes, ask:
- Did subgroup performance change?
- Did subgroup proportions change?
- Did both happen?
This decomposition prevents an overall number from being mistaken for the mechanism.
Clinic 7: repeated averages should preserve comparable definitions
A monthly average is difficult to compare across time if the inclusion rules, measurement method or population changed. Before interpreting movement, check that the statistic is built from comparable objects.
Clinic 8: the same label can hide a changing denominator
A “success rate” may be calculated over all attempts one month and only completed attempts the next. The percentages are not directly comparable if the denominator rule changed.
Statistical continuity requires definitional continuity.
Clinic 9: aggregation across categories can create a meaningless average
Do not average quantities that do not share a coherent scale merely because they are numerical. Averaging temperatures, ranks, category codes or unrelated percentages can produce a number without a useful interpretation.
A calculation can be legal in arithmetic and meaningless statistically.
Clinic 10: rank data require care
The difference between rank 1 and rank 2 is not necessarily the same substantive difference as rank 10 and rank 11. Treating ranks as ordinary interval-scale measurements can overstate what the numbers support.
Clinic 11: the mean of coded categories is often not meaningful
If categories are encoded 1, 2 and 3 for convenience, the arithmetic mean of those codes does not automatically represent a real central category unless the coding carries quantitative meaning.
Clinic 12: missing values should be counted
Before reporting any average, record how many observations were expected and how many were available. Missingness can alter the distribution and may itself carry information.
Clinic 13: sample selection should match the reader job
A dataset collected from volunteers, one location or one time window may answer a local descriptive question while remaining weak evidence for a broader population.
The statistical summary cannot be more general than the sampling process permits.
Clinic 14: uncertainty belongs next to estimates
When moving from sample description to population inference, point estimates should be interpreted with uncertainty. At more advanced levels this can involve standard errors, confidence intervals or hypothesis procedures.
The broader principle is simple: a sample statistic is not the population itself.
Clinic 15: visual scale can distort perception
A graph with a truncated vertical axis can make small differences look dramatic. An overly wide axis can make meaningful changes look flat.
Always inspect axis limits before interpreting visual magnitude.
Clinic 16: sorting data can help and hide
Sorting reveals median, quartiles and extremes. It removes time order. Preserve the original sequence if timing or progression matters.
Clinic 17: rounding can change comparisons near a threshold
If two means are 69.96 and 70.04, rounding to one decimal gives 70.0 for both. A reporting rule can therefore erase small differences.
Choose precision based on data quality and decision needs rather than display convenience.
Clinic 18: very precise averages can create false authority
A mean reported to six decimal places may look scientific even when the underlying measurements are coarse. Numerical precision should not exceed evidential precision without a reason.
Clinic 19: robust summaries do not eliminate the need to inspect the data
Median and IQR protect against some extreme-value influence. They still cannot show every cluster, gap, time pattern or data-quality problem.
Robust does not mean complete.
Clinic 20: standard deviation needs context
A standard deviation of 5 can be small or large depending on the scale and purpose of the variable. Compare it with the mean, allowable range, decision threshold or another relevant reference.
Clinic 21: coefficient of variation has conditions too
At advanced levels, relative spread measures can help compare variables on different scales, but ratios involving a mean near zero can become unstable or misleading. Every summary has a jurisdiction.
Clinic 22: extreme values can be the reader job
If the goal is safety, reliability or worst-case planning, the tail may matter more than the average. A low average failure rate can coexist with rare but severe failures.
Do not let central tendency erase the exact part of the distribution the decision cares about.
Clinic 23: ordinary experience and total burden can require different statistics
Median may describe a typical individual experience while mean may better represent total resource burden when every value contributes to the total.
The right summary follows the question.
Clinic 24: publication standard
A strong statistical statement should make clear:
- what was measured;
- how many observations were included;
- which summary was used;
- why that summary fits the reader job;
- what important distribution information remains outside the summary;
- whether subgroup, time-order or sampling limitations matter.
Final transfer task
A report states: “Average response time improved from 22 minutes to 18 minutes.”
Before concluding that the system improved for most users, ask for:
- sample sizes in both periods;
- median and spread;
- distribution or percentile information;
- any changes in case mix;
- the proportion of very long delays;
- whether measurement rules changed.
The average may be accurate and still insufficient for the intended interpretation.
Release standard
The learner is statistically ready when they stop treating the average as the answer and start treating it as one piece of evidence. They can pair centre with spread, recover denominators, inspect distribution shape, respect subgroup weights, preserve time order when it matters, and state what the available data cannot establish.

