Application of Mathematics in Real-World Usage · Guide 13 · BTT Mathematics Hub
Two machines produce parts with the same average diameter. One produces almost every part close to the target. The other alternates between parts that are too small and parts that are too large. Their averages agree, but their usefulness does not. A single number has compressed away the difference that matters.
Manufacturing gives school Mathematics a demanding real-world job: distinguish an acceptable item from a dependable process. Measurement creates observations. Inequalities express requirements. Statistics describes variation. Probability explains why an inspection sample can miss defects. Multiplication connects the yields of successive stages. None of these tools works well when its denominator, assumptions or decision boundary is left unclear.
This guide builds a fictional parts-production case from a specification through inspection, process monitoring and a defensible report. Every dimension, count, cost and production history below is an original teaching construction, not a measurement of a real factory. The examples do not certify equipment, establish an industrial acceptance plan or authorise the release of safety-critical products. Their purpose is to make the mathematical reasoning visible.
Start with specifications and measurements, continue to yield and material flow, then examine process capability, sampling risk and the twenty-question practice set.
A specification describes the allowed result, not the observed process
Suppose an educational model requires a diameter from 9.80 mm to 10.20 mm, inclusive, with target 10.00 mm. The specification interval is 9.80 ≤ d ≤ 10.20. Its width is 0.40 mm. The target sits at the midpoint, but meeting the target exactly is not the only permitted outcome.
A reading of 10.16 mm lies inside the numerical interval; a reading of 10.23 mm lies outside it. That comparison is an inequality test. It does not yet establish what a measuring instrument can resolve, whether the measurement is biased or how many future items will pass.
Write down the inclusion rule for the endpoints. In this example, equality at 9.80 and 10.20 is permitted. Another requirement might use a strict inequality. A rounded display at the boundary also needs interpretation: a displayed 10.20 does not necessarily identify an underlying value of exactly 10.20000. Representation and physical measurement are not identical objects.
A measured pass is not automatically a guaranteed pass
Consider an explicitly bounded measurement model: the true diameter is assumed to lie within 0.03 mm of the reading. For a reading of 10.16 mm, possible values run from 10.13 to 10.19 mm. Every value in that interval satisfies the specification, so this bounded model supports a guaranteed numerical pass.
For a reading of 10.19 mm, the possible interval is 10.16–10.22 mm. Some values pass and some fail. The central reading passes, but the stronger statement that all permitted true values pass is unsupported. For a reading of 10.25 mm, the interval is 10.22–10.28 mm, entirely outside the upper limit.
Under this particular hard-bound assumption, requiring the complete measurement interval to fit inside the specification produces an acceptance region for readings of 9.83–10.17 mm. This is a conservative classroom decision rule, not a prescribed industrial guard band. Actual measurement uncertainty often has a probability-based interpretation rather than an absolute guarantee, so its stated coverage and decision rule must be respected.
Repeatability, bias and resolution are different questions
Imagine repeated readings of a reference object: 10.11, 10.12, 10.11, 10.12 and 10.11 mm. The readings are tightly grouped. If the reference value is 10.00 mm and is trusted at the required accuracy, however, the group is displaced from it. Repeated agreement does not remove a systematic offset.
A second instrument might show only one decimal place. Its readings may look identical because differences below its display increment are hidden. Identical displayed numbers therefore do not prove zero variation. Extra digits, conversely, do not prove that the digits are accurate.
For a learner, the useful habit is to separate what was recorded from what was inferred. “The five displays were equal” is an observation about representation. “The five parts were identical” is a stronger claim. The earlier guide on Measurement, Units, Scale and Estimation develops these distinctions before they are used in a production decision.
The mean can be right while individual parts are wrong
Machine A produces the fictional readings 9.98, 10.00 and 10.02 mm. Machine B produces 9.70, 10.00 and 10.30 mm. Both means are 10.00 mm. Under the stated 9.80–10.20 interval and ignoring measurement uncertainty for this comparison, all three A readings pass, while two B readings fail.
The range for A is 0.04 mm. The range for B is 0.60 mm. This simple spread measure reveals something the mean hides. It does not provide enough evidence to estimate long-run process capability, because three constructed observations are not a production history.
The same mistake occurs outside factories. An average waiting time can conceal occasional very long waits. Average marks can hide uneven topic performance. The remedy is not to discard the mean, but to ask which additional feature of the distribution the decision needs.
Calculate spread with the correct formula and unit
Take five readings: 9.96, 9.98, 10.00, 10.02 and 10.04 mm. Their mean is 10.00 mm. Deviations from that mean are −0.04, −0.02, 0, 0.02 and 0.04 mm. Squared deviations sum to 0.004 mm².
Using the sample-variance convention, divide by n − 1 = 4 to obtain 0.001 mm². The sample standard deviation is √0.001, approximately 0.03162 mm. If the task instead describes these five values as the complete finite population of interest, dividing by five gives a different descriptive population variance. State the convention.
Variance carries squared units, while standard deviation returns to the measurement unit. This is why standard deviation can be compared with a tolerance width measured in millimetres. A number without its unit can conceal a mistake even when the calculator operations are correct.
Count nonconforming items separately from defects
Suppose 100 items are inspected. Four items each have one defect, and two other items each have three defects. There are six nonconforming items, but ten defects. The proportion of nonconforming items is 6/100 = 6%; the average number of defects per inspected item is 10/100 = 0.10.
Those measures use different numerators. Treating ten defects as ten defective items double-counts items with several defects. Conversely, counting only six items hides how many individual defect observations were recorded.
NIST’s proportions-control-chart guidance defines a nonconforming item by whether at least one inspected requirement fails. That definition supports item-level proportion calculations. It does not imply that a part with three defects represents three independent items. Choose the counting object before choosing the probability model.
First-pass yield and final yield tell different production stories
A fictional batch begins with 1,000 items. Nine hundred pass immediately. Of the remaining 100, seventy pass after rework and thirty are scrapped. First-pass yield is 900/1,000 = 90%. Final acceptable yield is 970/1,000 = 97%. Scrap proportion is 3%.
The final yield looks encouraging, but it does not erase the extra work applied to seventy items. A report that shows only final yield hides that operational burden. A report that shows only first-pass yield understates the final quantity available for use.
The conservation check is simple: 900 immediate passes + 70 recovered passes + 30 scrapped = 1,000 starts. Reworked items must not be counted as new starts in the same denominator unless the metric explicitly measures processing attempts rather than original items. Count identities are part of the mathematics.
Successive stage yields multiply when their bases are conditional
Imagine 100,000 units entering a three-stage process. Stage 1 passes 95%, leaving 95,000. Stage 2 passes 98% of those survivors, leaving 93,100. Stage 3 passes 99% of its incoming survivors, leaving 92,169. Overall first-pass yield is 0.95 × 0.98 × 0.99 = 0.92169, or 92.169%.
Adding the percentages would be meaningless. Averaging them would also answer a different question. The whole route requires every stage to succeed, and each stage’s denominator is the group that reached it.
This multiplication does not require the stages to be statistically independent when the stated percentages are conditional on reaching each stage. Multiplying separate unconditional probabilities, however, would require additional justification. The distinction is subtle but valuable: a correct-looking product can rely on very different assumptions depending on what its inputs mean.
Pooling batches requires pooling the underlying counts
One batch contains 10 nonconforming items out of 100; another contains 9 out of 900. The batch rates are 10% and 1%. Their simple average is 5.5%, but the pooled item-level rate is 19/1,000 = 1.9%.
The simple average gives equal weight to the two batches. The pooled rate gives equal weight to each item. Either can be defined, but only the second answers “what fraction of all inspected items was nonconforming?” Naming the averaging unit prevents a misleading comparison.
Keep batch identifiers even after pooling. A low combined rate can conceal a problematic line, shift or material lot. Aggregation is useful for totals, while subgrouping is useful for locating a pattern. The two views should support rather than replace one another.
Capability compares process spread with a specification
NIST’s process-capability chapter gives Cp = (USL − LSL)/(6σ) and Cpk = min[(USL − μ)/(3σ), (μ − LSL)/(3σ)]. Here USL and LSL are specification limits, μ is process mean and σ is process standard deviation. Interpretation requires an appropriate stable process model; these formulas do not certify a process merely because a few readings are available.
For our fictional specification, assume a stable normal process with μ = 10.00 mm and σ = 0.05 mm. Cp is 0.40/(6 × 0.05) = 1.3333. Both distances from the mean to the specification edges are 0.20 mm, so Cpk is also 0.20/0.15 = 1.3333.
These are calculations from stipulated population parameters, not estimates from the earlier five measurements. The distinction prevents a common evidential shortcut: using a tiny example to make a precise claim about long-run production. The six-standard-deviation width is a model comparison, not a statement that every observation lies within three standard deviations of the mean.
The same Cp can hide a shift toward one limit
Keep σ = 0.05 mm but move the assumed mean to 10.12 mm. The specification width and process spread are unchanged, so Cp remains 1.3333. The upper allowance is now only 10.20 − 10.12 = 0.08 mm; the lower allowance is 0.32 mm.
Cpk becomes min(0.08/0.15, 0.32/0.15) = 0.5333. The minimum selects the more restrictive side. The change exposes the loss of centring that Cp alone cannot show. If the assumed mean moves outside the allowed interval, one term can become negative; that is not a calculator error.
There is no universal release threshold asserted here. A real requirement depends on the application and its agreed quality system. Our purpose is to understand why two formulas can give different information about the same distribution. A larger collection of indices does not remove the need to understand what each measures.
A normal-model tail is a prediction, not an observed defect count
In the centred hypothetical normal model, both specification edges are four standard deviations from the mean because 0.20/0.05 = 4. The two-sided probability outside those edges is approximately 0.00006334, or about 63.34 per million. This follows from the normal distribution assumed for the example.
In the shifted model, the upper edge is only 1.6 standard deviations above the mean, while the lower edge is 6.4 below it. The predicted upper-tail probability is about 5.48%, with a negligible lower tail in this model. The same σ can therefore accompany very different modelled failure probabilities.
Do not convert the prediction into a claim that a million items were inspected. Do not apply these tail values to an arbitrary non-normal distribution either. A distributional assumption is part of the answer. NIST’s capability guidance specifically discusses the complications introduced by non-normal data and by estimating parameters from samples.
Specification limits are not control limits
A specification says what output is acceptable. A control limit is part of a rule for detecting unusual behaviour relative to a process model. A process can be statistically stable yet consistently produce too many out-of-specification items. It can also produce individually acceptable items while beginning to shift away from its historical pattern.
For a fictional stable process with known mean 10.00 and standard deviation 0.05 mm, the mean of 25 independent readings has standard deviation 0.05/√25 = 0.01 mm. A three-standard-error illustration would place mean-monitoring limits at 9.97 and 10.03 mm. Those are limits for subgroup means, not for individual diameters.
Comparing an individual part with those mean limits would mix two different random quantities. NIST’s control-chart overview provides the process-monitoring context. The classroom calculation shows why the sample size and the statistic being charted matter.
A proportion chart needs a model and a sample size
Assume a baseline nonconforming probability p = 0.04 and independent item outcomes. For a sample of n = 100, the standard deviation of the sample proportion is √[0.04 × 0.96/100], approximately 0.019596. The usual three-standard-deviation upper line is about 0.098788; the negative lower result is truncated at zero for the plotted proportion.
A subgroup with 11 nonconforming items has proportion 0.11 and crosses that illustrative upper line. It calls for investigation under the chosen monitoring rule. It does not prove a particular cause, nor does it mean that every item in the subgroup is defective.
NIST derives this p-chart structure from a stable binomial model. With small expected counts, the distribution is discrete and asymmetric, so a normal-style three-sigma rule does not automatically have the familiar exact two-tail probability. Variable sample sizes also change the line width. A wider chart is not evidence that the product requirement became more lenient.
A signal suggests where to look; it does not establish causation
Suppose several high readings occur after a material change. The sequence is worth investigating, but timing alone does not establish that the material caused the shift. An instrument adjustment, a temperature change, a different operator or a changed sampling method might also explain it.
A useful next observation distinguishes those explanations. Rechecking a trusted reference tests the measuring system. Comparing contemporaneous runs under controlled conditions can help test a material hypothesis. Combining all readings into a single monthly mean may hide the very transition that needs explanation.
This is a reasoning principle rather than a production procedure: a pattern earns investigation, not an invented mechanism. Mathematics is strongest when it narrows the next question instead of making uncertainty disappear rhetorically.
A clean sample does not prove a clean population
Assume independent sampling from a large population with true nonconforming probability 5%. The probability that twenty sampled items all pass is 0.95²⁰, approximately 35.85%. A process with a nonzero defect probability can therefore produce a completely clean small sample quite often.
For sixty samples, the same probability is 0.95⁶⁰, about 4.61%. More observations reduce this particular chance of missing every defect, but do not turn sampling into certainty. The calculation also assumes the selected items represent the process rather than being chosen because they look good.
NIST distinguishes acceptance sampling from estimating exact lot quality. An acceptance plan makes a decision with risks. This example demonstrates one such risk; it is not a recommended sample-size rule for an actual product.
Finite-lot sampling changes the probability calculation
Suppose a known finite teaching lot contains twenty items, exactly two of them nonconforming, and five distinct items are sampled uniformly without replacement. The probability of selecting no nonconforming items is C(18,5)/C(20,5), which equals 21/38, approximately 55.26%.
The independent-trial expression 0.9⁵ gives about 59.05%, a different result. Without replacement, the composition of the remaining lot changes after each draw. Counting valid combinations or multiplying the changing conditional probabilities respects that structure.
The difference is particularly visible when the sample is a substantial fraction of the lot. The denominator should describe the actual selection process, not whichever probability formula was most recently practised in class.
Zero observed failures still leaves an uncertainty interval
Under a constant independent binomial model, observing zero failures in 100 trials gives the one-sided exact 95% upper confidence limit 1 − 0.05^(1/100), approximately 2.95%. The equation sets the probability of zero failures at that boundary equal to 0.05.
This is not a statement that there is a 95% frequentist probability that the fixed true rate lies below the realised bound. It is a property of the confidence procedure under its assumptions. Nor does the interval account for biased sampling, missed defects or changing process behaviour.
NIST’s confidence-interval guidance for proportions explains why interval methods are needed rather than treating the sample proportion as the exact population rate. For school learning, the central insight is already powerful: “none found” is not synonymous with “none exist.”
Evaluate an improvement using both quantity and consequence
Return to the 1,000-item batch with seventy reworked and thirty scrapped. Assign fictional extra costs of $2 per reworked item and $5 per scrapped item, with no double counting of these categories. Their combined extra cost is 70 × 2 + 30 × 5 = $290.
A proposed change costs $120 per batch and is assumed to reduce rework to twenty items and scrap to ten. The new total is $120 + 20 × $2 + 10 × $5 = $210, a modelled reduction of $80. This is a conditional comparison, not an observed saving or a complete business case.
The new process must still satisfy its product requirements. A lower inspection rejection count achieved by weakening the specification is not the same improvement as producing better parts under the original requirement. Keep the target definition fixed when comparing results, or disclose that the question has changed.
Practice: twenty manufacturing decisions
Use the fictional assumptions exactly as stated. The first questions establish quantity control; later questions test probability and interpretation. Keep the question section separate from the answers while working.
- A specification is 24.7–25.3 mm inclusive. Does 25.28 mm pass the numerical specification?
- A reading is 25.28 mm with a hard error bound ±0.04 mm. Is a pass guaranteed under question 1?
- Find the mean and range of 9.9, 10.0 and 10.1 mm.
- A 200-item sample contains twelve defects spread across eight items. Find the nonconforming-item percentage.
- For question 4, find defects per inspected item.
- Of 500 starts, 450 pass immediately, 35 pass after rework and 15 are scrapped. Find first-pass and final yield.
- Two successive conditional stage yields are 0.90 and 0.95. Find overall yield.
- Samples contain 4 nonconforming out of 50 and 6 out of 150. Find the pooled proportion.
- For limits 9.8 and 10.2 mm and stipulated σ = 0.05 mm, calculate Cp.
- Add stipulated μ = 10.10 mm to question 9. Calculate Cpk.
- Individual standard deviation is 0.06 mm. Under independence, find the standard deviation of a mean of 36 readings.
- For p = 0.04 and n = 100, calculate the illustrative three-sigma p-chart upper limit.
- Does 11/100 cross that limit? What does a crossing not prove?
- Under independent trials with p = 0.05, calculate the probability of no defects in twenty samples.
- A lot contains ten items, two defective. Two are drawn without replacement. Find the probability that both pass.
- Explain why zero failures in a sample does not prove a zero population failure rate.
- Find the sample standard deviation of 2, 4 and 6 using denominator n − 1.
- Thirty reworks cost $3 each and eight scrapped units cost $7 each under disjoint cost categories. Find the extra cost.
- A sample mean crosses a process-monitoring line while every measured item meets specification. Is this logically possible?
- A factory reports a higher final yield but omits rework counts. What additional information is needed before calling the change more efficient?
Worked answers and interpretation
- Yes. The reading satisfies 24.7 ≤ 25.28 ≤ 25.3. This answers the stated numerical comparison, not a separate question about measurement uncertainty.
- No guarantee. Possible values run from 25.24 to 25.32 mm, crossing the upper limit. Some allowed values pass and some fail.
- Mean 10.0 mm; range 0.2 mm. The mean gives location, while maximum minus minimum gives one measure of spread.
- 4%. Eight distinct items fail among 200 inspected. The twelve defect observations are not twelve different items.
- 0.06 defects per item. Divide twelve defects by 200 inspected items. State this different numerator explicitly.
- 90% first-pass; 97% final. Immediate passes are 450; acceptable output after rework is 485. The fifteen scrapped units complete the count balance.
- 85.5%. Multiply 0.90 × 0.95 because the second yield is conditional on reaching stage 2.
- 5%. Pool ten nonconforming items across 200 inspected. Averaging 8% and 4% would weight batches equally instead.
- Cp = 1.3333 approximately. Divide the 0.40 mm specification width by 6 × 0.05 mm. Units cancel.
- Cpk = 0.6667 approximately. The smaller mean-to-limit distance is 0.10 mm; divide it by 0.15 mm.
- 0.01 mm. Use 0.06/√36. This is the variability of the mean, not the variability of each part.
- About 0.098788. Calculate 0.04 + 3√(0.04 × 0.96/100), about 9.879%.
- Yes. A signal does not identify a cause or imply that every item fails. It prompts investigation under the model and rule used.
- About 35.85%. Each independent pass has probability 0.95, so all twenty pass with probability 0.95²⁰.
- 28/45, about 62.22%. Use 8/10 × 7/9. The second probability changes because the first item is not replaced.
- Nonzero defect rates can generate clean finite samples. Sampling uncertainty and representativeness remain even when the recorded numerator is zero.
- 2. Mean is 4; squared deviations sum to 8; divide by 2 and take the square root.
- $146. Add 30 × $3 and 8 × $7. The stipulated categories prevent counting the same extra cost twice.
- Yes. Monitoring a process statistic and checking individual specification compliance answer different questions with different limits.
- Compare starts, rework effort, scrap, cost, time and unchanged product requirements. A higher final yield alone does not establish a lower resource cost per acceptable item.
How this develops mathematical judgement
Teach the sequence in layers. Begin with an interval and a measured value. Add a measurement bound. Next compare two tiny datasets with the same mean and different spread. Only after those ideas are secure should a learner encounter capability indices or monitoring limits. Otherwise the formulas arrive before the questions they answer.
For a stronger learner, require a two-column report: “established by the data or assumptions” and “not established.” The first column may contain a count balance, an observed sample proportion or a conditional model prediction. The second may contain future yield, cause of a shift or safety of product release. This prevents correct arithmetic from becoming an unsupported conclusion.
The most useful final question is not whether the learner remembers Cp. It is whether the learner can distinguish a part, a sample, a process and a future claim. Those are different objects. Manufacturing Mathematics becomes dependable when the representation keeps them separate.
Sources and connected applications
Technical reference routes are the NIST Engineering Statistics Handbook sections on process capability, control charts, proportion charts, acceptance sampling and confidence intervals for proportions. The worked batches and exercises are original educational examples.
Continue with Cooking, Recipes, Scaling, Timing and Unit Conversion for yields and production quantities; Sports, Pace, Scoring, Rankings and Performance Data for fair comparisons; and Data Networks, Storage, Bandwidth and Transfer Time for another system where averages and limits answer different questions. Return to the BTT Mathematics Hub.
