Bukit Timah Tutor Mathematics

A connected Mathematics learning system from school foundations to examinations, applications and advanced study. Use the Mathematics Hub to move between levels, concepts, diagnosis, examinations, applications and world routes.

Real-World Mathematics: Reliability Engineering, Failure Rates, Redundancy and Availability

Three learners review open books together at a classroom table, with stacks of textbooks, stationery and a whiteboard in the bright room.

Application of Mathematics in Real-World Usage · Guide 29 · BTT Mathematics Hub

A component can have a low failure rate and still make a system unreliable if many such components must all work. Two identical components can improve system reliability dramatically when either one is sufficient, but only if their failures are sufficiently independent. A machine can be very reliable yet have poor availability if repairs take a long time. Reliability Mathematics separates these ideas so that a percentage has a clear operational meaning.

This guide uses fictional components, operating times, failure rates and repair times for Mathematics teaching. It is not engineering certification, maintenance advice or a safety case. Real systems may have common-cause failures, changing environments, ageing, maintenance policies, imperfect detection and dependencies that simple textbook models omit.

Reliability is a survival probability

For a non-repairable item, reliability R(t) is the probability the item survives beyond time t under the stated conditions. If R(500)=0.90, the model says 90% of comparable items are expected to survive past 500 time units.

Reliability is tied to a mission time. Saying “90% reliable” without specifying the time and operating conditions can be incomplete.

Failure probability complements reliability

If F(t) is the probability of failure by time t, then F(t)=1−R(t). A reliability of 0.92 at a declared mission time corresponds to cumulative failure probability 0.08 by that time.

This does not mean an 8% failure rate per hour. A cumulative probability over a mission and an instantaneous hazard rate are different mathematical objects.

Hazard rate is conditional on survival so far

NIST defines the hazard or failure rate as the instantaneous conditional failure rate among units that have survived to time t. It can change with age.

A component population may have a high early-life hazard, a flatter middle-life hazard and an increasing wear-out hazard. A single constant-rate model should therefore be treated as an assumption, not a universal property.

The exponential model assumes constant hazard

Under the exponential lifetime model, R(t)=e^(−λt), where λ is a constant hazard rate. NIST notes that the mean lifetime in this model is 1/λ.

For λ=0.001 per hour, reliability at 500 hours is e^(−0.5)≈0.6065. The mean lifetime is 1,000 hours.

The median lifetime is ln2/λ≈693.1 hours. Mean and median differ because the exponential distribution is right-skewed.

Conditional survival reveals the memoryless property

For an exponential model, P(T>t+s | T>t)=e^(−λs). Once survival to t is given, the additional survival probability depends only on s.

This memoryless property is mathematically convenient but often unrealistic for ageing hardware. It belongs to the model, not to reliability in general.

Series systems multiply component reliabilities

If a system works only when every independent component works, system reliability is the product of component reliabilities. Three independent components each with mission reliability 0.90 give R=0.9³=0.729.

Each component looks strong individually, yet the all-must-work system is only 72.9% reliable. Adding required components can lower system reliability unless their individual reliabilities improve.

Parallel redundancy can increase reliability

NIST’s parallel model assumes the system works as long as at least one independent component remains operating. Two independent components each with reliability 0.90 fail together with probability 0.1²=0.01, so parallel system reliability is 0.99.

The improvement depends on independence. If both components share the same power supply, environment or design defect, common-cause failure can make the independence calculation optimistic.

k-out-of-n reliability uses the binomial model when components are independent and identical

A 2-out-of-3 system works if at least two of three components work. With independent component reliability p=0.90, system reliability is C(3,2)p²(1−p)+p³.

That is 3×0.9²×0.1+0.9³=0.972. The redundancy is weaker than a 1-out-of-2 parallel system’s 0.99 because two working units are required rather than one.

Availability includes repair

Reliability asks whether failure has occurred during a mission. Availability asks whether a repairable system is operational at a randomly chosen time under a stated steady-state model.

For a simple two-state model, availability A=MTBF/(MTBF+MTTR). If MTBF=1,000 hours and mean time to repair MTTR=10 hours, A=1000/1010≈99.01%.

Expected unavailable fraction is about 0.9901%. Across 8,760 hours, that corresponds to about 86.7 hours of expected downtime under the simplified steady-state interpretation.

Reliability and availability can rank systems differently

System A fails rarely but takes 100 hours to repair. System B fails twice as often but is restored in 1 hour. Depending on the mission and the repair model, A can have better mission reliability while B has better long-run availability.

The decision metric should match the real question: uninterrupted mission success, fraction of time operational, expected failures, or repair burden.

Weibull models allow changing hazard

A common reliability form is R(t)=exp[−(t/η)^β]. The scale η controls time scale and the shape β controls hazard behaviour.

When β=1, the model reduces to an exponential form. When β>1, hazard increases with time; when β<1, it decreases in the idealised Weibull model.

For η=1,000 hours, β=2 and t=500 hours, R=e^(−0.25)≈0.7788.

Expected failures use a process model

For a repairable homogeneous Poisson process with constant rate λ, NIST gives expected cumulative failures M(t)=λt.

If λ=0.002 failures/hour over 5,000 operating hours, expected failures are 10. The realised count need not equal 10; ten is the model expectation.

A reliability block diagram is logic written as probability

Series means logical AND: every required branch must succeed. Parallel means logical OR: at least one branch must succeed. Mixed systems can therefore be reduced section by section if the independence assumptions are appropriate.

For two parallel 0.9 components followed in series by one 0.95 component, reliability is [1−0.1²]×0.95=0.9405.

Redundancy has diminishing returns under independence

For n identical 0.90 components in a 1-out-of-n parallel arrangement, system failure probability is 0.1ⁿ. Reliability is 1−0.1ⁿ.

One component gives 90%, two 99%, three 99.9%, four 99.99%. Each added independent component removes another factor of ten from this particular failure probability, but practical cost and common-cause risks are not represented.

Testing without failure does not prove perfect reliability

Observing zero failures during a finite test provides evidence, but it cannot prove that the true failure probability is zero. Reliability inference requires a statistical model, test duration and confidence statement.

The absence of observed failures is therefore data, not proof of immortality.

A complete reliability report states the mission and dependence assumptions

State mission time, operating conditions, whether the system is repairable, the lifetime distribution assumed, the component dependency structure and whether the metric is reliability, failure intensity, MTBF or availability.

The mathematical return path is component behaviour → system logic → mission or repair model → system-level metric.

Practice: twenty reliability Mathematics questions

  1. If R=0.93, find cumulative failure probability.
  2. For λ=0.001/h, find exponential mean lifetime.
  3. Find exponential reliability at 500 h.
  4. Find exponential median lifetime for λ=0.001/h.
  5. Three independent series components each have reliability 0.9. Find system reliability.
  6. Two independent parallel components each have reliability 0.9. Find system reliability.
  7. For a 2-out-of-3 system with p=0.9, find reliability.
  8. MTBF=1000 h and MTTR=10 h. Find availability.
  9. Using question 8, estimate unavailable hours in 8,760 h.
  10. For Weibull η=1000, β=2, find R(500).
  11. If λ=0.002 failures/h for 5,000 h in an HPP model, find expected failures.
  12. Two parallel 0.9 components are in series with a 0.95 component. Find total reliability.
  13. Find reliability of three independent 0.9 components in 1-out-of-3 parallel form.
  14. Find reliability of four such parallel components.
  15. If reliability at a mission time is 0.8, what is failure probability by that time?
  16. If availability is 0.999, what fraction of time is unavailable?
  17. Why does adding required series components generally reduce reliability?
  18. Why can common-cause failure invalidate a simple parallel calculation?
  19. Why is MTBF not the same as guaranteed lifetime?
  20. Why does zero observed failures not prove failure probability zero?

Worked answers

  1. 0.07.
  2. 1,000 h.
  3. About 0.6065.
  4. About 693.1 h.
  5. 0.729.
  6. 0.99.
  7. 0.972.
  8. About 0.9901, or 99.01%.
  9. About 86.7 h.
  10. About 0.7788.
  11. 10.
  12. 0.9405.
  13. 0.999.
  14. 0.9999.
  15. 0.2.
  16. 0.001, or 0.1%.
  17. Because every required component creates another condition that must simultaneously succeed.
  18. Because the independence assumption fails when components share a failure mechanism.
  19. Because MTBF is an average under a model, not a minimum guaranteed operating time.
  20. Because finite observation cannot rule out a small nonzero failure probability.

Sources and connected applications

For failure-rate and exponential-model definitions, see the NIST Engineering Statistics Handbook: Failure Rate and Exponential Life Distribution. For redundant-system structure, see NIST: Parallel or Redundant Model.

Continue with Control Systems, Feedback, Sensors, Error and Stability; Signal Processing, Sampling, Filtering, Noise and Fourier Analysis; and Urban Planning, Land Use, Density, Accessibility and Networks. Return to the BTT Mathematics Hub.