Operational risk mathematics asks how often a process fails, how much each failure costs, and how those losses combine over time. Frequency models count incidents. Severity models describe their financial consequences. Aggregate loss models join the two while preserving reporting thresholds, dependence, uncertainty and recovery. This is the quantitative foundation behind questions about processing errors, failed controls, fraud losses, technology outages and disrupted banking services.
A useful operational risk model does more than multiply an incident rate by an average loss. It distinguishes expected annual loss from a severe annual outcome, observed incidents from unreported events, gross losses from delayed recoveries, and financial loss from interruption of a critical service. Poisson and negative-binomial counts, lognormal and tail severity models, compound distributions, Value at Risk and Expected Shortfall each answer a different part of that problem. Their formulas become reliable only when the event definition and data boundary are clear.
This Banking And Finance Mathematics guide develops the subject through original worked examples rather than treating a statistical distribution as a ready-made capital requirement. The Basel operational-risk framework uses a standardised capital methodology; the internal frequency–severity models developed here are analytical tools, not an assertion that a bank may replace its applicable regulatory rules with our calculations. The Basel principles for operational resilience also address continued delivery of critical operations, a question that cannot be answered by an annual loss quantile alone.
The 50-second route
Start with a defined incident and a reconciled loss record. Estimate frequency using the correct exposure period or transaction volume. Model the size of a loss conditional on an incident, correcting for what the database omits. Combine the count and severity distributions without assuming away common causes. Then test the intervention: does it prevent an event, detect it earlier, limit its size, transfer part of the cost or restore service sooner? Keep each improvement in its own units before comparing the total result.
Route 1 — Understand the data before fitting a model
Read the event and loss boundary, then the observation process. These sections explain why a database row is not automatically an independent incident, why recoveries need their own dates and why fewer reported losses can mean several different things.
Route 2 — Learn frequency and severity mathematics
Begin with frequency models and continue to loss severity. Follow the numerical examples for Poisson rates, overdispersion, Bayesian updating, lognormal losses and threshold correction. Every probability has a horizon and every severity has a currency unit.
Route 3 — Calculate the annual loss distribution
Go to aggregation. A complete finite example derives every possible annual loss before calculating the mean, variance, quantiles and tail average. The dependence examples then show why the same marginal losses can produce very different system-level outcomes.
Route 4 — Evaluate controls, insurance and recovery
Use control effectiveness and insurance and recoveries. The purpose is to distinguish fewer incidents, smaller incidents, transferred losses and shorter disruption. A control that improves one of these measures need not improve all of them equally.
Route 5 — Challenge a model or teach the subject
Read validation and regulatory boundaries, then the application workshop. Finish with reader questions and the teaching guide. A sound explanation identifies both the result and the assumption that could change it.
Contents and source routes
I. Data and event definitions · II. Frequency · III. Severity · IV. Aggregate loss · V. Controls and recovery · VI. Validation and boundaries · VII. Applications · Sources.
Conventions. Unless stated otherwise, monetary examples use hypothetical Singapore dollars or explicitly labelled thousands of Singapore dollars. No example describes an actual bank, insurer, customer or loss database. Adrian, Jo, Aisha, Ryan, Ben, Mira, Clara and Ethan are recurring fictional learners. The guide is educational, not financial, actuarial, insurance or regulatory advice. A date attached to an external source is not a claim that its findings predict the next incident.
I. The mathematics begins with what happened
1. A loss event is not the same thing as a bad financial result
Suppose a bank correctly executes an authorised bond purchase and the bond price subsequently falls. That is a market-loss story under the stated circumstances. Suppose instead a processing failure sends the wrong quantity or duplicates the purchase. The resulting financial consequences have an operational cause. The price movement may affect their size, but the event being investigated is the failure to execute the intended instruction correctly.
The Basel definition centres on losses caused by failures involving processes, people, systems or external events. It includes legal risk while distinguishing strategic and reputational risk from the defined category. That classification does not mean reputation is unimportant. It means a particular regulatory or analytical loss measure has a boundary, and consequences outside it need another clearly labelled measure.
Adrian wants one universal list of operational events. Jo asks what decision the list supports. A customer-service team may count every failed transaction to measure reliability. A finance team records financial charges. An incident-response team groups thousands of failed transactions under one technology disruption. These views can all be useful. Problems arise when one count is divided by another team’s denominator and presented as a single event rate.
Our first requirement is therefore a complete sentence: an event in this model means a distinct processing incident that meets these inclusion rules over this horizon. The definition should identify whether a recurring problem constitutes one continuing event or several new events, whether near misses are included, and what evidence closes the case. This is not administrative detail added after modelling. It determines the random variable whose distribution we intend to estimate.
A precise boundary also prevents false comparisons. A model of all customer inconvenience cannot be compared directly with a model of booked operational charges above a threshold. One may include a large outage with no immediate accounting loss, while the other may include a legal settlement recognised years after the original conduct. Before asking which model has the larger average, ask whether their averages refer to the same thing.
2. Count the incident once and record its consequences completely
A software release causes the same fee error in 10,000 accounts. The bank makes 10,000 customer adjustments, hires a specialist team and later pays a settlement. Counting each adjustment as an independent incident can produce a huge apparent frequency of tiny losses. Counting only the release error as one event produces low frequency and a large severity. Both representations can preserve the same total cash, but they imply different dependence and forecasting models.
For an event-based annual loss model, a sensible teaching convention is to group consequences by the common incident and keep the underlying transactions as linked detail. Frequency then counts incidents; severity sums the included consequences of each incident. A separate transaction-quality model can still count each incorrect charge. We have not discarded detail. We have assigned it to the level at which it answers the question correctly.
The event identifier should remain stable as new charges arrive. If a case costs 80,000 when discovered and another 20,000 when resolved, it need not become a second independent event. The event’s estimated ultimate severity has changed. A record version can show the original estimate, subsequent development and final payments. Otherwise the latest year’s frequency may rise merely because accountants processed another instalment of an older problem.
Grouping also needs limits. Two separate failures caused by the same weak control are not necessarily one event forever. There may be repeated incidents with a common causal factor. A useful hierarchy has a transaction ID, incident ID and root-cause ID. The incident defines count and severity, while the root-cause grouping supports dependence and control analysis. This permits investigation of recurring weaknesses without collapsing an entire business line into a single permanent loss.
Aisha tests the database by selecting a known incident and following every related charge. The sum should reconcile to the stated event total under the chosen inclusion convention. She then asks whether those same charges appear under another incident ID. This duplicate test can improve a model more than adding a sophisticated tail distribution to an unreconciled table.
3. Gross loss, net loss and temporary cash need are three different amounts
A mistaken transfer removes 100,000 from a bank’s account. It recovers 70,000 after two days and spends 5,000 on directly associated external recovery work. Under an explicitly chosen economic-event convention, gross cash damage including that recovery expense is 105,000, recovered funds are 70,000 and net loss is 35,000. The maximum temporary cash need may nevertheless have been 100,000 or more before the recovery arrived.
The database should not erase the original outflow when a recovery becomes likely. An expected recovery, a legally established receivable and collected cash are different states. A valuation model may recognise a receivable according to its own rules, but a liquidity model cannot spend it before collection. Recording gross consequences, recovery amounts, recovery type and dates preserves the information needed for both analyses.
Netting all records into a single signed amount can also damage severity modelling. If a recovery is entered as a separate negative event, a positive-loss distribution such as the lognormal no longer matches the raw variable. The repair is not to delete recoveries. Link them to their originating events and state whether the model fits gross severity, ultimate net severity or period cash flows. Each is a legitimate variable when defined properly.
Suppose insurance reimburses a further 20,000 six months later. Ultimate net event cost falls, but the bank still had to operate during the six-month cash gap and meet any coverage conditions. Reporting that the incident cost only 15,000 without its gross size and recovery timing conceals the operational burden. Conversely, reporting the original 105,000 forever without acknowledging recovery overstates the ultimate economic cost.
The precise entries in a regulatory loss database must follow the applicable rules, not our teaching convention. Our general lesson is broader: retain enough information to reconcile the event, its financial recognition and its cash path separately. A loss measure becomes misleading when its name implies one of these views but its data quietly combine all three.
4. The year of occurrence is not always the year of recognition
A control failure begins in one year, is detected in the next and produces a settlement two years later. Which year contains the event? An occurrence-rate study may use the date the failure started. A detection study may use the discovery date. An accounting-loss report may use recognition dates. A cash forecast uses payment dates. None should be changed silently to make a trend look smoother.
Recent occurrence cohorts may be incomplete because some incidents have not yet been discovered. Recent detection cohorts may contain unresolved severities. Comparing their currently recorded costs with fully developed older cohorts can make the newest control environment look artificially safe. This is a development problem: the observation available today may be only part of the event’s ultimate consequence.
Keep several dates rather than forcing one timestamp to answer every question. A practical record contains occurrence or estimated start, discovery, first recognition, subsequent financial adjustments, recoveries and closure. When dates are uncertain, store the uncertainty or interval. A guessed exact date can create spurious seasonality or a misleading claim that an intervention preceded every affected event.
For a simple example, suppose two annual cohorts each ultimately contain ten incidents costing 10,000 each. The old cohort is fully recorded at 100,000. Only six events from the recent cohort have been found, giving 60,000. Calling this a 40% improvement assumes away delayed discovery. A model of reporting delay or a like-for-like maturity comparison is needed before the operational conclusion follows.
This is why our synthetic examples declare that counts and severities are complete unless stated otherwise. Real data rarely deserve that assumption automatically. The model should show what is observed at the analysis date and what additional development it expects, rather than treating incomplete recent losses as final outcomes.
5. Observed incidents are produced by an observation process
Let the true number of relevant incidents in a year have expected value λ. Suppose each is detected and reported with probability q, independently of the others and of its size for this first illustration. The expected observed count is qλ. A database of reports therefore estimates the product of incident generation and observation, not automatically the underlying incident rate.
Take λ = 100 and q = 0.4. Expected reports are 40. A new monitoring programme raises q to 0.8 while the true rate stays 100. Expected reports rise to 80. A dashboard that labels the change a doubling of underlying failures has confused better visibility with deteriorating operations. The increase is real in the observed count; its cause is the changed observation probability.
The reverse is equally important. Reports can fall because incidents become less common, because staff stop reporting, because thresholds rise or because a business closes. The count alone does not identify which explanation is correct. Maintain exposure volumes, reporting rules, detection coverage and business changes alongside the incident series. Where observation is uncertain, use a range rather than presenting a precise corrected rate as known.
Size-dependent reporting complicates the model further. Large losses are more likely to be discovered and to cross a collection threshold. The observed severity distribution then overrepresents large events relative to all incidents. Dividing total observed loss by observed count can be a valid average for reported losses, but it is not the average loss of every underlying event. The later truncation example derives that difference exactly.
Near misses are useful evidence about control operation and exposure even when they produce no booked loss. They should not automatically be assigned an imagined full loss and mixed with realised severities. Record the prevented event and its hypothetical consequence separately. The distinction allows a control team to learn from successful interception without converting every warning into a fabricated financial loss.
6. A shared reporting language supports, but does not replace, a loss model
The Financial Stability Board’s Format for Incident Reporting Exchange, published on 15 April 2025, provides common information items for operational and cyber incident reporting. Its purpose includes reducing fragmentation and improving communication across jurisdictions. A consistent reporting language helps connect incident facts, but it does not prescribe the probability distribution that fits a particular bank’s events.
Ryan draws that boundary explicitly. A report can describe affected services, time, status and consequences. The statistical model still needs an exposure period, inclusion rule and population. A report updated three times should not become three observations merely because the reporting format permits updates. Interoperability of fields and independence of observations are separate qualities.
For this guide, the core modelling record is an incident identifier linked to complete consequences, dates, exposure context, recovery and control information. A compact model can then answer a precise question: what is the distribution of annual net financial loss from this defined population under stated assumptions? The next part constructs the count distribution before we ask how large its individual losses are.
II. Frequency: how many incidents occur in the chosen exposure?
A count model begins with an integer-valued variable N. N might count incidents in a year, errors in a million transactions or failures during a defined operating period. The denominator is part of the variable. Before fitting a distribution, establish whether two observations represent comparable exposure. A busy month with twice the transaction volume should not automatically be treated as the same opportunity for failure as a quiet month.
7. A rate is a count divided by the exposure that generated it
A processing team records 24 incidents across six million transactions. Its observed rate is four incidents per million transactions. The following period has eight million transactions and 28 incidents, a rate of 3.5 per million. The incident count has risen by approximately 16.7%, but the rate per transaction has fallen by 12.5%. Both statements are arithmetically correct. They answer different operational questions.
The total number of cases matters for staffing and total expected loss. The exposure-adjusted rate matters for assessing failure propensity under a comparable transaction mix. Neither metric replaces the other. If volume grows faster than control improvements reduce the rate, the organisation may need more investigation capacity even though each transaction is less likely to encounter a problem.
Exposure is not always transaction count. A system outage may be more naturally related to operating time, release events or a particular dependency. Employee-related events might need headcount or staff-years. A legal event can depend on historical customer populations rather than this year’s revenue. Choosing an exposure variable is a causal modelling decision, not an automatic instruction to divide every loss count by sales.
Suppose a million simple domestic transfers and a million complex cross-border transactions have different failure opportunities. Combining them under one overall count can hide a changing mix. If the share of complex transactions rises, a constant process within each category can produce a higher aggregate rate. Stratify by relevant exposure types or model their effects explicitly before concluding that controls deteriorated.
A basic model writes expected count in period t as θ times exposure Et. Here θ is incidents per exposure unit. This implies proportional scaling: doubling exposure doubles the expected count when everything else stays fixed. Capacity constraints, learning and congestion can break that assumption. The proportional model is a benchmark to test, not a guarantee that a system can scale without changing its risk.
8. The Poisson model is a disciplined starting point
The NIST Poisson reference gives the count probability P(N = n) = exp(−λ)λⁿ/n! for nonnegative integers n. In this parameterisation, λ is the expected count over the specified interval. The mean and variance both equal λ. This equality is a property of the model, not a requirement that every operational dataset must satisfy.
N ~ Poisson(λ) P(N = n) = exp(−λ) λⁿ / n! E[N] = λ Var(N) = λ P(N ≥ 1) = 1 − exp(−λ)
Take λ = 4 incidents per year. The probability of no incident is exp(−4), approximately 1.832%. The probability of one or more is approximately 98.168%. That does not mean exactly four incidents will occur. Four is an average across hypothetical repeated years governed by this model; the realised count remains random.
A homogeneous Poisson process adds assumptions about independent increments and constant intensity through time. A Poisson distribution for one annual count alone does not establish those temporal assumptions. If incidents cluster after a release or during a peak processing window, a uniform arrival model may misrepresent staffing pressure even when its annual average is reasonable.
The distinction between average and occurrence probability matters for rare events. A rate of 0.02 per year gives annual probability of at least one event equal to 1 − exp(−0.02), about 1.980%, not exactly 2%. For small rates the difference is small; for larger rates it is substantial. A claim of a 10% annual event probability corresponds to a Poisson annual mean of −ln(0.9), about 0.10536.
Thus a workshop phrase such as once every ten years must be clarified. It might mean an average waiting time of ten years in a homogeneous Poisson process, implying rate 0.1. It might instead mean exactly a 10% probability of at least one event in a year. Those are not identical mathematical statements. Translate the expert’s meaning before using it as a parameter.
9. Estimating the rate requires the right exposure total
Suppose independent period counts follow Poisson distributions with means θEt, where exposures Et are known and θ is constant. Ignoring terms that do not depend on θ, the log likelihood is total count times ln θ minus θ times total exposure. Differentiating gives the maximum-likelihood estimate θ̂ = total count / total exposure. The formula follows from the observation model rather than from an arbitrary averaging rule.
Two periods illustrate why the denominator matters. The first has ten incidents in one million transactions; the second has ten in nine million. Their rates are ten and approximately 1.111 per million. The simple average of the two rates is about 5.556. The pooled estimate is 20 incidents in ten million transactions, or two per million. The simple average gives the small exposure period the same weight as the large one.
The pooled estimate is appropriate only for the common-rate assumption. If the first period used a different system or transaction mix, pooling may conceal a genuine change. The correct response is not to choose whichever average looks preferable. Model separate regimes or covariates and state why the exposures are or are not comparable.
For equal one-year exposures, the rate estimate reduces to the arithmetic mean of annual counts. Under the model, an approximate standard error is the square root of total count divided by total years. With very small counts, symmetric normal approximations can produce implausible negative lower limits or misleading precision. Exact Poisson intervals or an explicitly specified Bayesian analysis may be more informative.
A confidence interval is about uncertainty in a fixed unknown rate under a repeated-sampling procedure. It is not the same as a prediction interval for next year’s realised count. Even with a perfectly known λ, next year remains random. Predictive uncertainty includes incident randomness and, when the rate is estimated, uncertainty about that rate as well.
10. Zero observed incidents do not establish a zero underlying rate
A team observes no relevant incidents in three complete years. The Poisson maximum-likelihood estimate of the annual rate is zero. It would be a mistake to treat that point estimate as proof that the process cannot fail. Several positive rates can generate zero observations with nonzero probability, especially when the observation window is short.
A simple one-sided 95% upper confidence bound solves exp(−3λ) = 0.05. The result is λ = −ln(0.05)/3, approximately 0.99858 events per year. Under a rate near one per year, observing zero over three years is unusual but still occurs with probability 5%. The calculation demonstrates how little a short zero-event history can establish about rare failure risk.
The interpretation is procedural: a confidence-bound method has a specified repeated-sampling coverage under its assumptions. It is not a statement that the rate has a 95% probability of lying below the bound unless a corresponding Bayesian model has been specified. It also assumes complete observation and a constant Poisson rate. Unreported incidents or regime changes create additional uncertainty that the formula does not address.
Exposure units can make a zero history more informative. Zero errors in a very large number of genuinely comparable independent opportunities can constrain a per-opportunity probability. Zero major platform outages across three years of one platform is a different amount of evidence. Avoid translating a large transaction count into many independent trials for a common-cause failure that could affect all transactions simultaneously.
Clara uses the result to separate an evidence statement from a control claim. The evidence is that no qualifying incidents were observed during the stated exposure. The control claim is that the process now has a lower failure rate. Establishing the second may require tests of the controls, comparable observations, changes in exposure and a causal explanation. The absence of an incident is useful information, but it is not a certificate of impossibility.
11. Overdispersion can reveal a missing source of variation
The Poisson model fixes variance equal to mean. Suppose a sequence of comparable periods has a mean count near four but much greater variability. Possible explanations include changing risk intensity, common-cause clusters, unmodelled exposure variation or a few unusual periods. A negative-binomial count model is one way to accommodate greater variance, but choosing it should not end the investigation into why the variability exists.
A transparent construction lets the annual rate Λ itself vary across hypothetical years. Conditional on Λ, count N is Poisson with that rate. Suppose Λ has a gamma distribution with shape a and rate b. Then E[Λ] = a/b and Var(Λ) = a/b². Applying conditional expectation and conditional variance gives E[N] = a/b and Var(N) = a/b + a/b². The extra term is variation in the incident-generating environment.
Choose a = 2 and b = 0.5. The expected count is four, matching Poisson(4), but its variance is twelve rather than four. The probability of no incidents is (b/(b + 1))ᵃ = (1/3)² = 1/9, about 11.111%. It can have more zero years and more high-count years than the fixed-rate Poisson model while preserving the same mean.
That result is not contradictory. Mixing low-intensity and high-intensity years creates both quiet and busy periods. A single mean rate blurs the environmental variation. The mixture’s distribution is negative binomial under this parameterisation, but software packages use different definitions of their probability and shape parameters. Always verify the implied mean and variance before copying parameter values between implementations.
There is also an identification issue. A gamma mixture can represent genuine year-to-year rate variation or uncertainty about a fixed unknown rate. The predictive count distribution can look similar while the interpretation differs. More data can reduce parameter uncertainty; it does not necessarily eliminate actual environmental variability. Label which role the mixture plays before concluding that a better estimate will make the business inherently more stable.
12. Bayesian updating joins evidence with a declared prior
Suppose an unknown constant annual incident rate λ has a gamma prior with shape a = 2 and rate b = 1. Its prior mean is two events per year. Observe three incidents over two complete years under the Poisson model. The posterior distribution has shape a + 3 = 5 and rate b + 2 = 3. Its mean is 5/3, approximately 1.6667 events per year.
The posterior combines the prior rate information with the observed rate of 1.5. The prior’s parameters cannot be chosen merely to produce a preferred answer. They should represent documented prior knowledge or a deliberately stated weak-information assumption. A reader should be able to repeat the analysis with plausible alternative priors and see whether the substantive decision changes.
Next year’s predictive count integrates over the posterior rate rather than plugging in only its mean. The predictive probability of zero incidents is (3/4)⁵, approximately 23.730%. A plug-in Poisson model at rate 5/3 gives exp(−5/3), approximately 18.888%. The difference comes from rate uncertainty, which adds predictive dispersion. Reporting only the posterior mean rate does not capture the entire predictive distribution.
The predictive mean is 5/3. Its variance is 5/3 + 5/9 = 20/9, approximately 2.2222. The second term is posterior uncertainty about λ. A future observed count can differ substantially from the mean even if the posterior is well estimated. The model makes that distinction explicit instead of treating a single rate forecast as a deterministic workload plan.
Shevchenko and Wüthrich’s research on combining loss data and expert opinions provides a historical operational-risk application of Bayesian reasoning. Its regulatory discussion belongs to its period. Our conjugate example is a statistical teaching calculation, not an instruction that this posterior is a present-day regulatory capital method. The enduring lesson is to expose the prior, likelihood and resulting uncertainty.
13. Separate slow exposure changes from incident clusters
A holiday processing peak can increase incident opportunities through predictable volume. A faulty release can create a burst of related incidents through a common cause. A monitoring improvement can create a burst of discoveries from old failures. All three can produce a high count in one month, but they call for different models and responses.
A nonhomogeneous Poisson model lets the intensity λ(t) change through time while retaining independent increments conditional on that intensity. Its expected count over an interval is the integral of λ(t) across that interval. This can represent a documented daily volume pattern. It does not automatically represent self-excitation, where one event changes the probability of subsequent events, or a cluster containing multiple consequences of the same underlying incident.
A regression can model log expected count as a sum of an exposure offset and explanatory variables. For instance, log E[Nt] = log Et + β₀ + β₁zt, where z might mark a system regime. The exponential keeps the predicted mean positive. But the coefficient on a control variable is not automatically causal: riskier processes may receive stronger controls precisely because they were risky.
Do not treat a good statistical fit as a complete explanation of the incident mechanism. An overdispersion parameter can absorb several missing effects at once. A time indicator can track an intervention and a simultaneous change in customers. Useful model development links the count pattern to exposure, process evidence and observation changes, then tests whether the explanation predicts data not used for fitting.
Frequency also interacts with severity through regimes. A stressed processing period may create more events and make each event harder to contain. Fitting count and size independently without examining that common state can understate annual tail loss. We will calculate a regime example after deriving the aggregate model so that the difference can be seen numerically rather than hidden in the word correlation.
14. The count model must answer an operationally useful question
A manager choosing next year’s investigation budget needs an expected workload and a range of plausible counts. A team staffing tomorrow’s response desk needs within-day arrival patterns and service times. A capital analyst needs the count distribution joined to loss sizes. A control owner needs a comparison that adjusts for changing exposures and detection. One fitted annual count distribution cannot automatically serve all four purposes.
Ben prepares a one-page frequency specification: what counts as an incident, what exposure generated the observations, whether reporting is complete, which rate or regime is assumed, and what horizon the prediction covers. He then supplies the observed counts and the predicted mean, dispersion and zero-event probability. This lets another reader see whether a strange result is a data issue, an assumption or a numerical mistake.
Once those foundations are settled, the next question is not how to make the count model more elaborate. It is what happens financially when one of the counted events occurs. That is the severity variable, and it requires equal care about inclusion, units, thresholds and the distinction between typical and extreme outcomes.
III. Severity: what one incident can cost
Let X be the included financial loss from one event, measured under a stated gross or net convention. Unlike a count, X can take fractional currency values. A severity model describes the distribution of X conditional on a relevant event occurring. It does not, by itself, say how many such events occur in a year. That distinction will prevent us from mistaking a large individual-event quantile for an annual aggregate-loss quantile.
15. A typical loss and an average loss need not be close
Consider five hypothetical event losses, in thousands: 2, 3, 5, 10 and 180. Their mean is 40 and their median is 5. Four of the five events cost less than the mean. Describing 40 as the typical incident would give the wrong impression of the ordinary case; ignoring the 180 because it is unusual would give the wrong impression of total financial exposure.
The mean matters for expected aggregate loss because it weights every possible loss by its probability. The median divides the distribution into two halves. A high quantile describes a point in the upper tail. A tail average describes the severity beyond a probability threshold. These are different summaries of the same variable, and their disagreement is information about the distribution rather than a reason to choose whichever number feels more representative.
Outliers deserve investigation, not automatic deletion. The 180 could be a genuine severe event, a duplicated charge, a currency conversion error or an incident belonging to a different population. Each explanation calls for a different treatment. Removing a valid large loss solely because it drives the model upward discards precisely the kind of outcome a tail-risk analysis is intended to consider.
Conversely, keeping an erroneous 180 after discovering that it should be 18 does not make the model conservative in a meaningful way. It makes the data incorrect. Data correction and risk conservatism are separate decisions. Preserve the audit trail, correct demonstrable errors and then apply explicit sensitivity or stress assumptions to genuine uncertainty.
The sample also illustrates instability. Removing the largest observation changes the mean from 40 to 5. A small dataset can therefore support a reasonably clear description of common small events while providing very weak evidence about extreme losses. The model should communicate that difference instead of using the same number of decimal places for every statistic.
16. The lognormal model separates multiplicative scale from dispersion
The NIST lognormal reference defines a positive variable whose natural logarithm is normally distributed. Write ln X = μ + σZ, where Z is standard normal and σ is positive. This gives X = exp(μ + σZ). The median is exp μ, while the mean is exp(μ + σ²/2). The extra variance term makes the mean exceed the median when dispersion is positive.
ln X ~ Normal(μ, σ²) Median(X) = exp(μ) E[X] = exp(μ + σ²/2) Var(X) = [exp(σ²) − 1] exp(2μ + σ²) Quantile_q(X) = exp[μ + σ Φ⁻¹(q)]
Measure X in thousands of Singapore dollars and choose median 20 with σ = 1. Then μ = ln 20. The mean is approximately 32.974, the standard deviation 43.224 and the 99% individual-event quantile 204.809. These are not forecasts from a bank’s data. They show how a distribution with a moderate median can still permit much larger losses.
Units matter inside the logarithm. Changing from thousands of dollars to dollars multiplies X by 1,000 and adds ln 1,000 to μ. It does not multiply σ by 1,000. A model that copies μ unchanged after changing the currency scale will be wrong by a factor of 1,000 in its severity outputs while still producing perfectly smooth probability curves.
Given a desired arithmetic mean m and standard deviation s, one can solve σ² = ln[1 + (s/m)²] and μ = ln m − σ²/2. For m = 50 and s = 150 in thousands, σ is approximately 1.51743 and μ approximately 2.76073. The resulting median is about 15.811 and the 99% quantile about 539.581. Matching a mean and variance does not mean the median should also equal the mean.
The lognormal is one candidate, not a universal law of operational loss. Its positive support suits a gross-loss variable, but that alone does not establish the correct tail. A process with a hard contractual cap or several distinct loss mechanisms may need another model. Fit, plausibility and sensitivity should all be examined before using any named family as the basis of a high quantile.
17. Matching averages does not identify the shape of the tail
Two severity distributions can share the same mean and variance while placing different probabilities on very large losses. Moment matching compresses a distribution into a few numbers. It cannot generally recover all of its quantiles. A model chosen only because its first two moments match the sample has not yet demonstrated that its extreme tail is appropriate.
A gamma distribution with shape k and scale θ has mean kθ and variance kθ². An exponential distribution is the special case k = 1. These can be useful for positive losses with particular dispersion patterns. Their upper tails differ from a lognormal or a Pareto-type model, even when parameters are selected to give a similar average. A tail comparison should therefore look at the actual probability assigned to decision-relevant thresholds.
Suppose one proposed model gives a 1% chance of a loss above 100 and another gives 0.1%. If the business question concerns whether a 100 financial buffer is adequate for an event, that tenfold difference is material. A slightly better fit around the many small observations may not settle the disagreement about the sparse upper tail. The chosen objective should guide the diagnostic.
A useful procedure compares several defensible families under the same dataset, reporting rule and estimation method. Record their implied median, mean, upper quantiles and probability of falling below the reporting threshold. Reject candidates that imply impossible behaviour or a implausible missing-event population. Do not select the smallest tail estimate merely because it improves a capital or budget number.
Hadley, Joe and Nolde’s severity-selection study examines instability arising from distribution selection and the treatment of reporting thresholds. Its historical regulatory discussion is not a substitute for today’s rules. The useful research lesson here is narrower: distribution choice, truncation and out-of-sample evaluation can materially affect an annual loss estimate, even when a fitted curve looks satisfactory.
18. A collection threshold changes what the sample represents
Suppose the underlying event loss X is exponential with mean 10, measured in thousands. The database records only losses above 5. The probability an event is recorded is P(X > 5) = exp(−0.5), approximately 0.60653. The observed sample contains about 60.653% of the underlying events under this model; the rest are missing because of the inclusion threshold, not because they cost zero.
For an exponential distribution, the average recorded loss is 5 + 10 = 15. That is larger than the unconditional mean 10. If the underlying incident rate is 100 per year, the expected observed count is approximately 60.653. Multiplying that observed count by the observed mean gives expected recorded loss of approximately 909.796. Expected total loss from all events is 1,000. The omitted small losses account for the difference.
Two naïve corrections go wrong in different directions. Using the full count of 100 with the observed mean of 15 gives 1,500, overstating the true model mean. Using the observed count with an unconditional severity fitted without acknowledging the threshold can distort both shape and frequency. The correct treatment keeps the population, observed count and conditional severity consistent.
If X has density f and cumulative distribution F, the density of a recorded loss above threshold u is f(x)/[1 − F(u)] for x > u. The denominator normalises the conditional distribution. Omitting it from the likelihood treats the observed values as though small losses could have appeared in the dataset but happened not to. That is not the observation process we specified.
A threshold can change over time. Then each observation or period needs the relevant inclusion rule, and count exposure must be adjusted consistently. Simply combining ten years of events collected at different minimum amounts can produce an artificial trend in both frequency and mean severity. Correcting for thresholds is not an optional technical refinement when the threshold defines which data exist.
19. Censoring preserves information that truncation removes
Truncation means events outside the observation rule do not appear in the sample. Censoring means we know an event exists but know only a bound or interval for its size. If a database records that ten losses were below 5 without retaining their exact values, those ten events provide information about F(5). If they vanish entirely, their count must be inferred or obtained elsewhere.
For independent exact observations above a known lower threshold and a known count m below it, a censored-data likelihood can include a factor F(u)ᵐ for the subthreshold events, multiplied by densities for the exact events. The exact expression depends on the sampling design. The point is that a known number of small events changes what can be estimated, even without their precise amounts.
Insurance limits produce an upper-bound version. A payment recorded as 100 because the policy limit is 100 does not establish that the underlying loss was exactly 100. The loss may have been much larger. Fitting an uncensored loss distribution to capped payments can make the tail appear artificially thin. The model must distinguish underlying loss, insurer payment and retained loss.
An unresolved incident may have a lower bound from charges already paid and an uncertain remaining amount. Treating the current paid total as final severity introduces another form of incompleteness. A development or interval model may be appropriate, or the analysis may restrict itself to comparably mature events. The chosen treatment should be stated rather than quietly mixing final and provisional amounts.
Research on operational-risk truncation and parameter uncertainty shows why assumptions about omitted observations can affect extreme annual-loss estimates. Our exponential example makes the mechanism visible without relying on an estimated bank dataset. The lesson is to model how observations were selected before interpreting the fitted distribution as the distribution of all events.
20. Heavy tails require checking whether the moments exist
A Pareto-type teaching model can use survival probability P(X > x) = (xm/x)ᵃ for x at least xm. The parameter a controls how quickly the tail decays. Its mean is finite only when a > 1, and its variance is finite only when a > 2. These conditions follow by integrating the tail or density; a finite sample mean does not override them.
Choose xm = 10 and a = 1.5. The theoretical mean is 30, but the theoretical variance is infinite. A simulation with a finite number of draws will still produce a finite sample variance. Printing that number with a narrow-looking standard error does not make it a stable estimate of a finite population variance. Rare larger draws can continue to change the estimate dramatically.
This matters when later formulas require E[X²]. The compound-loss variance formula is valid as a finite quantity only if the required moments exist. A model with an infinite severity variance may still have finite loss quantiles below probability one, but normal approximations based on a finite variance are not justified by merely forcing a sample variance into the equation.
Real organisations have finite resources and many event types have practical or legal limits. An unbounded distribution can still be a useful approximation over a range, but the assumed far tail needs scrutiny. A cap introduced because there is a genuinely defined exposure limit differs from an arbitrary numerical cap added to make the model converge. Explain the boundary and test the sensitivity of important outputs to it.
Tail terminology also needs precision. A lognormal has all positive polynomial moments finite, even though its upper tail can be much heavier than an exponential tail. Not every right-skewed distribution shares the same asymptotic properties. The mathematical condition relevant to the chosen calculation is more useful than simply labelling every large-loss model fat-tailed.
21. A tail model is conditional on exceeding its threshold
The NIST generalized-Pareto reference describes a model for excess Y = X − u conditional on X exceeding a sufficiently high threshold u. In a common parameterisation, its survival function is [1 + ξy/β]−1/ξ, with β positive and the expression inside the brackets positive. The ξ = 0 limit is exponential. The tail model does not automatically describe the many events below u.
Take u = 100, β = 50 and ξ = 0.25, in thousands. The 95% conditional quantile of total loss given that the loss already exceeds 100 is approximately 322.949. The 99% conditional quantile is approximately 532.456. These are quantiles within the exceedance population. Calling the latter the 99% quantile of all incidents would ignore the probability of ever entering that population.
Suppose only 2% of all events exceed 100. To find the 99.9% quantile of all event losses using this tail approximation, divide the unconditional upper-tail probability 0.1% by the exceedance probability 2%. The required conditional upper-tail probability is 5%, so the relevant conditional quantile is 95%, giving about 322.949 rather than 532.456.
The 99.9% quantile of an individual event is still not the 99.9% quantile of annual aggregate loss. Annual loss depends on the number of incidents and their joint occurrence. This sequence of conditioning levels is important: all events, threshold exceedances, individual severity and annual totals are different populations. A probability label should always name which population it describes.
A higher threshold may improve the plausibility of a tail approximation while leaving fewer observations for estimation. A lower threshold supplies more observations but may fit the body poorly with a tail family. Examine parameter stability and predictive performance across reasonable thresholds. Selecting the threshold solely because it gives a favourable risk number turns an estimation decision into hidden outcome manipulation.
22. Put historical severities on a comparable basis without inventing precision
An old loss recorded as 100,000 and a recent loss recorded as 100,000 need not represent the same purchasing power or business exposure. Currency conversion, price changes, transaction volumes and the nature of the event can all matter. Preserve the original amount and currency, then record each adjustment separately so a reader can reconstruct the analytical value.
For a purely illustrative index calculation, if the chosen cost index rises from 100 to 125, an old 100,000 cost becomes 125,000 in the new index units. This is not a statement about actual inflation. It is the arithmetic of a declared adjustment. The appropriate index for specialist legal or technology costs may differ from a broad consumer-price index, so the choice requires justification.
Scaling an external loss by the ratio of two banks’ revenues is another assumption, not a neutral conversion. Some losses may grow roughly with customer count; others have fixed investigation costs, capped exposures or network-wide common causes. A smaller organisation can suffer a disproportionately large loss if its controls or insurance differ. External events are valuable for scenario design without being mechanically interchangeable statistical observations.
Changing legal environments, technology and business models can also make an event more or less relevant. Retain a documented reason for including, adjusting or excluding an observation. The goal is neither to keep every historical number untouched nor to reshape history until it matches current preferences. It is to make comparability assumptions explicit and test whether they materially change the decision.
We now have two objects: the number of events and the size of each event. The next part combines them. Keeping the objects separate until this point lets us see exactly which uncertainty creates a large annual result and which intervention could reduce it.
IV. Aggregate loss: join the count and the consequence
Let annual loss S equal the sum of the losses from all incidents in that year. If no incident occurs, S is zero. If three occur, S contains three severities. The number of terms is random, which is the defining feature of a compound-loss model. A model of one event’s size and a model of annual total loss are related, but they are not the same distribution.
23. Derive the mean and variance instead of memorising them
Begin with independent identically distributed positive severities X₁, X₂ and so on, each with mean m and variance v. Assume the count N is independent of the entire severity sequence and has finite mean and variance. Conditional on N = n, the total loss has mean nm and variance nv. We can now average over the count distribution.
Conditional expectation gives E[S] = E[N]m. Conditional variance separates two sources of uncertainty: E[Var(S|N)] + Var(E[S|N]). The first is E[N]v, representing variation in individual sizes. The second is Var(N)m², representing variation in how many average-sized incidents occur. Both terms are needed; treating the event count as fixed drops the second.
S = X₁ + … + X_N, with S = 0 when N = 0 E[S] = E[N] E[X] Var(S) = E[N] Var(X) + Var(N) [E[X]]² For N ~ Poisson(λ): E[S] = λ E[X] Var(S) = λ E[X²]
The Poisson expression follows because Var(N) = E[N] = λ and E[X²] = Var(X) + [E[X]]². It is therefore not λ times the severity variance alone. Even when every incident costs exactly the same amount, the annual total is random because the number of incidents varies.
The assumptions are part of the formula. If a busy year also changes the severity distribution, conditional means and variances may depend on the environment. If several consequences share a common incident, the severities may not be independent. If a required moment is infinite, the corresponding variance is not a finite planning number. Use conditional calculations suited to the actual model rather than applying the compact expression outside its domain.
Units provide a useful check. Expected count is dimensionless for a fixed horizon, m has currency units, and v has squared-currency units. E[S] is in currency and Var(S) in currency squared. Standard deviation returns to currency only after taking the square root. Mixing dollars and thousands in different terms can produce a plausible-looking number that is wrong by orders of magnitude.
24. The same annual mean can conceal very different volatility
Suppose the annual count is Poisson with λ = 4. Let each severity have mean 50 and standard deviation 150, measured in thousands. Expected annual loss is 4 × 50 = 200. The second severity moment is 150² + 50² = 25,000. Annual variance is 4 × 25,000 = 100,000, and annual standard deviation is approximately 316.228.
The standard deviation exceeds the mean. That is possible for a nonnegative right-skewed distribution; it does not imply negative annual losses are allowed. A normal distribution chosen solely to match these two moments would assign substantial probability below zero. That is a warning that a symmetric approximation can be poorly suited to this compound-loss shape.
Now keep the same Poisson count but make every severity exactly 50. The annual mean remains 200, while variance becomes 4 × 50² = 10,000 and standard deviation 100. Uncertain event size contributed the difference. A budget based only on expected loss would treat the two systems identically, even though their annual financial variability differs substantially.
Alternatively keep the variable severity and replace the fixed-rate Poisson count with the gamma-mixture count having mean four and variance twelve. Annual variance is 4 × 22,500 + 12 × 2,500 = 120,000. The mean stays 200. The extra count variability increases the variance without changing average incident size or average count.
None of these moments gives an exact 99% quantile. Distributions with the same moments can have different tails. The calculations are valuable as checks, diagnostic comparisons and inputs to approximations whose limitations are acknowledged. To obtain a particular quantile, calculate or approximate the distribution itself and validate that step separately.
25. A complete finite model reveals every annual outcome
Use a small count distribution: no incident with probability 0.5, one with probability 0.3 and two with probability 0.2. Each incident independently costs 10 with probability 0.8 or 50 with probability 0.2, in thousands. Counts and severities are independent. Because there are at most two incidents and two severity values, we can enumerate the full annual distribution without simulation.
No incident gives annual loss zero with probability 0.5. One incident gives loss 10 with probability 0.3 × 0.8 = 0.24, or loss 50 with probability 0.3 × 0.2 = 0.06. With two incidents, two small losses give 20 with probability 0.2 × 0.8² = 0.128. One small and one large gives 60 with probability 0.2 × 2 × 0.8 × 0.2 = 0.064. Two large losses give 100 with probability 0.2 × 0.2² = 0.008.
| Annual loss S | Probability | Cumulative probability |
|---|---|---|
| 0 | 0.500 | 0.500 |
| 10 | 0.240 | 0.740 |
| 20 | 0.128 | 0.868 |
| 50 | 0.060 | 0.928 |
| 60 | 0.064 | 0.992 |
| 100 | 0.008 | 1.000 |
The probabilities sum to one. Expected annual loss is 10(0.24) + 20(0.128) + 50(0.06) + 60(0.064) + 100(0.008) = 12.6. The second annual moment is 535.6, so variance is 535.6 − 12.6² = 376.84. These quantities can be checked directly from the table.
Now check through the compound formulas. E[N] = 0.7 and Var(N) = 0.61. E[X] = 18 and Var(X) = 256. The mean is 0.7 × 18 = 12.6. Variance is 0.7 × 256 + 0.61 × 18² = 376.84. Agreement between enumeration and conditional formulas verifies both the probabilities and the moment calculation.
The finite model is intentionally simple, not a claim that real operational losses stop after two incidents. Its value is that every result is auditable. Before trusting a simulation or recursive calculation on a large model, test that implementation on a small distribution whose complete answer is already known.
26. A quantile is a boundary; Expected Shortfall describes the selected tail
Define the q quantile as the smallest loss whose cumulative probability is at least q. In the finite table, the 90% quantile is 50 because cumulative probability rises from 0.868 at 20 to 0.928 at 50. The 95% and 99% quantiles are both 60 because cumulative probability at 60 is 0.992. The same VaR value can therefore serve two different confidence levels in a discrete distribution.
Expected Shortfall at 99% averages the worst 1% of outcomes, including the appropriate fraction of probability at the boundary. The worst outcomes contain all 0.8% probability at loss 100 and another 0.2% from the mass at loss 60. Their weighted sum divided by 1% is [100(0.008) + 60(0.002)]/0.01 = 92.
At 95%, the worst 5% consists of 0.8% at 100 and 4.2% at 60, giving Expected Shortfall 66.4. At 90%, it includes all mass at 100 and 60 plus 2.8% at 50, giving 60.4. These numbers are larger than or equal to their corresponding quantiles because they average losses in the selected upper tail.
Simply calculating E[S|S ≥ VaR] can be wrong for the intended tail level when there is a probability atom at VaR. That conditional event may contain much more than the worst 1% or 5%. The quantile-integral definition, or an equivalent fractional-boundary calculation, preserves the desired tail probability. Software conventions should be tested on a discrete example before their results are interpreted.
VaR is not the maximum possible loss. In the table, 99% VaR is 60 even though 100 occurs with probability 0.8%. Expected Shortfall is not an absolute maximum either; it is an average of a specified tail. Both require the horizon, confidence level and loss convention. The risk-measure mathematics guide develops their wider uses and limitations.
27. Rare disasters can fall beyond the chosen quantile boundary
Consider an annual model with zero loss at probability 99.95% and a loss of one million at probability 0.05%. Its 99.9% quantile is zero under the stated smallest-cumulative-boundary definition. Its expected annual loss is 500. Expected Shortfall at 99.9% is 500,000 because the worst 0.1% contains the 0.05% million-loss mass and 0.05% at zero.
The zero quantile does not prove that no buffer is needed or that the severe event is irrelevant. It says the chosen probability boundary sits below that rare outcome. A decision can supplement quantiles with severe scenarios, exceedance probabilities at particular loss amounts and service-continuity constraints. These additions address different questions rather than correcting a mathematical error in the quantile.
Suppose a control halves the rare event’s loss but does not change its probability. The 99.9% quantile remains zero, while expected loss and Expected Shortfall halve. A performance measure based only on that VaR would miss the improvement. Conversely, a tiny change in event probability near the quantile cutoff can make the reported VaR jump dramatically without an equally dramatic change in the underlying process.
This is especially relevant to low-frequency operational events because the choice of horizon and confidence level can determine whether a particular scenario enters the summary. Do not choose those settings after seeing which result is most attractive. Define the decision and measure first, then report complementary outputs that reveal what the main metric leaves out.
28. Transforms and recursion provide checks, not magic
For a nonnegative severity X, define its Laplace transform LX(t) = E[exp(−tX)] for t at least zero. Conditional on N = n independent severities, the transform of their sum is [LX(t)]ⁿ. For a Poisson count, averaging over n gives LS(t) = exp{λ[LX(t) − 1]}. This identity follows by summing the exponential series.
The negative sign is useful. For nonnegative X and t ≥ 0, exp(−tX) lies between zero and one, so the transform exists. A positive-argument moment-generating function need not exist for every heavy-tailed severity, including a lognormal. Writing a transform formula without checking its existence can turn a formal-looking manipulation into an invalid calculation.
For severity values placed on a nonnegative integer grid, let fk be the probability of severity k grid units and gn the aggregate probability. The probability-generating function relation is G(z) = exp{λ[F(z) − 1]}. Differentiating and comparing coefficients yields g₀ = exp[λ(f₀ − 1)] and gn = (λ/n)Σk=1…n kfkgn−k.
The g₀ expression matters when the discretised severity has mass at zero. Rounding small positive losses down can create that mass. Blindly setting g₀ to exp(−λ) then ignores zero-sized model events. The recursion is exact for the stated discrete compound-Poisson distribution, not necessarily for the continuous model that was approximated by the grid.
Use the existing Panjer-recursion article for implementation detail. Here the essential audit is to compare total probability, mean and variance with analytical values, vary the grid, and track probability outside the computed range. A beautiful distribution plot is not evidence that omitted tail mass is negligible.
29. Simulation noise is different from uncertainty about the model
A Monte Carlo calculation draws a count, draws that many severities and sums them. Repeating the experiment gives a sample from the specified annual-loss model. It does not create new evidence about the true incident rate or tail family. Increasing the number of simulations improves numerical precision conditional on the assumptions; it cannot repair an incorrect observation rule or missing common cause.
At a 99.9% threshold, 100,000 independent simulated years contain an expected 100 observations above the population quantile. The count of such exceedances has a standard deviation close to ten under a fixed 0.1% exceedance probability. A high-quantile estimate can therefore remain noisy even when the simulation count sounds very large. Heavy-tailed severities add variation in how large the few selected observations are.
Use repeated seeds or numerical error assessments to study simulation precision. Separately vary rate parameters, severity families, reporting assumptions and dependence structures. A narrow simulation interval under one fixed model should not be presented as the full uncertainty interval for the business. These experiments answer different questions and should appear in separate columns.
Parameter uncertainty can be propagated by drawing a parameter set from a justified sampling or posterior distribution, then drawing incidents conditional on it. The outer parameter draw should preserve dependencies among fitted parameters. Randomising μ and σ independently when their estimates are related can create unrealistic combinations. Document how the uncertainty distribution was obtained rather than calling any arbitrary parameter jitter a confidence analysis.
Reproducibility requires the model specification, data version, numerical settings and random seed. Reproducibility is not truth: another analyst can reproduce a wrong model exactly. It is nevertheless an essential condition for diagnosing where two results differ and for distinguishing a substantive assumption change from accidental simulation variation.
30. Common incidents change aggregation across business lines
A common platform incident affects two business lines. Each occurrence costs line A 100 and line B 200, in thousands. The common annual incident count Z is Poisson with rate 0.2. Then line losses are 100Z and 200Z, and total loss is 300Z. Expected total loss is 60 and its variance is 300² × 0.2 = 18,000.
If an analyst instead gives each line an independent Poisson(0.2) count, the marginal count distributions and expected line losses remain unchanged. Total expected loss is still 60. But variance becomes 100² × 0.2 + 200² × 0.2 = 10,000. The independent model misses covariance of 4,000 between the line losses; twice that covariance explains the missing 8,000 in total variance.
This is a direct example of why independent business-line models can understate a shared operational event. The repair is to represent the common incident and its consequences across lines, not merely to increase each line’s mean. Means do not encode whether losses occur together. The incident identifier and root-cause hierarchy introduced earlier provide useful evidence for that dependence.
A different issue arises if both lines record the full 300 as their own financial loss. Adding those records would double the economic cost. Dependence modelling should preserve correlated distinct consequences, while consolidation should remove duplicate descriptions of the same consequence. These are not the same correction, and applying one does not automatically solve the other.
For a larger portfolio, aggregation requires the joint distribution or a structural event model. A correlation coefficient can help describe second moments when they exist, but it does not uniquely determine joint tail behaviour. A common-cause scenario and a statistical dependence model can complement one another if their boundaries are clear and the same event is not included twice.
31. A stressed year can raise both frequency and severity
Construct two annual regimes. With probability 0.9 the year is calm: incident count is Poisson with mean two and every incident costs 10. With probability 0.1 it is stressed: count is Poisson with mean eight and every incident costs 100. Amounts are thousands. The entire year shares its regime, so incidents within a stressed year have both higher frequency and greater size.
Expected annual loss is 0.9 × 2 × 10 + 0.1 × 8 × 100 = 98. Average annual count is 2.6. Taking the unweighted-by-events regime average severity, 0.9 × 10 + 0.1 × 100 = 19, and multiplying by 2.6 would give only 49.4. That severity average weights years, not observed events. Stressed years generate more events and should contribute more to the event-population average.
The event-weighted mean severity is 98/2.6, approximately 37.6923. Multiplying that by average count recovers the correct mean, but it still does not recover the annual distribution. If the event sizes are then drawn independently from the pooled event mix, the common annual state is lost. The model would scatter severe events across years differently from the specified regime system.
Conditional annual variance is 200 in calm years and 80,000 in stressed years. Its average is 8,180. The conditional annual means are 20 and 800, whose between-regime variance is 0.9 × 0.1 × (800 − 20)² = 54,756. Total annual variance is therefore 62,936. The between-regime term is much larger than the average within-regime variance.
This example explains a common modelling failure without relying on vague warnings about correlation. Pooling can preserve a correct overall mean while erasing a state that concentrates losses. The choice between an independent compound model and a regime model should be informed by process evidence and data, with sensitivity reported where the evidence is weak.
The next part asks what a control actually changes. Once frequency, severity and dependence are visible, the effect of prevention, containment, insurance and restoration can be calculated without pretending that they are all the same kind of risk reduction.
V. Controls: identify the part of the loss process that changes
A control can prevent an incident, reduce the consequences after it starts, improve observation or speed recovery. Those are different interventions. A model that represents every control as a percentage reduction in one final capital number loses the causal distinction and makes verification difficult. Begin with the operational mechanism, then identify the count, severity, dependence or duration parameter it could reasonably affect.
32. Prevention and containment do not add their percentage benefits
Use the annual Poisson count with rate four and severity mean 50, standard deviation 150, in thousands. Baseline expected annual loss is 200. Suppose a preventive control independently removes 20% of potential incidents without favouring small or large events. The remaining count has rate 3.2 under the thinning assumption. Suppose a separate containment improvement multiplies every remaining loss by 0.75.
New expected annual loss is 3.2 × 37.5 = 120. The combined reduction is 40%, not 45%. The prevention factor 0.8 and severity factor 0.75 multiply because containment applies only to incidents that remain. Adding the two advertised percentage reductions would count some of the same baseline loss twice.
Variance also changes differently. Scaling severity by 0.75 scales its second moment by 0.75². New compound-Poisson variance is 3.2 × 0.75² × 25,000 = 45,000, with standard deviation approximately 212.132. Expected loss falls by 40%, while standard deviation falls by about 32.9%. A single percentage cannot summarise all aspects of the improvement.
The thinning assumption is important. A control that mostly prevents small incidents can reduce count substantially while leaving much of the severe tail. A control that specifically prevents a rare catastrophic mechanism might barely change total count but greatly reduce the upper tail. Evidence about which incidents are affected is therefore more useful than a headline percentage effectiveness with no event population.
If a containment measure only caps losses above a threshold, it does not multiply every severity by the same factor. The retained distribution changes shape. Calculate the transformation event by event or from its distribution, then aggregate. The model should reflect the actual control rule rather than force every intervention into a proportional reduction because that is easy to calculate.
33. Two controls are not independent merely because they have different names
Suppose an incident produces loss only if two barriers fail. In an ordinary operating state, barrier A fails with probability 0.1 and barrier B with probability 0.05, independently conditional on that state. Their joint failure probability in that state is 0.005. It would be tempting to present this multiplication as the overall protection of the system.
Now add a common configuration failure that occurs with probability 0.01 and defeats both barriers. In the remaining 99% of cases, the ordinary-state probabilities apply. Overall joint failure probability becomes 0.01 + 0.99 × 0.005 = 0.01495. This is almost three times the ordinary-state product. The small common-cause component dominates much of the residual risk.
The individual unconditional barrier-failure probabilities have also changed: they are 0.109 and 0.0595. Multiplying those unconditional probabilities still does not recover the joint probability because their failures share the common state. The relevant calculation is the conditional event tree, not an assumption that different control labels imply independent outcomes.
Common causes can include shared data, shared permissions, the same validation logic or the same operational dependency. The example does not identify a vulnerability in any actual system. It explains why design reviews should ask whether a backup or second check depends on the condition it is supposed to protect against. Independence requires evidence about mechanisms, not simply separation on a diagram.
A model can test the benefit of reducing the common-cause probability separately from improving an ordinary-state barrier. That comparison may reveal a better use of resources than repeatedly strengthening a component whose failures are already rare outside the common state. The numerical answer remains conditional on the stated event probabilities and should be challenged with alternative plausible values.
34. Detection quality must be read with the base rate
A hypothetical monitoring system reviews one million transactions, of which 100 genuinely require intervention. It detects 90% of those cases, giving 90 true alerts. Its false-alert rate on the 999,900 ordinary transactions is 0.1%, producing an expected 999.9 false alerts. The system’s alert queue therefore contains approximately 1,089.9 cases, of which only about 8.258% are genuinely problematic.
This does not contradict the 90% detection rate. Detection sensitivity conditions on a true problem. The fraction of alerts that are true conditions on an alert. Bayes’ rule connects them through the base rate of true problems and the false-alert rate. In a low-prevalence setting, even a small false-alert probability can create many more alerts than true cases.
Reducing the false-alert rate to 0.01% while preserving sensitivity gives an expected 99.99 false alerts and 90 true alerts. The true fraction rises to approximately 47.371%. The absolute number of detected true cases is unchanged, but investigation workload falls substantially. This can be valuable because overloaded reviewers may take longer to handle the cases that matter.
Optimising only detection rate can therefore be incomplete. A lower threshold may detect more true cases but send too many ordinary cases into a constrained review process. Optimising only precision can miss serious events. A decision should account for missed-event severity, review cost, delay, customer friction and the capacity to investigate the resulting queue.
An alert is not automatically an avoided loss. It must be reviewed and acted on before the relevant damage becomes irreversible. A model should distinguish true-event detection, timely intervention and financial containment. Counting the full potential loss of every alert as a benefit would overstate effectiveness and ignore the events that were detected only after their losses had already occurred.
35. A review queue can turn a detection improvement into a delay problem
Use a deliberately simple queue with Poisson arrivals at rate λ and exponentially distributed service times with rate μ, one server, first-in-first-out service and an unlimited waiting room. Under these assumptions and λ < μ, the stationary probability of n cases in the system is (1 − ρ)ρⁿ, where ρ = λ/μ. Summing the geometric distribution gives an average of ρ/(1 − ρ) cases.
Dividing the average number in the system by the arrival rate gives average total time 1/(μ − λ). If nine cases arrive per operating hour and ten can be served per hour on average, the mean total time is one operating hour. The mean service time is six minutes, so most of that hour is waiting. High utilisation can create substantial delay even before average arrivals exceed average capacity.
Increase arrivals from nine to 9.5 per hour while service remains ten. Mean total time doubles to two hours. The input grows by only about 5.6%, but the gap between capacity and demand halves. This nonlinear congestion effect is one reason a small increase in alerts can have a large operational consequence near capacity.
The formula is not a staffing recommendation for a real bank. Real queues can have multiple reviewers, priorities, batch work, non-exponential service times and time-varying arrivals. The teaching model isolates a mechanism: variability and limited spare capacity produce waiting. A more realistic model should preserve the actual prioritisation and service-time evidence rather than assuming average throughput alone guarantees timely review.
Link delay to severity only through a stated mechanism. For example, an unresolved processing error might continue to affect transactions until correction. Shorter time to intervention can reduce the number affected. The count of root incidents need not change. The control benefit appears in incident severity and service delay, not necessarily in a lower initiating-event frequency.
36. Availability is an average, not a maximum outage guarantee
In an alternating renewal model with finite mean uptime U and mean downtime D, long-run availability is E[U]/(E[U] + E[D]) under the usual regenerative assumptions. Suppose mean uptime is 1,000 hours and mean downtime four hours. Availability is 1,000/1,004, approximately 99.6016%. This summarises the long-run fraction of time operating; it does not bound the length of any one outage.
Two systems can have the same mean downtime but different downtime distributions. One may usually recover in four hours. Another may recover almost immediately in most cases but occasionally take several days. Their average availability can match while their probability of breaching a critical recovery deadline differs greatly. A resilience assessment needs the distribution of disruption, not only a ratio of means.
Reducing mean repair time can improve availability without reducing the rate at which incidents begin. Conversely, preventing frequent short outages can raise availability while leaving a rare long outage unchanged. These effects mirror frequency and severity distinctions in financial loss. The appropriate performance measure follows the service requirement.
The BIS explanation of operational resilience focuses on maintaining critical operations through disruption. That objective motivates measures such as the time a service is unavailable, the transactions affected and the time needed to restore normal performance. None is automatically equivalent to the annual financial loss used in our compound distribution.
When reporting availability, state whether scheduled maintenance, partial degradation, unsuccessful transactions and downstream delays are included. A service can be technically reachable while failing to complete a critical customer task. The measurement boundary should follow the operation the organisation promises to deliver, not merely whether one server responds to a test.
37. Restoring the system does not instantly clear the backlog
A service normally receives 1,200 tasks per hour. It stops processing for two hours, accumulating 2,400 tasks in a simple fluid approximation with constant arrivals and no abandonment. When it returns, it can process 1,500 tasks per hour while new tasks keep arriving at 1,200. The net backlog-clearing rate is only 300 per hour, so clearing the accumulated work takes another eight hours.
The platform is restored after two hours, but the backlog is not fully cleared until ten hours after the disruption began. A recovery measure that stops at technical restart misses much of the customer-facing effect. The distinction matters particularly for payment, reconciliation and reporting tasks with deadlines that continue to run while the system is unavailable.
If recovery capacity equals the incoming rate, the backlog does not shrink in this simplified model. If it is lower, the backlog grows even though the system is technically online. Extra capacity, rerouting, prioritisation or reduced incoming demand may be needed. A plan that assumes automatic clearance after restart contains an unstated capacity improvement.
Not every delayed task has the same consequence. Critical payments may be prioritised while ordinary records wait. That changes which users experience delay, not merely the total count. A useful model represents priority classes and measures deadline breaches for each. Summing all tasks into one average waiting time can conceal the service that matters most.
The financial severity can be built from separately justified components: direct remediation expense, compensation, failed-settlement charges or other included consequences. Avoid multiplying every delayed transaction by its face amount and calling the result a loss. A delayed transfer of 1,000 is not automatically a destroyed 1,000. The causal route from disruption to financial damage must be stated.
38. Insurance transforms retained loss rather than preventing the event
For a simplified policy with deductible d and maximum payment l per event, define insurer payment Y = min(max(X − d, 0), l), assuming the event is covered and the insurer pays as promised. Retained event loss is R = X − Y. Actual policies require their own terms, exclusions and timing; the formula is a teaching contract, not a statement of available coverage.
Use severity X = 10 with probability 0.8 and X = 50 with probability 0.2, in thousands. Set deductible 15 and payment limit 25. A small event receives no payout and leaves retained loss 10. A large event receives 25 and leaves retained loss 25. Expected gross loss is 18, expected insurer payment five and expected retained loss 13.
If the illustrative premium is six per exposure period for this one-event experiment, expected retained loss plus premium is 19, greater than the gross expected loss of 18. That does not automatically make insurance undesirable: the contract reduces the large retained outcome and may transfer a risk the buyer is poorly placed to bear. A decision can value risk reduction and liquidity protection in addition to expected cost.
The premium is not an incident loss reduction. Keep it as the cost of risk transfer in the decision ledger. Otherwise the comparison can mix loss datasets and expenditure in ways that make frequency and severity estimates inconsistent. Likewise, an insurer payout transfers part of a loss to another entity; it does not imply that the underlying operational failure never happened.
Where recovery is uncertain or delayed, model those features. A promised payout may not be available before a customer compensation deadline. Some costs may not meet coverage conditions. A financial model that assumes immediate certain recovery from every large incident can understate both retained loss and short-term cash needs.
39. An annual policy limit must be applied after the relevant event aggregation
Keep the deductible of 15 and per-event payout limit of 25. Suppose two large events of 50 occur in the same year. Without an annual aggregate limit, the insurer pays 25 for each, totalling 50, and the buyer retains 50 of the gross annual loss of 100. Now add an annual payout cap of 30. Total insurer payment becomes 30 and annual retained loss 70.
Applying a 25 payout to each event and forgetting the annual cap understates retained loss by 20 in this scenario. Applying the annual cap independently to each event also misreads the contract. A policy can include both occurrence-level and annual-level rules, and their order matters. Event grouping under the wording can matter as well if several consequences are treated as one occurrence.
The resulting retained-loss distribution is not generally a compound model with independent identically distributed retained severities. After the annual cap is partly used, the protection available to a later event changes. A year-level simulation or exact state calculation should track remaining coverage. The dependence is contractual even if the original incidents were independent.
Reinstatements, coinsurance, waiting periods and exclusions add other transformations. They should be coded as the stated contract and checked on small boundary cases: loss below deductible, exactly at attachment, exactly at the limit, and multiple events exhausting annual capacity. A model that passes only average-case examples can still fail where the policy’s nonlinear features matter most.
Insurance data also describe a selected population. Paid claims may exclude uncovered incidents and losses below deductibles. They cannot automatically be treated as a complete sample of gross operational severity. Retain the distinction between the event-generating process, the insurance reporting process and the payment transformation.
40. Recovery timing and counterparty performance belong in the same scenario
A hypothetical incident creates an immediate loss of 1,000. There is an 80% probability of collecting a recovery of 400 exactly one year later and a 20% probability of no recovery. With an illustrative annual discount rate of 5%, expected present value of recovery is 0.8 × 400 / 1.05, approximately 304.762. Expected present-value net cost is approximately 695.238 before other expenses.
The immediate cash requirement is still 1,000 if that full amount must be paid now. Expected discounted recovery is not money already in the account. A cash plan needs a financing route for the intervening period. A profitability or economic-value model can use discounted expected amounts while a liquidity model tracks the realised timing and amount separately.
Recovery probability may depend on the same environment that causes the loss. A widespread disruption can affect the operational entity and its service providers or risk-transfer counterparties simultaneously. Assuming independence could overstate protection in the most severe scenario. This is a modelling possibility to examine, not an assertion about the credit quality of any particular insurer or provider.
A useful recovery model therefore records coverage, expected amount, collection lag and uncertainty, together with the event conditions. It should also prevent double counting the same reimbursement under an insurer payment and a vendor indemnity. The combined recovery cannot exceed the amounts permitted by the relevant contracts and actual losses.
When comparing controls, distinguish reductions in gross event damage from improvements in recoverability. Both can reduce retained loss, but only the first necessarily reduces the underlying harm. A strategy relying entirely on later compensation may leave critical services interrupted, customers inconvenienced and immediate liquidity strained even when ultimate financial recovery is substantial.
41. Evaluate control investment with a complete decision ledger
The prevention-and-containment example reduced expected annual loss from 200 to 120, a benefit of 80 in thousands. Suppose implementing the controls costs 100 immediately and operating them costs 30 at each year-end for three years. Assume the expected annual loss benefit is also realised at year-end, all figures are nominal and the illustrative discount rate is 5%. Net annual benefit is 50.
The three-year present value is 50/1.05 + 50/1.05² + 50/1.05³ − 100, approximately 36.162. This is a positive expected-value result under the stated assumptions. It is not a guarantee of savings in each year, nor does it value every resilience or customer benefit. The frequency and severity reductions are model inputs that should themselves be tested.
If effectiveness takes a year to develop, the timing changes. If volume grows, the exposure and potential savings change. If maintenance costs rise or the control creates new errors, the net benefit changes. A complete decision analysis includes these possibilities rather than assuming that a one-time installation permanently delivers the maximum advertised reduction.
Some controls may be necessary to meet a service, legal or governance requirement even when a narrow expected-loss calculation is negative. Others may be worthwhile because they reduce severe outcomes not adequately represented by the mean. Those reasons should be stated directly. Do not conceal a policy constraint or ethical judgement inside an invented probability chosen to make the financial calculation positive.
The corporate-finance mathematics route develops discounting and investment comparison. The distinctive operational-risk task is to supply a defensible change in the event process, not merely a spreadsheet row labelled savings. A control becomes measurable when the model states what it changes and the evidence that would confirm or contradict that change.
VI. Validation: what the calculation establishes, and what it cannot establish alone
A loss model can be numerically correct and still unsuitable for the decision. The arithmetic may accurately aggregate a population that does not match the organisation’s current exposure. A well-fitted severity curve may ignore a reporting threshold. A carefully simulated annual quantile may be irrelevant to a payment that must complete in the next hour. Validation therefore needs several separate questions, not one approval box labelled model checked.
Start with the data and the meaning of the variables. Then check the statistical model, its numerical implementation and its proposed use. Each stage has different evidence. A successful software test does not establish the completeness of the loss database. A good historical fit does not establish that a new control caused an improvement. The following sections make those boundaries explicit enough for another reader to challenge them.
42. Regulatory capital and an internal loss distribution are different objects
The Basel OPE navigator, checked on 20 September 2026, identifies a standardised approach to operational-risk capital. It also separates current chapters from forthcoming versions, including an OPE25 version effective on 1 January 2027. A published future version is not yet the effective version merely because it appears first in a search result. Application to a particular bank also requires the relevant jurisdiction’s implementation and scope.
Our frequency–severity model instead describes a hypothetical distribution of financial consequences. It can support internal investigation, control comparison, scenario analysis or a defined economic-risk measure. A fitted Poisson rate multiplied by a severity distribution does not, by itself, create a legal capital requirement. Conversely, complying with a standardised calculation does not prove that every operational vulnerability has been understood or every critical service can survive disruption.
Older operational-risk research often discusses the Advanced Measurement Approaches and high annual confidence levels within its historical regulatory setting. Those papers can still contain valuable statistical methods. Read the methods and the regulatory statements separately. A sound Bayesian update does not become obsolete merely because a capital framework changes, but a historical description of regulatory permission should not be presented as a current authorisation.
A useful report gives each number an unambiguous label: expected financial loss, simulated annual loss quantile, stressed cash requirement, standardised capital charge or service interruption. A conversion between them needs a stated rule. For example, subtracting an expected-loss budget from a tail quantile is a chosen definition of an additional buffer, not an automatic identity establishing required regulatory capital.
This distinction protects readers from a common error in financial education: believing that the most advanced-looking statistical formula is always the governing rule. Sometimes the governing rule is deliberately standardised. Sometimes the important management question requires a more detailed internal model than that rule supplies. The analyst should explain which problem is being solved rather than treating complexity as a source of authority.
43. A data reconciliation can falsify a model before a statistical test does
Suppose an incident register contains thirty cases, while the finance reconciliation identifies thirty-three distinct events with included charges. The missing three might be late records, events below the collection threshold or genuine omissions. Until the discrepancy is explained, a distribution fitted to the thirty cannot honestly be described as a model of all qualifying events. Statistical precision cannot compensate for an undefined population.
Reconcile identifiers before totals. Two databases can agree on total loss while assigning charges to different incidents. That preserves the annual sum but changes frequency, severity and dependence. If one system splits a common outage into a hundred independent events and another groups it once, the compound model built from each can predict a different future tail even when both reproduce last year’s booked amount.
Follow an event through its lifecycle. The operational report identifies what happened; financial records show recognition and payments; recovery records show collected or expected reimbursements. A reconciliation should explain why each amount is included, excluded or still provisional. Original currencies and adjustments should remain accessible. A corrected dataset should preserve an audit trail rather than overwriting history so thoroughly that nobody can explain why the model changed.
Missing and zero need separate codes. Zero may mean a recorded near miss, a fully recovered event under a net convention or a period with no qualifying incidents. Missing may mean the amount has not been estimated or the period was not monitored. Converting every blank to zero can lower both apparent frequency and severity. Dropping every zero can remove genuine information about the observation process or insurance transformation.
Clara makes the reconciliation useful by selecting several cases rather than relying only on an aggregate balance. She checks a large unresolved incident, a small fully recovered one, an event with multiple charges and a period of zero reports. Each probes a different boundary. Passing these checks does not prove a tail model is correct, but failing them identifies a concrete repair before more modelling effort is spent.
44. Test the distribution that generated the observations
If an exact continuous observation X truly follows cumulative distribution F with known parameters, the transformed value F(X) is uniform between zero and one. This follows because P(F(X) ≤ u) = u for the relevant inverse mapping. It provides a useful diagnostic foundation: a model that assigns too many observations to its own extreme tail is not matching the observations as claimed.
For losses observed only above collection threshold a, however, the appropriate observed cumulative distribution is [F(x) − F(a)]/[1 − F(a)] for x above a. Applying the unconditional F directly to the truncated observations cannot yield a uniform distribution over the full unit interval. The missing lower region is a consequence of the sampling rule, not evidence that the underlying loss family is necessarily wrong.
Estimated parameters complicate formal testing because the same observations helped choose the fitted distribution. The transformed values are not a fresh independent validation sample simply because they have been passed through a formula. A calibrated testing procedure, simulation under the fitted process or separate evaluation sample may be needed. The precise method should match the estimation and observation design.
Inspect more than one diagnostic. A model can match the central histogram while badly missing upper quantiles. It can fit the observed tail by implying an implausibly large population below the reporting threshold. It can reproduce marginal severities while failing to describe their clustering by root cause or business regime. A distributional check should therefore include the body, relevant tail, implied omitted population and dependence assumptions.
Numerical fitting also needs boundaries. A solver can find a parameter combination on the edge of an admissible region or a local optimum far from another reasonable solution. Report convergence diagnostics, alternative starting values and the financial implications of the fitted parameters. A software success flag confirms that a procedure stopped according to its rules, not that the resulting incident mechanism makes operational sense.
45. Preserve time and incident identity when evaluating predictions
Randomly splitting database rows into training and test samples can leak information when several rows belong to the same incident. If early charges from a legal case enter training while later charges from that case enter testing, the supposed test is not fully independent evidence about a new event. The same issue arises when related consequences of a technology outage are scattered across the split.
For a next-year prediction problem, a time-ordered evaluation is often more faithful to the information actually available at the decision date. Fit using records and estimates known then, and evaluate against later observations at a comparable development stage. Do not use a final recovery discovered years later as though it had been known when the original forecast was issued.
Thresholds, currency adjustments and event grouping should also be determined without looking ahead to favourable test results. Choosing a tail threshold on the full dataset and then describing the later period as untouched validation understates the amount of information used. An honest model comparison documents which choices were fixed before the evaluation and which were revised afterward.
Business changes can make an old forecast miss for an understandable reason. A merger adds exposure, a service closes or a new collection policy identifies more incidents. That does not mean the error should be ignored. Decompose it into exposure changes, observation changes, parameter error and omitted mechanisms where evidence permits. This turns evaluation into a learning process rather than a contest over whether one number was close.
A useful prediction archive stores the original forecast distribution, not only its mean. Later reviewers can then test count ranges, expected losses and threshold exceedances separately. Replacing an old forecast with an updated model destroys the ability to learn from the original decision. Versioning is a practical statistical control: it preserves the difference between what was known then and what became clear later.
46. A short history cannot strongly validate an extreme annual quantile
Suppose a correctly specified continuous annual-loss model has a 99.9% quantile, and annual outcomes are independent. The probability of seeing no exceedance in ten years is 0.999¹⁰, approximately 99.0045%. Therefore ten years with no breach is unsurprising under the model. It is weak evidence for the precise location of that extreme quantile because the evaluation contains very few opportunities to observe the relevant tail.
For comparison, a threshold with true exceedance probability 1% has probability 0.99¹⁰, approximately 90.4382%, of seeing no breach over ten independent years. A zero-breach history can occur under both a 0.1% and a 1% exceedance model. It does not distinguish them sharply. The absence of a breach should not be advertised as verification to three decimal places.
More frequent observations do not automatically create more independent annual tail tests. Monthly losses aggregated into overlapping twelve-month windows share data. Thousands of transactions affected by one outage share a cause. Treating these as independent repetitions of an annual catastrophe can create false confidence. The effective evidence depends on the event mechanism and sampling structure, not just the number of rows.
Validation can still make progress. Test more central parts of the distribution where data are informative. Examine whether the fitted model reproduces observed threshold exceedances, event counts and severity development. Challenge the extreme scenarios through exposure limits, causal analysis and alternative tail families. A tail estimate remains uncertain, but the uncertainties become specific enough to manage rather than hidden behind an untestable precision claim.
External events can help reveal missing mechanisms even when they cannot be pooled mechanically with internal observations. A well-documented outside outage may identify a shared dependency absent from the model. Its value can be conceptual rather than a direct addition to the likelihood. Keeping these uses separate permits learning from rare events without pretending that every organisation supplies an exchangeable statistical sample.
47. A change in the mix can overwhelm improvements within every category
Consider two transaction categories. In the first period, a low-complexity category processes 900,000 transactions at one incident per thousand, producing 900 incidents. A high-complexity category processes 100,000 at ten per thousand, producing 1,000. The total is 1,900 incidents per million transactions. These are constructed counts for a comparison, not observed bank data.
In the second period, both category rates improve by 20%: to 0.8 and eight incidents per thousand. But the transaction mix reverses. Only 100,000 low-complexity transactions produce 80 incidents, while 900,000 high-complexity transactions produce 7,200. The overall rate rises to 7,280 per million, despite improvement within both categories.
A dashboard using only the aggregate rate would describe a large deterioration. A dashboard using only within-category improvement would ignore the much larger total workload and loss opportunity. Both views are incomplete alone. Standardising the second-period rates to the first-period mix gives 0.9 × 0.8 + 0.1 × 8 = 1.52 incidents per thousand, correctly showing a 20% rate improvement at a constant mix.
The observed overall rate remains 7.28 per thousand, which is the quantity needed for the actual period’s workload. The standardised 1.52 is a comparison statistic, not a replacement for the real total. Reporting both separates process performance from exposure composition. This is the same discipline used throughout the article: a statistic must retain the population and question that define it.
Even the within-category rate change does not prove that a particular control caused the improvement. Other changes may have occurred within each category, and counts themselves are random. A causal assessment needs a credible comparator, timing and a mechanism linking the control to the affected incidents. Good adjustment removes one source of confusion; it does not remove every possible alternative explanation.
48. Bounds are useful when they are not mistaken for forecasts
For nonnegative annual loss S with mean m, Markov’s inequality gives P(S ≥ B) ≤ m/B for positive B. The reasoning is simple: on the event S ≥ B, loss contributes at least B, so the overall mean cannot be smaller than B times that event’s probability. This is a distribution-free upper bound given the mean, not a claim that the exceedance probability equals the bound.
Using the compound model with mean 200 and variance 100,000, all amounts in thousands, Markov’s bound for loss at least 1,000 is 20%. With finite variance, the one-sided variance bound gives P(S − 200 ≥ 800) ≤ 100,000/(100,000 + 800²), approximately 13.514%. This tighter bound uses more information, but it can still be much larger than the probability under a particular fitted distribution.
These bounds can help check an obviously inconsistent result or show what limited moment information alone permits. They cannot identify the exact tail. If a decision requires an exceedance probability below 1%, a 13.514% upper bound does not establish failure of that requirement; it merely fails to demonstrate that the requirement is met using this information. Distinguish evidence of a breach from insufficient evidence of safety.
Moment uncertainty matters too. A bound computed from estimated m and v inherits uncertainty in those estimates. If the severity variance is not finite under the chosen model, a finite-variance inequality cannot be applied by substituting a convenient sample variance. Mathematical conservatism depends on the assumptions actually holding. A bound with an unsupported input is not automatically conservative for the real process.
The same caution applies to stress scenarios described as worst case. A largest loss among one thousand simulations is only the largest simulated loss, not the maximum possible outcome. A contractual exposure cap may support a genuine upper bound; an arbitrary simulation range does not. Use language that distinguishes a tested scenario, a quantile, an observed maximum and a proven bound.
49. Service validation and financial validation need different pass conditions
A hypothetical payment service has an internal requirement that critical instructions complete within two hours under a defined disruption scenario. Its annual expected operational loss falls after an insurance purchase. That change does not establish compliance with the two-hour service objective. The insurance pays money after a covered event; it may do nothing to restore the service while the event is unfolding.
A separate recovery test should trace detection, decision, failover, data integrity, resumed processing and backlog clearance. The test passes only if the required operation is delivered under its stated conditions. A technically successful failover that loses necessary records or leaves priority instructions unprocessed may fail the operational objective even when a server is running.
The Basel operational-resilience guidance makes interdependencies and delivery of critical operations central to the assessment. Our teaching implication is to preserve a service ledger alongside the financial ledger: what stopped, who depended on it, what substitute existed and when delivery actually resumed. This is a distinct dimension of resilience, not another name for a loss quantile.
A combined report can therefore say that a control reduces expected loss, improves a particular recovery-time distribution and leaves a common dependency unresolved. That is more actionable than calling the organisation resilient without qualification. The remaining dependency becomes a specific test or repair, while the verified improvements retain their value. The final workshop puts this habit into decisions that a reader can reproduce with a calculator.
VII. Application workshop: ten decisions with complete numerical checks
These exercises combine the components rather than repeating a formula with new labels. Each changes a meaningful part of the model: the annual insurance contract, the question about extreme events, the targeted incident class, the observation process or the relationship between disruption and cost. All numbers are hypothetical. Amounts are in thousands of Singapore dollars unless a question explicitly uses transaction counts, hours or probabilities.
For each exercise, write the decision before the equation. A question about the largest event needs a different calculation from a question about the total year. A question about available cash needs dates as well as ultimate losses. After calculating, identify the first assumption that would change the conclusion. This final step turns a worked answer into a usable model rather than a memorised result.
50. Build the full retained annual-loss distribution
Return to the finite annual model: N is zero, one or two with probabilities 0.5, 0.3 and 0.2; each event costs 10 or 50 with probabilities 0.8 and 0.2. Add the hypothetical insurance contract with deductible 15, per-event payout limit 25 and annual payout cap 30. Assume every event is covered and payouts are certain; timing and premium are excluded from this particular loss-distribution calculation.
For each possible year, calculate gross loss, the sum of eligible per-event payouts and the annual capped payout. Then subtract the actual annual payout from gross loss. The probabilities of the underlying incident combinations do not change because insurance does not prevent those incidents in our model. It changes the financial amount retained by the buyer.
| Year’s incidents | Probability | Gross loss | Actual annual payout | Retained loss |
|---|---|---|---|---|
| No incident | 0.500 | 0 | 0 | 0 |
| One small incident | 0.240 | 10 | 0 | 10 |
| Two small incidents | 0.128 | 20 | 0 | 20 |
| One large incident | 0.060 | 50 | 25 | 25 |
| One small and one large | 0.064 | 60 | 25 | 35 |
| Two large incidents | 0.008 | 100 | 30 | 70 |
Expected retained annual loss is 9.26. Expected payout is 3.34, and the two add to the original gross mean of 12.6. Retained variance is 144.5524. The 99% retained-loss quantile is 35, and its Expected Shortfall is [70(0.008) + 35(0.002)]/0.01 = 63. The annual cap matters most in the two-large-event outcome.
Without the annual cap, that outcome would retain 50 rather than 70. Removing the cap therefore lowers expected retained loss by 20 × 0.008 = 0.16, but lowers the 99% Expected Shortfall by 16, from 63 to 47. The same contractual feature has a modest effect on the mean and a much larger effect on the selected tail. A premium comparison should not treat these two benefits as interchangeable.
The table also provides an implementation test. Every row satisfies gross loss equals retained loss plus payout. The probabilities sum to one. The cap is applied to the year’s combined eligible payouts, not independently to each incident. A numerical routine that fails any of those checks is misrepresenting the contract regardless of how sophisticated its simulation engine may be.
51. Calculate the probability of at least one large event
Suppose incident count is Poisson with annual rate four, and each independent severity has a 2% probability of exceeding a specified amount B. Conditional on n incidents, the probability none exceeds B is 0.98ⁿ. Averaging over the Poisson count gives exp(−4 × 0.02). The probability of at least one event above B is therefore 1 − exp(−0.08), approximately 7.6884%.
This is not generally the probability that annual total loss exceeds B. Two losses of 60 can make the annual total exceed 100 even though neither individual loss exceeds 100. For nonnegative severities, an individual loss above B guarantees the total is above B, but the reverse implication fails. The large-event probability is consequently a lower bound on that total-loss exceedance probability in this setting.
Let M be the largest individual loss in the year, with M = 0 in a zero-incident year. For nonnegative x, its cumulative distribution is P(M ≤ x) = exp{−λ[1 − F(x)]}. This follows from the same conditional argument. At a 99% annual-maximum quantile with λ = 4, the corresponding severity cumulative probability is 1 + ln(0.99)/4, approximately 0.9974874.
Thus the annual maximum’s 99% boundary corresponds to approximately the 99.7487% individual-event severity quantile in this model, not the individual 99% quantile. The conversion depends on the expected count and its distribution. If the requested maximum quantile lies within the zero-event probability mass, it is zero instead; the inverse formula must respect that boundary.
This exercise is useful when planning for a single large incident, setting an investigation threshold or understanding how exposure growth changes the chance of encountering an extreme event. It does not replace aggregation when the decision concerns total annual expenditure. State whether the object is an event, the maximum event or the sum before interpreting the probability.
52. A tiny change in incident count can remove a large share of expected loss
Construct two independent incident classes. Small events have Poisson rate 100 and cost one each. Large events have Poisson rate 0.2 and cost 1,000 each. Expected annual incident count is 100.2. Expected annual loss is 100 × 1 + 0.2 × 1,000 = 300. Most incidents are small, but the rare large class contributes two thirds of the mean loss.
Control A halves the small-event rate and leaves the large-event process unchanged. Total expected count falls to 50.2, an impressive-looking reduction of almost half. Expected loss falls to 250, a reduction of 50. Under the independent compound-Poisson model, annual variance falls only from 200,100 to 200,050 because the large-event class still dominates squared loss.
Control B leaves small events unchanged but halves the large-event rate from 0.2 to 0.1. Total expected count falls only from 100.2 to 100.1. Expected loss nevertheless falls to 200, a reduction of 100, and variance to 100,100. A dashboard ranking controls only by the number of incidents prevented would strongly favour A and miss B’s effect on severe financial exposure.
The arithmetic does not establish that B should always be purchased. Its cost, feasibility and evidential support may differ. Small incidents may also create customer burdens not included in these financial severities. The exercise establishes a narrower point: count reduction is not an adequate substitute for risk reduction when incident classes have very different consequences.
For an actual control comparison, identify the incident class affected and justify the change in its rate or size. Preserve a separate service measure where frequent small problems matter to users. A combined decision can then compare expected cost, severe outcomes and service reliability openly, rather than allowing the easiest count statistic to decide the result by default.
53. More reporting history cannot by itself identify an unknown detection rate
Assume underlying incidents form a Poisson count with rate λ and each is independently reported with constant probability q. Conditional on N incidents, the number reported is binomial with N trials and probability q. Summing over the Poisson count produces a Poisson reported count with rate λq. This is the thinning result used earlier, now applied as an identification test.
Compare world A with λ = 100 and q = 0.4 to world B with λ = 50 and q = 0.8. Both generate exactly the same distribution of annual reported counts: Poisson with mean 40. They do not merely share an average. Every probability for the observed count is identical under the specified model, even though their true incident rates differ by a factor of two.
A longer series of reported counts can estimate the product λq more precisely, but it cannot separate λ and q without additional information or restrictions. This is structural non-identifiability. The repair is not simply to collect more of the same variable. Independent audits, controlled detection tests or other evidence about reporting completeness can help identify q, while their own sampling assumptions must also be examined.
Suppose independent evidence only supports q between 0.5 and 0.8 and the observed mean is approximately 40. Ignoring estimation uncertainty in that mean for the illustration, the implied true rate lies between 50 and 80. Reporting a single corrected rate of 65 would introduce an unstated choice inside that range. A sensitivity analysis should show whether decisions differ across the admissible values.
If reporting depends on severity, the thinning model needs marks or categories rather than one q. Larger events may be more visible. An audit of easily discovered large incidents may say little about small failures. The lesson is to obtain evidence about the observation process relevant to the population, not merely to improve the precision of the reported total.
54. Compare two incident rates with their sampling uncertainty
A hypothetical before period has forty incidents in two million transactions. An after period has thirty incidents in three million. The observed rates are twenty and ten incidents per million, so the after-to-before rate ratio is 0.5. This calculation adjusts for exposure; comparing only counts would report a 25% reduction rather than a 50% rate reduction.
Under independent Poisson counts with known exposures and sufficiently large counts for the approximation, the standard error of the log rate ratio is approximately √(1/30 + 1/40), or 0.24152. The expression follows by approximating the variance of the log of each count and adding the independent contributions. An approximate 95% interval is exp[ln(0.5) ± 1.96 × 0.24152], about 0.311 to 0.803.
This interval concerns the ratio under the stated observation and sampling model. It does not incorporate an unmeasured change in detection probability, transaction mix or event definition. It also does not establish that a particular control caused the rate change. A before–after comparison can provide evidence of association while leaving alternative explanations unresolved.
For small or zero counts, the logarithmic approximation needs a different treatment; it cannot take the logarithm of zero and remain valid. An exact or otherwise justified model-based interval may be appropriate. Adding an arbitrary small count without explanation changes the estimator and its interpretation. State the method and test it against the actual sample size.
Aisha writes the practical conclusion in two parts. The exposure-adjusted observed rate fell by half, and the specified Poisson comparison supports a rate ratio below one within its approximate interval. Attribution to the control requires process and comparison evidence beyond that calculation. Keeping the statistical and causal claims separate makes the conclusion stronger, not weaker.
55. Join a severity body and tail without losing the probability weights
Suppose 98% of event losses lie at or below 100, with conditional mean 20. The other 2% exceed 100. For the excess above 100, use the hypothetical generalized-Pareto parameters β = 50 and ξ = 0.25. Its finite mean excess is β/(1 − ξ), or approximately 66.6667, so mean total severity among exceedances is approximately 166.6667.
The overall mean is 0.98 × 20 + 0.02 × 166.6667, approximately 22.9333. The tail supplies approximately 3.3333 of that mean, or about 14.535%, despite containing only 2% of the incidents. At annual Poisson frequency ten independent of severity, expected annual loss is approximately 229.333. The calculation uses both the tail shape and the probability of entering the tail.
A proper cumulative distribution can be written using a conditional body distribution B(x) and excess distribution G(y). At x ≤ 100, F(x) = 0.98B(x). Above 100, F(x) = 0.98 + 0.02G(x − 100). The total probability reaches one, and the two pieces meet at the threshold under the stated continuous-tail convention. The pieces are not two complete distributions to be added without weights.
Changing the exceedance probability changes expected loss even when the fitted excess distribution remains unchanged. Changing β or ξ changes the size of tail events even when their frequency remains unchanged. These are separate uncertainties. An analyst should not attribute a higher annual tail estimate solely to heavier severity if the rate of entering the tail also changed.
The body distribution must also be compatible with its stated support and mean. A separate uncapped lognormal used as the body could still assign probability above 100 unless explicitly conditioned or transformed. The splice should be defined mathematically, not created by pasting the left half of one chart beside the right half of another.
56. Two budgets can have the same breach frequency but different shortfalls
Use the retained annual-loss distribution from exercise 50, excluding premium. A hypothetical financial reserve of 35 is exceeded only by the retained loss of 70, which has probability 0.008. A reserve of 50 is also exceeded only by that same outcome. The two reserves therefore have the same 0.8% breach probability in this deliberately discrete model.
The expected uncovered amount differs. For reserve 35 it is 0.008 × (70 − 35) = 0.28. For reserve 50 it is 0.008 × (70 − 50) = 0.16. Increasing the reserve reduces the severity of a breach even though it does not change its frequency. The excess-loss measure E[(R − B)⁺] captures a distinction that the breach indicator does not.
At reserve 34, both retained outcomes 35 and 70 breach the budget. Breach probability jumps to 0.064 + 0.008 = 7.2%, while expected uncovered loss is 0.064 × 1 + 0.008 × 36 = 0.352. Moving from 34 to 35 therefore sharply changes the indicator but changes expected uncovered loss much more smoothly. This is the effect of probability mass at a boundary.
These calculations do not recommend a reserve level. Holding funds has an opportunity cost, and other sources of loss, cash timing and required capital lie outside the example. The exercise shows why a budgeting decision should inspect both how often a threshold is crossed and what happens after it is crossed. A pass–fail frequency alone may miss an economically important improvement.
The same logic applies to service thresholds. Two recovery plans can have the same probability of missing a deadline but different durations of delay when they miss it. Report the tail consequence alongside the threshold indicator. The question is not simply whether a promise fails, but how severely and for whom it fails in the scenarios that remain.
57. Recovery reliability should be conditional on the loss-producing state
Construct a two-state annual experiment. With probability 95% there is no loss event. With probability 5% a disruption causes loss 1,000 and creates an eligible recovery of 800 from a hypothetical counterparty. In the disruption state the counterparty pays with probability one half. In the ordinary state it would be able to pay, but no recovery claim arises. This is a teaching dependence model, not an assessment of an actual insurer.
The correct net-loss outcomes are zero with probability 0.95, loss 200 with probability 0.025 when recovery succeeds, and loss 1,000 with probability 0.025 when it fails. Expected net loss is 200 × 0.025 + 1,000 × 0.025 = 30. Gross expected loss is 50, and expected recovery is 20.
Across all years, the counterparty’s modelled ability to pay is 0.95 × 1 + 0.05 × 0.5 = 97.5%. Using that unconditional percentage as though it applied independently whenever the loss occurs would predict expected net loss of 0.05 × (1,000 − 0.975 × 800) = 11. The prediction is too low because payment reliability is worse precisely in the state that generates the claim.
A high overall reliability statistic can therefore be the wrong input for contingent protection. The relevant quantity is performance conditional on the event and contractual conditions. A model should examine shared exposures, timing and dependencies rather than assume that a recovery provider’s ordinary-state performance fully describes the stressed claim.
The repair is not to assume every recovery fails. That would discard useful protection without evidence. Use a conditional scenario or a sensitivity range, show how much the result depends on it, and identify the evidence that could narrow the uncertainty. The model then preserves both the benefit of protection and the conditions under which it may be weakest.
58. The cost of average downtime can be smaller than average downtime cost
Suppose an incident lasts one hour with probability 0.8 and ten hours with probability 0.2. Its mean duration is 2.8 hours. Define an illustrative event-cost function C(t) = 20 + 5t + 2[max(t − 2, 0)]², in thousands. The fixed term represents response mobilisation; the linear term operating expense; the squared term represents an explicitly assumed escalation after two hours. This is a chosen teaching function, not a measured industry cost curve.
A one-hour incident costs 25. A ten-hour incident costs 20 + 50 + 128 = 198. Expected event cost is therefore 0.8 × 25 + 0.2 × 198 = 59.6. Plugging mean duration 2.8 into the cost function instead gives only 35.28. Nonlinearity prevents the cost of the mean from replacing the mean cost.
At incident frequency 0.25 per year, independent of these durations, expected annual financial loss is 14.9. The average-duration shortcut would produce 8.82. That difference has nothing to do with a new event count or a mathematical error in taking the mean duration. It comes from ignoring the way long outages disproportionately increase severity under the specified function.
A recovery improvement reduces the ten-hour outcome to four hours while leaving the one-hour outcome and frequency unchanged. Four-hour cost is 48. Expected severity becomes 0.8 × 25 + 0.2 × 48 = 29.6, and expected annual loss 7.4. The control’s main value comes from truncating the long-duration consequence, not from reducing initiating incidents.
To use such a model in practice, each cost component needs an evidence-based interpretation and a clear inclusion boundary. Do not invent a quadratic because it produces a persuasive investment case. The example demonstrates the need to model nonlinear consequences; the actual shape and scale require separate justification, especially where customer harm cannot be reduced to a simple financial charge.
59. Solve the recovery-capacity requirement from the deadline
A service receives a constant a tasks per hour, stops processing for d hours and must clear the resulting backlog by H hours after the incident begins, where H is greater than d. In the fluid model with no abandonment, the backlog at restart is ad. During the remaining H − d hours, recovery capacity μ must handle both new arrivals and that backlog.
The requirement is (μ − a)(H − d) ≥ ad. Rearranging gives μ ≥ aH/(H − d). With arrivals 1,200 per hour, downtime two hours and a six-hour total clearance deadline, required recovery capacity is at least 1,800 tasks per hour. A plan offering 1,500 per hour cannot meet that deadline under the assumptions, even though it eventually clears the backlog.
If downtime rises to three hours while the same six-hour deadline remains, the required rate rises to 2,400 per hour. The available recovery window has shortened while the backlog has grown. As downtime approaches the deadline, the required catch-up rate becomes extremely large. This exposes why a recovery plan cannot compensate indefinitely for slow detection or delayed restart merely by promising to work faster afterward.
The fluid requirement is a deterministic capacity benchmark, not a guarantee under variable arrivals and service. Real plans need room for variability, retries, priorities, reconciliation and dependencies. Some tasks may also expire or breach individual deadlines before the entire backlog clears. A whole-backlog target should therefore be supplemented with completion measures for critical task classes.
Ethan finishes the workshop by asking what has actually been demonstrated. The calculation gives the minimum constant recovery throughput for a precisely defined simplified schedule. A test of the real service must establish that the required throughput, data integrity and prioritisation are attainable in the same disruption scenario. An equation can specify the demand on a plan; it cannot make untested capacity exist.
VIII. Turn the calculation into an accountable decision
A numerical result becomes useful when another person can identify what would make it change. That person might be a student checking an equation, a manager considering a control or a reviewer questioning a tail estimate. The final task is therefore to preserve the connection between the model and the decision. Do not compress everything into a number labelled operational risk and discard the incident, horizon, recovery and service assumptions that gave it meaning.
60. Five years are not obtained by multiplying every annual statistic by five
Suppose annual losses are independent and identically distributed, each with mean 200 and variance 100,000, measured in thousands and squared thousands respectively. The undiscounted five-year total has mean 1,000 and variance 500,000. Its standard deviation is approximately 707.107, which is the annual standard deviation multiplied by the square root of five. The mean scales linearly because expectation adds; the variance calculation additionally uses independence.
Neither calculation establishes that a particular five-year quantile is five times, or the square root of five times, its annual counterpart. Quantiles depend on the full distribution of the sum. In a rare-event model, changing the horizon changes the probability of encountering any event at all. An annual quantile may lie in a zero-loss probability mass while a longer-horizon quantile does not. A convenient scaling rule should not replace that distributional calculation without justification.
A persistent weak control can also make annual losses dependent. As a mathematical comparison, suppose each pair of annual losses has covariance 20,000 while annual variances remain 100,000. There are ten distinct pairs across five years. Total variance is 5 × 100,000 + 2 × 10 × 20,000 = 900,000, giving standard deviation approximately 948.683. The covariance is an explicit hypothetical assumption, not an estimate. It shows why repeating the same risk environment may provide less time diversification than independent years suggest.
Discounting introduces another distinction. For fixed discount weights wt, present-value mean is the sum of wtE[St], while variance contains wt²Var(St) and twice the weighted cross-year covariances. Cash needed on an actual payment date is not reduced by discounting its economic value. A multiyear investment appraisal and a cash-resilience plan may use the same scenarios but answer different questions about them.
Jo therefore writes three horizons on the model sheet: the period in which incidents begin, the period over which their financial consequences develop and the period over which the decision is evaluated. They need not coincide. A three-year control programme can affect incidents whose settlements occur later. A model that truncates every consequence at the programme’s end may understate the cost of the incidents it claims to cover.
61. A control benefit should survive a break-even calculation
Consider a hypothetical control that costs 40 per year, with no initial expense, and independently prevents fraction r of a defined incident population. Suppose annual frequency is λ and mean severity is m, in the same thousands-of-dollars units. Ignoring other benefits, implementation effects and risk preferences, its expected annual net benefit is λmr − 40. Break-even prevention is r = 40/(λm), provided the denominator is positive. This expresses what the control must achieve; it does not establish that it can achieve it.
At λ = 4 and m = 50, prevention must reach 20% to break even on this narrow expected-cost basis. If the credible annual rate were only two with the same mean severity, the required fraction would be 40%. If the rate were six, it would be approximately 13.333%. A decision that changes across those plausible rates deserves investigation of the frequency evidence. More precise estimation of an unrelated parameter would contribute less to resolving this particular choice.
The prevention assumption must remain attached to the incident class. If a control prevents only low-severity cases, the relevant m is the mean of the prevented population, not necessarily the overall mean. If prevention probability depends on loss size, expected avoided loss is E[X times the conditional prevention probability], multiplied by the appropriate event rate. Replacing that expression with average severity times an unweighted prevention percentage can misprice the benefit.
There may be reasons to implement the control despite a negative mean-cost result: a mandatory requirement, a critical service objective or a strong aversion to a severe remaining outcome. Those reasons should be represented as constraints or explicitly chosen objectives, not hidden inside an inflated loss assumption. Equally, a positive expected benefit does not override evidence that the control disrupts legitimate transactions or depends on unavailable operating capacity.
Mira asks what evidence would change the decision. A test showing that prevention targets the severe mechanism could be valuable. A revised reporting audit could change λ. A realistic workload trial could reveal operating costs absent from the proposal. The break-even equation gives these investigations a purpose. It helps the team distinguish an uncertainty that determines the decision from an uncertainty that only changes an unimportant decimal.
62. Report the model as a chain of claims that can each be checked
A complete model statement can begin with the population: qualifying processing incidents in specified businesses and exposure periods, grouped by a documented event rule. It then states the observation process: collection thresholds, discovery delays and exclusions. Next come the count and severity assumptions, including dependence and the treatment of recovery. Only after those choices should the report present an annual distribution or a control comparison.
The result should contain both a central estimate and a description of what makes it uncertain. For a small discrete example, exact probabilities and a complete outcome table may be appropriate. For a fitted model, ranges across defensible parameter values or tail families may be more informative. Label numerical simulation error separately from uncertainty in the business model. An accurate computation of an uncertain assumption remains uncertain about the business.
Attach an action interpretation to each important output. Expected count informs investigation workload. Expected gross loss informs a defined financial-cost baseline. Retained tail loss informs the consequence left after the specified protection. The cash path identifies financing deadlines. A service interruption measure identifies whether users receive the operation promised. An improvement in one column should not automatically be copied into all the others.
Finally, record a condition for review. A change in transaction mix, reporting coverage, vendor dependency, insurance wording or event development can invalidate a previous assumption. That is different from changing a model whenever an unfavourable result appears. A review trigger should refer to evidence about the process. Preserve the previous specification so the reason for the change can be understood instead of retrospectively erasing the old forecast.
Adrian’s final test is conversational: could a reader explain why the model predicts this result without saying that the computer said so? The answer should identify the incidents, their probabilities, their consequences and the relevant constraints. When the explanation reaches an assumption not supported by data, say so and test alternatives. That is where the useful work begins, not where the report should become more confident.
Reader questions: keep the decision and the probability together
Is frequency multiplied by severity the whole of operational risk mathematics?
It is an expected-loss calculation under a compatible population and model, not the whole annual distribution. It does not describe whether incidents arrive together, whether one event is exceptionally large, whether losses develop over several years or whether recovery arrives before the payment deadline. The compound variance and complete finite distribution showed how information beyond the mean changes the range of annual outcomes.
Start with the product because it is interpretable, then ask what decision remains unanswered. A staffing decision needs arrival and service information. A severe-loss question needs tail assumptions. A resilience decision needs a service and dependency model. Making the mean more precise will not answer a question that the mean was never designed to resolve.
Why did the incident count rise after controls improved?
Possible explanations include higher exposure volume, a more difficult transaction mix, better detection, changed reporting thresholds or actual deterioration in a separate process. The count alone cannot identify which explanation applies. The exposure and observation examples demonstrated that reported incidents can rise even when each comparable transaction is safer or the underlying incident rate is unchanged.
Compare consistent event definitions, exposure-adjusted rates and evidence about detection coverage. Preserve the actual workload count as well; investigators still have to handle the reports they receive. An increase in visibility can be beneficial while creating a genuine capacity requirement. It should not be dismissed as irrelevant or mislabelled as proof that the preventive control failed.
Should a fully recovered mistake disappear from the dataset?
Not from a dataset intended to understand incidents and their recovery paths. The initiating failure occurred even if ultimate net financial loss is zero. Record gross consequences and recovery separately, then derive the net variable needed for a particular analysis. A gross-severity model, a net-loss model and a process-quality count may include the record differently because they answer different questions.
Erasing the incident loses evidence about the control weakness, temporary cash demand and success of recovery. Counting its gross amount and later recovery as two independent loss events creates another problem. The stable incident identifier connects those entries and allows each analytical view to be produced without rewriting what happened.
Which severity distribution is best?
There is no distribution that becomes correct simply because the loss is operational. The choice must reflect support, event mechanisms, reporting rules and the available evidence. A positive distribution may fit gross monetary losses but not a signed cash-flow series. A tail family may be appropriate above a threshold without describing ordinary small events. A cap may be justified for one defined exposure but not for unrelated consequences.
Compare defensible candidates and inspect the consequences for the decision. Where the data cannot distinguish their extreme tails, present that uncertainty. Choosing the lowest estimate because it is commercially convenient or the highest because it sounds cautious does not substitute for an explanation of why the model describes the intended population.
Why is Expected Shortfall sometimes different from the average of losses above VaR?
For a continuous loss distribution with the appropriate regularity, the familiar conditional-tail expression can agree with the quantile-based definition. With probability mass at the boundary, conditioning on loss greater than or equal to VaR may select too much probability. Expected Shortfall defined as the average of the worst specified percentage uses only the necessary fraction of the boundary mass.
The finite example made this visible: the worst 1% included all 0.8% probability at the largest loss and only 0.2% from the next loss level. Using every observation at that next level would average a different portion of the distribution. State the definition and test the implementation on a distribution where the correct selected tail is visible.
Does a rare event deserve to be excluded because it has never happened internally?
An internal zero history is evidence about observed events over a stated exposure, not proof of impossibility. A relevant scenario can be informed by process analysis, dependency tests or carefully interpreted external incidents. That does not mean every imaginable event should be assigned an arbitrary probability and inserted into the fitted database. Separate realised observations from hypothetical scenarios and expert judgements.
The next question is whether the mechanism is possible within the organisation’s exposure and whether existing protections have been tested against it. Where probability is poorly known, a conditional consequence calculation can still be useful. It tells the reader what would happen under the stated event without pretending that its frequency has been statistically established.
Can several business lines simply add their operational-risk estimates?
Expected values of distinct financial consequences add, but the distribution of their sum also depends on joint occurrence. Shared platforms, processes or external conditions can cause losses across several lines at once. The common-incident example preserved each line’s marginal count while producing a larger aggregate variance than an independent model. Separate line estimates need an aggregation rule that represents that relationship.
First remove duplicates: two lines may have recorded the same total incident cost rather than different consequences. Then model dependence among the genuinely distinct amounts. Deduplication and dependence are separate tasks. Performing only one leaves either an inflated total or a misleading pattern of annual outcomes.
Does insurance make the operation itself safer?
In the contracts modelled here, insurance changes who bears a covered financial loss. It does not change the initiating event rate, repair speed or service availability unless a separate mechanism is specified. Coverage can be valuable because it reduces retained severity, but deductible, occurrence limit, annual cap, exclusions and collection timing determine the actual benefit.
Evaluate the service and cash path alongside ultimate retained loss. A bank may still need to compensate customers or restore processing before reimbursement arrives. The policy example showed that an annual cap can have a small effect on expected payout and a large effect on tail protection. Read the complete contract transformation rather than applying one generic insured percentage to every year.
What does a million simulated years prove?
It supplies a large numerical sample from the model that was programmed, subject to implementation and random-sampling checks. It is not a million years of observed banking history. It does not establish the true tail family, reporting completeness or stability of the operating environment. A large simulation can make an answer very precise conditional on assumptions that remain uncertain.
Use simulation size to control numerical error, and use separate experiments to challenge parameters and model structure. Test simple cases with known answers before trusting a large run. An exact reproduction of the finite workshop distribution is a stronger implementation check than a smooth histogram with no independent benchmark.
Why can the system be online while the critical operation is still disrupted?
A restarted platform may still have a backlog, incomplete reconciliations, unavailable dependencies or insufficient processing capacity. The operation is what the user needs completed, not merely the technical state of one component. The backlog calculation showed that two hours of downtime could require eight further hours of catch-up when recovery capacity only modestly exceeds new arrivals.
Define the recovery endpoint before testing it. For one service it may mean critical instructions complete within their deadlines; for another it may require verified data integrity and restored normal capacity. A financial reimbursement or a responding server does not automatically meet that operational promise. Measure the operation directly.
How much mathematics must a beginner know before using this guide?
Begin with rates, percentages, averages, conditional probability and the ability to keep units consistent. The finite annual-loss table requires multiplying and adding probabilities, not advanced numerical software. The formulas for compound mean and variance become understandable once the reader separates variation in the count from variation in event size. More advanced distributions can follow after those distinctions are secure.
A beginner should be able to explain why thirty incidents in three million transactions can be safer per transaction than forty in two million, why an annual cap is not a per-event cap, and why a recovery next year cannot pay today’s bill. These are substantive mathematical achievements. Sophisticated notation should compress understanding already gained, not conceal its absence.
What is the strongest evidence that the model needs revision?
A direct contradiction between its assumptions and the process is especially important: the event grouping is wrong, observation rules changed, a supposedly independent control shares the same failure source, or a recovery requires conditions the model omits. Numerical symptoms such as unstable tail parameters or persistent prediction errors can help locate the issue, but the repair should address its cause rather than simply adjusting the final output.
Preserve the old model and document what new evidence changed the specification. A revision can improve the model without making every earlier decision unreasonable; the relevant distinction is what was known at each point. Equally, a favourable historical result should not protect an assumption that has now been shown to be false. The model serves understanding of the operation, not the defence of its previous score.
Teaching guide: build from a record to a defensible explanation
Begin by giving learners descriptions rather than distributions. A duplicated payment, a correctly executed investment that loses market value, a near miss, a recovered transfer and a continuing outage form a useful invented set. Ask which records belong to the chosen incident population and which require a different classification or analytical view. The first learning objective is not naming a probability family. It is defining the object that will be counted and explaining why the definition is consistent.
Adrian can maintain the incident register while Jo maintains the financial ledger. Give one incident several charges and one later recovery. Their task is to agree on the incident count, gross severity, ultimate net amount and temporary cash need without deleting information. Aisha checks that each included consequence appears once. A successful group can produce several correct summaries from the same records and explain why the summaries differ. An unsuccessful group often tries to force every purpose into one net number.
Move next to exposure-adjusted frequency. Provide two periods with different volumes and ask learners to compare both workload and per-transaction risk. Then change the transaction mix while improving each category’s rate. Ryan should calculate the actual overall count; Ben should calculate a standardised comparison using fixed weights. Require each to write a sentence describing the quantity calculated. This reveals whether the learner understands why a standardised performance statistic can improve while the real investigation queue grows.
Use the finite severity and count example before introducing simulation. Ask each learner to enumerate one possible annual outcome, including all incident orders that lead to the same total. Combine the outcomes into a complete probability table and check that its probabilities sum to one. Calculate the mean twice: directly from the annual table and through expected count times expected severity. Calculate the variance through both routes as well. Agreement demonstrates a reproducible connection between the elementary probability model and the compact compound formulas.
Mira should identify the worst one per cent using the table rather than a software quantile command. She needs to include all probability at the largest loss and only the required fraction at the next level. Ask why using every observation at that next level selects too much probability. This makes the discrete-boundary issue observable. The learner is ready for continuous and simulated tails when the probability mass in this small example is no longer being treated as a collection of vaguely large numbers.
Clara can then act as the policy reviewer. Give her the deductible, occurrence limit and annual cap, and ask for the retained amount in each annual outcome. Ethan checks the identity gross loss equals payout plus retained loss. Change only the annual cap and ask which statistics respond most. The model should show a small mean difference alongside a larger tail difference. Learners should explain that this is a property of where the cap binds, not an inconsistency or evidence that one of the calculations is wrong.
For a service exercise, separate the technical restart time from clearance of the backlog. One learner calculates tasks accumulated during downtime; another calculates net processing capacity after restart; a third checks critical deadlines. Add a scenario in which the replacement service shares a dependency with the original. The group must identify that missing mechanism rather than simply increasing a downtime number. The aim is to connect arithmetic to the operational promise and to recognise when the current model lacks a necessary component.
Finish with a decision rather than another calculation. Supply a control cost and a plausible range for effectiveness, event frequency and severity. Ask learners to identify the break-even condition, the assumptions under which the decision changes and the evidence worth obtaining next. They should distinguish a positive expected financial benefit from satisfaction of a service requirement. A strong answer can recommend a conditional action while clearly stating which evidence supports it and which unresolved uncertainty prevents a stronger conclusion.
For independent transfer, ask each learner to design a new synthetic incident model with a clearly specified reporting threshold and a recovery delay. A partner must reproduce the event count, annual mean and one chosen tail or service measure using only the written specification. Do not reward a model simply for having more parameters. Reward complete units, coherent observation rules, balanced financial consequences, a correct probability calculation and an honest statement of limits. These are the habits that remain useful when the arithmetic becomes much larger.
Sources and the boundaries of their use
The Basel operational-risk framework and its definitions chapter supply the regulatory context. The Basel framework timeline distinguishes effective and forthcoming changes. These references are not permission to replace jurisdiction-specific requirements with the hypothetical loss models in this guide. Publication date, effective date and local implementation must remain separate.
The 31 March 2021 Principles for operational resilience and BIS executive summary support the distinction between financial-loss measurement and delivery of critical operations. The FSB’s 15 April 2025 FIRE report concerns structured incident information and interoperability. It is used here for reporting context, not as a prescribed statistical estimation method.
NIST references for the Poisson distribution, lognormal distribution and generalized-Pareto tail model provide mathematical reference points. Distribution formulas do not establish that a particular organisation’s observations follow those families. All numerical parameter selections, finite distributions, control comparisons and service examples in this article are original hypothetical teaching constructions.
Shevchenko and Wüthrich’s Bayesian operational-risk study, research on truncation and parameter uncertainty, and Hadley, Joe and Nolde’s severity-selection study provide research routes for deeper statistical analysis. Their historical regulatory discussions should be read in their original context. Source versions were checked for this edition on 20 September 2026. No hypothetical event count in this guide should be interpreted as a confidential or publicly estimated loss history of any named institution.
Continue through the connected learning routes
Return to Banking And Finance Mathematics for the complete route. Strengthen the statistical foundation through probability and financial-risk distributions. Use VaR and Expected Shortfall mathematics for risk-measure comparisons, and stress and reverse-stress mathematics when the starting question is which conditions breach a chosen outcome.
For system-wide dependencies, continue to financial networks, contagion and systemic risk. For the payment clock, use payments and intraday-liquidity mathematics. The existing Finance & Banking Algorithms library supplies implementation-focused routes, including the separate Panjer-recursion guide. These owners address different questions rather than offering interchangeable versions of the same article.
The broader eduKate perspective is available in Operational Risk: How a Healthy Balance Sheet Can Still Suffer an Operational Failure and How Financial Systems Work. Those explanations connect the mathematics to the services people rely on. A healthy balance sheet, a functioning process and an available payment service support one another, but none should be used as a substitute measure for the others.
The final principle: measure the failure, the loss and the recovery separately
Operational risk mathematics is the disciplined connection between an incident, its consequences and the conditions under which a control changes them. Count the right event. Preserve the observation rule. Model the size of the consequence in the correct units. Join counts and severities with their dependence intact. Then distinguish avoided harm, transferred cost, available cash and restored service. Each is worth measuring, and each can tell a different truth about the same incident.
A useful model does not need to pretend that its rarest outcomes are known precisely. It needs to make its assumptions visible enough to test and its calculations clear enough to reproduce. The small tables, threshold corrections, recovery examples and capacity equations in this guide provide that foundation. Their purpose is not to produce an impressive number in isolation. It is to help a reader recognise what failed, calculate what follows and identify a repair whose benefit can be checked against the operation it is supposed to protect.
