Quick answer: operational risk is the risk of loss from failed or inadequate processes, people, systems or external events. Banks model it differently from market or credit risk because many important events are sparse, irregular and difficult to infer from recent averages. A strong framework therefore combines observed loss-event data, exposure proxies such as the Basel Business Indicator, control and risk assessments, scenario analysis for severe events, and operational-resilience testing that asks whether critical services can continue through disruption.
A system can have enough capital to survive a loss and still fail its customers because the payment engine, data centre, people or controls stop working.
Why this belongs in mathematics
Operational risk combines counting processes, severity distributions, heavy tails, scenario probabilities, control effectiveness, dependency mapping and threshold design. Yet it also teaches where mathematics must stop pretending data are abundant. A bank can have millions of routine transactions and only a handful of truly catastrophic operational events. That makes the largest losses both important and statistically scarce.
The current Basel standard defines operational risk as loss resulting from inadequate or failed internal processes, people and systems, or from external events. Legal risk is included; strategic and reputational risk are not part of the Basel operational-risk definition. See the Basel operational-risk definitions.
1. Start with the loss event
An operational-risk dataset is built from events: a payment-processing failure, internal fraud, external fraud, legal settlement, system outage, cyber incident, execution error, damaged physical asset, third-party service failure or another event that causes measurable loss.
A useful event record normally needs more than the final money amount. It should preserve:
- event date and discovery date;
- business line and legal entity;
- event type and root cause;
- gross loss;
- recoveries and insurance where applicable;
- near misses and non-financial impacts where the framework captures them;
- control failures and remediation;
- links to related events that may really be one incident.
The European Banking Authority’s 2025 operational-risk-loss standards illustrate why taxonomy matters: if institutions classify events differently, aggregate statistics become difficult to compare and root causes disappear inside inconsistent labels.
See the EBA operational-risk loss taxonomy work.
2. Frequency and severity are different mathematical objects
Suppose a bank experiences many small processing errors and very few enormous legal or cyber losses. One way to describe the problem is:
Annual loss = sum of severities of all events occurring during the year.
A traditional statistical model might estimate:
- a frequency distribution for the number of events N, and
- a severity distribution for individual losses X1, X2, …, XN.
Total annual loss is then L = ΣXi. This is conceptually straightforward. The hard part is the tail: the event that occurs once in twenty years may matter more to solvency than thousands of everyday errors.
3. Why the largest operational losses are hard to estimate statistically
Suppose the bank has ten years of internal data and has never suffered a major cloud-provider outage. It cannot conclude that such an outage has zero probability. The absence of an event may mean the risk is low, but it can also mean the sample is short relative to the event’s frequency.
This is the operational-risk version of a familiar mathematical mistake: unobserved is not the same as impossible.
Historically, banks used internal loss data, external loss data, scenario analysis and business-environment/control factors in sophisticated internal models. Basel III replaced the earlier advanced measurement approaches for regulatory capital with a standardised approach, partly because operational-risk internal models produced excessive variability and comparability problems.
See the BIS executive summary of the operational-risk standardised approach.
4. The Basel Business Indicator
The current Basel regulatory capital framework uses a financial-statement-based Business Indicator (BI) as a proxy for the scale of operational risk. The BI combines three broad components calculated from financial statement information: interest/lease/dividend activity, services, and financial activity.
The BI is transformed into a Business Indicator Component (BIC) using regulatory marginal coefficients. The framework then uses an Internal Loss Multiplier (ILM) that can reflect historical loss experience, subject to the applicable implementation choices and rules.
At the conceptual level:
Operational Risk Capital = BIC × ILM
and operational-risk RWA equal 12.5 times the operational-risk capital requirement. See Basel OPE25.
5. Why a revenue proxy can make sense—and why it can fail
Larger and more complex banking activity generally creates more opportunities for operational failures. A business-volume proxy therefore has useful information. But BI is not a literal prediction of operational loss. Two banks of similar size can have radically different systems, outsourcing concentration, control quality, product complexity and cyber exposure.
That creates an important boundary:
Regulatory capital measurement is not the same thing as operational-risk management.
A bank can meet a capital requirement and still need detailed scenario analysis, control testing, resilience engineering and incident response.
6. Scenario analysis: model the event your data do not contain
Scenario analysis asks experienced business, technology and risk specialists to construct severe but plausible events and estimate their causes, frequency, impact and control response. Examples might include:
- a core-payment platform unavailable for many hours;
- a cyber incident corrupting customer or transaction data;
- a major third-party cloud or telecommunications dependency failing;
- a rogue employee bypassing limits over a long period;
- a legal or conduct event affecting a large customer population;
- simultaneous operational failures during a wider market shock.
The purpose is not imaginative pessimism. A useful scenario has a mechanism, evidence, assumptions, affected processes, recovery path and explicit controls. Old OCC material on advanced operational-risk measurement described scenario analysis as especially relevant where internal and external data are not sufficient to estimate rare, high-severity losses.
7. Control effectiveness must be connected to the event mechanism
Suppose a bank says that a four-eyes approval process reduces unauthorised-payment risk. The useful question is not merely “Does the control exist?” It is:
- Does the second approver receive independent information?
- Can one administrator change both the transaction and the approval log?
- Does the control operate during emergency processing?
- What happens when transaction volume spikes?
- Has the control ever detected the failure it is meant to detect?
A control that exists on paper but can be bypassed by the same failure mechanism provides false risk reduction.
8. Key risk indicators are early signals, not proof
Banks often monitor key risk indicators (KRIs): system downtime, failed transactions, access violations, staff turnover, unresolved audit issues, change failures, manual overrides, aged exceptions or vendor incidents.
A rising KRI can warn that control conditions are deteriorating. But it may not predict loss directly. Ten failed software changes can signal weak change management without implying that the eleventh will cause a S$100 million event.
This is why a good KRI framework connects indicator → mechanism → control weakness → potential event instead of ranking colourful dashboard numbers without causal meaning.
9. Operational resilience asks a different question
Operational-risk management asks how the bank identifies, controls and absorbs operational losses. Operational resilience asks whether the bank can continue delivering critical operations through disruption.
The Basel Committee defines operational resilience as the ability to deliver critical operations through disruption. Its principles emphasise mapping the people, processes, technology, facilities, information and third parties required to deliver critical services, then testing severe but plausible disruptions.
See Basel Principles for Operational Resilience and the revised 2026 Federal Reserve sound-practices guidance.
10. Dependency graphs reveal hidden single points of failure
A payment service can depend on an application, a database, an identity system, a telecommunications provider, a cloud region, several internal teams and one external vendor. Drawing only the application architecture can miss the human or third-party dependency that actually stops recovery.
A resilience graph can represent:
Critical service → process → application → data → infrastructure → people → third party → facility.
The mathematical value of the graph is not decoration. It allows the bank to identify nodes whose failure disconnects the service, shared dependencies used by several critical operations, and recovery paths that are not genuinely independent.
11. A worked scenario: payment outage
Imagine a bank processes 2 million payments per hour during peak periods. A database failure stops outbound payment processing for three hours. The immediate financial loss may be modest, but the operational state can deteriorate quickly:
- 6 million instructions accumulate.
- Customer service volume rises.
- Liquidity forecasts become inaccurate because expected payments have not left.
- Counterparties cannot distinguish delay from non-payment.
- Manual workarounds introduce duplicate-processing risk.
- Recovery creates a burst of queued payments that can overload downstream systems.
The loss is therefore not a single number at t=0. It is a dynamic process whose severity depends on queue growth, customer deadlines, liquidity effects and recovery speed.
12. Third-party concentration changes the risk surface
Modern banks rely heavily on external cloud, software, payment, telecommunications and data providers. Outsourcing a function does not outsource the bank’s need to understand whether the service can fail.
If several critical operations depend on one external provider, the bank may have a hidden concentration even if each internal system has redundancy. The current Federal Reserve operational-resilience guidance explicitly includes third-party risk among the interdependent disciplines required for resilience.
13. Creative-work lens: Rogue Trader and the danger of one person becoming the control system
Rogue Trader, based on the collapse of Barings, is not an operational-risk model. It is useful because it makes one control-design problem memorable: if the same person can create positions, influence records and conceal the independent evidence meant to challenge those positions, separation of duties has failed structurally.
The creative work helps us notice the human mechanism. A serious bank must then replace narrative with access logs, independent valuation, role permissions, reconciliation and audit evidence.
14. The operational-risk algorithmic pipeline
- Define the operational-risk taxonomy.
- Capture internal loss events and near misses.
- Reconcile losses to accounting and incident systems.
- Map critical processes and controls.
- Collect KRIs and control-test results.
- Use external events to challenge internal-data scarcity.
- Run severe but plausible scenarios.
- Calculate regulatory operational-risk capital under the applicable Basel implementation.
- Map critical operations and dependencies for resilience.
- Set disruption tolerances and recovery objectives.
- Test business continuity and technology recovery.
- Aggregate repeated root causes across apparently separate incidents.
- Escalate weaknesses that can propagate across critical services.
- Update controls and scenarios after every material incident.
15. Failure modes
- Data-only modelling. No historical catastrophe means the bank assumes no catastrophe is plausible.
- Capital=resilience confusion. Capital is treated as proof that services can continue during an outage.
- Taxonomy drift. Similar events are classified differently, hiding recurring root causes.
- Paper controls. A control is counted because it exists, not because it interrupts the failure mechanism.
- Scenario theatre. Workshops describe dramatic events without connecting them to data, dependencies or remediation.
- Third-party blindness. Internal systems are redundant but both depend on the same vendor.
- Recovery overload. The restart plan ignores the backlog created during downtime.
- Average-loss comfort. Routine losses dominate metrics while rare high-severity risks remain invisible.
16. Diagnostics and falsifiers
- Which critical operation has the largest number of shared dependencies?
- Which loss events were reclassified after root-cause review?
- What severe event has never occurred internally but has happened at a peer?
- How quickly does a transaction backlog grow during outage?
- Which control can be bypassed by the same administrator it is meant to constrain?
- Which third party supports more than one critical operation?
- Does the bank’s scenario loss change dramatically if recovery takes twice as long?
- Can an independent test reproduce the BI/BIC and operational-loss calculations?
Suppose someone claims, “We have never had a major cyber outage, so the model should assign it negligible risk.” A falsifier is credible external evidence showing that comparable institutions have suffered severe outages through a dependency the bank also uses. Internal absence alone cannot prove low tail risk.
17. Verification and update triggers
- reconcile operational-loss records with finance and legal records;
- independently review event taxonomy and root-cause coding;
- test KRIs against realised incidents rather than dashboard aesthetics;
- repeat scenarios after major technology or vendor changes;
- run resilience tests that actually interrupt services or credible test environments;
- update dependency maps after migrations and outsourcing;
- review whether remediation closes the cause rather than only the incident ticket;
- feed lessons from near misses into controls before a loss occurs.
Connections across the finance-and-banking algorithms lane
- Bank capital models — where operational-risk capital becomes part of total RWA.
- Transaction reconciliation — a control that can both prevent and detect operational failures.
- Payment systems and queues — useful for understanding backlog propagation during outages.
- A Formula Can Be Correct and Still Be the Wrong Model — the core warning when sparse tail data produce false confidence.
Research anchors
- Basel Framework — operational-risk standardised approach.
- Basel Committee — Principles for the Sound Management of Operational Risk.
- Basel Committee — Principles for Operational Resilience.
- Federal Reserve — Sound Practices to Strengthen Operational Resilience, revised June 2026.
- EBA — operational-risk loss taxonomy and data standards.
The deeper lesson
Operational risk is where finance discovers that not every important uncertainty is a price, interest rate or default probability. Sometimes the risk is whether the machine keeps working. Loss data tell the bank what has happened. Scenarios ask what the data have not yet seen. Controls try to interrupt the mechanism. Resilience asks whether the service survives anyway. The strongest model therefore connects numbers to the actual process that can fail.
Educational note: This article explains public banking and operational-risk concepts. It is not cybersecurity advice, legal advice, operational instructions for a specific bank, or regulatory guidance for any institution.
