Reader question: A credit model can rank borrowers from safer to riskier, but how does a bank turn those ranks into actual one-year probabilities of default without confusing ordering skill with probability accuracy?
This article owns the probability-of-default calibration problem: rating grades or score bands + observed one-year default experience + representativeness and conservatism rules → calibrated one-year PDs that can be checked against future outcomes. It does not own the downstream Basel capital formula. That separate step is covered by Basel IRB credit-risk algorithms.
The central mathematical lesson is simple but important: discrimination and calibration are different jobs. A model can rank every borrower correctly and still assign the wrong probabilities. A score that orders A above B above C tells us relative risk. A PD of 0.4%, 1.2% or 4.8% makes an absolute probabilistic claim that can be tested against default frequencies.
This is public mathematical and algorithmic education. It does not assess any named borrower, recommend lending decisions, or provide personalized financial advice.
What a PD means in this lane
In the current Basel Framework, probability of default for IRB purposes is built around a one-year horizon. The framework requires grade-level PD estimates to be grounded in observed historical one-year default experience and long-run evidence. For corporate, sovereign and bank exposures, Basel CRE36 states that PD estimates are based on the observed historical average one-year default rate for the rating grade, using obligor counts rather than exposure weighting. The same chapter requires relevant data, representativeness, empirical grounding and periodic review.
That definition immediately creates an algorithmic contract:
- the event definition must be stable;
- the observation window must be known;
- the denominator must represent borrowers actually at risk;
- the horizon must be exactly specified;
- the grade assignment must be frozen at the correct observation point;
- defaults must be counted consistently;
- changes in underwriting, populations or rating systems must be reconciled;
- the final PD must be validated as a probability, not merely as a rank.
The basic cohort calculation
Suppose grade G contains Nt non-defaulted obligors at the start of year t, and Dt of them enter default during the next 12 months under the chosen default definition.
The observed one-year default rate is:
DRt = Dt / Nt.
If the grade has five annual observed rates of 0.5%, 0.8%, 0.4%, 1.2% and 0.6%, the simple arithmetic average is 0.7%.
That 0.7% is evidence about central tendency. It is not automatically the final calibrated PD. Before committing to it, the algorithm must test whether the historical years are representative, whether the current grade definition matches the historical grade definition, whether the period contains a reasonable mix of economic conditions, and whether sparse-data uncertainty requires conservative adjustment.
Why exposure weighting can answer the wrong question
Suppose a grade contains ten small borrowers and one huge borrower. If the large borrower defaults, an exposure-weighted default rate could be enormous even though one of eleven obligors defaulted. If five small borrowers default, the exposure-weighted result could look modest even though nearly half the borrowers failed.
For grade-level IRB PD estimation, Basel CRE36 explicitly focuses on obligor-count-based default experience rather than EAD weighting. This makes sense because PD is a probability of an obligor default event, whereas EAD belongs to a different dimension of credit loss.
Loss mathematics later combines separate objects:
Expected loss ≈ PD × LGD × EAD.
Mixing exposure size into the probability calibration stage can therefore double-count economic size that belongs elsewhere.
Ranking is not calibration
Consider two models applied to 10,000 borrowers.
Model A assigns PDs from 0.1% to 20% and ranks nearly every future default toward the risky end. Model B produces the same ranking but multiplies every probability by four.
The two models have almost identical discrimination. Their AUC or rank-ordering statistics may be indistinguishable. But Model B is badly miscalibrated if observed default frequencies remain near Model A’s probabilities.
This is why validation needs at least two independent questions:
- Discrimination: does the model order risk?
- Calibration: do stated probabilities match observed frequencies closely enough for the intended use?
Step 1: freeze the event and observation clock
A probability is meaningless if the event moves while the model is being measured.
The calibration dataset therefore needs a versioned default definition. If the bank changes its treatment of delinquency, restructuring, unlikeliness-to-pay indicators or cure, historical observations may no longer be directly comparable.
A robust implementation stores:
- obligor identifier;
- grade assignment date;
- grade or pool version;
- 12-month observation start and end;
- default flag;
- default date;
- default-definition version;
- exclusion reason if an exposure leaves the sample;
- source-system timestamp.
This makes the cohort reproducible rather than merely reportable.
Step 2: establish the long-run central tendency
A one-year PD is forward-looking, but its calibration is anchored in historical one-year outcomes. The key challenge is that one year can be unusually benign or unusually stressed.
If calibration uses only a calm year, it may understate the long-run default tendency. If it uses only a crisis year, it may overstate the central tendency for a framework intended to represent a broader mix of conditions.
Basel therefore requires long-run experience and, where longer relevant histories are available, expects them to be used. The current CRE36 text requires at least five years for at least one relevant PD data source and asks for a representative mix of good and bad years where relevant.
The algorithmic question is not simply “how many years do we have?” It is:
Does this history represent the population and underwriting process we are calibrating today?
Step 3: map borrowers into grades without leakage
Each borrower must be assigned using information available at the rating date. Future delinquency, later financial statements or post-default information cannot leak backward into the score.
Leakage can make a model appear exceptionally discriminating and well calibrated in backtests while failing in real deployment.
A clean backtest therefore reconstructs the state as it existed at the decision date and uses the rating engine version that was actually available then.
Step 4: preserve monotonic risk ordering
If Grade 1 is intended to be safer than Grade 2, and Grade 2 safer than Grade 3, the calibrated PDs should normally preserve that order:
PD1 ≤ PD2 ≤ PD3.
Raw observed rates can violate monotonicity because of small samples. A safer grade might show two defaults out of 100 while the next grade shows one out of 120 merely by chance.
Possible responses include:
- pooling adjacent grades where economically justified;
- isotonic or monotone calibration;
- Bayesian shrinkage toward a portfolio-level central tendency;
- larger time windows;
- additional conservatism for low-default grades.
The correct response depends on the regulatory and modelling context. The important point is that a monotonicity repair must be documented rather than hidden.
Step 5: separate overrides from the underlying model
Human overrides can be useful when the statistical model misses material information. But overrides also change the population that ends up inside each grade.
That means calibration must measure the grades after the actual assignment process, not an imaginary world in which overrides never occurred. The European Banking Authority has specifically addressed the need to analyse how overrides affect observed default frequencies and risk quantification.
A useful diagnostic records:
- model grade before override;
- final grade after override;
- override direction;
- override reason;
- subsequent default outcome.
If overrides consistently move risky borrowers into safer grades, calibration drift will eventually reveal it.
Step 6: add conservatism for uncertainty, not as decoration
Suppose a very safe grade contains only 40 borrowers and has observed zero defaults over several years. Zero observed defaults does not prove the true PD is zero.
The sample is simply too small to establish such precision.
This is a classic low-default-portfolio problem. Conservative adjustments, pooling, external evidence, Bayesian methods or confidence-bound reasoning can be used to avoid treating “not observed” as “impossible”.
Conservatism should have a traceable cause:
- limited observations;
- missing downturn years;
- population mismatch;
- rating-system change;
- default-definition mismatch;
- data-quality defects;
- model uncertainty.
A flat add-on with no connection to an identified uncertainty is not a strong control.
A small calibration example
Imagine three grades with long-run observed default frequencies:
- Grade A: 0.3%;
- Grade B: 1.1%;
- Grade C: 4.2%.
The score model proposes 0.2%, 0.9% and 4.5%.
The first question is whether the ordering is correct. It is.
The second question is whether the level is correct. Grade A appears underpredicted relative to long-run experience; Grade B is also somewhat low; Grade C is slightly high.
A calibration layer might adjust the score-to-PD mapping so that grade-level predictions better align with observed central tendencies while preserving rank order. The exact method could be parametric, monotonic, Bayesian or rules-based depending on the model framework.
After recalibration, the system must be retested on data not used to fit the adjustment. Otherwise the apparent improvement may be only in-sample curve fitting.
Inputs and outputs
A production-shaped PD calibration engine can require:
- rating score or grade;
- grade-system version;
- obligor population definition;
- default definition and version;
- historical one-year cohorts;
- observed defaults;
- economic-period labels;
- override records;
- external or pooled data where justified;
- data-quality flags;
- calibration method;
- conservatism adjustments;
- minimum PD floors where applicable.
Outputs can include:
- grade-level calibrated PD;
- raw observed default rate;
- long-run central tendency;
- confidence interval or uncertainty band;
- conservatism component;
- override impact;
- calibration residual;
- monotonicity status;
- backtest status;
- data-representativeness flags.
Evidence polarity: what would increase confidence?
Evidence for confidence includes stable grade definitions, reproducible cohorts, sufficient history, reasonable representation of current underwriting, monotonic observed risk, predictions close to subsequent frequencies, consistent results across independent implementations, and conservative treatment of sparse grades.
Evidence against confidence includes frequent grade-definition changes, unexplained denominator shifts, default-definition drift, perfect in-sample fit with weak out-of-time calibration, unstable PDs under small sample changes, large override distortions, and repeated underprediction during stress periods.
Calibration diagnostics
- Calibration-in-the-large: compare total predicted defaults with total observed defaults.
- Grade-level expected/observed test: compare ΣPD with realized default counts by grade.
- Reliability plot: graph predicted PD bands against observed frequencies.
- Monotonicity test: confirm risk rises as grade quality deteriorates.
- Binomial uncertainty: compare observed defaults with a probability interval implied by the assigned PD and sample size.
- Time stability: inspect annual calibration rather than only the pooled average.
- Override backtest: compare outcomes before and after overrides.
- Population-stability test: detect changes in borrower mix or underwriting standards.
- Out-of-time validation: test a later period not used in calibration.
- Challenger calibration: compare against a second mapping method.
Counterexample: zero defaults does not imply zero PD
A grade with no defaults in a small sample can still have meaningful underlying risk. Treating empirical zero as probability zero would make the model infinitely confident from finite evidence.
The falsifier is straightforward: widen the observation window or combine comparable external evidence. If defaults appear, the zero-PD claim collapses immediately.
Counterexample: a perfect ranker can be a poor probability model
If every assigned PD is multiplied by ten, the borrower ordering is unchanged. Yet the absolute probabilities become unusable.
This is why AUC alone cannot validate PD calibration.
Counterexample: pooled averages can hide regime failure
Suppose a model predicts 1% every year. Actual defaults are 0.2%, 0.3%, 0.4%, 3.5% and 0.6%. The five-year average may look tolerable even though the model missed the stressed year badly.
Whether that is acceptable depends on the intended rating philosophy and framework, but the pooled average must not hide the time pattern. Annual and scenario-conditioned diagnostics are necessary.
Counterexample: a changed default definition breaks comparability
If Year 1 counts only hard payment default while Year 2 also captures unlikeliness-to-pay events, a higher observed default rate may reflect a definition change rather than a real deterioration in borrowers.
A versioned event definition is therefore a mathematical input, not an administrative footnote.
Counterexample: exposure weighting can distort grade PD
One very large obligor can dominate an EAD-weighted default frequency. That may be relevant for loss or concentration analysis but not for a grade-level obligor default probability. This is why the probability stage must preserve the unit of analysis.
Weak links
Denominator contamination. Borrowers not actually at risk for the full horizon are counted incorrectly.
Data leakage. Post-rating information enters the historical score reconstruction.
Grade drift. The same grade label changes meaning across model versions.
Override opacity. Manual changes are not preserved for backtesting.
Stress blindness. Long-run averages conceal systematic underprediction in adverse periods.
Low-default overconfidence. Zero or one observed default is treated as precise evidence.
Population mismatch. Historical borrowers are materially different from the current portfolio.
Double counting conservatism. Multiple add-ons compensate for the same uncertainty.
Alternatives and complements
Logistic recalibration can adjust intercept and slope when raw model probabilities are systematically too low or high.
Isotonic regression can preserve monotonicity without forcing a specific functional form.
Bayesian shrinkage is useful when grade-level data are sparse and information must be shared carefully across groups.
Survival models estimate timing as well as event probability. See the separate article in this series on Cox hazards and loan-default timing.
Rating-transition matrices model movements among discrete states and default rather than calibrating a single one-year PD directly. See credit-rating transition-matrix algorithms.
How this connects to the surrounding knowledge estate
This page supplies an upstream input to Basel IRB RWA mathematics. LGD estimation supplies a different risk component. Merton structural credit risk provides a market-structure challenger signal rather than a regulatory grade calibration. Gaussian-copula portfolio models require marginal PDs before dependence is added. The Basel output floor then constrains model-based RWA at a different layer.
What would falsify confidence?
Confidence should be withdrawn when reconstructed cohorts cannot be reproduced; when the grade system used in history is not comparable with the current system; when realized defaults repeatedly fall outside reasonable probability ranges without explanation; when monotonicity breaks persistently rather than randomly; when overrides cause systematic hidden drift; or when the model’s stated PDs fail out-of-time even though rank ordering remains strong.
Verification and update triggers
Preserve every calibration run with data extract date, cohort rules, default-definition version, rating version, observation window, borrower counts, defaults, calibration method, conservatism adjustments, override treatment and validation results. Re-open calibration after material underwriting changes, mergers, portfolio-mix shifts, definition-of-default changes, rating-model redevelopment, new downturn evidence, sustained calibration drift or material data repairs.
Primary and high-quality references
- Basel Committee on Banking Supervision, CRE36 — IRB approach: minimum requirements to use IRB approach, current Basel Framework chapter.
- Federal Reserve Board, Section 217.101 — Definitions, including the U.S. advanced-approaches definition of PD as an empirically based long-run average one-year default rate.
- European Banking Authority, Q&A 2019_5029, on PD calibration and the effect of rating overrides.
- Basel Committee on Banking Supervision, Stability of a through-the-cycle rating system during a financial crisis, for the distinction between point-in-time and through-the-cycle behaviour.
- Federal Reserve Board and OCC, SR 11-7 — Guidance on Model Risk Management, for validation, monitoring, limitations and independent challenge.
Educational boundary: This page explains how probability calibration works mathematically and operationally. It does not assign a PD to any real borrower, determine whether a loan should be made, or provide personal financial advice.
