
Contents
A borrower can remain in exactly the same relative risk position while the probability attached to that position changes materially.
The same score can carry a different PD
Consider a borrower whose score has not changed.
The borrower remains in the same risk band. Their position relative to every other borrower in the portfolio is unchanged. Nothing in the ordering has moved.
Yet the probability of default assigned to that position rises from 3.5% to 5.0%.
How can the same risk ranking produce a different probability of default?
The answer is not necessarily model failure. It is evidence that two analytical functions—often compressed into one mental model—are doing different jobs.
The ranking establishes where the borrower sits relative to others. Calibration determines what that position means in absolute risk terms.
Ranking and calibration answer different questions
Ranking asks: Who is riskier?
A ranking model establishes relative risk order. In a simple four-band structure, it might produce:
A < B < C < D
Borrowers in A are expected to be safer than borrowers in B; B safer than C; and C safer than D. Measures of discrimination evaluate how effectively that ordering separates subsequent good and bad outcomes.
That ordering matters, but it does not tell us whether band C represents a 3.5% PD, a 5.0% PD or something else. As the preceding analysis on model discrimination and decision quality shows, strong ranking performance cannot establish the accuracy—or economic suitability—of the absolute risk estimate.
Calibration asks: How risky are they?
Calibration translates positions in the ranking into probabilities. The same four bands can preserve exactly the same order while their absolute PD levels change.
- A0.8%
- B1.7%
- C3.5%
- D7.0%
- A1.2%
- B2.5%
- C5.0%
- D9.5%
The ranking can remain intact while calibration moves. That distinction is central to PD model calibration: relative risk stability and absolute risk stability are not the same property.
Calibration needs an anchor
If ranking supplies the relative shape of risk, calibration needs an anchor for the portfolio's overall level of default risk—its central tendency.
Observed historical defaults are essential evidence, but they cannot be interpreted mechanically. An observation period may not represent the portfolio or environment in which the calibrated PDs will be used.
Practitioners therefore need to consider whether the experience remains representative, including:
- portfolio composition and acquisition mix;
- structural changes in products or customers;
- economic conditions during and after the observation period;
- changes in underwriting standards or policy;
- available forward-looking information;
- the uncertainty surrounding every estimate.
This is not an invitation to adjust PDs until they feel comfortable. It is a requirement to make the calibration framework explicit, evidence-based and reviewable.
From borrower data to credit decision
The distinction becomes operational once PD enters the decision system.
- DataBorrower evidence
- RankingRelative risk
- Risk orderA < B < C < D
- CalibrationBridge
- PDAbsolute risk
- DecisionEconomics & policy
PD may feed expected loss, pricing, limits, approval policy, capital allocation, provisioning and portfolio strategy. A calibration error can therefore propagate across the economics of a credit decision even when the ranking model remains highly discriminatory.
The appropriate response depends on the institution's objectives and controls, but the analytical requirement is universal: relative position, absolute risk and decision consequence must remain distinguishable.
Entimema's Credit Risk work connects portfolio behaviour, model evidence and policy so those distinctions remain visible in the resulting decision architecture.
Ranking locates the borrower in the risk hierarchy. Calibration assigns the level of risk that the institution is prepared to use for decisions.
Monitor when the meaning changes
Ranking stability and calibration stability do not necessarily deteriorate together.
A modern monitoring architecture should therefore observe population drift, discriminatory performance, observed default behaviour, calibration drift and downstream decision outcomes as related but distinct signals.
An intelligent monitoring agent could continuously connect those signals and identify situations where the ranking still works but the meaning of the ranking has changed.
That role does not replace model validation or human judgement. It makes emerging inconsistencies easier to investigate before they propagate silently through pricing, limits, policy and portfolio outcomes.


