Model Calibration Drift: When a PD Model Still Ranks Risk Correctly but Predicts the Wrong Risk Level

Contents
A PD model can remain perfectly useful for ranking borrowers while becoming materially wrong about the absolute level of risk. Calibration drift appears when ordering survives but probabilities stop matching reality.
Discrimination and calibration are different model objects
Ranking asks whether PDa>PDb corresponds probabilistically to Riska>Riskb. Calibration asks whether borrowers assigned p default at approximately p over the governed horizon and population.
| Risk band | Predicted PD | Observed default |
|---|---|---|
| A | 1% | 2% |
| B | 2% | 4% |
| C | 4% | 8% |
| D | 8% | 16% |
Every band remains correctly ordered, so AUC/Gini can remain strong. Yet absolute risk is approximately doubled. A ranking metric cannot detect that error because monotonic probability transformations preserve ordering.
All bands move up or down while relative shape remains similar.
Risk separation becomes too steep, too flat or segment-dependent.
O/E finds level error; intercept and slope diagnose its form
O/E≈1 indicates broad alignment; 1.5 means observed risk is 50% above expected; 0.7 means the model overpredicts. Portfolio O/E can conceal offsetting segment errors and says little about shape.
Ideal: α=0 and β=1
α captures broad level bias conditional on the fitted calibration model. β captures whether the spread of predicted risk is appropriate. Under the stated convention, β<1 often signals overly extreme predictions and β>1 insufficient dispersion, but range, sample, estimation uncertainty and misspecification matter; never interpret slope as a slogan.
LogLoss = −(1/N)Σ[Yiln(PDi)+(1−Yi)ln(1−PDi)]
Brier and log loss evaluate probability quality but mix discrimination and calibration. Calibration curves compare predicted and observed rates by decile, score or PD band; grouping changes the picture. Hosmer–Lemeshow is grouping- and sample-size-sensitive: huge samples reject trivial deviations while small samples miss material ones.
A 250,000-account portfolio can rank well and understate risk
Observed defaults = 250,000×4.1% = 10,250
O/E = 10,250/7,000 = 1.464
| Band | Accounts | Predicted PD | Observed default | O/E |
|---|---|---|---|---|
| A | 62,500 | 0.8% | 1.2% | 1.50 |
| B | 62,500 | 1.6% | 2.3% | 1.44 |
| C | 62,500 | 3.0% | 4.4% | 1.47 |
| D | 62,500 | 5.8% | 8.5% | 1.47 |
AUC remains 0.76 versus 0.77 at validation, the estimated calibration slope is 0.98 and intercept is positive. The uniform band pattern supports an intercept-type challenger. It does not prove the cause or authorise production change: vintage, segment, policy, macro and data evidence still matter.
PSI and calibration answer orthogonal questions
Calibration: Pt(Y|Score) ≠ Pdev(Y|Score)
A stable score distribution and low PSI can coexist with rising defaults after unemployment, interest-rate or income stress. Conversely, a riskier applicant mix can raise PSI and average PD while each score band remains correctly calibrated. Population movement is not calibration failure.
Likewise, higher realised loss can originate from PD, LGD or EAD: EL=PD×LGD×EAD. Funding costs and margins can change lending economics while PD remains calibrated. Keep probability validity separate from decision economics.
Aggregate alignment can hide segment and maturity failure
| Segment | Predicted PD | Observed PD | Direction |
|---|---|---|---|
| Existing customers | 2.0% | 3.0% | Underpredicted |
| New customers | 6.0% | 4.5% | Overpredicted |
| Portfolio aggregate | 3.6% | 3.6% | Appears aligned |
Changing mix can make aggregate calibration appear correct while both segments are wrong—an intuitive Simpson's-paradox pattern. Diagnose product, channel, customer type, band, justified geography, risk grade and vintage, with uncertainty controls against noisy slicing.
A 12-month PD requires mature 12-month outcomes. Recent cohorts are censored; comparing annual PD with partial performance creates maturity bias. Use mature cohorts, disclose partial maturity, or use governed survival methods where appropriate. Low-default portfolios may require longer aggregation; high-risk short-tenor books can provide faster but more volatile feedback.
Recalibration method must match the diagnosed failure
| Method | Use | Risk |
|---|---|---|
| Intercept adjustment | Broad uniform level shift | Misses heterogeneous or shape drift |
| Logistic recalibration | Estimate α and β on logit(PD) | Can overfit temporary conditions |
| Segment-specific | Stable, justified segment differences | Fragmentation and weak samples |
| Isotonic regression | Flexible monotonic mapping | Stepwise overfit and extrapolation |
| Platt-style scaling | Parametric probability mapping | Functional-form mismatch |
| Prior/base-rate adjustment | Supported prevalence shift | Insufficient if P(Y|X) changed heterogeneously |
Intercept-only recalibration is plausible when ranking, slope and ordering remain stable, a broad level shift persists and mature outcomes are sufficient. Champion/challenger comparisons should include existing, intercept-only, full logistic and justified segmented mappings on out-of-time mature samples.
Frequent rolling recalibration is responsive but can chase noise, destabilise PDs and expand governance burden. Structural policy, product, channel or regime breaks can invalidate simple adjustment.
Recalibration is not a substitute for redevelopment
- Is ranking stable?
- Is calibration stable?
- Is population supported?
- Are variable relationships stable?
- Is drift segment-specific?
- Recalibrate, redevelop or monitor
- OOT validate
Falling Gini/AUC, unstable slope, variable-relationship drift, segment inversion, weak support, new product mechanics or structural policy change point beyond simple calibration. Data defects demand remediation, not statistical polishing.
Materiality=f(Deviation, Exposure, Duration, Decision sensitivity)
Large samples make tiny deviations statistically significant; small segments create wide uncertainty. Assess confidence intervals, absolute PD error, exposure, expected-loss impact and decision impact.
A stable score cut-off can conceal a changed risk appetite
A monotonic score may remain unchanged while recalibration updates Score→PD. If score 620 once implied 5% PD and now implies 7.5%, preserving the numeric cut-off accepts a different absolute risk level. SameScore ≠ SameRisk across calibration versions.
Score Scaling & PDO explains the mapping; Cut-Off Strategy connects it to economics and appetite. Direct-PD strategies can change approval immediately after recalibration; score strategies require explicit remapping and governance review.
Understated PD understates expected loss all else equal, distorting pricing, provisioning, portfolio forecasts and capital planning. Application-scoring PD, regulatory PD and IFRS 9 lifetime PD have different horizons and calibration architectures; never transfer a recalibration mechanically between uses.
A four-quarter case changes diagnosis as evidence matures
| Quarter | Predicted PD | Observed default | O/E | Gini | PSI | α | β | Response |
|---|---|---|---|---|---|---|---|---|
| Q1 | 2.9% | 3.0% | 1.03 | 0.47 | 0.03 | +0.03 | 1.01 | Monitor |
| Q2 | 3.2% | 3.4% | 1.06 | 0.47 | 0.18 | +0.05 | 0.99 | Diagnose mix; calibration stable |
| Q3 | 3.1% | 4.6% | 1.48 | 0.46 | 0.17 | +0.39 | 0.97 | Intercept challenger + vintage review |
| Q4 | 3.0% | 5.1% | 1.70 | 0.39 | 0.16 | +0.42 | 0.72 | Full relationship/model review |
Q2 is population movement with stable performance. Q3 adds persistent level drift while ranking and slope largely survive: recalibration becomes a credible challenger. Q4 introduces slope and Gini deterioration, so intercept adjustment alone would hide structural failure.
Operational calibration control joins predictions to mature outcomes
- Scores, PDs and mature outcomes
- O/E, intercept and slope
- Band, segment and vintage diagnosis
- PSI and materiality context
- Recalibration challengers
- OOT validation
- Governance and mapping update
- Production monitoring
Record model, calibration, score-scale and effective-date versions; test old and new score-to-PD maps; quantify cut-off, approval, expected-loss and segment impacts; preserve rollback and historical reconstruction. PD Model Monitoring provides the broader surveillance layer, and PD Ranking & Calibration develops the core validation distinction.
Fast-turning non-bank portfolios can learn sooner—and confound faster
Fintech, consumer-finance and instalment lenders may see faster turnover, higher default incidence and shorter outcome horizons, enabling responsive calibration monitoring. But frequent pricing, policy and channel changes alter selection and mix, making cause attribution harder. High-risk tails require granular evidence; low-default books require patience and pooled uncertainty.
Where calibrated parameters feed impairment, calibration drift can affect ECL. This does not make application PD equivalent to IFRS 9 lifetime PD; horizon, forward-looking scenarios and use-specific architecture remain distinct.
Seventeen failures turn recalibration into false assurance
| Failure | Why it fails |
|---|---|
| 1. Gini only | Ordering hides probability error |
| 2. Stable rank means fine | Absolute risk may be wrong |
| 3. Portfolio O/E only | Segment errors cancel |
| 4. Slope ignored | Shape deterioration is missed |
| 5. Segment drift ignored | Material local failure survives aggregation |
| 6. Immature cohorts | Censoring biases observed risk |
| 7. Wrong horizon | Outcome and PD are incomparable |
| 8. PSI confused with calibration | Inputs substitute for outcomes |
| 9. Loss drift equals PD drift | LGD/EAD are ignored |
| 10. No OOT validation | Noise is fitted and deployed |
| 11. Too frequent | PDs chase volatility |
| 12. Cut-off preserved blindly | Risk appetite changes silently |
| 13. Appetite impact ignored | Technical mapping bypasses economics |
| 14. One mapping for all | Heterogeneous segments remain wrong |
| 15. Significance equals materiality | Exposure and decision effect disappear |
| 16. Recalibrate bad ranking | Structural model failure is cosmetically masked |
| 17. Point estimates only | Uncertainty is hidden |
A PD Calibration & Drift Agent can make challenger evidence recurring
- Ingest predicted PDs
- Match mature outcomes
- Calculate O/E, α and β
- Build calibration curves
- Compare segments and vintages
- Join PSI and Gini
- Detect persistence
- Test challenger mappings
- Estimate cut-off impact
- Produce review evidence
Entimema's Credit Risk work connects PD calibration, validation, monitoring and redevelopment diagnostics. Decision Automation can operationalise controlled surveillance while accountable humans approve model changes.
Related foundations include Credit Scorecard Development, Logistic Regression for Credit Risk and Early Warning Indicators.


