Model Calibration Drift: When a PD Model Still Ranks Risk Correctly but Predicts the Wrong Risk Level

Entimema
Editorial artwork for Model Calibration Drift showing an ordered glass sequence intersected by a shifted luminous reference plane.
Contents

A PD model can remain perfectly useful for ranking borrowers while becoming materially wrong about the absolute level of risk. Calibration drift appears when ordering survives but probabilities stop matching reality.

Discrimination and calibration are different model objects

Ranking asks whether PDa>PDb corresponds probabilistically to Riska>Riskb. Calibration asks whether borrowers assigned p default at approximately p over the governed horizon and population.

Perfect ordering, materially wrong level
Risk bandPredicted PDObserved default
A1%2%
B2%4%
C4%8%
D8%16%

Every band remains correctly ordered, so AUC/Gini can remain strong. Yet absolute risk is approximately doubled. A ranking metric cannot detect that error because monotonic probability transformations preserve ordering.

Intercept-type driftBroad level shift

All bands move up or down while relative shape remains similar.

Slope-type driftSpread changes

Risk separation becomes too steep, too flat or segment-dependent.

The first can preserve ordering; the second challenges differentiation itself.

O/E finds level error; intercept and slope diagnose its form

O/E = Observed defaults / Expected defaults = Ȳ / PD̄
Calibration-in-the-large

O/E≈1 indicates broad alignment; 1.5 means observed risk is 50% above expected; 0.7 means the model overpredicts. Portfolio O/E can conceal offsetting segment errors and says little about shape.

logit(P(Y=1)) = α + β·logit(PDmodel)
Ideal: α=0 and β=1
Calibration regression

α captures broad level bias conditional on the fitted calibration model. β captures whether the spread of predicted risk is appropriate. Under the stated convention, β<1 often signals overly extreme predictions and β>1 insufficient dispersion, but range, sample, estimation uncertainty and misspecification matter; never interpret slope as a slogan.

Brier = (1/N)Σ(PDi−Yi
LogLoss = −(1/N)Σ[Yiln(PDi)+(1−Yi)ln(1−PDi)]
Proper scoring rules

Brier and log loss evaluate probability quality but mix discrimination and calibration. Calibration curves compare predicted and observed rates by decile, score or PD band; grouping changes the picture. Hosmer–Lemeshow is grouping- and sample-size-sensitive: huge samples reject trivial deviations while small samples miss material ones.

A 250,000-account portfolio can rank well and understate risk

Expected defaults = 250,000×2.8% = 7,000
Observed defaults = 250,000×4.1% = 10,250
O/E = 10,250/7,000 = 1.464
Portfolio level
Original mature score-band diagnosis
BandAccountsPredicted PDObserved defaultO/E
A62,5000.8%1.2%1.50
B62,5001.6%2.3%1.44
C62,5003.0%4.4%1.47
D62,5005.8%8.5%1.47

AUC remains 0.76 versus 0.77 at validation, the estimated calibration slope is 0.98 and intercept is positive. The uniform band pattern supports an intercept-type challenger. It does not prove the cause or authorise production change: vintage, segment, policy, macro and data evidence still matter.

PSI and calibration answer orthogonal questions

Population: Pt(X) ≠ Pdev(X)
Calibration: Pt(Y|Score) ≠ Pdev(Y|Score)
Different forms of drift

A stable score distribution and low PSI can coexist with rising defaults after unemployment, interest-rate or income stress. Conversely, a riskier applicant mix can raise PSI and average PD while each score band remains correctly calibrated. Population movement is not calibration failure.

Likewise, higher realised loss can originate from PD, LGD or EAD: EL=PD×LGD×EAD. Funding costs and margins can change lending economics while PD remains calibrated. Keep probability validity separate from decision economics.

Aggregate alignment can hide segment and maturity failure

Intuitive aggregation trap
SegmentPredicted PDObserved PDDirection
Existing customers2.0%3.0%Underpredicted
New customers6.0%4.5%Overpredicted
Portfolio aggregate3.6%3.6%Appears aligned

Changing mix can make aggregate calibration appear correct while both segments are wrong—an intuitive Simpson's-paradox pattern. Diagnose product, channel, customer type, band, justified geography, risk grade and vintage, with uncertainty controls against noisy slicing.

A 12-month PD requires mature 12-month outcomes. Recent cohorts are censored; comparing annual PD with partial performance creates maturity bias. Use mature cohorts, disclose partial maturity, or use governed survival methods where appropriate. Low-default portfolios may require longer aggregation; high-risk short-tenor books can provide faster but more volatile feedback.

Recalibration method must match the diagnosed failure

Practitioner challenger set
MethodUseRisk
Intercept adjustmentBroad uniform level shiftMisses heterogeneous or shape drift
Logistic recalibrationEstimate α and β on logit(PD)Can overfit temporary conditions
Segment-specificStable, justified segment differencesFragmentation and weak samples
Isotonic regressionFlexible monotonic mappingStepwise overfit and extrapolation
Platt-style scalingParametric probability mappingFunctional-form mismatch
Prior/base-rate adjustmentSupported prevalence shiftInsufficient if P(Y|X) changed heterogeneously

Intercept-only recalibration is plausible when ranking, slope and ordering remain stable, a broad level shift persists and mature outcomes are sufficient. Champion/challenger comparisons should include existing, intercept-only, full logistic and justified segmented mappings on out-of-time mature samples.

Frequent rolling recalibration is responsive but can chase noise, destabilise PDs and expand governance burden. Structural policy, product, channel or regime breaks can invalidate simple adjustment.

Recalibration is not a substitute for redevelopment

01Predicted PD02Mature outcomes03O/E04Calibration intercept05Calibration slope06Segment / vintage07PSI context08Ranking stability09Materiality10Recalibrate / redevelop / monitor
Separate level, shape, population support and ranking before choosing action.
ENTIMEMA FRAMEWORKRecalibration versus redevelopmentA governed diagnosis, not an automatic threshold tree.
  1. Is ranking stable?
  2. Is calibration stable?
  3. Is population supported?
  4. Are variable relationships stable?
  5. Is drift segment-specific?
  6. Recalibrate, redevelop or monitor
  7. OOT validate

Falling Gini/AUC, unstable slope, variable-relationship drift, segment inversion, weak support, new product mechanics or structural policy change point beyond simple calibration. Data defects demand remediation, not statistical polishing.

Persistencet(k)=ΣI(|O/Et−h−1|>c)
Materiality=f(Deviation, Exposure, Duration, Decision sensitivity)
Persistence and materiality

Large samples make tiny deviations statistically significant; small segments create wide uncertainty. Assess confidence intervals, absolute PD error, exposure, expected-loss impact and decision impact.

A stable score cut-off can conceal a changed risk appetite

A monotonic score may remain unchanged while recalibration updates Score→PD. If score 620 once implied 5% PD and now implies 7.5%, preserving the numeric cut-off accepts a different absolute risk level. SameScore ≠ SameRisk across calibration versions.

Score Scaling & PDO explains the mapping; Cut-Off Strategy connects it to economics and appetite. Direct-PD strategies can change approval immediately after recalibration; score strategies require explicit remapping and governance review.

Understated PD understates expected loss all else equal, distorting pricing, provisioning, portfolio forecasts and capital planning. Application-scoring PD, regulatory PD and IFRS 9 lifetime PD have different horizons and calibration architectures; never transfer a recalibration mechanically between uses.

A four-quarter case changes diagnosis as evidence matures

Original multi-period calibration case
QuarterPredicted PDObserved defaultO/EGiniPSIαβResponse
Q12.9%3.0%1.030.470.03+0.031.01Monitor
Q23.2%3.4%1.060.470.18+0.050.99Diagnose mix; calibration stable
Q33.1%4.6%1.480.460.17+0.390.97Intercept challenger + vintage review
Q43.0%5.1%1.700.390.16+0.420.72Full relationship/model review

Q2 is population movement with stable performance. Q3 adds persistent level drift while ranking and slope largely survive: recalibration becomes a credible challenger. Q4 introduces slope and Gini deterioration, so intercept adjustment alone would hide structural failure.

Operational calibration control joins predictions to mature outcomes

ENTIMEMA FRAMEWORKCalibration monitoring workflowObserve → diagnose → challenge → validate → govern → monitor.
  1. Scores, PDs and mature outcomes
  2. O/E, intercept and slope
  3. Band, segment and vintage diagnosis
  4. PSI and materiality context
  5. Recalibration challengers
  6. OOT validation
  7. Governance and mapping update
  8. Production monitoring

Record model, calibration, score-scale and effective-date versions; test old and new score-to-PD maps; quantify cut-off, approval, expected-loss and segment impacts; preserve rollback and historical reconstruction. PD Model Monitoring provides the broader surveillance layer, and PD Ranking & Calibration develops the core validation distinction.

Fast-turning non-bank portfolios can learn sooner—and confound faster

Fintech, consumer-finance and instalment lenders may see faster turnover, higher default incidence and shorter outcome horizons, enabling responsive calibration monitoring. But frequent pricing, policy and channel changes alter selection and mix, making cause attribution harder. High-risk tails require granular evidence; low-default books require patience and pooled uncertainty.

Where calibrated parameters feed impairment, calibration drift can affect ECL. This does not make application PD equivalent to IFRS 9 lifetime PD; horizon, forward-looking scenarios and use-specific architecture remain distinct.

Seventeen failures turn recalibration into false assurance

Calibration failure mechanisms
FailureWhy it fails
1. Gini onlyOrdering hides probability error
2. Stable rank means fineAbsolute risk may be wrong
3. Portfolio O/E onlySegment errors cancel
4. Slope ignoredShape deterioration is missed
5. Segment drift ignoredMaterial local failure survives aggregation
6. Immature cohortsCensoring biases observed risk
7. Wrong horizonOutcome and PD are incomparable
8. PSI confused with calibrationInputs substitute for outcomes
9. Loss drift equals PD driftLGD/EAD are ignored
10. No OOT validationNoise is fitted and deployed
11. Too frequentPDs chase volatility
12. Cut-off preserved blindlyRisk appetite changes silently
13. Appetite impact ignoredTechnical mapping bypasses economics
14. One mapping for allHeterogeneous segments remain wrong
15. Significance equals materialityExposure and decision effect disappear
16. Recalibrate bad rankingStructural model failure is cosmetically masked
17. Point estimates onlyUncertainty is hidden

A PD Calibration & Drift Agent can make challenger evidence recurring

ENTIMEMA FRAMEWORKPD Calibration & Drift AgentContinuous calibration surveillance + challenger analysis + governance evidence—not autonomous production recalibration.
  1. Ingest predicted PDs
  2. Match mature outcomes
  3. Calculate O/E, α and β
  4. Build calibration curves
  5. Compare segments and vintages
  6. Join PSI and Gini
  7. Detect persistence
  8. Test challenger mappings
  9. Estimate cut-off impact
  10. Produce review evidence

Entimema's Credit Risk work connects PD calibration, validation, monitoring and redevelopment diagnostics. Decision Automation can operationalise controlled surveillance while accountable humans approve model changes.

Related foundations include Credit Scorecard Development, Logistic Regression for Credit Risk and Early Warning Indicators.