
Contents
A PD model has deteriorated. The fastest response is to move its probabilities back towards observed default rates. That response is also capable of making a structurally obsolete model look repaired.
The opposite mistake is just as costly. If the model still separates relative risk well and only the absolute risk level has shifted, full redevelopment can add estimation uncertainty, implementation risk and governance burden without solving a deeper problem. The decision is not whether performance is poor. It is what kind of performance failed.
Recalibration and redevelopment repair different things
Recalibration preserves the model’s underlying score or ranking relationship while revising the mapping from that output to probability of default. It is plausible when ranking remains sufficiently useful, the population is relevant, structural relationships are stable enough, and the observed calibration shift can be estimated from representative outcomes.
Redevelopment revisits the model’s predictive architecture: variables, transformations, segmentation, functional form, interactions, estimation sample or modelling method. It becomes more plausible when relative ordering weakens, relationships no longer generalise, important populations sit outside the model’s support, or repeated adjustments cannot restore stable performance.
This is a continuum, not a clean binary. A concentrated segment may need a new submodel while the wider portfolio needs only recalibration. A decision-policy correction may improve outcomes without changing the model. A data defect must be repaired before either intervention is assessed. PD model monitoring detects and decomposes these signals; intervention begins only after that diagnostic work.
Four states organise the first decision
Discrimination and calibration answer different questions. Discrimination asks whether higher-risk borrowers still rank above lower-risk borrowers. Calibration asks whether predicted PD levels correspond to realised risk. Their combination provides a compact starting map.
No fundamental intervention indicated by these measures alone.
Ranking survives; test whether the probability mapping can be repaired.
Acceptable aggregate calibration may conceal deteriorating ordering.
Both relative and absolute risk performance have weakened.
Before action: verify data integrity · population relevance · segment behaviour · vintage maturity · decision-policy chronology
State 3 is easily missed. Aggregate predicted and observed default rates can align even while risk ordering deteriorates inside the portfolio. Offsetting errors may make the calibration total look acceptable. A level correction cannot restore lost separation, and acceptable O/E does not prove that account-level probabilities remain decision-useful. The distinction is developed further in PD ranking and calibration research.
Two similar symptoms, two different interventions
Consider a fictional consumer-credit model evaluated on mature, comparable twelve-month outcome windows. The figures are synthetic and demonstrate reasoning rather than universal action thresholds.
| Period | AUC | Predicted PD | Observed DR | O/E |
|---|---|---|---|---|
| Reference | 0.76 | 4.1% | 4.3% | 1.05 |
| Scenario A | 0.75 | 4.2% | 6.1% | 1.45 |
| Scenario B | 0.66 | 4.3% | 6.2% | 1.44 |
In Scenario A, AUC changes only modestly while observed risk rises far above predicted risk. The ranking relationship appears broadly intact; absolute risk has shifted. If the movement is persistent, outcomes are representative and calibration error is not hiding segment failure, recalibration is a credible candidate. It is not yet the decision.
Scenario B has almost the same aggregate calibration error, but AUC falls materially. A recalibration might improve portfolio-level O/E while preserving a weakened ordering underneath. The investigation now belongs at relationship, variable and segment level: rank ordering by band, stability of coefficients or feature effects, population support, missingness and policy selection. Redevelopment becomes materially more plausible.
Vintage evidence also matters. If deterioration is concentrated in recent originations after a cut-off or channel change, same-age vintage analysis can separate a new-book effect from broad model failure. Immature cohorts or temporary macro pressure can make an immediate intervention premature.
Recalibrate or rebuild? Follow the cause, not the alarm
Has calibration deteriorated?
Test persistence, outcome maturity, representativeness and error by segment.
RECALIBRATION MAY BECOME A CANDIDATEWhich relationships weakened?
Investigate variables, segmentation, functional form, population support and implementation.
STRUCTURAL INTERVENTION BECOMES MORE PLAUSIBLEIf deterioration is segment-specific, a portfolio-wide change may dilute the real problem. Investigate eligibility, overrides, feature coverage and a segment treatment before rebuilding everything. If a new cut-off, pricing strategy or acquisition campaign changed the booked population, assess the strategy and selection effect before blaming model methodology. If upstream data changed, restore implementation integrity before estimating any new parameters.
Governance escalation is not reserved for redevelopment. A material recalibration changes risk estimates and may affect approvals, pricing, provisions or capital. The required validation, approval and change classification depend on model use, materiality and the applicable regime. No universal statistical boundary can make that decision safely.
Choose the smallest intervention that repairs the actual failure
The model did not deteriorate in the abstract. Its data, population relevance, ranking, probability mapping, segment behaviour or operating environment deteriorated in a particular way. The defensible response is the smallest governed intervention that addresses that cause without concealing residual risk.
The recurring sequence is monitor → detect → diagnose → classify → prepare investigation → human decision. Calculation, historical comparison, exception prioritisation and evidence assembly can be systematically augmented. Model-change decisions, materiality judgements and accountability remain human-governed.
Diagnose before changing the model.
Entimema Credit Risk connects performance evidence, portfolio change and governance to a proportionate intervention.
Explore Credit Risk consulting →Make recurring diagnosis reproducible.
A future PD Model Monitoring Agent could prepare signals, classifications and investigation routes for human review. It is a product direction, not a live-agent claim.
Explore the AI Agents Library →Methodological boundary
The framework is original Entimema reasoning grounded in general model-risk principles. The 2026 US interagency model-risk guidance explicitly positions adjustment, recalibration and redevelopment as outcomes informed by validation; the EBA’s PD calibration guidance requires analysis of why realised and estimated default rates deviate; and the Basel Framework requires ongoing review of performance, stability, model relationships and outcomes. Institution-specific standards remain governed by portfolio, use and regulatory context.


