Entimema

PD Model Observation and Performance Windows

Entimema
Contents

A probability-of-default model can be statistically excellent and still answer the wrong risk question. The failure may have happened before regression, WoE or variable selection—when the development sample was placed on the wrong clock.

The model begins with a prediction clock

Every PD model needs two temporal definitions. The observation window determines the historical information available at the decision or reference point. The performance window determines the subsequent period over which default is observed. Between them sits time zero: the instant at which prediction is made.

The architecture is simple to state and consequential to design:

WHAT THE MODEL MAY KNOWWHAT THE MODEL MUST PREDICT
HISTORICAL DATAObservation window
TIME ZERODecision / snapshot
FUTURE OUTCOMEPerformance windowOUTCOME
LEAKAGE
The PD model prediction clock: information belongs to the past; the target belongs to the future. Leakage occurs when that boundary is breached.

For an application scorecard, time zero may be the application decision. For a behavioural model, it may be a month-end portfolio snapshot or review date. Account opening can be appropriate for another purpose. These choices are not interchangeable: they define the economic question the model learns.

Time zero is an information boundary

An ambiguous reference date creates more than untidy data. It permits different observations to contain different information sets, starts outcome measurement from inconsistent points and makes rows that look comparable answer different questions.

The core rule—the model must not know the future—is therefore stricter than removing an explicit default flag. Leakage can enter through a status field updated after decision, delinquency or restructuring information recorded inside the performance period, subsequent collections activity, or an aggregate whose calculation window quietly crosses time zero. A feature can be historically named yet temporally unavailable.

This is precisely why leakage is dangerous: it often improves validation statistics. The model appears powerful because it has been given early evidence of the event it is supposed to predict. Once deployed at the real decision point, that information does not exist.

AVALID
ObservationTime zeroFull performance

Known before decision · outcome fully matured

BLEAKAGE
ObservationTime zeroPerformance

Predictor information crosses the boundary

CIMMATURE
ObservationTime zeroPartial

Outcome horizon has not completed

Good sample, leaked sample and immature sample: three rows can look complete in a table while carrying fundamentally different evidential status.

A compact application example

A lender wants to predict 12-month default risk at application. An applicant applies on 31 January 2025. The permitted information set ends that day; performance runs from 1 February 2025 to 31 January 2026. If a behavioural aggregate uses transactions through March 2025, the model is no longer an application-time model. Better Gini would not repair the interpretation—it would evidence an easier, contaminated question.

Now consider another applicant accepted on 31 October 2025 when the dataset is extracted on 31 January 2026. Only three months of performance exist. Treating that account as non-default simply because no default has yet appeared converts incomplete observation into a “good” label. The resulting bias is temporal, not statistical.

The performance horizon defines what “risk” means

A shorter and a longer performance window do not produce alternative versions of the same target. They ask different questions. Horizon selection must connect the business use of the PD, the model purpose, default emergence, portfolio seasoning, censoring and the amount of mature data available.

A one-year horizon is common in particular PD and regulatory contexts, but it is not a universal methodological answer. The correct design is the horizon that matches the model's stated purpose and is supported by comparable, sufficiently mature outcomes. Mixing six-, nine- and twelve-month performance in one binary target does not enlarge evidence; it changes the meaning of “non-default” across rows.

Seasoning makes recent lending look safer

Suppose Vintage A has been observed for 12 months and Vintage B for four. Their observed default rates are not directly comparable: B has had less time for risk to reveal itself. This is the same maturity problem viewed from two analytical directions. In model development it governs target eligibility; in Credit Vintage Analysis it governs cohort comparison.

The issue extends to censoring. Accounts that close, leave the observable system or reach extraction before the horizon ends do not automatically become good outcomes. Their status depends on an explicit eligibility and censoring treatment consistent with the prediction question.

More rows are not necessarily more information

Monthly behavioural snapshots can increase observations rapidly, but the same borrowers, accounts and macro environment recur. Overlapping performance windows generate correlated outcomes; repeated rows create serial dependence; one concentrated period may dominate the sample. The nominal row count can therefore rise much faster than independent information. Development and validation design should recognise clusters and time structure rather than treating every row as a fresh experiment.

The development sample is a temporal model of the decision problem

ENTIMEMA FRAMEWORKPD Model Time ArchitectureThe sample construction sequence, governed by five temporal controls.
  1. Data history
  2. Observation window
  3. Reference date / time zero
  4. Performance window
  5. Outcome
  6. Development sample
Information availabilitySeasoningCensoringEconomic representativenessData consistency

This architecture determines which facts exist, which outcome is allowed to emerge and which observations become eligible. The resulting sample is not a neutral extract fed into a model. It is a temporal representation of the decision itself.

Recency competes with representativeness

A sample drawn mainly from benign conditions may encode a risk structure that weakens under stress. Extending history can improve cycle coverage, yet older observations may represent discontinued products, policies, channels, customer behaviour or data definitions. Neither maximum history nor maximum recency is automatically correct.

The design question is: what economic conditions does the sample contain, and are its risk relationships relevant to the environment in which the model will operate? Periods should be selected and weighted with that tension visible, then challenged through temporal and out-of-time validation.

Temporal sample-design failure modes
FailureWhy it happensWhat it distortsHow performance misleads
Unclear time zeroRecords are assembled around convenient datesFeature availability and outcomes vary between observationsPerformance looks portable when the prediction question is not consistent
Predictor–outcome overlapAggregates cross the reference boundaryFuture behaviour enters the predictor setDiscrimination is overstated and collapses in production
Immature vintagesRecent accounts are labelled before the horizon completesDefaults have not had time to emergeRecent lending appears artificially safe
One-regime sampleAvailability or recency dominates sample selectionRisk relationships reflect a narrow macro stateValidation outside that state reveals instability
Repeated observationsMonthly snapshots are counted as independent rowsBorrowers, accounts and macro conditions recurNominal sample size exaggerates independent information
Changing definitionsDefault or source fields change across periodsThe target or predictors cease to mean the same thingTemporal change is mistaken for risk differentiation

Mixing application and behavioural logic deserves particular attention. Application variables describe information available at origination; behavioural variables describe an account after experience has accumulated. Combining them without a single coherent time zero changes the model's use case inside the sample. Likewise, a default definition or source-field definition that changes across periods can create apparent signal that is actually measurement drift.

Define the clock before the probability

01What decision will the model support?
02When is that decision made?
03What information exists at that moment?
04What future outcome must be predicted?
05How long must performance mature?
06Which observations are eligible?
07Are periods and definitions comparable?
08Does the sample represent the operating environment?

Only after this sequence should modelling transformations begin. ranking and calibration depend on a target whose horizon is coherent; PD model monitoring depends on development and current outcomes being temporally comparable. Default definition, WoE and IV, logistic regression, scorecard development and model validation all inherit the clock established here.

Dataset construction is also a recurring control problem. Analytical workflow automation can enforce reference dates, test feature availability, flag aggregates that cross time zero, check outcome maturity and vintage eligibility, identify repeated observations, and produce development-dataset QA. The AI Agents Library is a contextual route for such controlled workflows—not a substitute for deciding model purpose, representativeness or acceptable evidence.

In practice, observation and performance architecture requires alignment across model purpose, data, default definition, portfolio behaviour, validation and the business decision. Entimema's Credit Risk consulting connects those elements when development or redevelopment moves from methodology into implementation.