Weight of Evidence & Information Value: What They Measure, Where They Fail, and How to Use Them Properly

Contents
Weight of Evidence and Information Value are useful because they impose structure on messy borrower data. They become dangerous when practitioners mistake that structure for truth.
Weight of Evidence begins with a modelling decision: the bins
WoE is calculated only after the analyst divides a continuous variable or groups categories. Change the boundaries and the Goods/Bads distribution changes; so do WoE and IV. WoE is partly a property of the variable and partly a property of the binning architecture chosen by the modeller.
Binning can capture non-linearity, isolate missingness, reduce extreme-value influence, group categories, make risk shape interpretable and support stable production rules. It also discards granularity and local structure.
Fine classing discovers; coarse classing decides what to preserve
| Stage | Purpose | Questions |
|---|---|---|
| Fine classing | Create granular initial partitions | Where do bad rates move, reverse, become sparse or expose data defects? |
| Coarse classing | Combine groups into defensible production bins | Are counts, bads, risk direction, economic meaning and future stability adequate? |
Continuous variables may begin with quantiles, equal-width bands, supervised splits or domain boundaries. Categorical variables require attention to order, rare levels, high cardinality and business meaning. No method is universally best, and categories should not be combined merely because one grouping maximises IV.
WoE compares class distributions, not merely the bad rate inside a bin
Under this convention, positive WoE means the bin contains a greater share of all Goods than of all Bads; negative WoE means the opposite. The reverse ln(DistBad/DistGood) convention is equally workable. Consistency across target coding, coefficient direction, points and production matters more than which sign is chosen.
Bad rate describes risk among observations inside the bin. WoE locates that bin relative to the total Goods and total Bads across the sample. They usually move together, but answer different questions. WoE's connection to relative log odds helps explain its natural use in traditional logistic scorecards.
An original Debt-to-Income WoE example
This fictional accepted-account sample contains 10,000 Goods and 1,200 Bads. The Missing group is retained because absence may represent a separate information mechanism. Percentages and WoE are rounded.
| DTI bin | Goods | Bads | DistGood | DistBad | Bad rate | WoE |
|---|---|---|---|---|---|---|
| <20% | 2,700 | 90 | 27.00% | 7.50% | 3.23% | 1.281 |
| 20–35% | 3,000 | 200 | 30.00% | 16.67% | 6.25% | 0.588 |
| 35–50% | 2,300 | 330 | 23.00% | 27.50% | 12.55% | −0.179 |
| 50–65% | 1,200 | 300 | 12.00% | 25.00% | 20.00% | −0.734 |
| >65% | 500 | 220 | 5.00% | 18.33% | 30.56% | −1.299 |
| Missing | 300 | 60 | 3.00% | 5.00% | 16.67% | −0.511 |
The observed relationship deteriorates monotonically across known DTI bands. Missing DTI is riskier than 35–50% but less risky than 50–65%, showing why missingness should not be automatically placed at either extreme. The economic question is why DTI is missing and whether that process will persist.
IV aggregates separation created by this population, target, period and partition
IV = Σj IVj
The synthetic DTI example produces total IV of approximately 0.615. That describes substantial univariate separation in this constructed sample. It does not establish that DTI is lawful, stable, incremental, correctly timed or production-ready.
IV measures how differently Goods and Bads are distributed across the chosen bins. It depends on the population, default definition, observation period, binning and data quality. It is not a permanent intrinsic property of “DTI”. Common weak/moderate/strong bands can be orientation aids, never laws.
Target leakage, later delinquency, collection activity, post-origination status, tiny cells, granular optimisation and historical policy can all generate impressive IV. A variable may describe what the previous decision strategy allowed into the portfolio rather than pure applicant risk. Reject Inference develops that selective-observation problem.
Infinite or extreme WoE is often a warning about evidence quantity
If Goodj=0 or Badj=0, raw WoE is undefined or infinite. A bin with three Goods and zero Bads can look extraordinarily protective even though its evidence is negligible.
Possible responses include merging economically similar bins, imposing minimum observations and bads, or applying a documented continuity adjustment. Conceptually:
Smoothing prevents numerical infinity; it does not manufacture evidence. Arbitrary ε can conceal structural sparsity and materially change IV. Report the adjustment, sensitivity and effective counts.
Missing values deserve causal diagnosis
Missing may mean no credit history, unavailable bureau data, a new customer, a different channel or process failure. A distinct bin can be predictive, but a model built around a temporary system defect deteriorates when operations improve. Ask not only whether missingness predicts, but why it occurs and whether its mechanism will continue.
Monotonicity is a governance preference, not an economic law
Monotonic bad rates or WoE can improve interpretation, reduce local reversals and simplify governance. But real borrower relationships can be U-shaped, threshold-driven, segmented or interaction-dependent. Forcing monotonic bins can erase genuine structure.
Automated binning may maximise IV subject to monotonicity, minimum size and maximum bins. That is a constrained in-sample optimisation, not proof of future validity. A numerically optimal split can be economically implausible, driven by one period, or unstable under a small boundary change. Compare candidate partitions out of time and prefer boundaries that production can implement and explain.
Stable moderate information can be more valuable than unstable high information
For development, validation and out-of-time samples, compare direction, magnitude, bin rank order, population and bad rate. A high-IV variable whose strongest bin reverses risk order may be less defensible than a moderate-IV variable with consistent structure.
| Variable | Development IV | Temporal evidence | Selection implication |
|---|---|---|---|
| A — acquisition device signal | 0.54 | 0.60 → 0.51 → 0.12; channel break | High separation, weak durability |
| B — verified debt burden | 0.24 | 0.23 → 0.25 → 0.22; stable bins | Moderate, repeatable evidence |
| C — relationship tenure | 0.07 | Stable; adds value with utilisation interaction | Low standalone IV, useful incrementally |
Ranking only by IVA>IVB>IVC would favour the least durable variable and discard interaction value.
Test time, vintage and segment—not only one aggregate sample
Track IVt, WoEj,v and IVv across months and origination vintages. Credit Vintage Analysis helps separate cohort and underwriting regimes. Repeat by product, channel, new/existing customer and relevant risk segment. Aggregate predictiveness can conceal opposite subgroup relationships—a form of Simpson's paradox.
Univariate separation is not incremental multivariate value
Binning can approximate non-linearity, limit extreme-value influence and give categorical variables a common log-odds-related representation. It often supports stable coefficients, but cannot guarantee the correct logistic specification.
Total debt, debt-to-income, monthly obligations and utilisation may each have respectable IV while encoding overlapping capacity information. Review correlation, variable families, coefficient stability and incremental likelihood or validation performance. A low-IV variable can matter through an economically sensible interaction: Risk=f(X,Z), not merely f(X)+f(Z).
This article is the specialist binning and screening node beneath Credit Scorecard Development. The later multivariate stage must still govern coefficients, likelihood, interactions, calibration, stability and validation. Entimema's existing logistic scorecard engineering research develops that translation.
Population drift and risk-relationship drift are different diagnoses
More applicants can move into a high-risk bin while its WoE remains stable; the score distribution and approvals still change. Conversely, bin populations can stay stable while bad rates and WoE deteriorate. Monitor bin population, missing rate, bad rate, WoE and IV together. PSI may summarise population movement but does not diagnose relationship drift.
PD Model Monitoring should connect these characteristic signals to ranking, calibration, overrides and strategy outcomes. Rebinning may be justified by structural population change, new categories, products or unstable economic relationships—but it changes the model's transformation and often its effective ranking. It therefore requires versioning, validation, calibration review and governance.
The Entimema variable assessment framework
- Predictive separation
- Stability
- Interpretability
- Incremental value
- Data reliability
- Production availability
- Governance
- Raw variable
- Data quality
- Fine classing
- Risk inspection
- Coarse classing
- WoE
- IV
- Stability testing
- Redundancy analysis
- Multivariate testing
- Production review
WoE and IV are attractive—and especially fragile—in fast-changing non-bank portfolios
Consumer finance, digital and short-tenor lending, fintechs and other NBFIs may combine large application volumes, small modelling teams, rapid strategy changes and strong explainability needs. WoE/IV offers transparent variable diagnostics and deterministic deployment.
The same speed can destabilise static bins as channels, policy, product design and applicant mix move. Use shorter stability intervals where performance matures quickly, retain strategy chronology, compare vintages and channels, and refuse apparent precision where bads per bin remain small.
Seventeen failure modes that turn structure into false confidence
| Failure mode | Why it fails |
|---|---|
| Universal IV thresholds | Context-dependent separation is treated as law |
| Maximising IV only | In-sample fit replaces stability |
| Too many bins | Sparse partitions overfit local noise |
| Tiny high-WoE bins | Magnitude disguises weak evidence |
| Zero counts ignored | Infinite estimates enter the model |
| Arbitrary smoothing | A numerical fix hides structural sparsity |
| Monotonicity forced | Real non-monotonic economics is erased |
| Missingness mechanism ignored | Operational defects become borrower risk |
| Variables selected by IV alone | Suitability and incrementality disappear |
| Correlation ignored | Duplicated information destabilises coefficients |
| Interactions ignored | Useful conditional signal is discarded |
| Historical selection ignored | Old policy is mistaken for risk |
| Post-decision leakage | Future outcomes masquerade as prediction |
| OOT stability omitted | Temporary relationships reach production |
| Sign convention assumed universal | Transform and points logic reverse |
| Production rebinning treated casually | The effective model changes without validation |
| Separation confused with meaning | A statistical pattern lacks economic credibility |
A Credit Scorecard Development Agent should accelerate evidence—not mechanically choose variables
A specialised workflow could profile variables and missingness, propose fine and supervised classing, calculate WoE and IV, diagnose zero counts, test monotonicity, compare OOT and segment stability, screen redundancy and prepare candidate-variable reports.
Its role is analytical acceleration + diagnostics + evidence generation. Human modellers must judge economic meaning, lawful use, interactions, stability trade-offs and final model architecture.
Entimema's Credit Risk practice connects variable architecture to scorecard development, validation, redevelopment and production monitoring.


