Score Scaling & Points to Double the Odds: Turning PD Models into Operational Credit Scores

Entimema
Editorial artwork for Score Scaling and Points to Double the Odds showing continuous glass risk curves resolving into calibrated metallic score markers.
Contents

A model probability is not yet an operational credit score. Score scaling is the mathematical and governance layer that converts model odds into a stable, interpretable and executable score architecture.

From a statistical probability to an operational language

logit(PDi) = zi = β0 + ΣjβjXij
PDi = 1/(1+exp(−zi))
Logistic model

A model may return PD=3.7% or logit(PD)=−3.26 while a decision system expects Score=642. The translation is not cosmetic: direction, odds convention, rounding and mapping determine whether development and production make the same decision.

Oddsbad = PD/(1−PD)
Oddsgood:bad = (1−PD)/PD = 1/Oddsbad
Name the odds convention

Both conventions are valid. Mixing them is catastrophic. In the traditional higher-score-is-lower-risk architecture, Score↑ implies PD↓. A higher-score-is-higher-risk system is also possible, but model, scale, documentation, cut-off and monitoring must all preserve the same direction.

01Logistic model02Log-odds03Good:Bad odds04Base score / odds / PDO05Scaled score06Variable points07Score-to-PD mapping08Cut-off09Production decision
Scaling carries one model signal through an explicit operational coordinate system.

PDO determines spacing; base score and odds anchor the scale

Score = Offset + Factor × ln(Oddsgood:bad)
Traditional Good:Bad scale

If adding PDO points doubles Good:Bad odds, then Score(2O)−Score(O)=Factor·ln(2)=PDO. Therefore:

Factor = PDO/ln(2)
Points to Double the Odds

At chosen base score S0 and base odds O0, S0=Offset+Factor·ln(O0), so:

Offset = S0 − Factor × ln(O0)
Base anchor

Original scale: 600 at 20:1 Good:Bad, PDO 50

Factor = 50/ln(2) = 72.134752
Offset = 600 − 72.134752ln(20) = 383.903595
Numerical constants
PDO and probability verification
ScoreGood:Bad oddsPD
55010:19.0909%
60020:14.7619%
65040:12.4390%
70080:11.2346%

PDO=20 compresses the same odds change into fewer points; PDO=100 expands it. Neither changes discrimination. Choice of scale should follow continuity, interfaces, reporting and strategy usability—not a universal ideal.

A score has no risk meaning without its inverse mapping

ln(O)=(S−Offset)/Factor
O=exp((S−Offset)/Factor)
PD(S)=1/[1+exp((S−Offset)/Factor)]
Inverse scale

The same model can be represented on 0–1000, 300–850 or 1–999. If S2=a+bS1 with b>0, borrower ranking is preserved. Equal numbers from different models, products or bureaus do not imply equal risk.

PDGood:Bad oddsScoreRisk grade
Each layer must remain traceable to the calibrated probability beneath it.

Grades and bands simplify reporting, pricing and strategy, but add discretisation. A common numerical scale across products is not automatically a common risk meaning.

Logistic coefficients become an additive scorecard

Because ln(Oddsgood:bad)=−logit(PD), the central implementation identity is:

Scorei = Offset − Factor(β0 + ΣjβjXij)
Model to score
Pointsij = −Factor·βjXij
BaseContribution = Offset − Factor·β0
Variable and base points

For a WoE scorecard, Xij=WoEij, so every governed bin receives points. The base contribution may remain central or be divided across n variables as BaseContribution/n. Distribution is an implementation convention, not a statistical necessity; changing conventions breaks reconciliation.

Original fictional points table—selected bins
Variable / binWoEββ×WoEPoints
DTI: ≤35%+0.80−0.65−0.520+37.51
DTI: 35–50%+0.30−0.65−0.195+14.07
DTI: >65%−0.70−0.65+0.455−32.82
Utilisation: <30%+0.70−0.55−0.385+27.77
Utilisation: 60–80%−0.40−0.55+0.220−15.87
Utilisation: >80%−0.85−0.55+0.468−33.73
History: clean+0.65−0.80−0.520+37.51
History: mild arrears−0.20−0.80+0.160−11.54
History: serious delinquency−0.90−0.80+0.720−51.94
Tenure: <1 year−0.55−0.35+0.193−13.89
Tenure: 2–4 years+0.10−0.35−0.035+2.52
Tenure: 4+ years+0.435−0.35−0.152+10.98

One borrower reconciles from raw inputs to a score of 612

Use β0=−3.20 and the coefficients above. The base contribution is 383.903595−72.134752(−3.20)=614.7348.

Spreadsheet-style implementation reconciliation
VariableRaw valueBinWoEββ×WoERaw pointsRounded
DTI44%35–50%+0.300−0.65−0.1950+14.0663+14
Utilisation72%60–80%−0.400−0.55+0.2200−15.8696−16
Credit historyMild arrearsMild arrears−0.200−0.80+0.1600−11.5416−12
Relationship tenure5.2 years4+ years+0.435−0.35−0.1523+10.9815+11
z = −3.20−0.195+0.220+0.160−0.15225 = −3.16725
PD = 1/(1+e3.16725) = 4.0417%
Good:Bad odds = e3.16725 = 23.7421:1
Score = 383.903595 + 72.134752ln(23.7421) = 612.3724
Full-precision chain

Raw variable points sum to −2.3624; adding base 614.7348 produces 612.3724 subject to displayed precision. Rounded components sum to −3; rounded base 615 gives the operational integer score 612. High utilisation and weak history are the largest adverse drivers; long tenure contributes positively.

Rounding becomes decision architecture near a hard threshold

Retaining 37.46, rounding to 37, rounding to nearest integer and truncating are different rules. Rounding each component can differ from rounding the continuous total.

Original boundary example
ComponentContinuousIndividually rounded
Base500.49500
A40.4940
B39.4939
C39.9340
Total620.40619

At an approval cut-off of 620, rounding the final continuous score approves; summing rounded parts declines or refers. The institution must document the stage, method and tolerance.

Score ≥ 620 ⇔ PD ≤ 1/[1+exp((620−383.903595)/72.134752)] = 3.6509%
Cut-off equivalence on this scale

Scaling represents risk; decision strategy chooses the boundary. The correct sequence is Model → Calibration → Scaling → Economics → Cut-Off → Decision. See Credit Cut-Off Strategy.

Calibration and redevelopment can break familiar score meanings

If ranking survives but calibration changes, an institution can preserve score and update its score-to-PD lookup, or rescale the model. Neither is automatically superior. Governance must identify which artefact changed and how legacy rules remain valid.

Scoreold=650 and Scorenew=650 do not guarantee equal PD. New-to-old alignment can use PD-equivalent, percentile or odds-equivalent mapping; naïvely matching numeric ranges disguises migration. PD Ranking & Calibration separates order from absolute risk.

Every result should retain model version, scale version, base score, base odds, PDO, calibration version and effective date. Caps and floors may fit interfaces but can conceal extreme-risk differences.

Production validity is demonstrated record by record

Scoredevelopment = Scoreproduction within documented tolerance
Reconciliation standard
Minimum implementation controls
ControlRequired evidence
Odds and directionNamed Good:Bad/Bad:Good convention; monotonicity test
TransformationsExact bin inclusivity, WoE version and categorical mappings
BoundariesFor every b, test b−ε, b and b+ε through bin, points and final score
Missing / unseenDedicated missing points; governed Other, fallback or exception
UnitsAnnual/monthly, currency, percentage/decimal and timing contracts
ArithmeticIntercept architecture, Factor, Offset, precision and rounding stage
Test populationMinimum/maximum risk, every bin, cut-off neighbourhood and exceptions

Development income of 36,000 annual is not production income of 3,000 monthly merely because the borrower is economically identical. A mathematically correct scale cannot rescue a wrong input contract. Never let missing values silently become zero or an unseen category acquire arbitrary points.

Monitor the score as an operating system, not a static number

Track mean, median, distribution, bands, approval by score, observed bad rate by score, score-to-PD mapping, overrides and population drift. For behavioural scores, St→St+1 migration may reveal deterioration before default. Population concentration can compress operational differentiation even when scale mathematics is unchanged.

PD Model Monitoring connects these signals to intervention. Low point contributions can support internal explanation, but customer-facing reason codes require governed business and legal mapping. Score≠FinalDecision: affordability, fraud, policy and manual review remain separate override layers.

Transparent scales help non-bank lenders—but tradition is not the objective

Fintech, instalment and consumer-finance lenders often benefit from understandable bands, lightweight systems and fast reconciliation between modelling and strategy teams. Yet a modern engine may consume score, calibrated PD, risk grade or expected loss. If direct PD decisioning is clearer, unnecessary legacy-style scaling adds governance surface without decision value.

The representation should serve the decision architecture. A score remains useful for familiarity, monotonic simplicity, banding, interfaces and explanation—not because 612 is intrinsically more informative than its governed 4.0417% PD.

Eighteen ways a mathematically valid model becomes a wrong score

Failure mechanisms
FailureMechanism
1. Odds convention mixedReciprocal is treated as the original; score direction flips
2. Direction reversedApprove rule selects higher risk
3. PDO formula wrongNumerical spacing no longer doubles odds
4. Base odds wrongEvery score receives the wrong risk anchor
5. Intercept mismatchCentral and distributed bases are double-counted or omitted
6. Different roundingDevelopment and engine diverge
7. Component roundingBorderline records cross the cut-off
8. Wrong boundariesBorrower enters a different bin
9. Wrong WoE versionPoints no longer represent fitted coefficients
10. Missing mismatchDesigned missing risk is silently replaced
11. Unseen categoryArbitrary points create uncontrolled decisions
12. Unit mismatchCorrect formula processes the wrong magnitude
13. Equal scores equatedDifferent models/products are assigned false common meaning
14. Score treated as PDAbsolute risk is inferred without the scale mapping
15. Score treated as decisionEconomics, policy and affordability disappear
16. Recalibration ignoredOld mapping misstates current PD
17. Legacy cut-off preservedRedeveloped model changes the boundary’s risk meaning
18. Distribution unmonitoredCompression, drift and strategy erosion remain invisible

A Scorecard Implementation & Reconciliation Agent can make control recurring

ENTIMEMA FRAMEWORKScorecard Implementation & Reconciliation AgentImplementation control + reconciliation + model-governance support—not borrower approval.
  1. Read coefficients and WoE specification
  2. Validate odds convention
  3. Calculate Factor and Offset
  4. Generate points tables
  5. Reconcile score and PD
  6. Test boundaries and rounding
  7. Challenge missing and unseen handling
  8. Produce validation evidence

Entimema's Credit Risk work connects scorecard development, scaling, calibration, implementation validation and monitoring. Decision Automation connects that governed risk representation to controlled production decision architecture.

For the upstream model mechanics, read Credit Scorecard Development, Weight of Evidence & Information Value and Logistic Regression for Credit Risk.