Reject Inference: Learning Credit Risk from the Customers You Never Approved

Entimema
Editorial artwork for Reject Inference showing glass applicants separated by an opaque decision plane into observable and unobservable trajectories.
Contents

Reject inference is not primarily a technique for inventing outcomes for rejected applicants. It is a problem of learning under selective observation.

The lender must learn about 100,000 applicants from the 40,000 it chose to observe

Suppose a lender receives 100,000 applications. Historical strategy approves 40,000 and rejects 60,000. Repayment becomes observable for booked borrowers; for most rejects it never does. Yet the next scorecard and credit cut-off strategy must make decisions across the next full applicant population. That is a fundamental information problem, not a routine label-completion exercise.

Ai = 1 if approved; Ai = 0 if rejected
Yi = 1 if default; Yi = 0 otherwise
Approval and outcome indicators
Yi observed when Ai = 1; generally unobserved when Ai = 0
Selective observation

Development data therefore describe P(X,Y | A=1), while the intended model may need to describe P(X,Y). The two distributions coincide only under restrictive conditions. Approval changes the distribution of scores, income, affordability, channel, documentation, policy flags and latent risk that reaches the outcome window.

01Applicant population02Historical policy / score / rules03Accepted population04Observed outcomes05Future model
The previous decision strategy becomes embedded in the next model's training data.

The development dataset is not merely “history”. It is partly an artefact of historical decisions. When the resulting model drives the next policy, selection becomes a feedback loop.

Missingness determines which claims are defensible

Missing-outcome mechanisms in rejected-applicant modelling
MechanismMeaningCredit-risk implication
MCAROutcome observation is unrelated to observed or unobserved informationRarely credible: lending decisions are deliberately selective
MAR conditional on XOnce recorded variables are known, approval contains no further information about YCan support adjustment if X reconstructs selection sufficiently and overlap holds
MNARSelection still depends on unrecorded information or latent outcome riskObserved accepts alone cannot identify reject outcomes without stronger assumptions or evidence
Is P(A=1 | X,Y) adequately represented by P(A=1 | X)?
The practical missingness question

If approval used variables absent from the modelling extract—manual judgement, documents, fraud signals, affordability details, policy exceptions or an earlier score—conditioning on available X does not reconstruct the selection mechanism. Even rich data do not prove missing-at-random. They make the assumption more or less plausible.

Approval is not repayment performance

Rejected ⇏ Bad    and    Approved ⇏ Good
A decision is not an outcome

A customer may be rejected because of a conservative cut-off, affordability, eligibility, missing documentation, suspected fraud, operational capacity or manual underwriting. Approval merely records what a strategy did. Default records what happened after credit was extended under particular terms. Treating all rejects as bad confuses a policy label with an outcome label.

Approval propensity reveals where accepted outcomes can—and cannot—carry evidence

e(Xi) = P(Ai=1 | Xi)
Approval propensity

Applicants with e(X) near one are well represented among accepts. Applicants with e(X) near zero contribute little or no booked-loan evidence. Common support—also called positivity or overlap—asks whether accepted and rejected applicants coexist at comparable observed X.

P(A=1 | X=x) ≈ 0 ⇒ outcome estimation at x is assumption-dominated
Weak support
SUPPORTEDAccepts and rejects coexist

Adjustment may be estimable under explicit conditional assumptions.

THIN SUPPORTFew comparable accepts

Variance, model dependence and sensitivity rise.

UNSUPPORTEDPolicy observed almost nobody

Extrapolation cannot create empirical information.

This is why accepted-only modelling may remain useful inside the historical acceptance region yet fail when a lender opens a new channel, changes population, or lowers an approximate score cut-off from 620 to 580. The decisive question is: where does reliable evidence about applicants scoring 580–619 come from?

An original 120,000-applicant selection diagnostic

This fictional portfolio contains 48,003 observed accepted accounts and 71,997 rejects. Values are synthetic and illustrate support, not a recommended score policy.

Synthetic applicant distribution by score band
Score bandApplicationsApproval rateObserved acceptsBad rate among acceptsReject shareEmpirical support
720+15,00096%14,4001.0%4%Strong
680–71921,00086%18,0601.9%14%Strong
640–67925,00051%12,7503.8%49%Moderate
600–63927,0009%2,4307.6%91%Weak
560–59919,0001.5%28511.2%98.5%Very weak
<56013,0000.6%78Highly uncertain99.4%Minimal

The top bands offer abundant outcome evidence, but mainly about historically acceptable risks. Below 640, accepts become sparse. The 11.2% accepted bad rate in 560–599 is an estimate for just 285 unusually selected booked cases—not proof of the rejected majority's bad rate. Below 560, reporting a precise population rate would conceal the absence of support.

A model can still calculate a prediction there. The calculation's existence does not make the prediction identified, calibrated or decision-safe. A lender considering expansion should show the support profile beside every simulated approval, loss and value result.

Reject-inference methods trade different assumptions—not uncertainty for truth

Main reject-inference method families
MethodWhat it doesPrimary weakness
Hard classificationAccepted-only model labels each reject good or badCircular and falsely deterministic
ParcelingAllocates inferred goods and bads within bands using assumed deteriorationResult can be driven by the multiplier
Augmentation / reweightingWeights accepts by inverse approval propensityRequires conditional selection and overlap; unstable near zero
Fuzzy augmentationAdds probabilistic rather than hard synthetic outcomesStill inherits model and missingness assumptions
ExtrapolationExtends accepted outcome relationships into rejected regionsUnsupported regions are highly model-dependent
External / bureau performanceUses later performance elsewhere where lawful and availableProxy outcome may differ from performance on the lender's loan and terms
Controlled explorationApproves a governed sample near the boundary to create evidenceConsumes risk appetite and requires customer, capital and regulatory control

Parceling makes its assumption visible

BRR = λBRA;   BRA=6%;   λ=1.5 ⇒ BRR=9%
Illustrative parceling scenario

The multiplier λ is not learned from missing reject outcomes. It is a scenario assumption. A credible analysis compares plausible values, documents their basis and carries each into calibration and cut-off economics.

Reweighting cannot solve a lack of support

wi = 1 / e(Xi)
Inverse probability weight
e(Xi) → 0 ⇒ wi → ∞
Instability at the boundary

Large weights let a handful of accepts represent many unlike applicants, increasing variance and sensitivity to propensity misspecification. Trimming or stabilising weights can improve numerical behaviour but changes the target estimand. Report the weight distribution, effective sample size, trimmed population and unsupported regions—not merely a fitted model.

Synthetic labels often return the model's own beliefs

AcceptsAccepted-only modelPredict rejectsSynthetic outcomesRetrain
Retraining on labels generated by the accepted-only model may add less information than the larger row count suggests.

If an accepted-only model predicts rejects and those predictions become training labels, the enlarged dataset is not enlarged ground truth. Improved development AUC, Gini or fit may mostly demonstrate consistency with the original model. It does not prove improved knowledge of rejected applicants.

Validation is necessarily asymmetric

Useful evidence can come from out-of-time accepts, policy changes that later admitted marginal bands, external performance, controlled exploration samples and simulations. Each answers a narrower question. Accepted-only out-of-time validation tests transport within observed policy; it does not directly validate the reject labels. Validation must therefore challenge robustness of assumptions as well as observable performance.

Make uncertainty decision-relevant through sensitivity analysis

θ ∈ {θ1, θ2, …, θk}   ⇒   M(θ)
Scenario family
ENTIMEMA FRAMEWORKEntimema sensitivity architectureVary assumptions, then trace their consequences through the lending decision—not model metrics alone.
  1. Inference assumption
  2. Alternative model
  3. Ranking
  4. Calibration / PD
  5. Expected loss
  6. Cut-off economics
  7. Approval boundary
  8. Governance conclusion

Vary parceling deterioration, propensity specification, weight trimming and extrapolation form. Compare not only AUC but score ordering, calibrated PD, approval rate, expected loss, expected value and marginal approval bands. A method can alter population bad-rate calibration without changing rank; alternatively, synthetic labels can change rank with little observable support.

Reject inference → PD → EL → Pricing → Cut-off → Approval
Propagation into lending economics

If three defensible approaches create effectively the same strategy, methodological differences may have limited economic importance. If small assumption changes move the cut-off sharply, that instability is itself a governance result. A point estimate must not hide it.

Credit strategy is an endogenous learning system

Historical strategySelectionObserved dataModelNew strategyNew selectionNew data
Strategy controls which future outcomes become observable, so every policy is also an information policy.

Exploitation approves applicants currently believed to create value. Exploration collects information about uncertain populations. Lending cannot behave like unrestricted online experimentation: any marginal approval programme must respect affordability, regulation, fair customer treatment, capital, expected economics, risk appetite and human governance. But a policy that never explores outside historical boundaries can preserve uncertainty indefinitely.

Why consumer finance, fintech and non-bank lenders should care

Selection risk becomes especially material when policies change frequently, channels open rapidly, scorecards evolve, applicant mix shifts, thin-file borrowers matter or risk appetite expands. Fast feedback can help, but short cycles do not remove selective observation. Operational teams need a simple discipline: preserve decision lineage, map overlap before expansion, test bounded scenarios, and monitor newly booked marginal vintages separately.

When reject inference is not the right intervention

Do not add inference by default when future and historical acceptance regions are similar, overlap is effectively absent, data quality or rejection reasons are poor, external evidence contradicts the inferred pattern, or complexity does not change a decision. Sometimes the correct conclusion is: we do not know enough about this population. That is more defensible than false precision.

The Entimema reject-inference decision framework

ENTIMEMA FRAMEWORKFrom target population to decision impactMethod choice follows diagnosis of selection, support and assumptions.
  1. Target population
  2. Historical selection mechanism
  3. Observed / missing outcomes
  4. Overlap diagnosis
  5. Assumption set
  6. Inference method
  7. Sensitivity analysis
  8. Validation evidence
  9. Decision impact

Operational evidence required

Minimum practical data architecture
Data layerRequired evidence
ApplicationTimestamped applicant variables exactly as available at decision time
DecisionApprove, reject, review, override, offer and final booked status
Policy lineageScorecard version, cut-off, rule versions, manual route and channel
ReasonStructured rejection, fraud, affordability, eligibility and documentation reasons
OutcomeGoverned default definition, performance window, exposure and censoring
External evidenceLawful bureau/proxy outcomes with provenance and outcome differences
Application history → Decision + reason → Score / PD → Performance → Selection + overlap → Inference scenarios → Model comparison → Strategy simulation → Governance
Operational workflow

This connects upstream to Default Definition and PD Ranking & Calibration, then downstream to Cut-Off Strategy, Decision Automation, PD Model Monitoring and portfolio outcomes by vintage.

Fifteen failure modes that create false confidence

Reject-inference failure mechanisms
Failure modeWhy it fails
All rejects treated as badPolicy decisions are substituted for repayment outcomes
Missingness assumed randomDeliberate underwriting selection is ignored
Historical policy ignoredThe mechanism producing the sample is omitted
Rejection reasons ignoredRisk, fraud, affordability and eligibility selection are mixed
Extrapolation beyond supportPredictions are mistaken for empirical evidence
Extreme propensity weightsA few accepts dominate estimates and variance
Arbitrary parceling multiplierAn assumption becomes a concealed result
Circular synthetic labelsThe original model's beliefs are recycled as truth
Inferred labels evaluated as observedInternal consistency masquerades as validation
Ranking and calibration mixedA change in PD level is misreported as better ordering
Strategy change ignoredThe target population differs from the historical one
Population drift ignoredOld conditional relationships are transported unchallenged
Complexity treated as correctnessTechnical machinery substitutes for identification
Metrics optimised without decisionsNo approval, loss or value consequence is demonstrated
Point estimates hide uncertaintyGovernance never sees assumption sensitivity

A Credit Model Development & Strategy Agent should expose uncertainty—not invent borrower outcomes

A future agent could reconstruct historical acceptance policy, compare accepted and rejected populations, estimate approval propensity, diagnose common support, identify unsupported regions, run alternative inference scenarios, compare ranking and calibration, simulate cut-off implications and prepare sensitivity evidence for human review.

Its role is research automation + scenario analysis + methodological diagnostics + decision support. It should not invent borrower outcomes or autonomously approve or reject applicants. Deterministic calculations, versioned assumptions, evidence provenance, risk authority and human judgement must remain explicit.

Entimema's Credit Risk practice connects development-sample design, model calibration and strategy. Where a governed strategy is ready for production, Decision Automation can make its execution traceable without pretending that automation resolves missing evidence.