Reject Inference: Learning Credit Risk from the Customers You Never Approved

Contents
Reject inference is not primarily a technique for inventing outcomes for rejected applicants. It is a problem of learning under selective observation.
The lender must learn about 100,000 applicants from the 40,000 it chose to observe
Suppose a lender receives 100,000 applications. Historical strategy approves 40,000 and rejects 60,000. Repayment becomes observable for booked borrowers; for most rejects it never does. Yet the next scorecard and credit cut-off strategy must make decisions across the next full applicant population. That is a fundamental information problem, not a routine label-completion exercise.
Yi = 1 if default; Yi = 0 otherwise
Development data therefore describe P(X,Y | A=1), while the intended model may need to describe P(X,Y). The two distributions coincide only under restrictive conditions. Approval changes the distribution of scores, income, affordability, channel, documentation, policy flags and latent risk that reaches the outcome window.
The development dataset is not merely “history”. It is partly an artefact of historical decisions. When the resulting model drives the next policy, selection becomes a feedback loop.
Missingness determines which claims are defensible
| Mechanism | Meaning | Credit-risk implication |
|---|---|---|
| MCAR | Outcome observation is unrelated to observed or unobserved information | Rarely credible: lending decisions are deliberately selective |
| MAR conditional on X | Once recorded variables are known, approval contains no further information about Y | Can support adjustment if X reconstructs selection sufficiently and overlap holds |
| MNAR | Selection still depends on unrecorded information or latent outcome risk | Observed accepts alone cannot identify reject outcomes without stronger assumptions or evidence |
If approval used variables absent from the modelling extract—manual judgement, documents, fraud signals, affordability details, policy exceptions or an earlier score—conditioning on available X does not reconstruct the selection mechanism. Even rich data do not prove missing-at-random. They make the assumption more or less plausible.
Approval is not repayment performance
A customer may be rejected because of a conservative cut-off, affordability, eligibility, missing documentation, suspected fraud, operational capacity or manual underwriting. Approval merely records what a strategy did. Default records what happened after credit was extended under particular terms. Treating all rejects as bad confuses a policy label with an outcome label.
Approval propensity reveals where accepted outcomes can—and cannot—carry evidence
Applicants with e(X) near one are well represented among accepts. Applicants with e(X) near zero contribute little or no booked-loan evidence. Common support—also called positivity or overlap—asks whether accepted and rejected applicants coexist at comparable observed X.
Adjustment may be estimable under explicit conditional assumptions.
Variance, model dependence and sensitivity rise.
Extrapolation cannot create empirical information.
This is why accepted-only modelling may remain useful inside the historical acceptance region yet fail when a lender opens a new channel, changes population, or lowers an approximate score cut-off from 620 to 580. The decisive question is: where does reliable evidence about applicants scoring 580–619 come from?
An original 120,000-applicant selection diagnostic
This fictional portfolio contains 48,003 observed accepted accounts and 71,997 rejects. Values are synthetic and illustrate support, not a recommended score policy.
| Score band | Applications | Approval rate | Observed accepts | Bad rate among accepts | Reject share | Empirical support |
|---|---|---|---|---|---|---|
| 720+ | 15,000 | 96% | 14,400 | 1.0% | 4% | Strong |
| 680–719 | 21,000 | 86% | 18,060 | 1.9% | 14% | Strong |
| 640–679 | 25,000 | 51% | 12,750 | 3.8% | 49% | Moderate |
| 600–639 | 27,000 | 9% | 2,430 | 7.6% | 91% | Weak |
| 560–599 | 19,000 | 1.5% | 285 | 11.2% | 98.5% | Very weak |
| <560 | 13,000 | 0.6% | 78 | Highly uncertain | 99.4% | Minimal |
The top bands offer abundant outcome evidence, but mainly about historically acceptable risks. Below 640, accepts become sparse. The 11.2% accepted bad rate in 560–599 is an estimate for just 285 unusually selected booked cases—not proof of the rejected majority's bad rate. Below 560, reporting a precise population rate would conceal the absence of support.
A model can still calculate a prediction there. The calculation's existence does not make the prediction identified, calibrated or decision-safe. A lender considering expansion should show the support profile beside every simulated approval, loss and value result.
Reject-inference methods trade different assumptions—not uncertainty for truth
| Method | What it does | Primary weakness |
|---|---|---|
| Hard classification | Accepted-only model labels each reject good or bad | Circular and falsely deterministic |
| Parceling | Allocates inferred goods and bads within bands using assumed deterioration | Result can be driven by the multiplier |
| Augmentation / reweighting | Weights accepts by inverse approval propensity | Requires conditional selection and overlap; unstable near zero |
| Fuzzy augmentation | Adds probabilistic rather than hard synthetic outcomes | Still inherits model and missingness assumptions |
| Extrapolation | Extends accepted outcome relationships into rejected regions | Unsupported regions are highly model-dependent |
| External / bureau performance | Uses later performance elsewhere where lawful and available | Proxy outcome may differ from performance on the lender's loan and terms |
| Controlled exploration | Approves a governed sample near the boundary to create evidence | Consumes risk appetite and requires customer, capital and regulatory control |
Parceling makes its assumption visible
The multiplier λ is not learned from missing reject outcomes. It is a scenario assumption. A credible analysis compares plausible values, documents their basis and carries each into calibration and cut-off economics.
Reweighting cannot solve a lack of support
Large weights let a handful of accepts represent many unlike applicants, increasing variance and sensitivity to propensity misspecification. Trimming or stabilising weights can improve numerical behaviour but changes the target estimand. Report the weight distribution, effective sample size, trimmed population and unsupported regions—not merely a fitted model.
Synthetic labels often return the model's own beliefs
If an accepted-only model predicts rejects and those predictions become training labels, the enlarged dataset is not enlarged ground truth. Improved development AUC, Gini or fit may mostly demonstrate consistency with the original model. It does not prove improved knowledge of rejected applicants.
Validation is necessarily asymmetric
Useful evidence can come from out-of-time accepts, policy changes that later admitted marginal bands, external performance, controlled exploration samples and simulations. Each answers a narrower question. Accepted-only out-of-time validation tests transport within observed policy; it does not directly validate the reject labels. Validation must therefore challenge robustness of assumptions as well as observable performance.
Make uncertainty decision-relevant through sensitivity analysis
- Inference assumption
- Alternative model
- Ranking
- Calibration / PD
- Expected loss
- Cut-off economics
- Approval boundary
- Governance conclusion
Vary parceling deterioration, propensity specification, weight trimming and extrapolation form. Compare not only AUC but score ordering, calibrated PD, approval rate, expected loss, expected value and marginal approval bands. A method can alter population bad-rate calibration without changing rank; alternatively, synthetic labels can change rank with little observable support.
If three defensible approaches create effectively the same strategy, methodological differences may have limited economic importance. If small assumption changes move the cut-off sharply, that instability is itself a governance result. A point estimate must not hide it.
Credit strategy is an endogenous learning system
Exploitation approves applicants currently believed to create value. Exploration collects information about uncertain populations. Lending cannot behave like unrestricted online experimentation: any marginal approval programme must respect affordability, regulation, fair customer treatment, capital, expected economics, risk appetite and human governance. But a policy that never explores outside historical boundaries can preserve uncertainty indefinitely.
Why consumer finance, fintech and non-bank lenders should care
Selection risk becomes especially material when policies change frequently, channels open rapidly, scorecards evolve, applicant mix shifts, thin-file borrowers matter or risk appetite expands. Fast feedback can help, but short cycles do not remove selective observation. Operational teams need a simple discipline: preserve decision lineage, map overlap before expansion, test bounded scenarios, and monitor newly booked marginal vintages separately.
When reject inference is not the right intervention
Do not add inference by default when future and historical acceptance regions are similar, overlap is effectively absent, data quality or rejection reasons are poor, external evidence contradicts the inferred pattern, or complexity does not change a decision. Sometimes the correct conclusion is: we do not know enough about this population. That is more defensible than false precision.
The Entimema reject-inference decision framework
- Target population
- Historical selection mechanism
- Observed / missing outcomes
- Overlap diagnosis
- Assumption set
- Inference method
- Sensitivity analysis
- Validation evidence
- Decision impact
Operational evidence required
| Data layer | Required evidence |
|---|---|
| Application | Timestamped applicant variables exactly as available at decision time |
| Decision | Approve, reject, review, override, offer and final booked status |
| Policy lineage | Scorecard version, cut-off, rule versions, manual route and channel |
| Reason | Structured rejection, fraud, affordability, eligibility and documentation reasons |
| Outcome | Governed default definition, performance window, exposure and censoring |
| External evidence | Lawful bureau/proxy outcomes with provenance and outcome differences |
This connects upstream to Default Definition and PD Ranking & Calibration, then downstream to Cut-Off Strategy, Decision Automation, PD Model Monitoring and portfolio outcomes by vintage.
Fifteen failure modes that create false confidence
| Failure mode | Why it fails |
|---|---|
| All rejects treated as bad | Policy decisions are substituted for repayment outcomes |
| Missingness assumed random | Deliberate underwriting selection is ignored |
| Historical policy ignored | The mechanism producing the sample is omitted |
| Rejection reasons ignored | Risk, fraud, affordability and eligibility selection are mixed |
| Extrapolation beyond support | Predictions are mistaken for empirical evidence |
| Extreme propensity weights | A few accepts dominate estimates and variance |
| Arbitrary parceling multiplier | An assumption becomes a concealed result |
| Circular synthetic labels | The original model's beliefs are recycled as truth |
| Inferred labels evaluated as observed | Internal consistency masquerades as validation |
| Ranking and calibration mixed | A change in PD level is misreported as better ordering |
| Strategy change ignored | The target population differs from the historical one |
| Population drift ignored | Old conditional relationships are transported unchallenged |
| Complexity treated as correctness | Technical machinery substitutes for identification |
| Metrics optimised without decisions | No approval, loss or value consequence is demonstrated |
| Point estimates hide uncertainty | Governance never sees assumption sensitivity |
A Credit Model Development & Strategy Agent should expose uncertainty—not invent borrower outcomes
A future agent could reconstruct historical acceptance policy, compare accepted and rejected populations, estimate approval propensity, diagnose common support, identify unsupported regions, run alternative inference scenarios, compare ranking and calibration, simulate cut-off implications and prepare sensitivity evidence for human review.
Its role is research automation + scenario analysis + methodological diagnostics + decision support. It should not invent borrower outcomes or autonomously approve or reject applicants. Deterministic calculations, versioned assumptions, evidence provenance, risk authority and human judgement must remain explicit.
Entimema's Credit Risk practice connects development-sample design, model calibration and strategy. Where a governed strategy is ready for production, Decision Automation can make its execution traceable without pretending that automation resolves missing evidence.


