## 6. A Neural Network with Four Input Variables (Features), Two Outcomes, and ݊ Hidden Layers.

## Source details

**Canonical URL:** [6. A Neural Network with Four Input Variables (Features), Two Outcomes, and ݊ Hidden Layers.](https://www.imf.org/-/media/files/publications/wp/2019/wpiea2019109.pdf)

## Other formats

- [Markdown version](/-/media/files/publications/wp/2019/wpiea2019109.pdf.md)
- [Structured JSON version](/-/media/files/publications/wp/2019/wpiea2019109.pdf.json)

---

### I. Introduction — purpose and contributions
- Paper objective: nontechnical analysis of how ML could enhance assessing credit risk of borrowers and implications for financial inclusion.
- Two steps taken:
  - Discusses fundamental challenges in credit risk assessment confronting FinTech and traditional lenders.
  - Provides a nontechnical primer on key ML concepts and common ML techniques used in credit analysis.
- Core framing: credit rating summarized by Probability of Default (PD) and Loss Given Default (LGD).
- Emphasis on the five Cs of credit (capacity, capital structure, coverage, character, conditions) as the practitioner framework ML can augment.

### II. The Five Cs of Credit — data sources and implications for ML
- Capacity
  - Key indicator: Debt-to-Income (DTI) ratio.
  - Alternative: model future path of business income using business model, infrastructure, key expertise, competitiveness, customer base, potential growth, technological advantage.
- Capital Structure
  - Higher capital ratios imply larger owner stake and cushion against default.
- Coverage
  - Loan coverage via pledged collateral or guarantors; challenge when assets lack formal market prices — need for asset-specific pricing models.
- Character
  - Track record of missed payments, fraud, legal expenses; often summarized in credit scores (e.g., FICO).
- Conditions
  - Macroeconomic and industry factors (growth rate, unemployment rate, inflation rate, interest rate, tariffs, regulations, financial cycle, geography); reflect systematic/nondiversifiable risk.

### III. Machine Learning (ML) methods — fundamentals and practice
- Fundamental concepts
  - ML objective: improve performance on a task through training experience; common loss metrics include mean squared errors (MSE) and mean absolute errors.
  - ML models are largely nonparametric (flexible, require large training data) versus parametric econometric models (parsimonious, interpretable, require smaller samples).
  - Main concern for nonparametric models: overfitting.
- Cross-validation and model selection
  - Train/test split to assess out-of-sample predictive power; overfitting identified by low train error and high test error.
  - k-fold cross-validation: example fivefold cross-validation for a sample with 120 observations; average MSE across partitions used as cross-validation MSE.
  - Hyper-parameters (tree size, learning rate, number of NN layers/nodes) are calibrated to minimize cross-validation prediction error; alternative three-way split (train, test, prediction) can be used to keep a final evaluation sample independent.
- Supervised vs. unsupervised learning
  - Supervised: labeled outcome available (most credit models, e.g., PD classification).
  - Unsupervised: no outcome labels; used for clustering and discovering similarities.
  - Relevancy condition: labeled dataset must be representative of loans the model will assess.
- Classification vs. regression
  - PD models = classification (binary default/no default).
  - LGD models = regression (quantitative outcome).

### IV. Prominent ML model classes — descriptions and numerical illustrations
- Tree-based models (Decision Trees, Random Forest, Gradient-Boosting Decision Trees)
  - Decision Trees
    - Example for LGD with features LTV and DTI: first split LTV threshold 0.80; when LTV<0.8, DTI threshold 0.50; when LTV>0.8, DTI threshold 0.36; for LTV<0.8 and DTI>0.5 further split LTV at 0.3; predicted LGD for LTV<0.30 and DTI>0.50 is 4 percent. Tree depth in example: 3.
    - Advantages: interpretability for small trees; handle multi-output. Disadvantages: prone to overfitting, greedy/local optimization, instability, bias when classes dominate.
  - Random Forest
    - Two principles: bagging (bootstrap subsamples) and decorrelation (consider subset of features at each node).
    - Rule of thumb: include about √p features at each split. Example numbers preserved: if there are 105 features, the method randomly chooses 10 (√105) features at each node.
    - Advantages: reduced overfitting relative to single trees, better generalization, handles correlated features. Disadvantage: complex, hard to interpret without feature importance metrics.
  - Extremely randomized trees (extension)
    - Random thresholds for splits and use full train dataset (no bootstrapping) to reduce computation time while keeping comparable performance.
  - Gradient-Boosting Decision Trees (GBDT)
    - Builds trees sequentially; each tree trained on residuals of previous tree. Final prediction is weighted sum of tree predictions with a learning rate parameter.
    - Empirical guidance: small learning rate (for example, less than 0.1) combined with a large number of trees improves out-of-sample performance relative to learning rate equaling one with fewer trees; cost: higher computation.
    - Stochastic gradient boosting: use random subsample of training data at each iteration (fraction between 0.5 and 0.8 recommended) to improve performance and reduce computation.
    - Tools: feature importance and Partial Dependence Plots (PDPs) to interpret marginal relationships.
- Support Vector Machines (SVMs)
  - Objective: find best separating hyperplane maximizing margin between classes; uses support vectors.
  - Kernels: linear, polynomial, radial basis function (RBF), or custom kernels to allow nonlinear boundaries.
  - Advantages: memory efficient, scale well to high-dimension, less prone to overfitting, work well with unstructured data (text, images). Disadvantage: difficult to interpret; do not directly provide probability estimates for PD.
- Neural Networks (NNs) and Deep Learning
  - Structure: input layer (features X1..X4 in illustrative figure), hidden layers producing nodes Z, activation functions (sigmoid, tanh, softmax, rectifiers).
  - Numerical illustration of parameter counts:
    - To evaluate each Z in layer 1: five coefficients (four weights and a bias) per node.
    - Five nodes in layer 1 → layer 1 has 25 parameters.
    - If layer n is the only other hidden layer with three nodes → 18 more parameters.
    - Two-layer NN total parameters in example: 40 parameters (compare to five coefficients for an ordinary linear regression).
  - Advantages: flexibility to learn complex nonlinear patterns, performance gains with large data. Disadvantages: black-box (hard to interpret), require large datasets and computational power, difficult to encode tacit/business knowledge into network structure.

### V. How ML differs from econometrics — practical contrasts
- ML focus: maximize out-of-sample predictive accuracy; include variables that improve prediction even if not statistically significant in traditional sense.
- Econometrics focus: causal identification, interpretability, statistical inference (treatment of selection bias, endogeneity).
- ML parameters (number of trees, margin parameter, NN depth) inform model flexibility but often lack business-meaningful interpretation.
- Relevance and generalizability issues: ML predictive validity depends on dataset representativeness for intended prediction tasks.

### VI. Strengths of ML-based lending (implications for financial inclusion)
- Cost-effective credit assessment for small borrowers: automation reduces sunk costs of underwriting and enables frequent small loans and monitoring.
- Access to alternative data sources (BigTech transaction/e-commerce data; payment transaction histories) can substitute for missing credit registries and financial reports.
- ML hardens soft information: large-scale inclusion of varied data types (text, image, location, social media) can convert soft signals into usable predictors.
- Better capture of nonlinearities and local patterns: tree-based partitioning and ensemble methods can identify creditworthy borrowers overlooked by traditional indicators.
- Mitigates information asymmetry: enhanced monitoring (text, image, video) and frequent screening can reduce adverse selection and moral hazard.
- Empirical mentions/examples from literature: studies documenting superior ML performance in certain lending contexts (Jagtiani & Lemieux (2017); Schweitzer and Barkley (2017)).

### VII. Weaknesses and risks of ML-based lending
- Risk of digital exclusion and redlining:
  - Training data may underrepresent historically excluded groups; using raw data can perpetuate or exacerbate unfair exclusion.
  - Mitigation: exclude discriminatory variables, monitor input feature sets, use feature-importance diagnostics.
- Consumer protection, ethics, and privacy concerns:
  - Opaqueness of ML decisions, difficulty detecting dominant factors among many risk drivers.
  - Supervisory actions: monitor input variables; require feature selection (e.g., LASSO) and oversight of most significant drivers.
- Limited handling of structural change:
  - ML models trained on historical data may fail when rapid structural changes occur; ML algorithms themselves cannot easily incorporate non-data intuition or tacit knowledge.
  - Need for analyst judgment to validate relevance of training data and detect structural breaks.
- Gaming and data manipulation:
  - Borrowers can fake indicators once they learn which variables drive scores (e.g., inflate social media connections), degrading predictive validity.
- Endogeneity and selection bias remain relevant concerns; practitioners should check samples and treatment effects.

### VIII. Policy and operational recommendations (explicit and implied)
- Data governance and quality
  - Ensure legal and technological feasibility to gather digitalized data reliably and reduce noisy/biased inputs.
  - Implement cybersecurity measures given sensitivity of credit information.
- Model governance and monitoring
  - Ongoing monitoring of ML models to detect structural changes, faked indicators, and unjustified drivers.
  - Use feature-importance tools and PDPs to assess and justify top drivers of credit ratings against business knowledge.
  - Consider three-way data splits (train/test/prediction) and robust cross-validation practices when tuning hyper-parameters.
- Regulatory and supervisory actions
  - Enforce exclusion of discriminatory variables from inputs and require disclosure of key model drivers where feasible.
  - Monitor adoption of ML in credit markets to guard against systemic risks while enabling financial inclusion.
- Market and infrastructure
  - Develop technological infrastructure for big-data decision-making and make it accessible to FinTech credit providers.
  - In Emerging Market Economies (EMEs), maximize benefits by building high-quality data and digital infrastructure; require sufficient capital buffers for lenders taking on higher-risk underserved borrowers.

### IX. Concluding remarks — net assessment
- ML and FinTech credit present promise to increase financial inclusion by lowering costs, hardening soft information, capturing nonlinearities, and mitigating information asymmetry, particularly for small borrowers and in markets with weak credit registries.
- Important weaknesses: potential for noisy-driven exclusion, difficulty handling structural change, susceptibility to gaming, and persisting econometric concerns (endogeneity/selection bias).
- EMEs can benefit substantially from FinTech credit if accompanied by reliable data collection, cybersecurity, infrastructure, and disciplined model oversight; lenders should hold sufficient buffers when expanding credit to underserved populations.
- Data requirement note: training ML models requires loan performance data over a full business cycle and a credit cycle to appropriately capture risk dynamics.

*Source: IMF working paper section titled "6. A Neural Network with Four Input Variables (Features), Two Outcomes, and ݊ Hidden Layers."*

### References .............................................................................................................

### References

### Figures

- 1. Fivefold Cross-Validation of a Sample with 120 Observations. ........................................ 12
- 2. Left: An Illustrative Decision Tree Model for Estimating LGD. Right: Partitioning of the Features Space Implied by the Estimated Tree. ...................................................................... 15
- 3. The Random Forest Model ................................................................................................. 17
- 4. An SVM Model for Predicting Default Outcome Based on Debt-To-Income (DTI) and Loan-To-Value (LTV) Ratios. ................................................................................................ 20
- 5. SVM With Linear and Nonlinear Kernels .......................................................................... 21

*Source: wpiea2019109 - References*

### 6. A Neural Network with Four Input Variables (Features), Two Outcomes, and ݊ Hidden

### 6. A Neural Network with Four Input Variables (Features), Two Outcomes, and ݊ Hidden Layers.

### I. Introduction — purpose and contributions
- Paper objective: nontechnical analysis of how ML could enhance assessing credit risk of borrowers and implications for financial inclusion.
- Two steps taken:
  - Discusses fundamental challenges in credit risk assessment confronting FinTech and traditional lenders.
  - Provides a nontechnical primer on key ML concepts and common ML techniques used in credit analysis.
- Core framing: credit rating summarized by Probability of Default (PD) and Loss Given Default (LGD).
- Emphasis on the five Cs of credit (capacity, capital structure, coverage, character, conditions) as the practitioner framework ML can augment.

### II. The Five Cs of Credit — data sources and implications for ML
- Capacity
  - Key indicator: Debt-to-Income (DTI) ratio.
  - Alternative: model future path of business income using business model, infrastructure, key expertise, competitiveness, customer base, potential growth, technological advantage.
- Capital Structure
  - Higher capital ratios imply larger owner stake and cushion against default.
- Coverage
  - Loan coverage via pledged collateral or guarantors; challenge when assets lack formal market prices — need for asset-specific pricing models.
- Character
  - Track record of missed payments, fraud, legal expenses; often summarized in credit scores (e.g., FICO).
- Conditions
  - Macroeconomic and industry factors (growth rate, unemployment rate, inflation rate, interest rate, tariffs, regulations, financial cycle, geography); reflect systematic/nondiversifiable risk.

### III. Machine Learning (ML) methods — fundamentals and practice
- Fundamental concepts
  - ML objective: improve performance on a task through training experience; common loss metrics include mean squared errors (MSE) and mean absolute errors.
  - ML models are largely nonparametric (flexible, require large training data) versus parametric econometric models (parsimonious, interpretable, require smaller samples).
  - Main concern for nonparametric models: overfitting.
- Cross-validation and model selection
  - Train/test split to assess out-of-sample predictive power; overfitting identified by low train error and high test error.
  - k-fold cross-validation: example fivefold cross-validation for a sample with 120 observations; average MSE across partitions used as cross-validation MSE.
  - Hyper-parameters (tree size, learning rate, number of NN layers/nodes) are calibrated to minimize cross-validation prediction error; alternative three-way split (train, test, prediction) can be used to keep a final evaluation sample independent.
- Supervised vs. unsupervised learning
  - Supervised: labeled outcome available (most credit models, e.g., PD classification).
  - Unsupervised: no outcome labels; used for clustering and discovering similarities.
  - Relevancy condition: labeled dataset must be representative of loans the model will assess.
- Classification vs. regression
  - PD models = classification (binary default/no default).
  - LGD models = regression (quantitative outcome).

### IV. Prominent ML model classes — descriptions and numerical illustrations
- Tree-based models (Decision Trees, Random Forest, Gradient-Boosting Decision Trees)
  - Decision Trees
    - Example for LGD with features LTV and DTI: first split LTV threshold 0.80; when LTV<0.8, DTI threshold 0.50; when LTV>0.8, DTI threshold 0.36; for LTV<0.8 and DTI>0.5 further split LTV at 0.3; predicted LGD for LTV<0.30 and DTI>0.50 is 4 percent. Tree depth in example: 3.
    - Advantages: interpretability for small trees; handle multi-output. Disadvantages: prone to overfitting, greedy/local optimization, instability, bias when classes dominate.
  - Random Forest
    - Two principles: bagging (bootstrap subsamples) and decorrelation (consider subset of features at each node).
    - Rule of thumb: include about √p features at each split. Example numbers preserved: if there are 105 features, the method randomly chooses 10 (√105) features at each node.
    - Advantages: reduced overfitting relative to single trees, better generalization, handles correlated features. Disadvantage: complex, hard to interpret without feature importance metrics.
  - Extremely randomized trees (extension)
    - Random thresholds for splits and use full train dataset (no bootstrapping) to reduce computation time while keeping comparable performance.
  - Gradient-Boosting Decision Trees (GBDT)
    - Builds trees sequentially; each tree trained on residuals of previous tree. Final prediction is weighted sum of tree predictions with a learning rate parameter.
    - Empirical guidance: small learning rate (for example, less than 0.1) combined with a large number of trees improves out-of-sample performance relative to learning rate equaling one with fewer trees; cost: higher computation.
    - Stochastic gradient boosting: use random subsample of training data at each iteration (fraction between 0.5 and 0.8 recommended) to improve performance and reduce computation.
    - Tools: feature importance and Partial Dependence Plots (PDPs) to interpret marginal relationships.
- Support Vector Machines (SVMs)
  - Objective: find best separating hyperplane maximizing margin between classes; uses support vectors.
  - Kernels: linear, polynomial, radial basis function (RBF), or custom kernels to allow nonlinear boundaries.
  - Advantages: memory efficient, scale well to high-dimension, less prone to overfitting, work well with unstructured data (text, images). Disadvantage: difficult to interpret; do not directly provide probability estimates for PD.
- Neural Networks (NNs) and Deep Learning
  - Structure: input layer (features X1..X4 in illustrative figure), hidden layers producing nodes Z, activation functions (sigmoid, tanh, softmax, rectifiers).
  - Numerical illustration of parameter counts:
    - To evaluate each Z in layer 1: five coefficients (four weights and a bias) per node.
    - Five nodes in layer 1 → layer 1 has 25 parameters.
    - If layer n is the only other hidden layer with three nodes → 18 more parameters.
    - Two-layer NN total parameters in example: 40 parameters (compare to five coefficients for an ordinary linear regression).
  - Advantages: flexibility to learn complex nonlinear patterns, performance gains with large data. Disadvantages: black-box (hard to interpret), require large datasets and computational power, difficult to encode tacit/business knowledge into network structure.

### V. How ML differs from econometrics — practical contrasts
- ML focus: maximize out-of-sample predictive accuracy; include variables that improve prediction even if not statistically significant in traditional sense.
- Econometrics focus: causal identification, interpretability, statistical inference (treatment of selection bias, endogeneity).
- ML parameters (number of trees, margin parameter, NN depth) inform model flexibility but often lack business-meaningful interpretation.
- Relevance and generalizability issues: ML predictive validity depends on dataset representativeness for intended prediction tasks.

### VI. Strengths of ML-based lending (implications for financial inclusion)
- Cost-effective credit assessment for small borrowers: automation reduces sunk costs of underwriting and enables frequent small loans and monitoring.
- Access to alternative data sources (BigTech transaction/e-commerce data; payment transaction histories) can substitute for missing credit registries and financial reports.
- ML hardens soft information: large-scale inclusion of varied data types (text, image, location, social media) can convert soft signals into usable predictors.
- Better capture of nonlinearities and local patterns: tree-based partitioning and ensemble methods can identify creditworthy borrowers overlooked by traditional indicators.
- Mitigates information asymmetry: enhanced monitoring (text, image, video) and frequent screening can reduce adverse selection and moral hazard.
- Empirical mentions/examples from literature: studies documenting superior ML performance in certain lending contexts (Jagtiani & Lemieux (2017); Schweitzer and Barkley (2017)).

### VII. Weaknesses and risks of ML-based lending
- Risk of digital exclusion and redlining:
  - Training data may underrepresent historically excluded groups; using raw data can perpetuate or exacerbate unfair exclusion.
  - Mitigation: exclude discriminatory variables, monitor input feature sets, use feature-importance diagnostics.
- Consumer protection, ethics, and privacy concerns:
  - Opaqueness of ML decisions, difficulty detecting dominant factors among many risk drivers.
  - Supervisory actions: monitor input variables; require feature selection (e.g., LASSO) and oversight of most significant drivers.
- Limited handling of structural change:
  - ML models trained on historical data may fail when rapid structural changes occur; ML algorithms themselves cannot easily incorporate non-data intuition or tacit knowledge.
  - Need for analyst judgment to validate relevance of training data and detect structural breaks.
- Gaming and data manipulation:
  - Borrowers can fake indicators once they learn which variables drive scores (e.g., inflate social media connections), degrading predictive validity.
- Endogeneity and selection bias remain relevant concerns; practitioners should check samples and treatment effects.

### VIII. Policy and operational recommendations (explicit and implied)
- Data governance and quality
  - Ensure legal and technological feasibility to gather digitalized data reliably and reduce noisy/biased inputs.
  - Implement cybersecurity measures given sensitivity of credit information.
- Model governance and monitoring
  - Ongoing monitoring of ML models to detect structural changes, faked indicators, and unjustified drivers.
  - Use feature-importance tools and PDPs to assess and justify top drivers of credit ratings against business knowledge.
  - Consider three-way data splits (train/test/prediction) and robust cross-validation practices when tuning hyper-parameters.
- Regulatory and supervisory actions
  - Enforce exclusion of discriminatory variables from inputs and require disclosure of key model drivers where feasible.
  - Monitor adoption of ML in credit markets to guard against systemic risks while enabling financial inclusion.
- Market and infrastructure
  - Develop technological infrastructure for big-data decision-making and make it accessible to FinTech credit providers.
  - In Emerging Market Economies (EMEs), maximize benefits by building high-quality data and digital infrastructure; require sufficient capital buffers for lenders taking on higher-risk underserved borrowers.

### IX. Concluding remarks — net assessment
- ML and FinTech credit present promise to increase financial inclusion by lowering costs, hardening soft information, capturing nonlinearities, and mitigating information asymmetry, particularly for small borrowers and in markets with weak credit registries.
- Important weaknesses: potential for noisy-driven exclusion, difficulty handling structural change, susceptibility to gaming, and persisting econometric concerns (endogeneity/selection bias).
- EMEs can benefit substantially from FinTech credit if accompanied by reliable data collection, cybersecurity, infrastructure, and disciplined model oversight; lenders should hold sufficient buffers when expanding credit to underserved populations.
- Data requirement note: training ML models requires loan performance data over a full business cycle and a credit cycle to appropriately capture risk dynamics.

*Source: IMF working paper section titled "6. A Neural Network with Four Input Variables (Features), Two Outcomes, and ݊ Hidden Layers."*

---


_Source: https://www.imf.org/-/media/files/publications/wp/2019/wpiea2019109.pdf_
