## 1. The Fundamental Problem of Causal Inference

## Source details

**Canonical URL:** [1. The Fundamental Problem of Causal Inference](https://www.imf.org/-/media/files/publications/wp/2019/wpiea2019228-print-pdf.pdf)

## Other formats

- [Markdown version](/-/media/files/publications/wp/2019/wpiea2019228-print-pdf.pdf.md)
- [Structured JSON version](/-/media/files/publications/wp/2019/wpiea2019228-print-pdf.pdf.json)

---

### Introduction
- ML adoption accelerated over the "past 10 years" for predictive capabilities, though its pedigree extends back "more than fifty years".
- Empirical economics often requires causal inference: estimating the unobserved counterfactual path that would have occurred absent a policy or event.
- Counterfactuals cannot be directly validated; this is the central challenge of causal inference (Holland, 1986).
- "Causal machine learning" aims to apply ML strengths to causal inference to produce more precise, less biased, and more reliable estimators of causal effects.
- Running example used throughout: assessing the impact of a hypothetical financial crisis on output growth, with focus on Causal Forests (Athey and others, 2018) and methods for interpreting complex models.

### Estimating Causal Relationships: Predicting the Counterfactual
- Randomized controlled experiments are the theoretical gold standard but often impractical, unethical, or impossible in economics.
- Observational data require constructing credible proxies for the unobservable counterfactual by conditioning on observed confounding factors.
- Conditional on observed confounders, treatment can be treated as "essentially as good as randomly assigned", allowing average differences to proxy causal effects.
- Common econometric strategies:
  - Standard regression by including pre-identified control variables (constructing a prediction of the counterfactual).
  - Propensity score approach (Rosenbaum and Rubin, 1983): estimate conditional probability of treatment given covariates; match on propensity score (Abadie and Imbens, 2006).
- Key limitation: counterfactual predictions can never be validated; emphasis on plausibility and persuasiveness.

- Box 1 summary (preserved exactly):
  - Causal inference entails arriving at a plausible estimate of an unobserved counterfactual.
  - Best possible construction: experiments with "comparable" subjects (e.g., identical twins as an idealized experiment).
  - Average outcome for untreated similar subjects can serve as counterfactual for treated subjects; differences give treatment effects.
  - Estimating heterogeneous effects requires matched experiments per subgroup; estimating average effects for heterogeneous populations requires random allocation (RCT).
  - In observational studies, must statistically account for systematic differences between groups.

### Machine Learning: Making Better Predictions
- ML focuses on predictive performance and methods to assess out-of-sample predictive accuracy (Box 2).
- ML strengths:
  - Discover complicated relationships not specified in advance, handling interactions and thresholds.
  - Help predict rare events by sifting many potential predictors.
- Random Forests (RF) highlighted for:
  - Handling binary, categorical, numerical predictors.
  - Implicit variable selection focusing on covariates with greatest predictive power.
  - Applicability to classification, regression, and cluster analysis.
  - Generally strong predictive performance and lower training effort in many settings.
- No single algorithm dominates all applications ("no free lunch theorem"); RF usually performs well.

- Box 2 summary (preserved exactly):
  - Overfitting: good in-sample fit may model idiosyncratic noise; out-of-sample performance matters.
  - Regularization prevents overfitting by adding a penalty dependent on coefficient magnitudes.
  - LASSO (Least Absolute Shrinkage and Selection Operator) can set some coefficients exactly to zero, effectively performing variable selection.
  - Tuning parameter lambda (λ) controls penalty strength; higher λ implies greater regularization.
  - Model validation methods:
    - Holdout validation: split data into training and testing sets; use test set to compute validation error.
    - Cross validation: divide data into K folds (example K=3); rotate test/training folds to obtain validation errors and choose hyperparameters (e.g., λ) that minimize cross-validation error.
  - Best practice often uses an additional quarantined evaluation set as a final gauge of likely future performance.

### An Emerging Synthesis: "Causal" Machine Learning
- Historical divide:
  - Prediction asks what usually happens given circumstances.
  - Causal inference asks what would happen under an intervention.
  - Econometrics emphasized explanation, parsimony, interpretability, and in-sample fit; ML emphasized predictive accuracy and out-of-sample performance.
- Convergence: many estimation problems decompose into prediction-like steps, enabling ML integration into causal workflows.
- Examples of ML aiding causal estimation:
  - IV problems with many potential instruments: use ML (e.g., LASSO) to select instruments (Belloni and others, 2012).
  - Exogenous treatment with many confounders:
    - Solution A: LASSO-style regression to select confounders predictive of the outcome.
    - Solution B: use ML to estimate propensity scores with covariates selected for predicting inclusion in treatment (Wyss and others, 2014).
  - Double-selection procedure (Belloni and others, 2014b): run LASSO for outcome predictors and separately for treatment predictors; include union in OLS.
- Primary focus of these ML modifications: estimating the average treatment effect (ATE).
- Interest in heterogeneous treatment effects (HTE) motivates ML use for flexible discovery of interactions and thresholds driving heterogeneity.

### Decision Trees and Random Forests (mechanics and modifications)
- Decision trees:
  - Structured yes/no rules to predict outcomes; algorithm considers every possible split on every predictor and chooses the split that best separates samples based on the predicted outcome.
  - Binary partitions continue recursively; regression trees approximate real-valued functions by group averages.
  - Strengths: computational efficiency; handle nonlinearities and interactions. Weakness: can underperform when the true relation is linear.
- Random Forest (RF) modifications to reduce overfitting:
  1. Bootstrap aggregation ("bagging"):
     - Each tree built on a random sample (~two thirds of observations); remaining one-third are out-of-bag (OOB) for accuracy assessment.
     - Repeat hundreds or thousands of times; aggregate predictions (e.g., average).
     - Rationale: averaging many weak, unpruned trees reduces noise.
  2. Random subset of predictors at each split:
     - At each split, consider only a random subset of predictors (usually total predictors divided by three) to promote diversity among trees.
- Implementation scale: RF typically builds hundreds or thousands of unpruned trees using bootstrap samples.

### Using Random Forests for causal inference: the causal forest
- Problem setup: binary treatment W affecting outcome Y with many confounders X; linear regression requires pre-specifying controls, interactions, thresholds—risking misspecification and biased τ estimates.
- Causal forest:
  - A random forest of honest causal trees.
  - Causal trees partition to maximize differences in treatment effects across leaves (average differences in outcomes between treated and non-treated within each leaf).
  - Each terminal leaf acts as a custom-made artificial experiment; leaf average treatment effect predicts individual effects for future observations with the same features.
- Honesty and inference:
  - Honest trees split data in two: half to determine splits; half to populate leaves and estimate treatment effects.
  - Though a single honest tree uses only half the data for estimation, RF builds thousands of trees with new bootstrap samples, ultimately exploiting all data for both splitting and estimation.
  - Alternative noted: train a propensity-classification tree to predict treatment assignment (W) so leaves include well-matched training observations by treatment likelihood (Wager & Athey, 2018).

### Case study: The impact of a financial crisis on growth — framing and data
- Conceptual distinction:
  - Crisis triggers (events that touch off crises) vs. vulnerabilities (structural weaknesses that propagate shocks).
  - Crisis risk combines likelihood of event and impact (vulnerability).
- Traditional approaches:
  - Early-warning models: logit/probit or signaling approaches—interpretable but challenged by out-of-sample prediction and many complex predictors.
  - Two empirical approaches to estimate crisis impact:
    1. Dummy-variable (cross-country panel) approach: crisis dummy coefficient indicates growth cost.
    2. Comparison of actual post-crisis growth to pre-crisis trend: output loss measured as gap between actual and expected trend; trend identification methods vary and can yield sensitive results.
  - Few studies examine determinants of cross-country variation in output loss severity or allow for nonlinearities/interactions; Bayesian Model Averaging (BMA) used as robustness check but is computationally intensive and typically worse at prediction than ensemble ML.
- Causal-forest framing:
  - Treat crisis as treatment; outcome is cumulative output growth in the two years immediately following the crisis (including crisis year).
  - Causal forest produces individual-country counterfactuals by drawing on experiences of similar countries; can estimate hypothetical crisis cost even if crisis is unlikely.
- Data overview:
  - Financial-crisis treatment variable from Laeven and Valencia (2018): requires evidence of significant financial distress and a significant policy intervention.
  - Crisis frequencies: averaging one crisis every 30 to 45 years across income groups; historical clustering varies by period and income group.
  - Final dataset: 46 variables over 1985–2017, covers 107 countries, total over 3300 observations (Annex reports over 3364 observations as dataset total).

### Dataset features, resampling, and estimation details (Annex findings)
- Crises represent around 2½ percent of all observations in the sample.
- Dataset was resampled to improve class balance, raising the overall proportion of crises from 2½ percent to almost 10 percent.
- Covariates are lagged by one period to avoid post-treatment contamination.
- Resampling method: synthetic oversampling technique (SMOTE) to create synthetic minority samples similar to existing crisis observations, minimizing overfitting risk.

### Individual estimates: estimated impact of a financial crisis
- Model estimate for average cumulative cost of a crisis: 7.2 percentage points of growth over two years.
- Comparisons with previous studies (preserved exactly):
  - Hutchison and Noy (2002) suggest a 2-year cost of 6–7 percentage points.
  - Demirgüç-Kunt and others (2006) estimate a 2 -year cost of 7½ percentage points.
  - Abiad and others (2009) baseline estimate: 4½ percentage points; when counterfactual trend uses growth projections of Fund country desks, Abiad et al. estimate is almost 8 percentage points.
- Individual Treatment Effects (ITEs) with confidence intervals provided per country; interval widths vary with availability of comparable peers.
- Example country ITEs (preserved exactly where given):
  - Australia: predicted output loss of only 4.6 percentage points (sample average 7.2 points). Mitigating factors: flexible exchange rate, low manufacturing share, high trade openness, moderate government consumption (past five years), relatively small output gap.
  - China: predicted to experience a more severe crisis (chart signals up to -11.1 percent in figure), driven by high manufacturing share, rapid GDP growth, sharp terms-of-trade growth, volatile inflation; partially offset by potential fiscal space.

### Explaining a “black box” model: Shapley values and interpretability
- Purpose: explain which covariates drive the predicted cost of a hypothetical crisis (not the likelihood of a crisis).
- Method: apply Shapley values to decompose the difference between an individual prediction and the mean prediction across covariates.
- Properties and interpretation:
  - Shapley values capture each covariate’s contribution to an individual prediction.
  - Values are consistent and additive; tailored to each instance.
  - Shapley values are the average contribution across all possible coalitions of variables; cited as having a solid theoretical basis.
  - Identical covariate values across countries can yield different Shapley values because importance depends on the country’s full context.

### Aggregate findings (sample-wide variable importance)
- Top drivers (12 most important variables ranked by mean absolute Shapley value) include:
  - Exchange-rate flexibility (most important overall): greater flexibility associated with smaller output loss.
  - Manufacturing share (importance suggests nonlinear/threshold effects).
  - Cumulative public consumption (5 years).
  - Trade openness.
  - Output gap.
  - CPI inflation.
  - Prev. GDP growth (5 years).
  - Broad Money (% GDP).
  - Financial Development Index.
  - Oil Prices.
  - Total Debt (% GDP).
  - Trading Partner Growth.
- Distributional patterns and nonlinearities:
  - Manufacturing share: skewed distribution—effect appears only above a threshold, where larger shares associate with more severe crises.
  - Trade openness: little effect except for countries below a threshold—relatively closed countries tend to have more costly crises.
  - Public spending (cumulative): skewed with a possible hump-shaped relationship—both excessively high and excessively low levels associated with more costly crises.

### Exploring exchange-rate flexibility: partial dependence and interactions
- Partial dependence plot (PDP) for exchange-rate flexibility shows greater flexibility is associated with a milder crisis impact.
- Variance-decomposition (Friedman and Popescu, 2008) indicates exchange-rate flexibility’s effect is strongly influenced by interactions with other variables.
- Key interaction identified:
  - Exchange-rate flexibility × Financial development: the shock-absorber role of the exchange rate is significantly reduced for countries with a lower level of financial development.
- Implementation note: PDPs, Shapley-value calculations, and variance decompositions implemented using the “iml” package in R (Molnar, 2018).

### Methodological extensions and further research
- Causal Forest algorithm applied to a macroeconomic dataset; approach adaptable to micro-level analyses (labor-market reforms, tax policy, etc.).
- Causal forest extensions:
  - Continuous treatments: in each leaf compute covariance between treatment and outcome; ITE estimates partial impact of a one-unit increase in treatment.
  - Instrumental variables: implement an IV approach by computing the IV estimator in each leaf ("instrumental forests").
- Paper positioned as an introduction to causal machine learning for policy economists rather than a comprehensive contribution to financial-crisis literature.

### Conclusion: contribution and policy relevance
- Machine-learning causal methods (Causal Forest, Shapley-value explanations, PDPs, variance decompositions) can:
  - Produce plausible estimates consistent with prior literature (average impact and roles of exchange-rate flexibility and financial development).
  - Enable richer exploration of thresholds, nonlinearities, and interactions often infeasible with traditional econometric methods.
  - Provide tailored assessments of individual country circumstances to complement existing policy-analysis techniques.

_Source: IMF working paper chapter "1. The Fundamental Problem of Causal Inference" from wpiea2019228-print-pdf_

### 1. The Fundamental Problem of Causal Inference  ________________________________5

### 1. The Fundamental Problem of Causal Inference

### Introduction
- Machine Learning (ML) has a pedigree that extends back "more than fifty years", but adoption accelerated over the "past 10 years" due to predictive capabilities.
- Prediction is important, but empirical economics often requires causal inference — estimating what would have happened in the absence of a particular policy or event (the counterfactual).
- Counterfactuals cannot be validated directly because we never observe the path not taken; this is the central challenge of causal inference.
- "Causal machine learning" seeks to apply ML strengths to causal inference to produce more precise, less biased, and more reliable estimators of causal effects.
- This chapter uses assessment of a hypothetical financial crisis on output growth as a running example, focusing on Causal Forests (Athey and others, 2018) and methods for interpreting complex models.

### Estimating Causal Relationships: Predicting the Counterfactual
- The theoretical gold standard is a randomized controlled experiment, but experiments are often impractical, unethical, or impossible in economics.
- Observational data are used widely, but estimating causal effects from them is problematic because the counterfactual is unobserved (Holland, 1986 calls this the "fundamental problem of causal inference").
- Strategy: construct a credible proxy for the unobservable counterfactual by conditioning on observed "confounding factors" (factors correlated with both the outcome and the likelihood of the event).
- Conditional on observed confounders, treatment can be treated as "essentially as good as randomly assigned", allowing average differences between treated and untreated to proxy causal effects.
- Standard regression conditions by including pre-identified control variables—this is essentially constructing a prediction of the counterfactual.
- Propensity score approach (Rosenbaum and Rubin, 1983): calculate conditional probability of treatment given covariates; propensity score matching predicts missing counterfactuals by using closest observations from the other group (Abadie and Imbens, 2006).
- Key limitation: counterfactual predictions can never be validated; important that such predictions are as plausible and persuasive as possible.

- Box 1 summary (exact points preserved):
  - Causal inference entails arriving at a plausible estimate of an unobserved counterfactual.
  - Best possible construction: experiments with "comparable" subjects (e.g., identical twins as an idealized experiment).
  - Average outcome for untreated similar subjects can serve as counterfactual for treated subjects; differences give treatment effects.
  - Estimating heterogeneous effects requires matched experiments per subgroup; estimating average effects for heterogeneous populations requires random allocation (RCT).
  - In observational studies, must statistically account for systematic differences between groups.

### Machine Learning: Making Better Predictions
- ML origins in computational statistics; chiefly concerned with algorithms to identify patterns within datasets (Kuhn and Johnson, 2016).
- Distinguishing feature: focus on predictive performance and designing experiments to assess how well a model trained on one dataset will predict new data (Box 2).
- ML excels at discovering complicated relationships not specified in advance, handling interactions and thresholds.
- Predicting rare events (e.g., financial crises) is difficult; ML helps by sifting through many potential predictors to find reliable relationships.
- Random Forests (RF) highlighted as a popular, successful general-purpose algorithm:
  - Handles binary, categorical, numerical predictors.
  - Implicit variable selection—focuses on covariates with greatest predictive power.
  - Applicable to classification, regression, cluster analysis.
  - Generally impressive predictive performance; often requires less time and effort to train than many alternatives.
- No single algorithm dominates all applications ("no free lunch theorem"), but Random Forests usually perform well.

- Box 2 summary (exact points preserved):
  - Overfitting: good in-sample fit may model idiosyncratic noise; out-of-sample performance matters.
  - Regularization prevents overfitting by adding a penalty dependent on coefficient magnitudes.
  - LASSO (Least Absolute Shrinkage and Selection Operator) can set some coefficients exactly to zero, effectively performing variable selection.
  - Tuning parameter lambda (λ) controls penalty strength; higher λ implies greater regularization.
  - Model validation methods:
    - Holdout validation: split data into training and testing sets; use test set to compute validation error.
    - Cross validation: divide data into K folds (example K=3); rotate test/training folds to obtain validation errors and choose hyperparameters (e.g., λ) that minimize cross-validation error.
  - Best practice often uses an additional quarantined evaluation set as a final gauge of likely future performance.

### An Emerging Synthesis: "Causal" Machine Learning
- Prediction and causal inference historically separate: prediction asks what usually happens given circumstances; causal inference asks what would happen under an intervention.
- Econometrics emphasized explanation, parsimonious and interpretable models, statistical significance, and in-sample fit. ML emphasized predictive accuracy and out-of-sample performance.
- The divide has narrowed: many estimation problems decompose into steps that resemble pure prediction problems, enabling ML integration.
- Examples where ML assists causal estimation:
  - IV problems with many potential instruments: use ML (e.g., LASSO) to select instruments optimally (Belloni and others, 2012); ML used in first-stage predictive regressions.
  - When treatment is exogenous but many potential confounders exist, ML can select relevant controls:
    - Solution A: LASSO-style regression to select confounders predictive of the outcome.
    - Solution B: use ML to estimate propensity scores with covariates selected for predicting inclusion in treatment (Wyss and others, 2014).
  - Double-selection procedure (Belloni and others, 2014b): run LASSO to select covariates correlated with the outcome and separately with the treatment; take the union of selected covariates and include them as controls in OLS to potentially improve treatment effect estimation.
- These ML modifications focus primarily on estimating the average treatment effect (ATE).
- Interest often lies in heterogeneous treatment effects (HTE): whether treatment effects differ across subjects (e.g., financial crisis impacts varying by country).
- ML's predictive flexibility is especially useful for estimating HTE, enabling exploration of factors and interactions that drive heterogeneous causal effects.

*Source: IMF working paper chapter "1. The Fundamental Problem of Causal Inference" from wpiea2019228-print-pdf*

### Box 3. Better Predictions Through Ensembles: Decision Trees and Random Forests (Cont.)

### Box 3. Better Predictions Through Ensembles: Decision Trees and Random Forests (Cont.)

### Decision trees: intuition and mechanics
- Decision trees provide a structured set of yes/no questions to predict an outcome; they reduce complex, non-linear problems with many predictors to an intuitive flowchart.
- Algorithmic split rule:
  - The algorithm considers every possible split on every possible predictor variable and chooses the one split on the one variable that best separates the sample into the two most dissimilar subsamples (based on the predicted outcome).
  - Binary partitions continue recursively; each subsequent split only considers the subsample under which it falls.
- Regression trees approximate a continuous real-valued function by sorting the dataset into groups of similar observations and using the group average as the non-parametric estimate of the expected outcome.
- Strengths and weaknesses:
  - Computationally efficient and suitable where there are important nonlinearities and interactions.
  - Tend not to work as well if the underlying relationship is linear, though they can reveal aspects not apparent from a traditional linear approach (Varian, 2014).

### Random Forest (RF) modifications to reduce overfitting
- Motivation:
  - Individual decision trees often overfit the training sample and perform poorly out-of-sample.
  - One traditional fix is pruning (imposing a penalty for overly long/complex trees, analogous to the penalty term (λ) in Box 2).
- Two key Random Forest modifications:
  1. Bootstrap aggregation ("bagging"):
     - An individual tree is built on a random sample of the dataset, roughly two thirds of the total observations.
     - The remaining one-third are out-of-bag (OOB) observations and can be used to gauge the accuracy of the tree.
     - This process is repeated hundreds or thousands of times.
     - Prediction for a new instance is obtained by feeding it through each of the thousands of individual trees and aggregating their predictions (e.g., taking the average).
     - Rationale: unpruned trees are weak models; by averaging a large ensemble, noise averages out and the underlying signal remains.
  2. Random subset of predictors at each split:
     - At each split, the algorithm considers only a random subset of the available predictors (usually the total number of predictors divided by three).
     - This reduces the risk that bagging simply produces multiple versions of the same tree when predictors are highly correlated or one predictor dominates.
     - Randomizing the predictor space promotes diversity among trees so that combining many (uncorrelated) weak models yields a strong aggregate prediction.
- Implementation scale:
  - RF typically builds hundreds or thousands of (unpruned) trees using bootstrap samples.

### Using Random Forests for causal inference: the causal forest
- Standard causal estimation problem:
  - Researcher studies impact of a binary treatment (W) on outcome (Y) with many potential confounders (X).
  - Linear regression requires specifying controls, interactions, thresholds, and higher-order terms in advance; misspecification risks biased estimates of τ (average treatment effect).
  - Interest may extend to heterogeneity of treatment effects across subgroups, but specifying all potential interactions a priori is impractical and risks overfitting.
- Causal forest approach:
  - A causal forest is a random forest composed of honest causal trees.
  - Causal tree vs. regression tree:
    - Regression trees partition data to maximize differences in outcomes across leaves.
    - Causal trees focus on average differences in outcomes between treated and non-treated observations within each leaf—i.e., where treatment effects differ most while still allowing accurate estimation.
    - Each terminal leaf represents a custom-made artificial experiment where subjects are as similar as possible; the leaf average treatment effect helps predict individual effects for future observations with the same features.
  - Honesty and inference:
    - To obtain valid confidence intervals and inference asymptotically, each tree must be “honest.”
    - Honest trees divide the data in two: half for determining splits; half for populating the tree and estimating treatment effects.
    - Although a single honest tree appears to throw away half the data, the RF algorithm builds thousands of trees with new bootstrap samples for each tree, so the algorithm ultimately exploits all the data for both splitting and estimation.
    - (Wager & Athey (2018) note an alternative: build honest trees without double splitting by training a propensity-classification tree to predict treatment assignment (W) instead of a regression tree for τ; leaves of the propensity tree include well-matched training observations by likelihood of treatment.)

### Case study: The impact of a financial crisis on growth — framing and data
- Conceptual distinction:
  - Crisis triggers vs. vulnerabilities: triggers (events that touch off the crisis) differ from vulnerabilities (structural weaknesses that propagate and amplify shocks).
  - Crisis risk combines likelihood of an event and impact (vulnerability). Financial crises reflect confluence of vulnerability and trigger.
- Traditional early-warning and cost-estimation approaches:
  - Early-warning models: discrete-choice (logit/probit) regression or signaling approaches; strengths include interpretability but weaknesses include large in-sample vs out-of-sample gaps and difficulty handling many complex predictors.
  - Two empirical approaches to estimate crisis impact:
    1. Dummy-variable (cross-country panel) approach: coefficient on crisis dummy indicates growth cost of crisis; counterfactual is average for sample.
    2. Comparison of actual post-crisis growth to a pre-crisis trend: measures output loss as gap between actual growth and expected trend; methods to identify trend vary (linear projection excluding pre-crisis years; HP filter; time-series forecasting) and results can be sensitive to method choice.
  - Few studies explore determinants of cross-country variation in output loss severity or allow for nonlinearities and interactions.
  - Bayesian Model Averaging (BMA) has been used (Abiad and others, 2009) to address variable selection among many confounders but is computationally intensive and typically used as robustness check; BMA typically performs less well at prediction than ensemble machine-learning (Minka, 2000).
- Causal-forest application to crisis costs:
  - Treat crisis as the treatment, outcome is short-run trajectory of growth immediately following the crisis.
  - Causal forest provides an individual-country counterfactual by drawing on experiences of similar countries and can estimate hypothetical crisis cost even if crisis is unlikely.
- Data used in the study:
  - Financial-crisis treatment variable sourced from Laeven and Valencia (2018); their crisis definition requires evidence of significant financial distress and a significant policy intervention.
  - Crisis frequencies: averaging one crisis every 30 to 45 years across income groups; historical patterns differ by period and income group (1980s–1990s concentrated in EM and LIC; since 2000 clustered during the GFC and mostly among AE).
  - Outcome variable: cumulative rate of output growth in the two years immediately following the crisis (including the crisis year).
  - Covariates: broad range across financial, external, fiscal, real sectors; final dataset includes 46 variables over 1985–2017 and covers 107 countries for a total of over 3300 observations.
  - Data limitations: some informative variables have short histories (financial soundness indicators) or limited country coverage (housing prices), constraining inclusion.

*Source: wpiea2019228-print-pdf - Box 3. Better Predictions Through Ensembles: Decision Trees and Random Forests (Cont.)*

### Annex 1. A key feature of our dataset, however, is that crises represent a very small minority

### wpiea2019228-print-pdf - Annex 1. A key feature of our dataset, however, is that crises represent a very small minority

### Dataset and estimation approach
- Crises represent around 2½ percent of all observations in the sample.
- To improve class balance the dataset was resampled, raising the overall proportion of crises from 2½ percent to almost 10 percent.
- All covariates are lagged by one period to ensure they are not themselves affected by the treatment.
- Resampling method: synthetic oversampling technique (SMOTE) is used to create synthetic samples similar to existing minority observations, minimizing overfitting risk.

### Individual estimates: estimated impact of a financial crisis
- The model estimates the average cumulative cost of a crisis at 7.2 percentage points of growth over two years.
- Comparisons with previous studies:
  - Hutchison and Noy (2002) suggest a 2-year cost of 6–7 percentage points.
  - Demirgüç-Kunt and others (2006) estimate a 2 -year cost of 7½ percentage points.
  - Abiad and others (2009) baseline estimate: 4½ percentage points; when counterfactual trend uses growth projections of Fund country desks, Abiad et al. estimate is almost 8 percentage points.
- The model provides Individual Treatment Effects (ITEs) and confidence intervals for each country; interval widths vary by availability of comparable peers.
- Example country estimates:
  - Australia: predicted output loss of only 4.6 percentage points (compared to a sample average of 7.2 points). Key mitigating factors include flexible exchange rate, low manufacturing share, high trade openness, moderate government consumption over the past five years, and relatively small output gap.
  - China: predicted to experience a more severe crisis (chart signals up to -11.1 percent in figure), driven mostly by high manufacturing share, rapid GDP growth, sharp terms-of-trade growth, and volatile inflation; partially offset by potential fiscal space.

### Explaining a “black box” model: Shapley values and interpretability
- Purpose: explain which covariates drive the predicted cost of a hypothetical crisis (not the likelihood of a crisis).
- Method: apply Shapley values (from cooperative game theory) to decompose the difference between an individual prediction and the mean prediction across covariates.
- Properties of Shapley values used:
  - Capture contribution of each covariate to an individual prediction.
  - Values are consistent and additive; tailored to each specific instance.
  - Shapley values are the average contribution across all possible coalitions of variables and are the only estimators cited as having a solid theoretical basis.
- Interpretation note: identical covariate values across countries can produce different Shapley values because the importance of a variable depends on the country’s full context.

### Aggregate findings (sample-wide variable importance)
- Top drivers (12 most important variables ranked by mean absolute Shapley value) include:
  - Exchange-rate flexibility (most important overall): greater flexibility associated with smaller output loss.
  - Manufacturing share (importance suggests nonlinear/threshold effects).
  - Cumulative public consumption (5 years).
  - Trade openness.
  - Output gap, CPI inflation, Prev. GDP growth (5 years), Broad Money (% GDP), Financial Development Index, Oil Prices, Total Debt (% GDP), Trading Partner Growth.
- Distributional patterns:
  - Manufacturing share: skewed distribution suggests effect appears only above a threshold, where larger shares associate with more severe crises.
  - Trade openness: little effect except for countries below a threshold—relatively closed countries tend to have more costly crises.
  - Public spending (cumulative): skewed with a possible hump-shaped relationship—both excessively high and excessively low levels associated with more costly crises.

### Exploring exchange-rate flexibility: partial dependence and interactions
- Partial dependence plot (PDP) for exchange-rate flexibility shows greater flexibility is associated with a milder crisis impact.
- Variance-decomposition (Friedman and Popescu, 2008) indicates exchange-rate flexibility’s effect is strongly influenced by interactions with other variables.
- Key interaction:
  - Exchange-rate flexibility × Financial development: the shock-absorber role of the exchange rate is significantly reduced for countries with a lower level of financial development.
- Implementation note: PDPs, Shapley-value calculations, and variance decompositions implemented using the “iml” package in R (Molnar, 2018).

### Methodological extensions and further research
- Causal Forest algorithm applied to a macroeconomic dataset; approach can be adapted to micro-level analyses for individual interventions (labor-market reforms, tax policy, etc.).
- Causal forest extensions:
  - Continuous treatments: in each leaf the algorithm computes the covariance between treatment and outcome; ITE estimates partial impact of a one-unit increase in treatment.
  - Instrumental variables: the algorithm can implement an IV approach by computing the IV estimator in each leaf (called “instrumental forests”).
- The paper is positioned as an introduction to causal machine learning for policy economists rather than a comprehensive contribution to the financial-crisis literature.

### Conclusion: contribution and policy relevance
- Machine-learning causal methods (Causal Forest, Shapley-value explanations, PDPs, variance decompositions) can:
  - Produce plausible estimates consistent with prior literature (average impact and roles of exchange-rate flexibility and financial development).
  - Enable richer exploration of thresholds, nonlinearities, and interactions that are often infeasible with traditional econometric methods.
  - Provide tailored assessments of individual country circumstances to complement existing policy-analysis techniques.

*Source: IMF VE Database, author’s calculations*

### REFERENCES

### REFERENCES

### Bibliographic references
- Abadie, Alberto, and Guido W. Imbens, 2006, “Large Sample Properties of Matching Estimators for Average Treatment Effects,” Econometrica, Vol. 74(1), pp. 235–267.
- Abiad, Abdul, Petya Koeva Brooks, Irina Tytell, Daniel Leigh, and Ravi Balakrishnan, 2009, “What’s the Damage? Medium-Term Output Dynamics After Banking Crises,” IMF Working Paper No. 09/245 (Washington: International Monetary Fund).
- Alessi, Lucia & Carsten Detken, 2018, “Identifying Excessive Credit Growth and Leverage,” Journal of Financial Stability, Vol. 35(C), pp. 215–225.
- Athey, Susan and Guido W. Imbens, 2016, “Recursive Partitioning for Heterogeneous Causal E ffects,” Proceedings of the National Academy of Sciences, Vol.113 (27), pp. 7353–7360.
- ———, 2017, “The State of Applied Econometrics: Causality and Policy Evaluation,” Journal of Economic Perspectives, Vol. 31 (2), pp. 3–32.
- Athey, Susan, Julie Tibshirani, and Stephan Wager, 2019, “Generalized Random Forests,” Annals of Statistics, Vol.47 (2), pp. 1148–1178.
- Balta, Narcissa and Plamen Nikolov, 2013, “Financial Dependence and Growth Since the Crisis,” Quarterly Report on the Euro Area, Directorate General Economic and Financial, European Commission, Vol. 12 (3), pp. 7–18.
- Belloni, Alexandre, Daniel Chen, Victor Chernozhukov, and Christian Hansen, 2012, “Sparse Models and Methods for Optimal Instruments with an Application to Eminent Domain,” Econometrica, Vol. 80 (6), pp. 2369–2429.
- Belloni, Alexandre, Victor Chernozhukov, and Christian Hansen, 2014a, “High-Dimensional Methods and Inference on Structural and Treatment Effects,” Journal of Economic Perspectives, 28 (2), pp. 29–50.
- ———, 2014b, “Inference on Treatment Effects after Selection among High-Dimensional Controls,” The Review of Economic Studies, Vol. 81 (2), pp. 608–650.
- Breiman, Leo, 2001a, “Statistical Modeling: The Two Cultures,” Statistical Science, Vol. 16 (3), pp. 199–231.
- ———, 2001b, “Random Forests,” Machine Learning, Vol. 45 (5), pp. 5–32.
- Cerra, Valerie, and Sweta Chaman Saxena, 2008, “Growth Dynamics: The Myth of Economic Recovery,” American Economic Review, Vol. 98 (1), pp. 439–57.
- Chawla, Nitesh V., Kevin W. Bowyer, Lawrence O. Hall, and W. Philip Kegelmeyer, 2002, “SMOTE: Synthetic Minority Over-sampling Technique,” Journal of Artificial Intelligence Research, Vol. 16, pp. 321–357.
- Demirguc-Kunt, Asli, Enrica Detragiache, and Poonam Gupta, 2006, “Inside the Crisis: An Empirical Analysis of Banking Systems in Distress,” Journal of International Money and Finance, Vol. 25 (5), pp. 702–718.
- Edwards, Sebastian, and Eduado Levy Yeyati, 2005, ”Flexible Exchange Rates as Shock Absorbers,” European Economic Review, Vol. 49 (8), pp. 2079–2105.
- Friedman, Jerome H, and Bogdan E Popescu, “Predictive Learning Via Rule Ensembles,” The Annals of Applied Statistics, Vol. 2 (3), pp. 916–54.
- Ghosh, Swati R., and Atish R. Ghosh, 2002, “Structural Vulnerabilities and Currency Crises,” IMF Working Paper No. 02/9 (Washington: International Monetary Fund).
- Holland, P. W., 1986, Statistics and Causal Inference, Journal of the American Statistical Association, Vol. 81 (396), pp. 945–960.
- Hutchison, Michael, and Ilan Noy, 2005, “How Bad Are Twins? Output Costs of Currency and Banking Crises,” Journal of Money, Credit and Banking, Vol. 37 (4), pp. 725–52.
- Imbens, Guido W., and Donald B. Rubin, 2015, Causal Inference for Statistics, Social, and Biomedical Sciences: An Introduction (Cambridge University Press).
- IMF, 2010, The IMF-FSB Early Warning Exercise: Design and Methodological Toolkit, International Monetary Fund.
- Kuhn, Max and Kjell. Johnson, 2013, Applied Predictive Modeling. New York: Springer.
- Laeven, Luc, and Fabian Valencia, 2018, “Systemic Banking Crises Revisited,” IMF Working Paper No. 18/206 (Washington: International Monetary Fund).
- Minka, Thomas P., 2000, “Bayesian Model Averaging Is Not Model Combination,” http://www.stat.cmu.edu/minka/papers/bma.html.
- Morgan, Stephen L., and Christopher Winship, 2015, Counterfactuals and Causal Inference: Methods and Principles for Social Research, 2nd ed. (Cambridge University Press).
- McGue, Matt, Merete Osler, and Kaare Christensen, 2010, “Causal Inference and Observational Research: The Utility of Twins,” Perspect Psychol Sci. Vol. 5 (5), pp. 546–556.
- Molnar, Christoph, 2018, “iml: Interpretable Machine Learning,” R package version 0.5.1. https://CRAN.R-project.org/package=iml.
- ———, 2019, Interpretable Machine Learning, https://leanpub.com/interpretable-machine-learning.
- Pearl, Judea, 2009, Causality: Models, Reasoning, and Inference. 2nd ed. (Cambridge University Press).
- Rajan, Raghuram G., and Luigi Zingales, 1998, “Financial Dependence and Growth," American Economic Review, Vol. 88 (3), pp. 559–586.
- Rosenbaum, Paul R., and Donald B. Rubin, 1983, “The Central Role of the Propensity Score in Observational Studies for Causal Effects,” Biometrika, Vol. 70 (1), pp. 41–55.
- Savona, Roberto and Vezzoli, Marika, 2015, “Fitting and Forecasting Sovereign Defaults Using Multiple Risk Signals,” Oxford Bulletin of Economics and Statistics, Vol. 77 (1), pp. 66–92.
- Tiffin, Andrew J., 2016, “Seeing in the Dark; A Machine-Learning Approach to Nowcasting in Lebanon,” IMF Working Paper No. 16/56 (Washington: International Monetary Fund).
- Varian, Hal R. 2014, “Big Data: New Tricks for Econometrics,” Journal of Economic Perspectives, Vol. 28 (2), pp. 3–28.
- Wager, Stephan, and Susan Athey, 2018, “Estimation and Inference of Heterogeneous Treatment Effects using Random Forests,” Journal of the American Statistical Association, Vol. 113 (523), pp. 1228–1242.
- Wyss, Richard, Alan R. Ellis, M. Alan Brookhart, Cynthia J. Girman, Michele Jonsson Funk, Robert LoCasale, and Til Stürmer, 2014, “The Role of Prediction Modeling in Propensity Score Estimation: An Evaluation of Logistic Regression, bCART, and the Covariate-Balancing Propensity Score,” American Journal of Epidemiology, Vol. 180 (6), pp. 645–655.

### ANNEX I — Dataset overview and features
- Dataset provenance:
  - Data are taken from the Fund’s Vulnerability Exercise Database, with variables chosen to ensure a broad coverage: across sectors, across countries, and across time.
  - The final dataset includes 46 variables over 1985-2017, and covers 107 countries from both emerging markets and advanced economies, for a total of over 3364 observations.

- External Variables / Financial Variables:
  - Import Growth (goods and services, annual percent change)
  - Current Account Balance (percent of broad money)
  - Export Growth (goods and services, annual percent change)
  - Broad Money (percent of GDP)
  - Terms of Trade ( annual percent change)
  - Private Debt, Loans, and Securities (percent of GDP)
  - Trading Partner GDP (annual percent change)
  - Federal Funds Rate (effective, percent)
  - Export Prices (annual percent change)
  - 3-month US T-bill Rate (percent)
  - Trade Openness (exports plus imports/GDP)
  - US Term Spread (bps)
  - Oil Prices (average Brent, WTI, Dubai)
  - Financial Development (index)
  - Financial Account, Net Financial Liabilities (percent of GDP)
  - Bank Credit to Private Sector (percent of GDP)
  - Oil Imports (percent of GDP)
  - Credit Gap (deviation from trend, percent of GDP)
  - Oil Exports (percent of GDP)

- Real Variables:
  - Goods and Services Imports (percent of GDP)
  - Commodity Exporter (dummy)
  - Goods and Services Exports (percent of GDP)
  - Consumer Price Inflation (annual average, percent)
  - Current Account Balance (percent of GDP)
  - GDP Growth (real, 5-year cumulative)
  - Net FDI (percent of GDP)
  - Output Gap (HP filter, percent of GDP)
  - International Reserves (percent of IMF ARA metric)
  - Consumer Price Inflation (annual percent, 5-year cumulative)
  - Gross Foreign Exchange Reserves (percent of GDP)
  - Inflation Volatility (rolling, 4-year)
  - Exchange Rate Flexibility (index)
  - Gross Fixed Capital Formation (real, 5-year cumulative growth, relative to GDP)

- Fiscal Variables:
  - Private Consumption (real, 5-year cumulative growth, relative to GDP)
  - Central Government Debt (percent of GDP)
  - Domestic Demand (real, 5-year cumulative growth, relative to GDP)
  - Public Consumption (real, 5-year cumulative growth, relative to GDP)
  - Income per capita (PPP, relative to USA)
  - External Debt (percent of GDP)
  - Gross Public Savings (percent of GDP)
  - Fiscal Balance (percent of GDP)
  - Gross Private Savings (percent of GDP)
  - GDP Growth Volatility (5-year rolling)
  - Services Value Added (percent of GDP)
  - Manufacturing Value Added (percent of GDP)
  - Agriculture Value Added (percent of GDP)

*Source: REFERENCES and ANNEX I, wpiea2019228-print-pdf*

---


_Source: https://www.imf.org/-/media/files/publications/wp/2019/wpiea2019228-print-pdf.pdf_
