## wpiea2021150-print-pdf

## Source details

**Canonical URL:** [wpiea2021150-print-pdf](https://www.imf.org/-/media/files/publications/wp/2021/english/wpiea2021150-print-pdf.pdf)

## Other formats

- [Markdown version](/-/media/files/publications/wp/2021/english/wpiea2021150-print-pdf.pdf.md)
- [Structured JSON version](/-/media/files/publications/wp/2021/english/wpiea2021150-print-pdf.pdf.json)

---

### II. Past attempts at predicting crises
- Historical context and scope
  - Crisis prediction literature revived after the Asian crisis of the late 1990s.
  - Early studies primarily examine currency and/or financial crises; EWS for fiscal crises are far less common and cover relatively small samples of mostly emerging market economies.
  - Only a few studies include LIDCs (e.g., Cerovic et al. (2018)).
  - Recent literature expands beyond external debt crises toward a broader notion of fiscal stress.

- Two broad methodological strands
  - Multivariate regressions
    - Rely on the Generalized Linear Model (GLM – typically the probit or logit version).
    - Common practice: predictor variables are manually pre-selected based on authors’ judgement.
    - A handful of papers use more agnostic approaches (Extreme Bound Analysis; selection algorithms based on bivariate correlations).
    - Probit models used in policy practice (e.g., IMF/World Bank debt sustainability framework for low-income countries).
  - Tree-based approaches
    - Classification trees capture non-linear structures and complex interactions via iterative binary splits.
    - Aggregation of several trees (ensembles) and short trees as the “signaling” approach are used in some studies.
  - Machine learning examples
    - Random forest: Savona et al. (2015), Jarmulska (2020).
    - Variable selection for random forest: Moreno Badia et al. (2020).
    - Other ML: artificial neural networks in Rodriguez and Rodriguez (2006); Fioramanti (2008).

### III. Defining fiscal crises
- Working definition
  - Fiscal crises described as periods of heightened budgetary distress; not all fiscal crises involve sovereign default.
  - IMF’s fiscal crises database (Medas et al., 2018) is used with four criteria; a country-year is classified as crisis if at least one criterion is met.

- Four criteria (a country-year classified as crisis if at least one criterion is met)
  1. Credit events: sovereign defaults to private or official creditors and debt restructurings; minimum size and accumulation requirements imposed to exclude small technical defaults.
  2. Exceptionally large official financing: any IMF financial arrangement with a fiscal adjustment objective and access above 100 percent of quota counts as a crisis episode; EU financial assistance programs also considered.
  3. Implicit domestic public debt default: (1) high inflation (thresholds vary by income group); (2) accumulation of domestic arrears proxied by other accounts payable.
  4. Loss of market confidence: (1) loss of market access (bond issuance halt); (2) large spikes in sovereign yields.

- Sample- and episode-level statistics
  - 418 crisis episodes for a sample of 188 countries over the period 1980–2016.
  - A country is in crisis in any year if at least one criterion is met; consecutive crisis years count as a single episode; at least two years of no crisis required to separate distinct events.
  - On average, countries have undergone two fiscal crises since 1980.
    - LIDCs have experienced more than three crises on average.
    - AEs have less than one crisis on average.
  - Crisis duration:
    - EMEs and LIDCs show the longest episodes—on average five years.
  - Historical waves:
    - Largest concentration in the 1990s with 93 countries at its peak.
    - Bunching also occurred in the early 1980s and in 2010.
  - Incidence by group:
    - About two-thirds of LIDCs are in fiscal crisis at any point in time.
    - EMEs have a frequency of about 40 percent.
    - Less than 15 percent of AEs experience a fiscal crisis in any given year.
  - Crisis triggers:
    - Credit events account for over 40 percent of episodes.
    - Exceptionally large official financing became the second-most common criterion in the last decade, accounting for almost a third of crisis episodes.

### IV. Empirical strategy and data
- Prediction objective
  - Map a vector of k observables at time t into the probability of a crisis start occurring sometime between the end of year t and the end of year t+h.
  - Prediction window chosen: h =2 to give policy makers time to act.

- Sample coverage and pooling approach
  - Analysis covers 188 countries; sample spans from 1979 to 2015.
  - Estimation pools all countries together; evaluation reports performance separately for AEs/EMEs and LIDCs.
  - Rationale: pooling leverages similarities across countries; tree-based models can endogenously sort clusters.

- Model evaluation: sample-splitting and rolling threshold
  - Out-of-sample evaluation via training and held-out test samples; avoid test-sample leakage from variable selection, cross-year overlap, and two-sided filtering.
  - Iterative rolling procedure:
    - Estimate model with data up to year t – h to make predictions using data observed at time t.
    - For each new test year t+j, move cutoff year for training to t+j-h.
  - Test sample: 2001-2015, includes 1,878 country-year observations, 301 of which include a crisis start in year t+1 or t+2.
  - For a test sample spanning 15 years, 15 models are estimated; hyperparameters retuned accordingly.

- Performance measures
  - Three loss functions for predicted probabilities:
    - log-likelihood.
    - weighted log-likelihood with weights 0.5 on crisis and 0.5 on non-crisis components.
    - mean squared error (MSE), also known as the Brier score.
  - Additional evaluation tools:
    - Calibration curves.
    - Receiver-operator characteristic curves and AUROC.
  - Statistical inference and robustness:
    - Standard errors via bootstrapping on the test sample.
    - Adjust standard errors for two-way clustering (across countries and years).
    - Compute standard errors for differences in performance to permit t-tests (Diebold-Mariano for log-likelihoods and MSE).

- Heuristic benchmarks (three rules of thumb)
  - Pooled global averages: empirical frequency at which all countries have entered a crisis between 1980 and year t-1; for a two-year prediction window the predicted probability is computed as 푝푝_t + (1−푝푝_t)푝푝_t.
  - Country specific averages: empirical frequency at which a specific country has entered a crisis between 1980 and year t-1.
  - Combination: simple average of the pooled global average and the country-specific average.

### Heuristic benchmarks: calibration and behavior by income group
- Global averages:
  - Show very little variation.
  - Overpredict crises for advanced and emerging markets.
  - Underpredict crises for low-income countries.
  - AUROC for pooled (AE+EM): 0.362 (0.031).
- Country-specific averages:
  - Show larger dispersion; lie relatively close to the 45-degree line.
  - AUROC for country-specific (AE+EM): 0.750 (0.03).
- Combination rule:
  - Reduces extreme mistakes while preserving discrimination gains.
  - AUROC combination: 0.744 (0.031) for market access countries and 0.695 (0.035) for low-income developing countries.
- Predictive-performance context:
  - Jorda et al. (2011, 2016): largest in-sample AUROC for financial crises is 0.719.
  - Cerovic et al. (2018): out-of-sample AUROC of 0.69 for market access countries and 0.68 for developing countries; commodity-exporting developing countries: AUROC as high as 0.78 reported.

### Estimation methods overview and tuning
- Penalized logit: elastic net
  - Penalized objective: max_β [ L(X,y;β) − λ ( α ∑ ||β_i||_k + (1−α) ∑ β_i^2 ) ].
  - Hyperparameters λ and α chosen by grid search and cross-validation.
  - Selected values (reported in Annex B): α = 1; λ = .0096.

- Classification trees
  - Tuned hyperparameters to reduce overfitting:
    - maximum tree depth: 4 levels.
    - minimum observations per node: at least 7 observations.
    - minimum Gini improvement for any split: at least 0.01.
    - additional splits attempted only if a node has at least 20 observations.
    - up to 16 leaves possible.
  - Limitations: locally optimal, path dependent, inherently unstable.
  - Aggregation rationale: standalone trees unreliable; ensembles enhance performance if trees are diverse.

- Ensemble tree methods
  - Gradient boosting (XG boost)
    - Trees grown sequentially to explain residuals.
    - Hyperparameters tuned; selected values in Annex B: number of trees = 70; maximum depth = 14; η = .05; γ = 1.5; minimum child weight = 11.
  - Random forest
    - Double randomization: bootstrap samples per tree and random subset m_try at each node.
    - Number of trees set to 3000.
    - Selected m (caret rf): 38.

- Cross-validation procedure (Annex B)
  - For each candidate hyperparameter, select evaluation year t* within training sample, drop observations in t*-1 to t*+1, estimate model, predict for t*, compute log-likelihood loss; average across years; select hyperparameter minimizing average loss.
  - Hyperparameter tuning and model estimation make no use of the test sample.

### Data, feature engineering, and missing values
- Predictors and feature engineering
  - Large database of predictors; permutations include levels, first differences, 3-year differences, lags, and global averages.
  - Resulting number of individual series used: 748.
  - Ten baseline predictors used for comparison: log(GDP per capita); real GDP growth; foreign exchange reserves (in months of imports); current account balance (in percent of GDP); trade openness (10-year average of sum of exports and imports in percent of GDP); real exchange rate (measured using PPP price levels); total external debt (in percent of GDP); general government interest expense (in percent of GDP); Public debt (in percent of GDP); Polity IV score.

- Treatment of collinearity and outliers
  - Many predictors likely collinear; ML methods used can utilize collinear variables.
  - To avoid undue influence on linear models, each variable is winsorized at the 1st and 99th percentile.

- Missing values and imputation policy
  - Objective: provide predictions for all countries; prefer keeping all observations.
  - Missing values replaced with the median over the entire training sample of the non-missing values for that variable.
  - Rationale: simple, conservative; enables demonstration of gains attainable even with unsophisticated imputation.
  - Most missing data is in the 1980s; since 2000 missingness is significantly less and similar across country groups.

### Predictive performance: baseline predictors (no imputation/pooling)
- AE + EM (Table 5.2.A) — methods using same 10 variables, training start 1979, test t = 2001-2015:
  - log(likelihood):
    - Logit: -0.302 (0.034)
    - Elastic net: -0.3 (0.033)
    - XG boost: -0.287 (0.034)
    - Random forest: -0.281 (0.027)
  - weighted log(likelihood):
    - Logit: -1.038 (0.062)
    - Elastic net: -1.053 (0.035)
    - XG boost: -0.902 (0.055)
    - Random forest: -0.906 (0.045)
  - MSE:
    - Logit: 0.086 (0.012)
    - Elastic net: 0.085 (0.012)
    - XG boost: 0.083 (0.012)
    - Random forest: 0.081 (0.011)
  - AUROC:
    - Logit: 0.697 (0.039)
    - Elastic net: 0.697 (0.033)
    - XG boost: 0.750 (0.033)
    - Random forest: 0.772 (0.024)

- LIDC (Table 5.2.A):
  - AUROC:
    - Logit: 0.553 (0.054)
    - Elastic net: 0.546 (0.028)
    - XG boost: 0.530 (0.041)
    - Random forest: 0.577 (0.043)
  - Conclusion: ML methods outperform logit for AE + EM; for LIDCs gains are smaller and not always statistically significant.

### Imputation and pooling experiments (Tables 5.3.A, 5.4)
- Two expansion strategies: (i) pooling all countries; (ii) imputing missing values with training-sample median.
- Logit with imputation and pooling (selected patterns):
  - For market access countries:
    - Imputation without pooling doubles the training sample and yields pronounced, statistically significant gains.
    - Once missing values are imputed, pooling adds no benefit.
  - For LIDCs:
    - Pooling and imputation combined improves performance across all criteria.
    - Improvements for weighted log-likelihood and AUROC have confidence of 78 percent and 84 percent respectively.
    - Imputation alone can add noise and deteriorate log-likelihoods.
- Table 5.4 (all methods, impute in training sample = yes; impute in test sample = no; training pooled across countries):
  - AE + EM:
    - AUROC: Logit 0.712 (0.036); Elastic net 0.716 (0.037); XG boost 0.761 (0.033); Random forest 0.777 (0.022).
    - weighted log(likelihood): Random forest -0.882 (0.036).
  - LIDC:
    - AUROC: Logit 0.556 (0.048); Elastic net 0.555 (0.048); XG boost 0.555 (0.047); Random forest 0.581 (0.038).
  - Conclusion: pooling + imputation improves performance for AE + EM; random forest remains top performer.

### Graphical comparisons: calibration and ROC
- AE + EM (Figures 5.1)
  - Logit models overconfident and poorly ranked in some quintiles.
  - Random forest predictions relatively well calibrated, monotonic, and close to 45-degree line; reliable at identifying low crisis risk observations.
  - ROC shapes reflect better ranking by random forest.
- LIDC (Figure 5.2)
  - Differences marginal; random forest overpredicts less than logit but both have non-monotonic calibration curves.

### Model-based variable selection and high-dimensional experiments
- Expanded variable set: all 748 variables; missing values imputed; sample pooled across income groups.
- Linear models and selection
  - Logit with stepwise forward selection maximizing BIC selects 25 variables on average.
  - Logit with BIC-selected variables performs worse than the 10 pre-selected variable model except marginal AUROC improvement.
  - Elastic net considers many predictors without deterioration; uses on average more than 116 variables out of 748 across the 15 rolling regressions.
  - Elastic net outperforms logit/BIC along all dimensions.
- Tree ensembles and feature pre-selection
  - Random forest potentially harmed by many noisy predictors; more than half of variables provide no valuable information.
  - Recursive feature elimination (RFE) selects 93 out of 748 variables for random forest.
  - Random forest with RFE outperforms logit with BIC on every dimension and is more robust (smaller standard deviations).
  - Random forest with 748 variables performs substantially better for LIDCs than with pre-selected baseline variables.
  - RFE improves statistical significance though AUROC and weighted log-likelihood can decline for AE + EM when using RFE.

### Predictors of fiscal crises: variable importance and top predictors (Table 7.1)
- Variable importance measures
  - Linear models: predictor slope coefficient scaled by variable standard deviation.
  - XG boost: improvement in fit measured by reduction in Gini.
  - Random forest: out-of-bag permuted predictor importance.
- Shrinkage regimes
  - Sparse estimators (elastic net with α = 1) reduce predictors used; dense estimators (random forest) distribute importance across collinear predictors.
  - Elastic net tends to concentrate importance on top variables; random forest spreads importance more evenly.

- Top predictors (summary of empirical findings)
  - Considerable overlap across ML methods.
  - Crisis history and GDP per capita are within top 3 for elastic net, boosted trees, and random forest.
  - Boosted trees and random forest agree on 7 variables in the top 10.
  - Confirmed important predictors from literature-based model: GDP per capita; debt service; external debt; current account balance.
  - For random forest:
    - Current account importance shared with external gross financing needs.
    - Foreign exchange reserves important but importance divided across related ratios (e.g., debt amortization to reserves).
  - Variables not identified as important by these algorithms: trade openness, real exchange rate, real GDP growth rate.
  - Demographics and governance:
    - Size of working-age population and age dependency ratio rank strongly.
    - Quality of bureaucracy ranks strongly.
  - Other important variables: gross domestic savings; inflation; volatility measures.
  - Many fiscal series absent from top 30: government revenues (apart from foreign development aid), expenditures, deficits, and interest-growth differentials.
  - Public debt matters mainly when external.
  - Global variables (oil prices, interest rates) often redundant or irrelevant.

- Representative ranks and examples (from Table 7.1)
  - log(GDP per capita), PPP level: ranks include 2, 2, 31 across methods.
  - Official reserves (in months of imports) level: rank 6 in at least one method.
  - Public external debt / GDP level: ranks include 7 and 11 in some methods.
  - Volatility: CPI inflation 10-year std. dev. rank 30 in at least one method.
  - Country characteristics: Island (dummy) rank examples include 15, 25.

### Predictive performance summaries and statistical significance
- Predictive performance of historical averages (selected metrics)
  - log(likelihood):
    - pooled: -0.353 (0.032)
    - country-specific: -0.484 (0.126)
    - combination: -0.314 (0.031)
  - weighted log(likelihood):
    - pooled: -0.972 (0.003)
    - country-specific: -1.778 (0.416)
    - combination: -0.886 (0.025)
  - MSE:
    - pooled: 0.098 (0.014)
    - country-specific: 0.092 (0.012)
    - combination: 0.089 (0.012)
  - AUROC:
    - combination (AE+EM): 0.744 (0.031)
    - combination (LIDC): 0.695 (0.035)

- Table 6.1.A baseline specification (with imputation and pooling) — selected out-of-sample series (t = 2001-2015, 15 rolling regressions)
  - AE + EM AUROC sample entries: 0.714 (0.033), 0.748 (0.032), 0.717 (0.038), 0.766 (0.032), 0.750 (0.03), 0.771 (0.031), 0.780 (0.031), 0.802 (0.026), 0.787 (0.029), 0.744 (0.031).
  - LIDC AUROC sample entries: 0.62 (0.039), 0.662 (0.041), 0.623 (0.039), 0.676 (0.042), 0.634 (0.044), 0.683 (0.042), 0.654 (0.034), 0.705 (0.04), 0.70 (0.041), 0.695 (0.035).

- Statistical significance tables (Tables 5.2.B, 5.3.B, 6.1.B)
  - Report degree of confidence (1 - p-value) that model in row i outperforms model/benchmark in column j using one-sided t-tests with bootstrapped standard errors adjusted for two-way clustering.
  - Many pairwise confidence values reported; random forest often achieves high confidence values (up to 0.99 or 1.00) when compared to alternatives in AE + EM blocks.

### Conclusions and high-level takeaways
- Crisis predictions based on established econometric approaches are of limited use; they often cannot outperform a naïve prediction rule.
- Performance ranking depends on the performance measure and is not always statistically significant.
- Random forest approaches systematically outperform heuristic benchmarks and other statistical models in many settings, especially for market access countries.
- Relying on a single model for all countries and imputing missing values does not lower predictive performance.
- Allowing for a large set of predictor variables and letting algorithms select variables improves accuracy, notably for low-income developing countries.

*Source: wpiea2021150-print-pdf*

### Section VII discusses the ranking of predictors by importance.

### Section VII discusses the ranking of predictors by importance.

### II. Past Attempts at Predicting Crises
- Historical context and scope
  - Crisis prediction literature revived after the Asian crisis of the late 1990s.
  - Early studies primarily examine currency and/or financial crises; EWS for fiscal crises are far less common and cover relatively small samples of mostly emerging market economies.
  - Only a few studies include LIDCs (e.g., Cerovic et al. (2018)).
  - Recent literature expands beyond external debt crises toward a broader notion of fiscal stress.

- Two broad methodological strands
  - (1) Multivariate regressions
    - Rely on the Generalized Linear Model (GLM – typically the probit or logit version).
    - Common practice: predictor variables are manually pre-selected based on authors’ judgement.
    - A handful of papers use more agnostic approaches (Extreme Bound Analysis; selection algorithms based on bivariate correlations).
    - Probit models used in policy practice (e.g., IMF/World Bank debt sustainability framework for low-income countries).
  - (2) Tree-based approaches
    - Models based on classification trees (Breiman et al. (1984)) capture non-linear structures and complex interactions via iterative binary splits.
    - Aggregation of several trees (ensembles) and short trees as the “signaling” approach are used in some studies.
  - Machine learning developments
    - Examples: random forest used in Savona et al. (2015), Jarmulska (2020); Moreno Badia et al. (2020) explore variable selection techniques for random forest.
    - Other ML: artificial neural networks in Rodriguez and Rodriguez (2006); Fioramanti (2008).

### III. Defining Fiscal Crises
- Working definition
  - Fiscal crises described as periods of heightened budgetary distress; not all fiscal crises involve sovereign default.
  - To capture broader fiscal stress, the IMF’s fiscal crises database (Medas et al., 2018) is used with four criteria.

- Four criteria (a country-year classified as crisis if at least one criterion is met)
  1. Credit events: sovereign defaults to private or official creditors and debt restructurings; minimum size and accumulation requirements imposed to exclude small technical defaults.
  2. Exceptionally large official financing: any IMF financial arrangement with a fiscal adjustment objective and access above 100 percent of quota counts as a crisis episode; EU financial assistance programs also considered.
  3. Implicit domestic public debt default: (1) high inflation (thresholds vary by income group); (2) accumulation of domestic arrears proxied by other accounts payable.
  4. Loss of market confidence: (1) loss of market access (bond issuance halt); (2) large spikes in sovereign yields.

- Sample- and episode-level statistics
  - Based on this definition, there are 418 crisis episodes for a sample of 188 countries over the period 1980–2016.
  - A country is classified as in crisis in any given year if at least one of the four criteria is met; consecutive crisis years count as a single episode; at least two years of no crisis required to separate distinct events.
  - On average, countries have undergone two fiscal crises since 1980, with large heterogeneity:
    - LIDCs have experienced more than three crises on average.
    - AEs have less than one crisis on average.
  - Crisis duration differs by group:
    - EMEs and LIDCs show the longest episodes—on average five years.
  - Historical waves:
    - Largest concentration in the 1990s with 93 countries (more than 50 percent of which were EMEs) undergoing a crisis at its peak.
    - Bunching also occurred in the early 1980s and in 2010.
  - Incidence by group:
    - About two-thirds of LIDCs are in fiscal crisis at any point in time.
    - EMEs have a frequency of about 40 percent.
    - Less than 15 percent of AEs experience a fiscal crisis in any given year.
  - Crisis triggers:
    - Credit events account for over 40 percent of episodes.
    - Exceptionally large official financing became the second-most common criterion in the last decade, accounting for almost a third of crisis episodes.

### IV. Empirical Strategy and Data
#### A. The prediction problem and sample
- Prediction objective
  - Map a vector of k observables at time t into the probability of a crisis start occurring sometime between the end of year t and the end of year t+h.
  - Prediction window chosen: h =2 to give policy makers time to act.

- Sample coverage and pooling approach
  - Analysis covers 188 countries; sample spans from 1979 to 2015.
  - While evaluation reports performance separately for AEs/EMEs and LIDCs, estimation pools all countries together (letting algorithms decide how to split data).
  - Rationale: pooling can leverage similarities across countries (e.g., some EMEs resemble some LIDCs); tree-based models can endogenously sort clusters.

#### B. Model evaluation
- Sample splitting and rolling threshold
  - Out-of-sample evaluation: split data into training and held-out test samples; no test-sample information may influence variable selection or estimation.
  - Sources of test-sample leakage to avoid:
    - Variable selection using all observations (including test sample).
    - Training/test samples not separated by year (e.g., Spain 2009 in training while Portugal 2009 in test).
    - Predictor variables constructed using two-sided filtering approaches that incorporate future information.
  - Iterative rolling procedure:
    - Estimate the model with data up to year t – h to make predictions using data observed at time t.
    - With each new year t+j in the test sample, move the cutoff year for training to t+j-h.
    - Test sample chosen: 2001-2015, includes 1,878 country-year observations, 301 of which include a crisis start in year t+1 or t+2.
    - For a test sample spanning 15 years, 15 models are estimated; hyperparameters are retuned accordingly.

- Performance measures (emphasis on different penalties and uses)
  - Three measures of accuracy for predicted probabilities (each penalizes errors differently):
    - log-likelihood (the logit/probit objective function).
    - weighted log-likelihood with weights 0.5 on crisis and 0.5 on non-crisis components:
      - weighted log-likelihood defined using multipliers 0.5 preceding the crisis-sum and the non-crisis-sum.
    - mean squared error (MSE), also known as the Brier score.
  - Practical implications:
    - Predictions are continuous between zero and one; loss functions penalize large-margin missed crises differently (e.g., log-likelihood assigns a score near -∞ when predicted probability is close to zero for an observation that is within two years of a crisis start; MSE assigns at most a loss of 1).
    - Unweighted log-likelihood typically larger than weighted log-likelihood because most observations are non-crisis and easier to predict.
  - Additional evaluation tools
    - Calibration curves: plot predicted probabilities against observed crisis frequencies to assess calibration and discrimination.
    - Receiver-operator characteristic curves and AUROC: AUROC captures how well rankings of probabilities correspond to observed outcomes (probability that an algorithm selects the crisis out of a randomly selected crisis and non-crisis case).
    - Note: AUROC assesses ranking but not absolute calibration—models with poor calibration and poor discrimination can still have an AUROC close to 1.

- Statistical inference and robustness
  - Standard errors for performance measures are reported using bootstrapping on the test sample.
  - Adjust standard errors for two-way clustering (across countries and across years), following Cameron et al. (2011).
  - Compute standard errors for differences in performance between methods to permit t-tests of differences (for log-likelihoods and MSE these correspond to Diebold-Mariano (2002) tests).

- Heuristic benchmarks
  - To put predictive performance into perspective, model predictions are compared against simple rules of thumb (three such rules are considered; details of the rules not included in the supplied excerpt).

*Italic: Source: wpiea2021150-print-pdf - Section VII discusses the ranking of predictors by importance.*

### 1.      Pooled global averages: this rule of thumb takes the empirical frequency at which all

### wpiea2021150-print-pdf - 1.      Pooled global averages: this rule of thumb takes the empirical frequency at which all

### Heuristic benchmarks for crisis prediction
- Three simple rules of thumb for predicted crisis probability for year t+1:
  - Pooled global averages: empirical frequency at which all countries have entered a crisis between 1980 and year t-1; assigns the same probability to all countries at a point in time (as in Hellwig (2018)).
    - For a two-year prediction window, the predicted probability is computed as 푝푝_t + (1−푝푝_t)푝푝_t.
  - Country specific averages: empirical frequency at which a specific country has entered a crisis between 1980 and year t-1.
  - Combination: simple average of the pooled global average and the country-specific average.

- Calibration and behavior by income group:
  - Global averages show very little variation and:
    - Overpredict crises for advanced and emerging markets.
    - Underpredict crises for low-income countries.
  - Country-specific averages show larger dispersion and lie relatively close to the 45-degree line, implying substantial information in a country’s crisis history.
  - The combination (third rule) is motivated by large likelihood-score penalties incurred when country-specific averages assign near-zero probabilities to countries that subsequently experience a crisis.

### Predictive performance and metrics
- Comparative performance observations:
  - Pooled averages: very low AUROC because they do not discriminate between high and low risk countries.
  - Country-specific averages: better ranking (higher AUROC) but prone to costly likelihood-score mistakes when a historically crisis-free country has a crisis.
  - Combination of historical averages reduces extreme mistakes while preserving discrimination gains.

- Reported AUROC figures:
  - Combination of historical averages delivers an AUROC of .744 for market access countries and .695 for low-income developing countries.
  - For context from literature:
    - Jorda et al. (2011) and Jorda et al. (2016): largest in-sample AUROC for model-based predictions of financial crises is 0.719.
    - Cerovic et al. (2018): out-of-sample AUROC of 0.69 for market access countries and 0.68 for developing countries.
    - For the subsample of commodity exporting developing countries, an AUROC as high as 0.78 is reported.

- Additional remark:
  - Predictive performance (log-likelihood, MSE, AUROC) based on historical experience is higher for market access countries than for developing countries because there are fewer crises in market access countries, reducing underlying uncertainty.

### Estimation methods overview
- Focus: methods commonly used in applied economics literature; descriptions kept relatively non-technical. Cross-validation procedure and tuning grids are in Annex B.

- Penalized logit: elastic net
  - Uses the elastic net (Zou and Hastie, 2005) to mitigate overfitting in the logit framework by adding a penalty term to the likelihood maximization problem:
    - max_β [ L(X,y;β) − λ ( α ∑ ||β_i||_k + (1−α) ∑ β_i^2 ) ].
  - Hyperparameters λ and α are chosen by grid search and cross-validation.
  - Special cases:
    - When λ is zero → estimated β corresponds to the logit estimator.
    - When λ→∞ → estimated β values converge to zero.
    - When α=1 (LASSO) and λ sufficiently large → some β elements set to exactly zero.
    - When α=0 (Ridge) → all β elements are non-zero.

- Classification trees
  - Trees split samples recursively using thresholds on the most informative variable (via Gini index).
  - Tuned hyperparameters to reduce overfitting:
    - maximum tree depth: 4 levels.
    - minimum observations per node: at least 7 observations.
    - minimum Gini improvement for any split: at least 0.01.
    - additional splits attempted only if a node has at least 20 observations.
    - with these constraints, up to 16 leaves are possible.
  - Limitations:
    - Finding globally optimal trees is infeasible; trees are locally optimal and path dependent.
    - Trees are inherently unstable; small sample variations can change the top split and propagate.
  - Aggregation rationale:
    - Stand-alone trees are unreliable; aggregating many trees often enhances performance if trees are diverse.

- Ensemble tree methods used
  - Gradient boosting (XG boost by Chen et al. (2015))
    - Trees grown sequentially; each new tree explains residuals from prior trees.
    - Tuning over: number of trees; maximum tree depth; minimum gain required for an additional split; minimum size of any leaf.
  - Random forest (Breiman, 2001)
    - Diversity via double randomization:
      - Each tree is estimated on a synthetic sample drawn at random from the original estimation sample.
      - At each node, the algorithm chooses splits from a random subset of m_try < k predictors.
    - m_try chosen via cross-validation.
    - As suggested by Breiman, no other restrictions on tree growing; each tree grown exhaustively and can perfectly separate crisis starts from non-crisis observations in its bootstrap sample.
    - Number of trees set to 3000.
    - Interpretation: each tree overfits different observations; averaging cancels overfitting errors as number of trees becomes large.

### Data, feature engineering, and missing values
- Data and predictors:
  - Large database of possible predictors of fiscal crises (detailed variable list in Annex Table 3).
  - Predictors include country-specific economic and institutional variables plus global variables (interest rates, commodity prices, financial market conditions).
  - Feature engineering: use of various permutations (levels, first differences, 3-year differences) and lags; inclusion of global averages for many series to reduce confounding.
  - Resulting number of individual series used: 748.

- Treatment of collinearity:
  - Many predictors likely collinear, but machine learning methods used here can utilize collinear variables (unlike OLS or logit), enabling investigation of relative attention to revenues vs. expenditure vs. overall fiscal balance, or public vs. private vs. total external debt, etc.

- Outlier handling:
  - To avoid undue influence of outliers in linear models, each variable is winsorized at the 1st and 99th percentile.

- Missing values and imputation policy:
  - Objective: provide predictions for all countries regardless of data availability; prefer keeping all observations in the sample.
  - Rationale:
    - Dropping observations with missing values reduces training sample size and exacerbates overfitting.
    - Crisis risk is likely correlated with data availability/quality, so dropping observations would bias performance assessment.
  - Imputation method:
    - Missing values replaced with the median over the entire training sample of the non-missing values for that variable.
    - Chosen deliberately as a simple, conservative approach to demonstrate gains are achievable even with unsophisticated imputation.
    - Expectation stated that future research into more involved imputation methods could yield additional gains.

*Italic: Source — wpiea2021150-print-pdf*

### 4.2   shows, most of the missing data is in the 1980s, while since 2000 it is significantly less of

### wpiea2021150-print-pdf - 4.2   shows, most of the missing data is in the 1980s, while since 2000 it is significantly less of

### Imputation and data coverage
- Most missing data is in the 1980s; since 2000 missingness is significantly less and similar across country groups.
- Imputation is particularly useful when using a large number of variables because, without imputation, additional variables would cause loss of observations.
- By imputing variables, the author avoids variable selection based on data availability; variables with smaller coverage are acknowledged as more noisy and the algorithm is left to trade off noise versus information.

### Baseline predictors used for comparison
- Ten baseline predictors: log(GDP per capita); real GDP growth; foreign exchange reserves (in months of imports); current account balance (in percent of GDP); trade openness (10-year average of sum of exports and imports in percent of GDP); real exchange rate (measured using PPP price levels); total external debt (in percent of GDP); general government interest expense (in percent of GDP); Public debt (in percent of GDP); Polity IV score.
- These predictors are drawn from the literature preceding the test sample and used to evaluate whether models can aggregate information in a fixed set of variables.

### Econometric approach (logit and classification tree)
- Training samples initially split by income group and missing values dropped (no imputation).
- In-sample vs out-of-sample:
  - In-sample likelihood fit for logit is best attainable on training data (market access country data from 2001-2015).
  - Out-of-sample performance deteriorates dramatically; reported AUROC of 0.697 (nearly identical to Cerovic et al (2018)).
  - Dropping the 1980s from training data further deteriorates performance, suggesting smaller sample causes more overfitting.
- Classification tree:
  - Outperforms logit in-sample, suggesting non-linearities may help.
  - Out-of-sample, the classification tree performs worse than logit — better in-sample fit comes at cost of worse out-of-sample performance.
- Comparison to heuristic benchmark (combination of country-specific and pooled averages):
  - In-sample, logit fit better than heuristic for all but one measure, but AUROC and MSE differences not statistically significant.
  - Tree in-sample fit significantly better than historical benchmark along all dimensions.
  - Out-of-sample, neither logit nor tree robustly outperform the historical benchmark; weighted log-likelihood is significantly worse for these models.
- Low-income countries (LIDCs):
  - Similar mixed picture; LIDCs experience more crises and crisis years are harder to predict.
  - Unweighted likelihood worse for LIDCs than for more developed countries; weighted likelihood higher because models assign higher crisis probabilities to LIDC observations.
  - Overall, commonly used predictors are not sufficiently useful for predicting fiscal crises in developing countries.

### Machine learning algorithms (random forest, XG boost, elastic net)
- For advanced and emerging economies:
  - Logit is outperformed by all three ML methods on nearly every measure.
  - Random forest dominates in level and stability (smaller standard deviations).
  - Tree ensemble methods (random forest and XG boost) outperform linear methods (logit / elastic net); improvements are statistically significant except for MSE.
  - Random forest outperforms the three heuristic benchmarks with high probability (somewhat lower probability for weighted log-likelihood).
- For low-income developing countries:
  - Random forest still outperforms logit on all criteria, but overall ranking is more ambiguous.
  - Not all differences between random forest and logit are statistically significant.
  - None of the four models systematically outperform the heuristic benchmarks for LIDCs.
- Summary: ML methods, especially random forest, improve performance for market-access countries; benefits for LIDCs depend on variable selection and sample construction.

### Imputation and pooling experiments
- Two ways to expand training sample: (i) pooling all countries into a single sample; (ii) replacing missing values with imputed value (sample median).
- Logit with imputation and pooling (Table 5.3.A highlights):
  - For market access countries (panel (i)):
    - Expanding training sample to include LIDCs (column (ii)) leads to marginally better performance except for weighted log-likelihood.
    - Imputation without pooling (column (iii)) doubles the training sample size relative to column (i) and yields pronounced, statistically significant gains.
    - Once missing values are imputed, pooling adds no benefit (compare columns (iii) and (iv)); imputation reduces overfitting enough that pooling LIDCs adds more noise than information.
  - For LIDCs (panel (ii)):
    - Pooling and imputation combined (column (iv)) improves performance relative to column (i) across all criteria.
    - Improvements for weighted log-likelihood and AUROC have confidence of 78 percent and 84 percent respectively.
    - Imputation alone (doubling sample in column (iii)) can add noise and deteriorate log-likelihoods relative to column (i).
    - Using information from advanced and emerging market economies (column (ii), tripling sample) marginally outweighs costs of heterogeneity.
- Pooling and imputing across all learning algorithms (Table 5.4):
  - Test sample restricted to complete observations for comparability.
  - For advanced and emerging economies: unweighted likelihood and AUROC systematically improve; random forest marginally better across measures and standard errors are smaller.
  - Logit exhibits largest gains relative to Table 5.2.A because it lacks built-in overfitting controls and benefits most from expanded training sample.

### Graphical comparisons (calibration and ROC)
- Advanced and emerging market economies (Figures 5.1):
  - Logit models are overconfident, with poor ranking (middle quintile has lower actual crisis frequency than 2nd lowest quintile).
  - Random forest predictions are relatively well calibrated, monotonic, and close to 45-degree line; particularly reliable in identifying observations with low crisis risk.
  - ROC shapes reflect better ranking by random forest.
- Low-income countries (Figure 5.2):
  - Differences in performance are marginal.
  - Random forest overpredicts less than logit but both have non-monotonic calibration curves.

### Model-based variable selection (full set of variables and imputation/pooling)
- Expanded variable set to all 748 variables; missing values imputed; sample pooled across income groups.
- Imputation applied in training and test samples so performance reflects all observations.
- Linear models:
  - Logit variable selection via stepwise forward selection maximizing BIC selects 25 variables on average (over 15 rolling regressions).
  - Logit with BIC-selected variables performs worse than the 10 pre-selected variable model except marginal AUROC improvement.
  - Elastic net can consider many predictors without performance deterioration; for LIDCs additional predictors improve performance.
  - Elastic net uses on average more than 116 variables out of 748 across the 15 rolling regressions.
  - Elastic net outperforms logit/BIC along all dimensions; differences unlikely spurious per Table 6.2.B.
  - BIC leads to sparse selection and maximal likelihood coefficients, which in small samples mechanically overfits; elastic net regularizes and cross-validates penalty parameter based on out-of-sample accuracy.
- Tree ensembles:
  - XG boost uses informative variables and ignores rest; marginal improvement over elastic net for market access countries; for LIDCs improvement over logit but not over elastic net.
  - Random forest potentially harmed by many noisy predictors; more than half of variables provide no valuable information (variable importance chart).
  - Recursive feature elimination (RFE) used to pre-select variables for random forest by iteratively dropping least important variables; RFE selects 93 out of the 748 variables.
  - Stand-alone random forest outperforms logit with BIC on every dimension and is more robust (smaller standard deviations); outperforms heuristic benchmarks with high confidence.
  - Random forest AUROCs and weighted log-likelihoods particularly large, suggesting strength at avoiding “missed” crises.
  - Random forest with 748 variables performs substantially better for LIDCs than with pre-selected variables — baseline set was too narrow.
  - Random forest combined with RFE yields qualitatively similar but more statistically significant results; performance with RFE improves relative to other methods though AUROC and weighted log-likelihood can decline for advanced and emerging markets when using RFE.

### Predictors of fiscal crises (variable importance and top predictors)
- Variable importance measures:
  - Linear models (logit, elastic net): predictor slope coefficient scaled by variable standard deviation.
  - Gradient boosted trees (XG boost): improvement in fit measured by reduction in Gini.
  - Random forest: out-of-bag permuted predictor importance.
- Shrinkage regimes and implications:
  - Sparse estimators (elastic net as LASSO with α = 1) reduce predictors used; dense estimators (random forest) keep variables and distribute importance across collinear predictors.
  - Elastic net tends to attribute greater importance to top variables; random forest distributes weights more evenly.
- Empirical findings on important predictors (Top 30 discussion):
  - Considerable overlap across machine learning methods in variables selected.
  - Crisis history and GDP per capita are within top 3 for elastic net, boosted trees, and random forest.
  - Boosted trees and random forest agree on 7 variables in the top 10.
  - Confirmed important predictors from literature-based model: GDP per capita; debt service; external debt; current account balance.
  - For random forest, current account importance is shared with external gross financing needs; foreign exchange reserves important but importance divided across related ratios (e.g., debt amortization to reserves).
  - Trade openness, real exchange rate, and real GDP growth rate are not identified as important by these algorithms.
  - Demographics (size of working-age population, age dependency ratio) and governance (quality of bureaucracy) rank strongly.
  - Other important variables: gross domestic savings; inflation; volatility measures.
  - Many fiscal series absent from top 30: government revenues (apart from foreign development aid), expenditures, deficits, and interest-growth differentials; public debt matters mainly when external.
  - Global variables (oil prices, interest rates) are often redundant or irrelevant for predicting crises.
- Interpretation: external sector variables, both stocks and flows, play a key role as early warning indicators; volatility of macro fundamentals increases likelihood of exercising a default-like option.

### Conclusion (high-level takeaways)
- Crisis predictions based on established econometric approaches are of limited use; they often cannot outperform a naïve prediction rule.
- Performance ranking depends on the performance measure and is not always statistically significant.
- Random forest approaches systematically outperform heuristic benchmarks and other statistical models.
- Relying on a single model for all countries and imputing missing values does not lower predictive performance.
- Allowing for a large set of predictor variables and letting algorithms select variables improves accuracy, notably for low-income developing countries.

*wpiea2021150-print-pdf - 4.2   shows, most of the missing data is in the 1980s, while since 2000 it is significantly less of*

### References

### References

### Cited Works
- Abbas, S. A., N. Belhocine, A. ElGanainy, and M. Horton. 2011. “A Historical Public Debt Database,” IMF Economic Review, vol. 59, issue 4, pp. 717–42.
- Abiad, A., 2003. “Early Warning Systems: A Survey and a Regime-Switching Approach,” IMF Working Paper No. 03/32 (International Monetary Fund).
- Arslanalp, S. and T. Tsuda, 2012. “Tracking Global Demand for Advanced Economy Sovereign Debt,” IMF Working Papers 12/284.
- Baldacci, E., I. Petrova, N. Belhocine, G. Dobrescu, and S. Mazraani. 2011. “Assessing Fiscal Stress,” IMF Working Paper No. 11/100 (International Monetary Fund).
- Basu, S.S., R. Perrelli, and W. Xin, 2019. „External Crisis Prediction Using Machine Learning: Evidence from Three Decades of Crises Around the World.”
- Belloni, A. and Chernozhukov, V., 2011. “High dimensional sparse econometric models: An

*Source: wpiea2021150-print-pdf - References*

### introduction,” In Inverse Problems and High-Dimensional Estimation (pp. 121-156).

### wpiea2021150-print-pdf - introduction,” In Inverse Problems and High-Dimensional Estimation (pp. 121-156).

### Predictive performance of historical averages
- Accuracy of predictions (pooled / country-specific / combination) — log(likelihood):
  - pooled: -0.353 (0.032)
  - country-specific: -0.484 (0.126)
  - combination: -0.314 (0.031)
  - pooled (LIDC): -0.655 (0.056)
  - country-specific (LIDC): -1.01 (0.231)
  - combination (LIDC): -0.578 (0.049)
- Weighted log(likelihood):
  - pooled: -0.972 (0.003)
  - country-specific: -1.778 (0.416)
  - combination: -0.886 (0.025)
  - pooled (LIDC): -0.97 (0.003)
  - country-specific (LIDC): -1.515 (0.424)
  - combination (LIDC): -0.82 (0.023)
- MSE:
  - pooled: 0.098 (0.014)
  - country-specific: 0.092 (0.012)
  - combination: 0.089 (0.012)
  - pooled (LIDC): 0.224 (0.024)
  - country-specific (LIDC): 0.193 (0.019)
  - combination (LIDC): 0.198 (0.02)
- AUROC (accuracy of rankings):
  - pooled (AE+EM): 0.362 (0.031)
  - country-specific (AE+EM): 0.750 (0.03)
  - combination (AE+EM): 0.744 (0.031)
  - pooled (LIDC): 0.379 (0.033)
  - country-specific (LIDC): 0.701 (0.035)
  - combination (LIDC): 0.695 (0.035)
- Sample sizes and crisis counts:
  - size of test sample: 132 313 132 313 555 555 555 (table lists these concatenated)
  - of which: crisis observations: 136 136 136 165 165 165
- Note: backward-looking averages used as predicted probability of a crisis start occurring in year t+1 or t+2; test sample covers t = 2001-2015. Bootstrapped standard deviations in parentheses (with two-way clustering by country and year).

### Baseline predictors: benchmarking in-sample and out-of-sample performance of econometric approaches (Table 5.1)
- Training sample coverage and start years reported for variants (AE + EM, LIDC, combinations) with start years: 2001, 1979, 1989.
- Accuracy of predictions (selected entries shown):
  - log(likelihood) examples:
    - -0.28* (0.037)
    - -0.302 (0.034)
    - -0.31 (0.036)
    - -0.231*** (0.035)
    - -0.332 (0.041)
  - log(likelihood), weighted examples:
    - -1.03*** (0.058)
    - -1.038*** (0.062)
    - -1.094*** (0.078)
    - -0.844 (0.071)
  - MSE examples:
    - 0.081 (0.013)
    - 0.086 (0.012)
    - 0.087 (0.012)
    - 0.064*** (0.012)
  - AUROC examples:
    - 0.756 (0.037)
    - 0.697 (0.039)
    - 0.682* (0.04)
    - 0.822** (0.027)
- Average and test sample sizes:
  - average size of training sample examples: 1148 1309 994 1148 1309 994 (concatenated entries)
  - of which: crisis observations examples: 112 146 108 112 146 108
  - size of test sample: 1148 (repeated)
  - of which: crisis observations: 112 (repeated)
- Low-income developing countries (LIDC) block examples:
  - log(likelihood): -0.638 (0.033), -0.692 (0.038), -0.705 (0.034)
  - weighted log(likelihood): -0.697*** (0.023), -0.741** (0.034), -0.733*** (0.033)
  - MSE: 0.223 (0.015), 0.246 (0.017), 0.253 (0.015)
  - AUROC: 0.627 (0.044), 0.553 (0.054), 0.547* (0.052)
  - average training sample: 355 276 218
  - crisis observations in training: 134 122 97
  - test sample size: 355; crisis observations: 134

### Baseline predictors: predictive performance of different models without imputation or pooling (Table 5.2.A)
- Advanced and emerging market economies (AE + EM):
  - Methods: Logit, Elastic net, XG boost, Random forest
  - log(likelihood):
    - Logit: -0.302 (0.034)
    - Elastic net: -0.3 (0.033)
    - XG boost: -0.287 (0.034)
    - Random forest: -0.281 (0.027)
  - weighted log(likelihood):
    - Logit: -1.038 (0.062)
    - Elastic net: -1.053 (0.035)
    - XG boost: -0.902 (0.055)
    - Random forest: -0.906 (0.045)
  - MSE:
    - Logit: 0.086 (0.012)
    - Elastic net: 0.085 (0.012)
    - XG boost: 0.083 (0.012)
    - Random forest: 0.081 (0.011)
  - AUC / AUROC:
    - Logit: 0.697 (0.039)
    - Elastic net: 0.697 (0.033)
    - XG boost: 0.750 (0.033)
    - Random forest: 0.772 (0.024)
  - average size of training sample: 1309 (all methods)
  - size of test sample: 1148; of which crisis observations: 112
- Low-income developing countries (LIDC):
  - log(likelihood):
    - Logit: -0.692 (0.038)
    - Elastic net: -0.666 (0.012)
    - XG boost: -0.68 (0.024)
    - Random forest: -0.681 (0.026)
  - weighted log(likelihood):
    - Logit: -0.741 (0.034)
    - Elastic net: -0.7 (0.004)
    - XG boost: -0.718 (0.016)
    - Random forest: -0.702 (0.022)
  - MSE:
    - Logit: 0.246 (0.017)
    - Elastic net: 0.237 (0.006)
    - XG boost: 0.243 (0.011)
    - Random forest: 0.244 (0.012)
  - AUROC:
    - Logit: 0.553 (0.054)
    - Elastic net: 0.546 (0.028)
    - XG boost: 0.530 (0.041)
    - Random forest: 0.577 (0.043)
  - training sample: 276; crisis observations: 122
  - test sample: 355; crisis observations: 134
- Note: All models use the same 10 variables to predict probability of a crisis start in t+1 or t+2; test sample covers t = 2001-2015; training samples start in 1979. Bootstrapped standard deviations in parentheses (two-way clustering by country and year).

### Significance of differences in model performance (Table 5.2.B and Table 5.3.B)
- Tables report degree of confidence (1 - p-value) that model in row i outperforms model or benchmark in column j, based on one-sided t-tests using bootstrapped standard errors adjusted for two-way clustering by country and year.
- Entries include many cell values indicating confidence levels; examples (AUC, AE + EM block):
  - logit vs elastic net: 0.83
  - logit vs XG boost: 0.97
  - elastic net vs XG boost: 0.92
  - random forest vs others: values up to 0.99 in some cells
- Tables cover metrics: log(likelihood), weighted log(likelihood), MSE, AUROC for AE + EM and LIDC samples.
- Table 5.3.B focuses on logit model variations (no pooling/no imputation; pooling/no imputation; no pooling/imputation; pooling/imputation) and reports confidence values for performance differences across global average, country average, mix of averages, for AE + EM and LIC.

### Baseline predictors: logit model performance with and without imputation and pooling (Table 5.3.A)
- Advanced and emerging market economies (examples):
  - impute in training sample: no / yes (variants)
  - log(likelihood):
    - -0.302 (0.034)
    - -0.298 (0.038)
    - -0.297 (0.033)
    - -0.295 (0.039)
  - weighted log(likelihood):
    - -1.038 (0.062)
    - -1.093 (0.06)
    - -0.978 (0.057)
    - -1.087 (0.059)
  - MSE: 0.086 (0.012), 0.084 (0.013), 0.085 (0.012), 0.083 (0.013)
  - AUROC: 0.697 (0.039), 0.705 (0.037), 0.714 (0.038), 0.712 (0.036)
  - average size of training sample examples: 1309 1585 1888 2679 (concatenated); crisis observations in training vary (e.g., 146, 268, 238, 471)
  - size of test sample: 1148; crisis observations: 112
- Low-income countries (examples):
  - log(likelihood): -0.692 (0.038), -0.695 (0.04), -0.709 (0.063), -0.672 (0.045)
  - weighted log(likelihood): -0.741 (0.034), -0.724 (0.051), -0.817 (0.059), -0.733 (0.049)
  - MSE: 0.246 (0.017), 0.250 (0.016), 0.246 (0.024), 0.238 (0.018)
  - AUROC: 0.553 (0.054), 0.536 (0.051), 0.572 (0.048), 0.556 (0.048)
  - average training sample sizes listed: 276 1585 791 2679 (concatenated); crisis observations in training: 122 268 233 471
  - size of test sample: 355; crisis observations: 134

### Baseline predictors: predictive performance with imputation and pooling (Table 5.4)
- Advanced and emerging market economies (all countries, impute in training sample = yes; impute in test sample = no):
  - Methods: Logit, Elastic net, XG boost, Random forest
  - log(likelihood):
    - Logit: -0.295 (0.039)
    - Elastic net: -0.294 (0.035)
    - XG boost: -0.281 (0.032)
    - Random forest: -0.283 (0.024)
  - weighted log(likelihood):
    - Logit: -1.087 (0.059)
    - Elastic net: -1.057 (0.043)
    - XG boost: -0.924 (0.044)
    - Random forest: -0.882 (0.036)
  - MSE:
    - Logit: 0.083 (0.013)
    - Elastic net: 0.083 (0.013)
    - XG boost: 0.082 (0.012)
    - Random forest: 0.082 (0.01)
  - AUROC:
    - Logit: 0.712 (0.036)
    - Elastic net: 0.716 (0.037)
    - XG boost: 0.761 (0.033)
    - Random forest: 0.777 (0.022)
  - average size of training sample: 2679 2679 2679 2679 (all methods); crisis observations: 471
  - size of test sample: 1148; crisis observations: 112
- Low-income developing countries (all countries, impute in training sample = yes; impute in test sample = no):
  - log(likelihood):
    - Logit: -0.672 (0.045)
    - Elastic net: -0.669 (0.046)
    - XG boost: -0.688 (0.049)
    - Random forest: -0.669 (0.03)
  - weighted log(likelihood):
    - Logit: -0.733 (0.049)
    - Elastic net: -0.737 (0.044)
    - XG boost: -0.778 (0.041)
    - Random forest: -0.734 (0.031)
  - MSE:
    - Logit: 0.238 (0.018)
    - Elastic net: 0.237 (0.018)
    - XG boost: 0.244 (0.021)
    - Random forest: 0.237 (0.013)
  - AUROC:
    - Logit: 0.556 (0.048)
    - Elastic net: 0.555 (0.048)
    - XG boost: 0.555 (0.047)
    - Random forest: 0.581 (0.038)
  - average size of training sample: 2679 (all methods); crisis observations: 471
  - size of test sample: 355; crisis observations: 134
- Note: All models use the same 10 variables to predict crisis start in year t+1 or t+2; test sample covers t = 2001-2015. Bootstrapped standard deviations in parentheses (with two-way clustering by country and year).

*Source: wpiea2021150-print-pdf - introduction,” In Inverse Problems and High-Dimensional Estimation (pp. 121-156).*

### 2015. Out-of-sample performance obtained from 15 rolling regressions. Training samples start in 1979.

### 2015. Out-of-sample performance obtained from 15 rolling regressions. Training samples start in 1979.

### Predictive performance (Table 6.1.A) — baseline specification (with imputation and pooling)
- Testing framework:
  - Models predict probability of crisis start occurring in year t+1 or t+2; test sample covers t = 2001-2015.
  - Out-of-sample performance obtained from 15 rolling regressions. Training samples start in 1979 and pool all countries.
  - Missing values in the test and training samples are imputed using the training sample median.
  - Bootstrapped standard deviations in parentheses (with two-way clustering by country and year, following Cameron et al. 2008).

- (i) Advanced and emerging market economies — accuracy of predictions:
  - log(likelihood): -0.305, -0.322, -0.304, -0.294, -0.293, -0.289, -0.289, -0.3, -0.293, -0.314 (0.038) (0.047) (0.047) (0.037) (0.032) (0.036) (0.039) (0.027) (0.029) (0.031)
  - log(likelihood), weighted: -1.081, -1.158, -1.051, -0.978, -0.94, -0.945, -0.875, -0.747, -0.799, -0.886 (0.052) (0.101) (0.059) (0.051) (0.041) (0.061) (0.044) (0.029) (0.041) (0.025)
  - MSE: 0.087, 0.090, 0.087, 0.087, 0.085, 0.085, 0.085, 0.088, 0.087, 0.089 (0.013) (0.013) (0.017) (0.013) (0.012) (0.013) (0.015) (0.01) (0.011) (0.012)
  - AUROC: 0.714, 0.748, 0.717, 0.766, 0.750, .771, 0.780, .802, 0.787, 0.744 (0.033) (0.032) (0.038) (0.032) (0.03) (0.031) (0.031) (0.026) (0.029) (0.031)

- (ii) Low-income developing countries — accuracy of predictions:
  - log(likelihood): -0.593, -0.645, -0.59, -0.576, -0.591, -0.575, -0.58, -0.557, -0.558, -0.578 (0.042) (0.062) (0.027) (0.046) (0.046) (0.049) (0.022) (0.036) (0.038) (0.049)
  - log(likelihood), weighted: -0.73, -0.771, -0.735, -0.715, -0.767, -0.74, -0.725, -0.665, -0.661, -0.82 (0.046) (0.08) (0.044) (0.047) (0.038) (0.055) (0.039) (0.028) (0.038) (0.023)
  - MSE: 0.203, 0.216, 0.202, 0.198, 0.202, 0.195, 0.199, 0.189, 0.190, 0.198 (0.017) (0.021) (0.011) (0.019) (0.019) (0.019) (0.009) (0.015) (0.016) (0.02)
  - AUROC: 0.62, 0.662, 0.623, 0.676, 0.634, 0.683, 0.654, 0.705, 0.70, .695 (0.039) (0.041) (0.039) (0.042) (0.044) (0.042) (0.034) (0.04) (0.041) (0.035)

- Note on methods and predictors:
  - Methods reported include logit, logit with BIC, random forest, random forest with rfe, elastic net, XG boost, combination of country-specific and pooled average, and others.
  - Number of predictors reported across specifications include 10, 74, 81, 07, 48, 10, 74, 81, 07, 48 (as shown in the table header).

### Statistical significance of performance differences (Table 6.1.B)
- Table purpose:
  - Indicates degree of confidence (1 - p-value) that the model in row i outperforms the model (or historical benchmark) in column j, based on a one-sided t-test using bootstrapped standard errors, adjusted for two-way clustering by country and year (following Cameron et al., 2008).
- Metrics reported in table annotations:
  - log(likelihood), AE + EM; log(likelihood), LIDC
  - weighted log(likelihood), AE + EM; weighted log(likelihood), LIDC
  - MSE, AE + EM; MSE, LIDC
  - AUC, AE + EM; AUC, LIDC
- (Table cells present many pairwise confidence values between model pairs such as logit with BIC, elastic net, XG boost, random forest, random forest with rfe and combinations of global average, country average, and mix of averages; values shown in the table include entries like 0.88, 0.95, 0.37, -0.049, 0.04, 0.19, 0.10, 0.59, 0.95, 0.04, -0.000, 0.02, 0.01, 0.01, 10.97, 0.96, 0.95, -0.250, 0.650, 0.450, 0.99, 0.97, 0.53, 1.00, -0.470, 0.14, 0.13, and many more as printed.)
- Note: the table contains a dense matrix of pairwise comparisons; users should consult the full table layout for exact row/column pairings and values as presented.

### Predictor importance and rankings (Table 7.1)
- Table lists all predictors that are ranked 30th or higher for at least one prediction method. Numbers indicate ranks in terms of variable importance (see Section 7).
- Examples of top-ranked predictors and their ranks (across methods: logit, elastic net, XG boost, random forest):
  - Crisis history:
    - Historical crisis frequency: 31 (+), 73
    - Years since last crisis: 8 (-), 1 (-), 24
    - Number of countries with banking crisis starts, current year: 21 (+), 486
  - Output, demand, prices:
    - log(GDP per capita), PPP level: 2 (-), 2 (-), 31
    - log(GDP per capita), PPP lag1: 12
    - log(GDP per capita), PPP 10-year change: 66 (-), 2789
    - log(GDP), US dollars level: 8 (-), 22 23
    - Real GDP per capita 3-year growth rate: 19 (-), 42 (-), 90 112
    - Gross domestic savings / GDP level: 16 (-), 23 56
    - Trading partner real GDP 3-year growth rate: 30 163
    - Natural resource rents / GDP level: 12 (-), 54121
  - Volatility:
    - CPI inflation 10-year std. dev.: 30 (+), 12 14
    - Terms of trade growth 10-year std. dev.: 26 (+), 56 (+), 20 55
  - External sector:
    - Current account balance / GDP level: 10 (-), 9 (-), 21 25
    - External gross financing needs / GDP level: 16 122
    - Official development assistance / GDP level: 129 26
    - Official reserves (in months of imports) level: 6 50
  - Population and demographics:
    - Population density level: 6 (-) 10 (-) 14 71
    - Working-age population, share of total level: 12 (-) 6
  - Fiscal, debt, and debt-service related ranks are reported across many indicators (e.g., Public external debt / GDP level: 7 (+) 11 (+) 48; Public external debt amortization / reserves level: 539; Total external debt amortization / reserves lag: 13 (+) 18 (+) 23330), among others.
- Country characteristics and global variables:
  - Island (dummy): 15 (-) 25 22 44
  - Currency union member (dummy): 14 (+) 5 (+) 65188
  - US Treasury bill rate lag: 29 (+) 270364
  - VIX, end of period 3-year change: 23 (+) 539
  - World nominal GDP (million USD) 5-year change: 3 (+) 46248

### Annex A — definitions, samples, and data description
- Fiscal Crisis: Definitions and Data Sources (Annex Table 1)
  - Criterion: Minimum two years gap between crises.
  - Default, restructuring, or rescheduling:
    - (i) of substantial size (in percent of GDP p.a.); AND
    - (ii) defaulted nominal amount grows by a substantial amount (in percent p.a)
  - High-access IMF financial arrangement with fiscal adjustment objective in place (in percent of quota); OR EU program; Baldacci and others (2011); IMF.
  - High inflation rate (in pct. of growth of annual average CPI p.a.) OR ≥ 35; sources: Baldacci and others (2011); Sturzenegger and Zettelmeyer (2006); and Fisher, Sahya and Vegh (2002); IMF (World Economic Outlook).
  - Steep increase in domestic arrears (in first difference of the ratio of 'other account payables (OAP)' to GDP in percentage points); sources include Checherita-Westphal, Klem, and Viefers (2015); Reinhart and Rogoff (2011a); Eurostat; OECD.
  - High price of market access (in basis points of sovereign spreads or CDS spreads) OR (a) Level of spreads (bps) — Sy (2004); Baldacci and others (2011); (b) Annual change in spreads (bps) ≥ 300 ≥ 650 na; sources: IMF (2015); Kose and others (2017); Guscina, Sheheryar, and Papaioannou (2017); Gelos, Sahay, and Sandleris (2004); Reuters Datastream; Bloomberg.
  - The database mainly contains external defaults on sovereign debt denominated in foreign currency; sources listed include BoC-BoE Sovereign Default Database complemented with information from IMF desks; Cruces and Trebesch (2013); World Bank; Detragiache and Spilimbergo (2001); Chakrabarti and Zeaiter (2014); Reinhart and Rogoff (2011b).
  - Thresholds and event labels in annex table include entries such as: (1) Credit Event >0.5; ≥ 10; Event (4); Loss of market confidence ≥ 1,000 bps when market access is lost (after maintaining market access for a 1/4 of the sample time and 2 consecutive years before the loss year); ≥ 100; (3) Implicit domestic public default ≥ 100 ≥ 1; (2) Exceptionally large official financing.

- Annex Table 2 — Sample of Countries:
  - The annex provides an extensive country listing grouped under Advanced Economies, Emerging Markets, and Low Income Developing Countries. (Country names are listed in full within the annex.)

- Annex Table 3 — Description and sources of data:
  - Provides variable-by-variable descriptions, primary data sources, and permutations used (t, t-1, fd_t, fd_t-1, 3-year change, 5-year change, 10-year change, W, L2, pc3_t, etc.).
  - Examples of variable categories and sources:
    - Country category variables: Dummy: Monetary union member — IMF WEO; Dummy: Island country — Wikipedia; Dummy: Landlocked country — CIA World Factbook; Dummy: Small state — Authors' calculations, IMF WEO, WB WDI; Dummy: Fragile state — WB (FY2017 list); Dummy: Commodity exporter — IMF WEO.
    - Contagion / Crisis history: Contagion: number countries with fiscal crisis start — Medas et al (2018); Years passed since last fiscal crisis — Medas et al (2018); Historical fiscal crisis frequency — Medas et al (2018); Dummy: Banking crisis start — Laeven and Valencia (2018).
    - External sector variables: Net official development assistance (% of GDP) — OECD; Current account balance (% of GDP) — IMF WEO; External gross financing needs — author's calculations; IMF WEO; WDI.
    - Fiscal variables: General government expenditures (% of GDP) — IMF WEO; General government revenues in percent of GDP — IMF WEO; Stock and flow adjustments to public debt — author's calculations.
    - Global variables: Percent change of crude oil price — IMF WEO (GAS Live); VIX — Bloomberg; US T-Bill rate — IFS.
    - Institutions / elections: Revised Combined Polity Score — Center for Systemic Peace; Checks and balances index — DPI; Bureaucracy Quality — PRS Group; Corruption — PRS Group; Political Stability and Absence of Violence/Terrorism — WB WGI.
    - Demographics: Population ages 15-64 — WDI; Urban population (% of total) — WDI; Population density — WDI.
    - Private debt variables: (One-sided) credit gap — GDD; Total Debt, loans and securities, (% of GDP) — GDD; Domestic credit to private sector by banks (% of GDP) — WDI.
    - Public debt and debt-service variables: Public external debt (% of GDP) — IMF WEO; Public debt (% of GDP) — GDD; General government short-term external debt (% of GDP) — IMF WEO; Debt service measures — IMF WEO; WDI; IFS.
    - Real sector variables: log(real GDP per capita) (PPP) — IMF WEO; Percent change of real GDP per capita — IMF WEO; Percent change of period average consumer price index — IMF WEO.
  - Note: t = current year value; t-1 = past year value; fd_t = first difference; fd_t = lagged of first difference; W = cross sectional weighted average for all the permutations listed. WEO = World Economic Outlook; DPI = Cruz et al. (2016); GDD = Mbaye et al. (2018); IFS = International Financial Statistics; WDI = World Development Indicators; WGI = Worldwide Governance Indicators.

*Italic final source attribution: Content derived from wpiea2021150-print-pdf (2015) as provided.*

### Annex B: Cross-validation and hyperparameter tuning

### Annex B: Cross-validation and hyperparameter tuning

### Cross-validation procedure
- To select (tune) an algorithm’s hyperparameter values, a grid of candidate values is searched and the value that delivers the smallest average log-likelihood loss when cross-validated is selected.
- Cross-validation algorithm:
  - From the years included in the training sample, select a year t* that serves as the evaluation fold.
  - For the given hyperparameter value, estimate the model on the training sample after dropping observations between years t*-1 and t*+1.
  - Using the model estimates from step 2, make predictions for year t* and compute the log-likelihood loss.
- The steps above are repeated for each year in the training sample and then the average log-likelihood loss is taken.
- The procedure is repeated for each candidate hyperparameter value, so that each hyperparameter value is associated with a different loss function value.
- The hyperparameter value that minimizes the loss function is selected and used to re-estimate the model on the training sample.
- Hyperparameter tuning and model estimation make no use of the test sample (steps 1-3 above are done within the training sample). The resulting model is then used to make predictions for observations in the test sample.

### Tuning parameter grids (Annex Table 4)
- Method: Logit
  - Hyperparameter: - - - -
  - Tuning grid: glm, glmStepAIC
  - Selected value (in Section VII): (none specified)
- Method: Classification tree
  - Hyperparameter: - - - -
  - Tuning grid: rpart2
  - Selected value (in Section VII): (none specified)
- Method: Elastic net
  - Hyperparameter: α
    - Tuning grid: 11 [0, 1]
    - Selected value (in Section VII): 1
    - Method reference in caret package: glmnet
  - Hyperparameter: λ
    - Tuning grid: 400 [.003, 10]
    - Selected value (in Section VII): .0096
    - Method reference in caret package: glmnet
- Method: XG boost
  - Hyperparameter: number of trees
    - Tuning grid: 6 [10, 85]
    - Selected value (in Section VII): 70
    - Method reference in caret package: xgbTree
  - Hyperparameter: maximum depth
    - Tuning grid: 5 [1, 14]
    - Selected value (in Section VII): 14
    - Method reference in caret package: xgbTree
  - Hyperparameter: η (learning speed)
    - Tuning grid: 4 [.05, .6]
    - Selected value (in Section VII): .05
    - Method reference in caret package: xgbTree
  - Hyperparameter: γ (minimum gain required for split)
    - Tuning grid: 5 [.005, 6]
    - Selected value (in Section VII): 1.5
    - Method reference in caret package: xgbTree
  - Hyperparameter: minimum child weight
    - Tuning grid: 5 [1, 40]
    - Selected value (in Section VII): 11
    - Method reference in caret package: xgbTree
- Method: Random forest
  - Hyperparameter: m
    - Tuning grid: try 15 dynamic
    - Selected value (in Section VII): 38
    - Method reference in caret package: rf

*Source: wpiea2021150-print-pdf - Annex B: Cross-validation and hyperparameter tuning*

---


_Source: https://www.imf.org/-/media/files/publications/wp/2021/english/wpiea2021150-print-pdf.pdf_
