## wpiea2020001-print-pdf

## Source details

**Canonical URL:** [wpiea2020001-print-pdf](https://www.imf.org/-/media/files/publications/wp/2020/english/wpiea2020001-print-pdf.pdf)

## Other formats

- [Markdown version](/-/media/files/publications/wp/2020/english/wpiea2020001-print-pdf.pdf.md)
- [Structured JSON version](/-/media/files/publications/wp/2020/english/wpiea2020001-print-pdf.pdf.json)

---

### I. Introduction and empirical strategy
- Research question: whether current high public debt levels are a bellwether of future fiscal crises with large economic costs.
- Motivation and background:
  - Standard debt sustainability frameworks imply governments should worry about high public debt because real economic growth dips in the immediate aftermath of a fiscal crisis and the loss of output is often permanent (Medas et al. 2018; Asonuma et al. 2019).
  - Argument that “public debt may have no fiscal cost” (Blanchard 2019) is reinforced by historically low interest rates and the global stock of negative-yielding debt being around $12 trillion by the end of 2019.
  - If interest rates are lower than the economic growth rate (the interest-growth differential is negative), issuing debt can be feasible without later increasing taxes.
- Empirical strategy:
  - Sample: 188 countries dating back to the 1980s; identified 418 crisis episodes over 1980–2016.
  - Predictor universe: 748 indicators across debt, economic activity, fiscal aggregates, external indicators, global factors, demographics, and institutions.
  - Two-step procedure:
    1. Fit flexible machine learning models to capture nonlinearities and interactions while avoiding overfitting.
    2. Apply feature selection to reduce 748 indicators to those containing signal rather than noise.
  - Post-selection interpretation: use variable importance, partial dependence plots (PDPs), H-statistics, Shapley values, and event studies to interpret results.

### II. Key empirical findings
- Predictive hierarchy:
  - Public debt is the most important group of predictors; public external debt has especially high predictive value.
  - Public debt service ranks closely behind public debt.
  - Interest-growth differential ("r-g") has low predictive value.
- Nonlinearities and interactions:
  - Beyond certain debt levels, the likelihood of fiscal crises increases significantly irrespective of whether the interest-growth differential is highly positive or negative.
  - Crisis probability rises steeply at high public debt levels and also at moderate debt levels when accompanied by high inflation or large current account deficits.
  - Public external debt exhibits the strongest cumulative interactions with other features.
- Robustness:
  - Results hold across all income groups (AEs, EMs, LICs).
  - Performance assessed via out-of-sample predictive accuracy and stability of variable selection with respect to sampling variation.

### III. Crisis definitions, sample statistics, and crisis composition
- Crisis types defined (a year is a fiscal crisis year if at least one criterion met):
  - Credit events: debt service not paid on due date or creditor incurs losses including through debt restructuring.
  - Exceptionally large official financing: large IMF or EU support.
  - Implicit domestic public debt default: (1) periods of high inflation, or (2) accumulation of domestic arrears.
  - Loss of market confidence: loss of market access or very large borrowing costs/yield spikes.
- Sample incidence and basic statistics:
  - Identified crisis episodes: 418 for 188 countries over 1980–2016.
  - Average crises per country since 1980: 2 fiscal crises (reported as Total 2.2 average per country elsewhere).
  - Share of countries with at least one crisis since 1980: more than three quarters.
  - Frequency by income group (at any point in time):
    - LICs: about two-thirds are in fiscal distress.
    - EMs: on average, 40 percent.
    - AEs: less than 15 percent.
  - Decade concentration: 1990s highest; early 1980s and around 2010 show bunching.
- Crisis-type composition (selected figures from Table 1 — crisis-start criteria percent):
  - Credit event: Total 72.7; AEs 20.0; EMs 68.3; LICs 84.3
  - Exceptionally large official financing: Total 33.7; AEs 56.0; EMs 34.7; LICs 29.8
  - Implicit domestic default: Total 9.8; AEs 24.0; EMs 11.4; LICs 6.3
  - Loss of market confidence: Total 25.1; AEs 84.0; EMs 32.2; LICs 9.9
  - Average duration (years): Total 5.2; AEs 3.2; EMs 5.3; LICs 5.3
  - Number of countries with no fiscal crisis: Total 36; AEs 19; EMs 15; LICs 2
  - Number of countries by group with crises starts: AEs: 25; EMs: 202; LICs: 191
  - Total number of crises starts: 418

### IV. Data, debt metrics, and r-g measurement
- Debt data:
  - Assembled comprehensive range: public debt, public external debt, private indebtedness, total debt; uses Global Debt Database and other sources.
  - Debt scaled by GDP and complements such as reserves or revenues used.
  - Public external debt defined by residency of holder; includes public and publicly guaranteed debt.
- Interest-growth differential (r-g) definition and calculation:
  - r-g defined as (r−g)/(1+g) where r is the effective interest rate and g is the GDP growth rate.
  - Effective interest rates calculated using consistent time series for the stock of public debt.
  - Data suggests r-g has been close to zero or negative since the 1980s on average across income groups, but with wide dispersion and positive r-g not uncommon.
  - Stock-flow adjustment (SFA) constructed: SFA_t = d_t − ( (1+r)/(1+g) ) d_{t−1} + p_t, where d_t and d_{t-1} are stocks of public debt and p_t is the primary balance.
- Predictor inclusion rule:
  - A variable is included if 70 percent of the data exists.
  - Missing values imputed with the training sample median of non-missing values.

### V. Methodology — machine learning workflow and evaluation
- Model of choice: Random Forest (RF, Breiman 2001) selected for ability to capture nonlinearities and interactions and deliver predictive improvements.
- Feature selection algorithms (RF-based) applied to 748 variables:
  - P-values via permutation importance (PIMP).
  - Recursive Feature Elimination (RFE).
  - Boruta (shadow variable comparison).
  - VSURF (two subsets: interpretative and predictive).
- Training and evaluation:
  - Prediction window: two years; ŷ ∈ [0,1] predictive probability of fiscal crisis over next two years.
  - Rolling cutoff procedure: 15 rolling regressions starting with training sample 1980–2000.
  - Hyperparameter tuning: m_try chosen by k-fold cross-validation where k equals number of years in training sample; tuning length 10; number of trees = 2000; trees grown exhaustively.
  - Evaluation metrics averaged across 15 test samples: AUROC (AUC), Log-likelihood, Mean squared error (MSE).
  - Out-of-sample performance comparison also against LASSO and SVM.

### VI. Variable selection results and model performance
- Feature set sizes by algorithm (Figure 8 bar values):
  - RF: 748; Boruta: 336; PIMP: 176; RFE: 68; VSURF: 8.
- Stability (Pearson index) when randomly dropping 5 percent of observations:
  - Boruta: 0.92
  - RFE: 0.63
  - PIMP: 0.81
- Out-of-sample performance (selected reported figures; SDs in parentheses):
  - Advanced and emerging market economies:
    - AUROC: Boruta 0.805 (0.025); RFE 0.793* (0.026); VSURF 0.734*** (0.031); PIMP 0.791** (0.026); RF 0.806 (0.026)
    - Log(likelihood): Boruta -0.289*** (0.027); RFE -0.286** (0.029); VSURF -0.34 (0.043); PIMP -0.29** (0.028); RF -0.299 (0.026)
    - MSE: Boruta 0.084** (0.011); RFE 0.085 (0.011); VSURF 0.096*** (0.011); PIMP 0.085 (0.011); RF 0.087 (0.01)
  - Low-income countries:
    - AUROC: Boruta 0.708 (0.042); RFE 0.712 (0.042); VSURF 0.688 (0.039); PIMP 0.707 (0.042); RF 0.706 (0.043)
    - Log(likelihood): Boruta -0.547* (0.043); RFE -0.544** (0.044); VSURF -0.601* (0.052); PIMP -0.548* (0.043); RF -0.556 (0.039)
    - MSE: Boruta 0.186 (0.018); RFE 0.185 (0.018); VSURF 0.207** (0.019); PIMP 0.187 (0.018); RF 0.189 (0.017)
- Given predictive power and stability, Boruta chosen as benchmark algorithm.
- RF vs alternatives (Appendix Table A5.1 — selected figures):
  - Advanced and emerging market economies:
    - AUROC: RF 0.806; Lasso 0.763**; SVM 0.767*** (SDs: (0.026), (0.03), (0.025))
    - Log(likelihood): RF -0.299; Lasso -0.299; SVM -0.295 (SDs: (0.026), (0.036), (0.032))
    - MSE: RF 0.087; Lasso 0.089; SVM 0.087 (SDs: (0.01), (0.013), (0.012))
  - Low-income countries:
    - AUROC: RF 0.706; Lasso 0.676; SVM 0.682 (SDs: (0.043), (0.043), (0.04))
    - Log(likelihood): RF -0.556; Lasso -0.581; SVM -0.587 (SDs: (0.039), (0.052), (0.056))
    - MSE: RF 0.189; Lasso 0.2*; SVM 0.201 (SDs: (0.017), (0.021), (0.022))

### VII. Variable importance, Shapley analysis, and PDP findings
- Group importance (Boruta RF out-of-bag permuted predictor importance):
  - Top groups: Public debt; Public debt service; Level of development (GDP per capita); Demographics; External (capital flows); Private debt; Economic activity; External debt; Current account; Exchange rate; Fiscal; Institutions.
  - Interest-growth differential and global conditions among the least relevant groups.
- Shapley differences (crisis vs non-crisis):
  - Public debt and public debt service most important for EMs and LICs; public debt among top 3 for AEs.
  - Private debt appears more important than public debt for AEs.
  - Public debt and public debt service top categories both pre- and post-2000, with public debt highest post-2000.
- Public external debt — univariate PDPs (selected threshold findings):
  - Positive, non-linear relationship with crisis probability.
  - AEs: probability increases substantially once public external debt ≈ 70 percent of GDP.
  - EMs: flat for debt < 30 percent of GDP; rises steeply above 30 percent.
  - LICs: predicted probabilities much higher from the start; steepening at lower debt levels.
- Crisis probability thresholds minimizing sum of type I and type II errors:
  - AEs: 8.5 percent — captures 80 percent of crises while false alarms at 20 percent.
  - EMs: 22 percent.
  - LICs: 28 percent.
- High debt levels where estimated probability breaches threshold irrespective of other factors:
  - AEs and EMs: 70 percent of GDP.
  - LICs: 80 percent of GDP (higher threshold possibly due to minimization trade-off and higher concessional borrowing share).

### VIII. Interactions: public external debt with r-g, inflation, current account, credit gap
- Interaction measures:
  - H-statistic ranks pairs by interaction strength; public external debt has highest cumulative interactions.
  - Top interacting variables with public external debt: external assets, inflation rate, GDP relative to US, public external amortization.
- Public external debt × interest-growth differential (bivariate PDPs):
  - AEs and EMs: if public external debt sufficiently high, estimated probability breaches threshold irrespective of r-g.
  - LICs: both highly negative and positive r-g can imply higher crisis probability for same debt level.
  - Event studies: r-g spikes at onset of crisis, reducing value as a pre-crisis signal.
- Inflation:
  - Univariate PDPs: strong non-linear relationship; crisis probability for AEs and EMs increases significantly when inflation > 20 percent.
  - LICs: high inflation reduces debt thresholds at which crisis probability breaches group thresholds.
  - Both increases and declines in inflation can associate with higher crisis probability, notably for EMs.
- Current account:
  - Univariate PDPs: external deficits between 3–7 percent of GDP substantially increase crisis probability for AEs and EMs.
  - Interaction: moderate debt levels yield steep crisis probability rises when current account deficits are high.
- Credit gap (private sector leverage):
  - Probability of crisis increases significantly in AEs and EMs when credit gap > 40 percent.
  - Interactions: for EMs, large credit gaps lower public external debt levels needed to breach crisis thresholds.
  - Private debt dynamics less relevant for LICs.

### IX. Interest-growth differential: interpretation and limits as an early warning
- r-g limited as a leading indicator:
  - Variable importance analysis: r-g among least relevant predictors.
  - PDPs: large variations in r-g produce barely any change in estimated crisis probability; never breaches crisis thresholds on its own.
  - Event study: r-g often remains low long before crisis and shoots up at onset (t=0), undermining pre-crisis signaling value.
- Interpretation caveat: r-g matters for debt dynamics but has limited predictive/signal value for crisis onset in this analysis.

### X. Conclusions and policy implications
- Empirical conclusions:
  - Public debt in its various forms is the most important predictor of fiscal crises; public external debt particularly discriminating.
  - Nonlinear thresholds and interactions mean crisis risk depends on combinations of vulnerabilities (debt together with high inflation, large current account deficits, high private leverage).
  - Interest-growth differential has limited signaling value as a pre-crisis indicator.
- Policy implications:
  - Complacency about high public debt would be ill-advised even if r-g remains low.
  - High public debt entails non-linear risks that can escalate quickly; by the time r-g signals distress, a crisis may already be underway.
  - Reducing debt is not universally the correct policy; debt used for countercyclical purposes, public investment, or structural needs can be appropriate.
  - Risk management should consider interacting vulnerabilities (external balances, inflation, private leverage, institutional strength) rather than single indicators alone.
- Limitations:
  - Machine learning models identify correlations and predictive associations; they do not establish causality.

*Source: wpiea2020001-print-pdf - References*

### References .............................................................................................................

### wpiea2020001-print-pdf - References

### I. Introduction: research question and approach
- Research question: whether current high public debt levels are a bellwether of future fiscal crises with large economic costs.
- Key motivations and background points:
  - Standard debt sustainability frameworks imply governments should worry about high public debt because real economic growth dips in the immediate aftermath of a fiscal crisis and the loss of output is often permanent (Medas et al. 2018; Asonuma et al. 2019).
  - The argument that “public debt may have no fiscal cost” (Blanchard 2019) is reinforced by historically low interest rates and the global stock of negative-yielding debt being around $12 trillion by the end of 2019.
  - If interest rates are lower than the economic growth rate (the interest-growth differential is negative), issuing debt can be feasible without later increasing taxes.
- Empirical strategy overview:
  - Use machine learning models to identify robust predictors of fiscal crises and assess whether public debt is a reliable leading indicator by itself or via interactions.
  - Address methodological challenges: crises are rare, debt data limitations, and difficulty capturing nonlinearities and interactions with classic econometric techniques.
  - Sample: broad sample of 188 countries dating back to the 1980s.
  - Predictor universe:
    - Start with a wide range of predictors commonly associated with crises.
    - Consider many permutations yielding a total of 748 indicators.
  - Two-step procedure:
    1. Fit flexible machine learning models to avoid overfitting while capturing complex functional forms.
    2. Apply feature selection to reduce the 748 indicators to those containing signal rather than noise.
  - Post-selection interpretation: use statistical measures to go beyond the black box and uncover variable importance and interactions.

### I.1. Key empirical findings (as reported)
- Public debt is the most important group of predictors.
- Relative importance details:
  - Some forms of debt are more important than others, in particular public external debt.
  - The interest-growth differential has low predictive value.
- Nonlinearities and interactions:
  - Beyond certain debt levels, the likelihood of fiscal crises increases significantly irrespective of whether the interest-growth differential is highly positive or negative.
  - Event studies show the interest-growth differential spikes only at the onset of the crisis, reducing its value for signaling prior to crises.
  - Public debt interactions matter: crisis probability rises steeply at high public debt levels and also at relatively moderate debt levels when accompanied by high inflation or large current account deficits.
- Robustness:
  - Results hold across all income groups.
  - Performance assessed via out-of-sample predictive accuracy and stability of variable selection with respect to sampling variation.

### II. Literature review: evolution and gaps
- Coverage: survey of 42 papers chosen from a pool of 63 based on empirical relevance and clear identification of predictors (details in Appendix 1).
- Periodization and main lessons:
  - 1970s–1990s: focus on capacity to repay
    - Emphasis on external variables: external debt service, size of external debt, foreign exchange reserves to imports; economic growth also important.
    - Crisis definitions mainly debt rescheduling or arrears on external debt.
    - Typical methods: logit models with few predictors or linear regression.
  - 2000–2010: continued emphasis on external debt
    - Crisis definitions broaden to include access to IMF programs.
    - Common predictors: external debt, real GDP growth, debt service, maturity of debt, exchange rate, default history.
    - Methods: logit remains popular; some use classification and regression trees or neural networks.
  - 2011 onwards: growing debate on public debt
    - Crisis definitions expand to include debt defaults (mainly external), IMF programs, implicit debt defaults (high inflation, domestic arrears), and loss of market access.
    - Broader set of methodologies but persistent use of parsimonious approaches due to overfitting concerns and data constraints.
    - Mixed evidence on the predictive importance of public debt:
      - Savona and Vezzoli (2015) and Bruns and Poghosyan (2018) do not find evidence public debt matters for predicting crises.
      - Cerovic et al. (2018) and Sumner and Berti (2017) find some evidence but not robust across specifications.
      - Reinhart and Rogoff (2011a) find changes in public debt significant for debt crises but not for the post-World War II period.
    - No prior studies explicitly analyze the interest-growth differential as a predictor of fiscal crisis.
- Methodological gap: limited exploration of complex nonlinearities and interactions in prior literature; limited use of machine learning techniques in macroeconomic crisis prediction.

### III. Data and definitions
- Crisis definition strategy:
  - Follow Medas et al. (2018) and identify a fiscal crisis in any given year if any of four criteria is met (details in Appendix 2).
- Data coverage and challenges:
  - Sample of 188 countries, covering decades back to the 1980s.
  - Debt dataset is more comprehensive than prior studies and includes debt characteristics such as creditor structure.
  - Limitations noted in previous datasets: many compiled datasets include only a few countries or use narrow/changing definitions of debt, limiting scope (e.g., Reinhart and Rogoff 2009; Abbas et al. 2011; Jordà, Schularick, and Taylor 2016).
- Predictor construction:
  - 748 indicators generated by considering many permutations and moments of variables.
  - Predictors grouped broadly (public debt, public external debt, interest-growth differential, inflation, current account, credit gap, etc.).

### IV. Methodological approach (high level)
- Machine learning workflow:
  - Fit flexible models to capture nonlinearities and interactions without overfitting.
  - Use feature selection algorithms to identify robust predictors from 748 indicators.
  - Evaluate models on out-of-sample predictive accuracy and stability of variable selection.
  - Use interpretability tools (variable importance, partial dependence plots, interaction strength, event studies) to unpack model results.
- Comparative algorithm assessment:
  - Alternative algorithms assessed not only for predictive accuracy but also for the stability of selected variables under sampling variation.

### V. Implications and interpretation
- Public debt matters for fiscal crisis prediction, particularly public external debt.
- Interest-growth differential is a poor pre-crisis signal despite its policy relevance in debates about debt sustainability.
- Nonlinear thresholds and interactions suggest policy attention should focus on combinations of vulnerabilities (e.g., debt together with high inflation or large current account deficits), not just single indicators.
- Machine learning methods provide useful tools to uncover complex dynamics and improve early warning systems beyond traditional econometric approaches.

*Source: wpiea2020001-print-pdf - References*

### 1. Credit events. A crisis is triggered when the debt service is not paid on the due date or

### wpiea2020001-print-pdf - 1. Credit events. A crisis is triggered when the debt service is not paid on the due date or

### Definitions of fiscal crisis types
- Credit events: A crisis is triggered when the debt service is not paid on the due date or the creditor incurs any other type of losses including through debt restructuring.
- Exceptionally large official financing: Episodes where the country receives large financial support from the IMF or the European Union.
- Implicit domestic public debt default: Two criteria are considered: (1) periods of high inflation (usually associated with monetary financing of the budget); or (2) accumulation of domestic arrears.
- Loss of market confidence: Episodes associated with extreme market pressures as proxied by: (1) loss of market access, capturing sovereign defaults or bond issuance coming to a halt; or (2) very large borrowing costs or sovereign yield spikes.

### Sample, incidence, and basic statistics
- Identified crisis episodes: 418 crisis episodes for a sample of 188 countries over the period 1980–2016.
- Average crises per country since 1980: 2 fiscal crises.
- Share of countries with at least one crisis since 1980: more than three quarters.
- Frequency by income group:
  - LICs: about two-thirds are in fiscal distress at any point in time.
  - EMs: on average, 40 percent.
  - AEs: less than 15 percent of them are in fiscal distress in any given year.
- Decade concentration:
  - 1990s: decade with the highest concentration of crises; at the peak, about half of the countries (EMs for the most part) were in fiscal distress.
  - Early 1980s: some bunching—reflecting the collapse of commodity prices and a surge in global interest rates.
  - 2010: some bunching—following the onset of the global financial crisis.

### Crisis type composition and co-occurrence
- Relative frequency of crisis types (overall): credit events account for close to two-thirds of episodes.
- Advanced economies (AEs) are outliers: most AE episodes associated with loss of market confidence and/or exceptional large official financing.
- Overlap with other crises:
  - About one-third of fiscal crises overlap with currency crises.
  - Synchronicity with financial (banking) crises is relatively low.
  - Triple crises (simultaneous fiscal, currency, and financial crises) account for 3 percent of events.

### Table 1 — Fiscal Crises Episodes (1980–2016) — selected reported figures
- Total number of crises starts: 418
- Number of countries by group with crises starts: AEs: 25; EMs: 202; LICs: 191
- Breakdown of crisis-start criteria (percent) 1/:
  - Credit event: Total 72.7; AEs 20.0; EMs 68.3; LICs 84.3
  - Exceptionally large official financing: Total 33.7; AEs 56.0; EMs 34.7; LICs 29.8
  - Implicit domestic default: Total 9.8; AEs 24.0; EMs 11.4; LICs 6.3
  - Loss of market confidence: Total 25.1; AEs 84.0; EMs 32.2; LICs 9.9
- Average per country: Total 2.2; AEs 0.7; EMs 2.1; LICs 3.2
- Number of countries with no fiscal crisis: Total 36; AEs 19; EMs 15; LICs 2
- Average duration (years): Total 5.2; AEs 3.2; EMs 5.3; LICs 5.3

Notes on Table 1:
- 1/ Crisis starts can be associated with more than one criterion. Therefore, the breakdown does not need to add up to 100. A year is considered to be a fiscal crisis year when at least one of the four criteria is met. To separate between crisis events, we require at least two years of no fiscal crisis between the distinct events.

### Debt data, metrics assembled, and relevance for prediction
- Coverage and scope:
  - The study assembles a comprehensive range of debt metrics: public debt, private indebtedness, and total debt in the economy.
  - Time series leverage: Global Debt Database (private, public, and total debt for 190 countries going as far back as 1950—see Mbaye, Moreno Badia, and Chae 2018b).
- Scaling and complementary indicators:
  - Debt scaled by GDP and by indicators that can proxy for available liquidity such as reserves or revenues; complements debt service ratios.
- Debt composition heterogeneity:
  - External debt is the main component among LICs but not in other income groups.
- Purpose: Exploit heterogeneity in debt characteristics to identify which debt metrics have higher discriminating value as predictors of fiscal crises.

### Interest-growth differential ("r-g") and fiscal measurement
- Definition provided in the paper:
  - r-g is defined as:
    (푟−푔
    1+푔
    )
  - r is the effective interest rate and g the GDP growth rate.
- Calculation details and caveats:
  - Effective interest rates are calculated using consistent time series for the stock of public debt.
  - Trade-off: if the interest bill refers to a broader perimeter of government than the stock of debt, the interest-growth differential may be over-estimated.
  - For LICs, reported interest bills typically refer to the same level of government as debt stocks, reducing this discrepancy.
  - Data suggests that, on average, the interest-growth differential has been close to zero or negative since the 1980s across all income groups, but with wide dispersion and positive r-g not uncommon.
  - For foreign-currency denominated debt, depreciation-adjusted interest-growth differentials may be important, but limited data on currency composition prevents adjustment; however, measures of exchange rate depreciation are included among predictors.
- Stock-flow adjustment (SFA) constructed to capture valuation effects and contingent liabilities:
  - 푆퐹퐴 푡 =푑 푡 − ( 1+푟 1+푔 ) 푑 푡−1 +푝 푡
  - where d_t and d_{t-1} are the stock of public debt in periods t and t-1 respectively, and p_t is the primary balance in period t.

### Predictors dataset and inclusion rules
- Scope of predictors:
  - Dataset covers 748 indicators across measures of debt, economic activity, level of development, prices, fiscal aggregates, external indicators, global factors, demographics, and institutions.
  - Several permutations used for each variable (levels, lags, differences at various horizons) and cross-sectional averages to capture global factors and spillovers.
- Missing data rule:
  - A variable is included if 70 percent of the data exists.
  - Missing values are imputed with the training sample median of the non-missing values.

### Methodology for predictor selection and crisis prediction
- Objective: identify a stable and robust set of predictors from a large number of variables, accounting for interactions and non-linearities.
- Model of choice: random forest (Breiman 2001).
  - Reasoning: able to deal with complexity and deliver significant improvements in predicting fiscal crises relative to standard econometric approaches typically used in the early warning literature and other machine learning algorithms.

*Source: wpiea2020001-print-pdf - 1. Credit events. A crisis is triggered when the debt service is not paid on the due date or*

### Appendix 5 for an empirical comparison of the out-of-sample performance of the random

### wpiea2020001-print-pdf - Appendix 5 for an empirical comparison of the out-of-sample performance of the random forest (RF) against other econometric approaches

### Variable selection
- Benchmark RF starts with a large set of variables: 748.
- Predictive model notation: ŷ = f(X), where X is a matrix with n (annual) observations and m variables and ŷ ∈ [0,1] is the predictive probability of a fiscal crisis over the next two years. y = 1 if there is a crisis and 0 otherwise.
- Feature selection algorithms used (all built around RF):
  - P-values computed with permutation importance (PIMP)
    - Repeated permutations of the outcome vector produce “null importances”.
    - Fit a probability distribution (normal, lognormal, or gamma) to null importances; parameters estimated by maximum likelihood; P-values computed as probability of observing an importance larger than the original importance under the fitted distribution; keep only significant predictors.
  - Recursive Feature Elimination (RFE)
    - Start with RF on all variables; remove a specific proportion of least important variables; regenerate RF recursively until out-of-bag (oob) predictive error is larger than the initial/previous oob error; select variable set with smallest oob error or within a small range of the minimum.
  - Boruta
    - Compare importance of real predictors to random “shadow” variables; for each real variable test its importance against the maximum of shadow variables; declare variables important/unimportant; remove unimportant and shadow variables; repeat until classification finished or pre-specified runs reached.
  - VSURF
    - Returns two subsets: (1) important variables including some redundancy (relevant for interpretation), (2) smaller subset focusing on prediction and avoiding redundancy; uses RF permutation-based importance ranking followed by stepwise forward variable introduction.
- Criteria for choosing among algorithms:
  - High predictive power: compare out-of-sample performance (main metric AUROC) of RFs estimated using features from each algorithm against RF with full set of 748 variables; statistical significance assessed via t-tests with standard errors adjusted for two-way clustering (Cameron, Gelbach, and Miller 2011). Other metrics reported include log likelihood and mean squared errors (MSE).
    - AUROC interpretation: perfect model AUROC = 1; no predictive power AUROC = 0.5.
  - Stability of feature selection: construct two samples by randomly dropping 5 percent of observations; compare overlap of features selected by each algorithm across samples using the Pearson Correlation Coefficient (range -1 to 1; 1 = perfect overlap).

### Assessing variable importance
- Methods used:
  - Out-of-bag permuted predictor importance
    - Importance estimated as increase in prediction error after permuting a feature; model errors calculated on the oob sample.
    - Note on oob: each tree uses a bootstrap sample; about one-third of cases are left out (oob sample).
  - Shapley values
    - Measure each variable’s contribution to an individual prediction’s deviation from the historical mean, constructed as the mean of marginal contributions across all combinations of other variables.
    - Used to rank variables by contribution to the probability of a crisis and to compute differences in Shapley values between crisis and non-crisis events.
    - Note: Shapley values measure contribution to deviation from the mean prediction, not the change in prediction after removing the feature from a trained model.

### Studying interactions and nonlinearities
- Partial dependence plots (PDPs)
  - PDPs show marginal effect of one or several features on the predicted outcome and identify whether relationships are linear, monotonic, or more complex.
  - Univariate PDP: line plot showing relationship between a feature and the predicted outcome (probability of a crisis).
  - Bivariate PDP: surface plot visualizing predicted outcome for a pair of features by marginalizing over other variables.
  - Intuition: partial dependence at a feature value represents the average prediction if all data points were forced to assume that feature value.
- H-statistics (measure of two-way interaction strength)
  - H-statistic uses PDPs and is defined as the variance of the difference between observed bivariate PDP and the sum of the two individual PDPs, normalized by the variance of the bivariate PDP.
  - Interpretation:
    - H-statistic = 0 if there is no interaction.
    - H-statistic = 1 if variance of PD_jk is fully explained by the interaction (each single PD function is constant; effect on prediction comes only through the interaction).
  - The H-statistic ranks all (n-1) pairs of the n-th variables by relative interaction strength.

### Results — Variable selection (empirical findings)
- Feature set sizes vary widely across algorithms: from less than 10 variables (VSURF) to more than 300 (Boruta).
- Out-of-sample performance comparisons (summary of findings):
  - Full RF model outperforms for AEs and EMs relative to LICs, but overall predictive power is higher than in previous studies.
  - AUC (full model): 0.81 for AEs and EMEs; 0.71 for LICs.
    - Comparisons cited: Cerovic et al. (2018) report maximum AUC of 0.69 and 0.68 respectively.
  - Among feature selection algorithms:
    - Boruta is always at least as good as the full model for both income groups across performance metrics.
    - Boruta is statistically significantly better than the full model across the board when assessed by the log likelihood and for AEs and EMEs when assessed by the MSE.
    - VSURF shows the worst performance throughout.
    - PIMP usually performs worse than Boruta or RFE.
    - RFE performs in between, underperforming relative to Boruta across some metrics.
- Note on evaluation details: out-of-sample performance obtained from 15 rolling regressions; bootstrapped standard deviations based on 100 random resamples of the test sample with replacement. Stars indicate confidence that a model outperforms the RF estimated with the full set of 748 variables: * 90%, ** 95%, and *** 99%.

*Source: Appendix 5, wpiea2020001-print-pdf - Appendix 5 for an empirical comparison of the out-of-sample performance of the random forest (RF) against other econometric approaches*

### Appendix 5).

### Appendix 5)

### Stability of feature selection and model performance
- Comparison of stability restricted to best performing algorithms: Boruta, RFE, and PIMP.
- Pearson index (stability across replicates):
  - Boruta: 0.92
  - RFE: 0.63
  - PIMP: 0.81
- Table of performance metrics (reported with standard errors in parentheses):
  - Advanced and emerging market economies
    - AUROC:
      - Boruta: 0.805 (0.025)
      - RFE: 0.793* (0.026)
      - VSURF: 0.734*** (0.031)
      - PIMP: 0.791** (0.026)
      - RF: 0.806 (0.026)
    - Log(likelihood):
      - Boruta: -0.289*** (0.027)
      - RFE: -0.286** (0.029)
      - VSURF: -0.34 (0.043)
      - PIMP: -0.29** (0.028)
      - RF: -0.299 (0.026)
    - MSE:
      - Boruta: 0.084** (0.011)
      - RFE: 0.085 (0.011)
      - VSURF: 0.096*** (0.011)
      - PIMP: 0.085 (0.011)
      - RF: 0.087 (0.01)
  - Low-income countries
    - AUROC:
      - Boruta: 0.708 (0.042)
      - RFE: 0.712 (0.042)
      - VSURF: 0.688 (0.039)
      - PIMP: 0.707 (0.042)
      - RF: 0.706 (0.043)
    - Log(likelihood):
      - Boruta: -0.547* (0.043)
      - RFE: -0.544** (0.044)
      - VSURF: -0.601* (0.052)
      - PIMP: -0.548* (0.043)
      - RF: -0.556 (0.039)
    - MSE:
      - Boruta: 0.186 (0.018)
      - RFE: 0.185 (0.018)
      - VSURF: 0.207** (0.019)
      - PIMP: 0.187 (0.018)
      - RF: 0.189 (0.017)
- Given predictive power and stability, Boruta chosen as benchmark algorithm.

### Variable importance (Boruta-based analysis)
- Initial set reduced by half but Boruta still selects 336 indicators (including permutations); individual indicators expected to have small predictive power.
- Indicators grouped into 23 categories for interpretation.
- Out-of-bag permuted predictor importance — main findings:
  - Public debt is the most important group of predictors, followed closely by public debt service.
  - Institutional slow-moving variables rank highly: level of development (GDP per capita), demographics, and to a lesser degree the quality of institutions.
    - These variables likely help discriminate countries' underlying propensity to crisis (AEs vs EMs vs LICs).
  - External variables are important, particularly external capital flows, and to a lesser extent external debt, current account, and the exchange rate.
    - Fiscal crises overlap with currency crises in a third of cases.
  - Fiscal flow variables (deficits, revenues, spending) are relevant but considerably less than debt or some external variables.
  - Interest-growth differential and global conditions are among the least relevant variables.
- Shapley differences (discriminating power between crisis and non-crisis):
  - Public debt and public debt service are most important for EMs and LICs; somewhat less for AEs (public debt still among top 3).
  - Private debt appears more important than public debt for AEs.
  - Pre- and post-2000 periods: public and public debt service top categories in both periods, with public debt highest in post-2000.

### Analysis of selected predictors — public debt and interactions
- Focus on public external debt as the individual debt measure with highest predictive value.
- Univariate PDPs for public external debt:
  - Positive, non-linear relationship between public external debt and predicted probability of entering a crisis across income groups.
  - AEs: probability increases substantially once debt is around 70 percent of GDP.
  - EMs: estimated probability relatively flat for debt < 30 percent of GDP but rises steeply above that level.
  - LICs: predicted probabilities much higher from the start; steepening of the curve takes place at lower debt levels than other groups.
- Crisis probability thresholds minimizing sum of type I and type II errors:
  - AEs: 8.5 percent — captures 80 percent of crises while false alarms at 20 percent.
  - EMs: 22 percent.
  - LICs: 28 percent.
- High debt levels at which estimated probability breaches the crisis threshold regardless of other factors:
  - AEs and EMs: 70 percent of GDP.
  - LICs: 80 percent of GDP.
  - Possible reasons for higher LIC level: minimization trade-off to avoid false alarms; higher share of concessional borrowing among LICs.
- Interaction analysis (H-statistics and 2-way interactions):
  - Public external debt has strongest cumulative interactions with all other variables.
  - Top interacting variables with public external debt: external assets, the inflation rate, GDP relative to US, and public external amortization.
  - Conclusion: probability of crisis may be high even for moderate debt if interacting factors are adverse.

### Interest-growth differential
- Variable importance analysis shows limited information content for interest-growth differential.
- PDPs: even for large variations of the interest-growth differential, estimated crisis probability barely changes and remains relatively low; it never breaches the crisis threshold.
- Event study dynamics: interest-growth differentials can remain low for long stretches and only shoot up at onset of crisis, making it irrelevant as a leading indicator.
- Bivariate PDPs (public external debt × interest-growth differential):
  - AEs and EMs: if public external debt is sufficiently high, estimated probability breaches crisis threshold irrespective of interest-growth differential.
  - LICs: both highly negative and positive interest-growth differentials can imply higher probability of crisis for the same level of debt.
    - Possible reasons: both extremes may signal imbalances; governments may respond to low interest-growth differentials by increasing deficits, negating benefits of low borrowing costs.
- Interpretation caveat: interest-growth differential matters for debt dynamics but has limited signaling value as a leading indicator in this analysis.

### Inflation
- Univariate PDPs indicate a strong relationship between inflation and estimated probability of crises, with strong non-linearities.
- For AEs and EMs, probability of crises increases significantly when inflation is above 20 percent.
- Both increases and declines in inflation can be associated with higher probability of crises, particularly for EMs (risk of deflation and snowball effects).
- LICs: levels of public external debt at which estimated probabilities breach crisis thresholds decrease with inflation — even moderate debt levels imply much higher crisis probability when inflation is high.
  - Possible mechanism: limited capacity to manage debt leads to monetization of deficits.

### External and financial imbalances
- Current account balance PDPs show non-linear pattern: once external deficits are between 3–7 percent of GDP, probability of crisis increases substantially for AEs and EMs.
  - Current account deficits less relevant for predicting fiscal crises in LICs.
- Interaction of public external debt with current account (notably for AEs):
  - Even moderate debt levels lead to steep rise in crisis probability when current account deficits are high.
  - Current account surpluses do not shield countries from crises if debt levels are high.
- Private sector leverage (credit gap = private debt as a share of GDP relative to the 10-year average):
  - Probability of crisis increases significantly in AEs and EMs when credit gap is above 40 percent.
  - Interactions with public external debt for EMs: crisis thresholds breached for lower debt levels if credit gap is large.
  - Private debt dynamics much less relevant for LICs, likely reflecting low financial deepening.

### Conclusion and policy implications
- Main empirical conclusion:
  - Public debt in its various forms is the most important predictor of fiscal crises; it matters broadly and robustly.
  - Interactions between public debt and other predictors materially affect crisis risk.
  - Interest-growth differential has limited signaling value; beyond certain debt levels, crisis likelihood surges regardless.
- Limitations:
  - Machine learning models do not establish causality; results speak to strong correlations and predictive associations.
- Policy implications:
  - Complacency about high public debt would be ill-advised, even if interest-growth differentials remain low.
  - High public debt entails non-linear risks that can escalate quickly; by the time interest-growth differentials signal distress, a crisis may already be underway.
  - Reducing debt is not universally the correct policy in all contexts; debt used for countercyclical purposes, public investment, or structural needs can be appropriate.
  - Evidence suggests that public debt is not costless and warrants careful risk management and consideration of interacting vulnerabilities (external balances, inflation, private leverage, institutional strength).

*Source: Appendix 5).*

### REFERENCES

### REFERENCES

### Major thematic areas covered in the references
- Fiscal stress, sovereign debt, and sovereign default:
  - Historical and modern treatments of public debt (e.g., Abbas et al. 2011; Mauro et al. 2015; Reinhart and Rogoff 2009; Reinhart, Reinhart, and Rogoff 2012).
  - Determinants, pricing, and costs of sovereign default and restructuring (e.g., Cruces and Trebesch 2013; Bocola, Bornstein, and Dovis 2019; Asonuma et al. 2019; Catão and Sutton 2002).
  - Debt sustainability and debt dynamics (Escolano 2010; Kraay and Nehru 2006; Messmacher and Kruger 2004).

- Early warning systems, crisis prediction, and fiscal-crisis indicators:
  - Development and assessment of early warning systems for sovereign or financial crises (Berg, Borensztein, and Pattillo 2005; Bussiere and Fratzscher 2006; Manasse, Roubini, and Schimmelpfennig 2003).
  - Empirical studies designing country-specific or region-specific fiscal stress indices (Berti, Salto, and Lequien 2013; De Cos et al. 2014; Pamies Sumner and Berti 2017).
  - Methodological contributions to predicting sovereign debt crises and loss of market access (Gelos, Sahay, and Sandleris 2004; Guscina, Sheheryar, and Papaioannou 2017; Medas et al. 2018).

- Machine learning, feature selection, and model interpretability applied to economic and financial prediction:
  - Core machine learning algorithms and ensembles: Random Forest (Breiman 2001), Gradient Boosting (Friedman 2001), Support-vector Networks (Cortes and Vapnik 1995).
  - Feature selection, variable importance, and stability: Boruta, VSURF, PIMP, and other RF-based selection methods (Kursa and Rudnicki 2010; Genuer, Poggi, and Tuleau-Malot 2015; Altmann et al. 2010).
  - Model explanation and interpretation methods: Partial dependence plots and PDP tools (Apley 2016; Greenwell 2017), Shapley values and unified interpretation approaches (Lundberg and Lee 2017; Strumbelj and Kononenko 2010).
  - Reviews and comparative evaluations of classifiers and feature selection stability (Fernandez-Delgado et al. 2014; Nogueira and Brown 2016; Degenhardt et al. 2019).

- Empirical econometric techniques and robustness:
  - Panel data, logit and discriminant analysis for debt-servicing capacity and rescheduling (Feder and Just 1977; Hajivassiliou 1987; Frank and Cline 1971).
  - Multiway clustering and robust inference (Cameron, Gelbach, and Miller 2011).
  - Variable-selection penalization (Tibshirani 1996; Ma and Huang 2008).

- Data sources and databases referenced:
  - IMF-related datasets and internal work: IMF Global Debt Database; IMF Working Papers (numerous references with paper numbers, e.g., No. 19/69, No. 18/181, No. 18/111, No. 17/246, No. 11/100, No. 03/221, No. 01/02, No. 02/149, No. 04/211, No. 18/141, No. 15/13, No. 18/206).
  - World Bank and other cross-country databases (World Development Indicators; Database of Political Institutions DPI2017; World Bank Development Prospects Group Policy Research Working Paper No. 8157).
  - Other data sources cited explicitly in figures: Bloomberg; Datastream; Eurostat; Gelos, Sahay, and Sandleris (2004); Guscina, Sheheryar, and Papaioannou (2017); IMF, International Financial Statistics; Laeven and Valencia (2018); OECD; Reuters; authors’ calculations.

### Methodological and technical emphases in the cited literature
- Feature selection and robustness:
  - Algorithms and packages: RF, Boruta, PIMP, RFE, VSURF, VSURF (Genuer et al. 2015; Kursa and Rudnicki 2010; Altmann et al. 2010).
  - Stability measures: Pearson correlation coefficient for feature set stability under random sample drops (note described procedure: alternative samples drawn randomly dropping 5 percent of observations).
- Model interpretability:
  - Use of Partial Dependence Plots and Shapley values for interpreting contribution to predicted crisis probability (Apley 2016; Lundberg and Lee 2017).
- Evaluation of predictive systems:
  - Receiver Operating Characteristic (ROC) area comparisons (Rodriguez and Rodriguez 2006), extreme-bound analysis (Bruns and Poghosyan 2018), and cross-validation of predictive classifiers (Fioramanti 2008; Dawood, Horsewood, and Strobel 2017).

### Figures and empirical displays (captions, notes, and data coverage)
- Figure 1. Predictors of Fiscal Crises in the Literature
  - Based on a literature review of 42 empirical papers; variables plotted are those statistically significant in at least a third of the papers during the reference period. (See Appendix 1 in source.)
- Figure 2. Most Common Predictors in the Literature (Share of surveyed papers, 1970–2018)
  - Chart based on literature review of 42 empirical papers.
  - Listed variables across panels include: Debt service; External Debt; FX reserves; GDP Growth rate; Exchange rate; Short-Term Debt; Default history; GDP levels; Public debt; Fiscal Balance; Current Account; Openness; US variables; Political variables; GDP per Capita; Inflation; Exports; Imports; Private debt; Commodity Prices; Government expenditures; Savings; Cost of rescheduling; Trade balance; IMF credit; Banking crisis history.
- Figure 3. Countries with Fiscal Crises, 1980–2016 (Number)
- Figure 4. Overlap with Other Crises, 1980–2016 (Number of crises episodes)
  - Note: Two crises are identified as overlapping if they start within two years of each other. Financial crises are banking crises episodes as defined in Laeven and Valencia (2018).
  - Panel categories shown: Advanced Economies; Emerging Market Economies; Low-income Countries.
  - Crisis-type counts visible in figure: 253; 111; 25; 29 (present in figure area).
- Figure 5. Debt Statistics: Country Coverage, 1980–2016 (Number of countries)
  - Data sources: IMF, Global Debt Database; IMF, World Economic Outlook; World Bank, World Development Indicators; U.S. Bureau of Economic Analysis; Haver; Arslanalp and Tsuda (2014); authors’ calculations.
  - Series include: Public debt; Public external debt.
- Figure 6. Public and Public External Debt, 1980–2016 (Weighted average, percent of GDP)
  - Series: Advanced Economies; Emerging Economies; Low-Income Countries.
  - Note: Public external debt refers to public and publicly guaranteed debt.
  - Sources: IMF Global Debt Database; IMF World Economic Outlook; World Development Indicators; U.S. Bureau of Economic Analysis; Haver; Arslanalp and Tsuda (2014); authors’ calculations.
- Figure 7. Interest-Growth Differential, 1980–2016 (Percent)
  - Series by country group: Advanced Economies; Emerging Market Economies; Low-Income Countries.
  - Data sources: IMF Global Debt Database; IMF World Economic Outlook; authors’ calculations.
  - Chart annotations: MinMaxAverage lines and axis ranges include values shown such as -150, -100, -50, 0, 50, 100, 150 on some panels.
- Figure 8. Feature Selection Algorithms (Number of variables selected)
  - Bar values shown: RF 748; Boruta 336; PIMP 176; RFE 68; VSURF 8.
  - Note: Chart shows number of variables selected by each feature selection algorithm and the full RF model estimated over the full sample.
- Figure 9. Robustness in Variable Selection (Pearson correlation coefficient)
  - Pearson index values shown: Boruta 0.92; PIMP 0.81; RFE 0.63.
  - Note: Pearson index measures stability of chosen feature set to variations in training data; alternative samples drop 5 percent of observations.
- Figure 10. Variable Importance by Group of Predictors
  - Variable importance computed via an in-built out-of-bag permuted predictor importance function in R based on RF estimated with variables selected by Boruta.
  - Predictor groups listed include: Public debt; Level of development; Public debt service; External (capital flows); Demographics; Private debt; Economic activity; External debt; Current account; Exchange rate; Fiscal; Institutions; Cross Sectional; Inflation; Natural resources; Crisis history; Total debt; External debt service; FX reserves; Foreign aid; r-g; Global.
- Figure 11. Contribution to Probability of a Crisis (Shapley Values)
  - Charts display mean Shapley value difference (crisis versus non-crisis observations) by income groups: Advanced Economies; Emerging Market Economies; Low-Income Countries.
  - Predictor list in Shapley charts includes: Level of development; Private debt; Public debt; External (capital flows); Current account; Inflation; External debt; Public Debt Service; Demographics; Economic growth; Institutions; Total debt; Cross Sectional; Fiscal; Natural resources; External Debt service; Crisis history; Exchange rate; FX reserves; Foreign aid; r-g; Global.
- Figure 12. Partial Dependence Plots 1/ and Event Studies
  - (Figure label appears; figure content referenced without further numerical detail in the references section.)

### Representative methodological citations with exact paper identifiers and details
- IMF Working Papers, examples:
  - Asonuma, T., M. Chamon, A. Erce, and A. Sasahara. 2019. IMF Working Paper No. 19/69.
  - Baldacci, E., I. Petrova, N. Belhocine, G. Dobrescu, and S. Mazraani. 2011. IMF Working Paper No. 11/100.
  - Cerovic, S., K. Gerling, A. Hodge, and P. Medas. 2018. IMF Working Paper No. 18/181.
  - Guscina, A., M. Sheheryar, and M. Papaioannou. 2017. IMF Working Paper No. 17/246.
  - Laeven, L. and F. Valencia. 2018. IMF Working Paper No. 18/206.
  - Mbaye, S., M. Moreno-Badia, M., and K. Chae. 2018a. IMF Working Paper No. 18/141.
  - Mbaye, S., M. Moreno-Badia, M., and K. Chae. 2018b. IMF Working Paper No. 18/111.
  - Manasse, P., N. Roubini, and A. Schimmelpfennig. 2003. IMF Working Paper No. 03/221.
  - Detragiache, E. and A. Spilimbergo. 2001. IMF Working Paper No. 01/02.
  - Catão, L. and B. Sutton. 2002. IMF Working Paper No. 02/149.

- Key non-IMF works and methodological references:
  - Breiman, L. 2001. Random Forest. Machine Learning. 45(1), 5–32.
  - Friedman, J. H. 2001. “Greedy Function Approximation: A Gradient Boosting Machine,” Annals of Statistics Vol. 29, No. 5, pp. 1189–1232.
  - Tibshirani, R. 1996. “Regression Shrinkage and Selection via the Lasso,” Journal of the Royal Statistical Society. Series B, Vol. 58, No. 1, pages 267–88.
  - Lundberg, S. M. and S. Lee. 2017. “A Unified Approach to Interpreting Model Predictions,” Advances in Neural Information Processing Systems, pp. 4765–74.
  - Apley, D. W. 2016. “Visualizing the Effects of Predictor Variables in Black Box Supervised Learning Models.” arXiv preprint arXiv:1612.08468 (2016).

*Source: REFERENCES section of wpiea2020001-print-pdf*

### 1.  Public External Debt, AEs and EMs

### 1.  Public External Debt, AEs and EMs

### Boruta random forest PDPs and probability thresholds
- Charts (1)-(3) display PDPs based on the Boruta random forest.
- Solid lines show the PDP curve which represents the average prediction across all levels of public external debt (charts 1 and 2) and the interest-growth differential (chart 3).
- Dotted lines show probability thresholds based on minimizing the sum of type I and type II errors (missed crises and false alarms).
- Probability axes and scales appearing in figures: 0 50 100 150 (Probability of a crisis (percent)).
- Country groupings shown: AE (Advanced economies), EM (Emerging markets), LIC (Low-income countries).
- Numeric thresholds and axis ticks visible in figures: 0, 2, 4, 6, 8, 10, 12; 17, 18, 19, 20, 21, 22, 23, 24, 25; 28, 29, 30, 31, 32, 33, 34; 17.0, 17.5, 18.0, 18.5, 19.0; 3.0, 3.5, 4.0, 4.5, 5.0; -20 -10 0 10 20.

### Interest-growth differential: event study framework and interpretation
- Chart (4) displays an event study based on the framework developed by Gourinchas and Obstfeld (2012) where t=0 is the start of the fiscal crisis.
- Estimated equation (as presented): 푦푖,푡 = 훼푖 + ∑훽푡+푗 5 푗=−5 퐷푖,푡+푗 + 휀푖,푡
  - y is the interest-growth differential.
  - D_{i,t} is a dummy equal to 1 when the country is j periods away from the start of a crisis in period t and zero otherwise.
- Each data point should be interpreted as the interest-growth differential at time t+k, relative to “non-crisis” times benchmark.
- Event window plotted: t-5, t-4, t-3, t-2, t-1, t=0, t+1, t+2, t+3, t+4, t+5.
- Separate curves displayed for AE (Advanced economies) and EM (Emerging markets).

### Interaction strength and top interactions with public external debt
- Figure 13: Overall interaction strength (H-statistic) for each feature with all other features for the Boruta RF.
  - Public external debt has the highest relative interaction effect with all other features.
- Figure 14: Top-10 2-way interaction strengths (H-statistic) between public external debt and each other feature.

### Bivariate PDPs: Public external debt with interest-growth differential, inflation, current account, credit gap
- Figure 15: Bivariate partial dependence plots for Public External Debt and r-g (interest-growth differential) across:
  - Advanced economies
  - Emerging Market Economies
  - Low-income Countries
- Notes on bivariate PDPs:
  - Cells highlighted in red depict combinations of public external debt and the interest-growth differential for which the estimated probability of a crisis is above the probability thresholds calculated for that income group (minimizing the sum of type I and type II errors).
  - The darker the blue color, the lower the probability of a crisis.
- Example axis and tick values shown in the panels:
  - Public external debt to export: 00.020.040.060.08
  - Working age population, Bureaucracy Quality, Public external debt, % of exports, lag; Log GDPpc relative to US (PPP, lag); Amortization of public external debt, (% of reserves); Age dependency ratio; GDPpc (PPP, lag); Time since crisis; Crisis history; Public external debt, % of exports.
- Figure 16: Inflation — Univariate PDPs
  - Axes include Inflation rate and Change in the inflation rate (Percent).
  - PDPs based on Boruta RF; estimated probabilities on vertical axis and inflation on horizontal axis.
- Figure 17: Inflation and Public External Debt — Bivariate PDPs
  - Advanced Economies, Emerging Market Economies, Low-Income Countries.
  - Red cells denote combinations above group-specific probability thresholds.
  - Example numeric axis ticks visible: 17.0, 18.0, 19.0, 20.0, 21.0, 22.0, 23.0; 3.0, 4.0, 5.0, 6.0, 7.0, 8.0, 9.0; -10 -5 0 5 10 30 50.
- Figure 18: Current Account — Univariate PDPs (Percent of GDP)
  - Panels for Advanced and Emerging Market Economies and Low-Income Countries.
  - Figure 19: Current Account and Public External Debt — Bivariate PDPs (Percent of GDP)
  - Red cells denote combinations above group-specific probability thresholds.
  - Axis ticks visible include -20 -10 0 10 20; 28.0, 28.5, 29.0, 29.5, 30.0; -40 -20 0 20 40.
- Figure 20: Credit Gap — Univariate PDPs (Percent of GDP)
  - Panels for Advanced and Emerging Market Economies and Low-Income Countries.
  - Figure 21: Credit Gap and Public External Debt — Bivariate PDPs (Percent of GDP)
  - Red cells denote combinations above group-specific probability thresholds.
  - Axis ticks include -200 20 40 60; -40 -20 0 20 40.

### Key methodological points and visualization conventions
- PDPs (partial dependence plots) show estimated probabilities on the vertical axis and the predictor (public external debt, inflation, current account, credit gap, r-g) on the horizontal axis.
- Bivariate PDPs highlight regions where the estimated probability exceeds income-group-specific thresholds calculated by minimizing the sum of type I and type II errors (missed crises and false alarms).
- Interaction strength measured by H-statistic to identify features that interact strongly with public external debt.

### Appendix 1 — Literature review (selected entries and recurring predictors)
- Literature surveyed spans studies defining crises as rescheduling, arrears, default, IMF large access (>100 percent of quota), yield spreads, sovereign debt restructuring, and other criteria.
- Recurring predictors and variables across reviewed studies include:
  - Debt measures: debt service to exports; debt/exports; public external debt/GDP; total external debt/GNP; short-term debt; amortization rates; debt service ratios.
  - Reserves and liquidity: reserves to imports; reserves to GDP; FX reserves measures; reserves plus IMF credits.
  - External accounts: current account/GDP; exports to imports; openness (exports+imports/GDP+imports); export growth; trade balance/GDP.
  - Growth and macro variables: GDP growth; per capita income (GNP or GDP per capita); inflation; volatility of growth; real exchange rate deviations/overvaluation.
  - Financial and credit variables: short-term debt to reserves; short-term debt share; credit to GDP; domestic credit; credit gap (in PDPs).
  - Political/institutional variables: political instability; index of political freedom; democracy measures; past defaults/rescheduling history.
  - Global and external conditions: US interest rates/T-bill rates; world commodity price changes; global interest; IMF access/IMF programs.
- Examples of studies and their emphasized determinants (selection from Appendix):
  - Feder (1977): Logit — Debt service/exports, imports/reserves, amortization/debt, income/capita, capital inflows/debt service, GDP growth, export growth.
  - Cline (1984): Logit — Debt service to exports; reserves to imports; per capita economic growth; current account squared; gross debt minus reserves to exports; amortization rate.
  - Manasse, Roubini, and Schimmelpfennig (2003): Logit and binary recursive tree — Ratio of short-term debt to reserves; ratio of debt services to reserves; ratio of current account balance to GDP; US treasury bills rate; GDP growth; dummy for inflation rate above 50%; dummy for past defaults.
  - IMF (2017): Probit on 80 LICs, 1970–2014 — Institutions, external debt to GDP, debt to exports, debt service to revenue, debt service to exports, domestic growth, reserves, remittances, world growth, GDP per capita, openness, FX income, risk premium, Dummy for conflicts.

*Source: wpiea2020001-print-pdf - 1.  Public External Debt, AEs and EMs*

### APPENDIX 1. LITERATURE REVIEW

### APPENDIX 1. LITERATURE REVIEW

### Summary of reviewed studies and key predictors
- Sumner and Berti (2017)
  - Reference definition of crisis: Fiscal crises
  - Sample: 35 countries; 1970–2015
  - Model: Signal extraction; logit
  - Importance/Significance:
    - For the logit1 (less data): gross public debt; Current account; real GDP growth; World GDP growth; at 10%: change in gross public debt; private sector credit flow.
    - For logit2 (larger sample): private sector credit flow; world GDP growth; at 5%: current account; real GDP growth; at 10% change in gross public debt.
    - For the S0 (signal): yield curve; private sector credit flow; net savings of households; current account; net international investment position; gross financing needs; cyclicality adjusted balance; share construction on GDP; GDP per capita in PPP; ST debt of households; ST debt of nonfinancial corporates; ST debt GG; Net debt; private sector debt; change in nominal unit labor costs; primary balance; gross debt; change in gross debt; change in expenditures GG; change in REER; Real GDP growth.

- Bruns and Poghosyan (2018)
  - Reference definition of crisis: Fiscal crises
  - Sample: 29 AEs and 52 EMEs, 1970–2015
  - Model: Extreme Bound Analysis; Logit
  - Importance/Significance: Output gap; External Current Account; FX reserves growth; FX reserves as share of GDP; Openness; Primary balance gap (share of GDP from VEE); Real GDP growth; Overall fiscal balance; Primary balance; FX debt (GG from VEE).

- Cerovic, Gerling, Hodge and Medas (2018)
  - Reference definition of crisis: Fiscal crises
  - Sample: 188 countries, 1970–2015 (2007–15 out of sample)
  - Model: Signals and logit
  - Importance/Significance:
    - Logit results (in sample): For AE/EM: Current Account, Primary expenditures growth, Output gap, interest expenses/revenue, public debt/revenue.
    - For LICs: FDI, World food prices, GDP growth.

- Ghulam and Derber (2018)
  - Reference definition of crisis: External sovereign default
  - Sample: 70 countries 1970–2010
  - Model: Survival analysis (general logit/specified logit incl. significant increase in hazard up to 15 years)
  - Importance/Significance: Volatilities of US treasury bills rates and USD-denominated LIBOR; Political uncertainty; Export/Import growth; Increase in inflation; Debt/GDP; Increase in external; GDP per capita; Previous banking; US treasury rate; Central government debt/GDP; High current account deficit and exchange rate volatility.

### Notes on terminology used in the literature
- AEs: Advanced Economies
- EMEs: Emerging Market Economies
- LICs: Low-income Countries

*Source: APPENDIX 1. LITERATURE REVIEW (extracted content).*

### APPENDIX 4. DATA: DEFINITION, SOURCES, AND PREDICTOR GROUPINGS

### APPENDIX 4. DATA: DEFINITION, SOURCES, AND PREDICTOR GROUPINGS

### Crisis history and financial crisis variables
- Source: Authors’ calculations based on Medas et al (2018).
- Variables and permutations listed:
  - Average crisis history
  - Number of crisis
  - Number of crisis by year, Advanced and Emerging Economies (AE+EM)
  - Number of crisis by year, Emerging and Low-Income Economies (EM+LIC)
  - Number of crisis by year, AEs
  - Number of crisis by year, EMEs
  - Number of crisis by year, LICs
  - Number of crisis by year (sum of the current and previous year)
  - Number of crisis by year (sum of the current and previous year), AE+EME
  - Number of crisis by year (sum of the current and previous year), EME+LIC
  - Number of crisis by year (sum of the current and previous year), AE
  - Number of crisis by year (sum of the current and previous year), EME
  - Number of crisis by year (sum of the current and previous year), LIC
  - Number of crisis by country, AE+EME; EME+LIC; AE; EM; LIC
  - Year since last crisis
- Banking crisis start year:
  - Source: Authors’ calculations based on Laeven and Valencia (2018).
  - Permutation indicators: L, L2

- Financial Crisis indicators (grouped by income/region): Financial Crisis, AE+EM; Financial Crisis, EM+LIC; Financial Crisis, AE; Financial Crisis, EME; Financial Crisis, LIC

### Level of development, output, and urbanization
- Urban population (% of total) — Source: World Bank, World Development Indicators.
  - Permutations: L, 5yr_L, 10yr_L, wavg
- Log of real GDP per capita (in PPP dollars, Units), relative to US — Source: IMF, World Economic Outlook.
  - Permutations: L, L2, fd_L, d3_L, 5yr_L, 10yr_L, wavg; Selected: L, L2, fd_L, d3_L, 5yr_L, 10yr_L, L_wavg, L2_wavg
- Log of nominal GDP in USD, relative to US — Source: IMF, World Economic Outlook.
  - Permutations: L, L2, fd_L, d3_L, 5yr_L, 10yr_L, wavg; Selected: L, L2, fd_L, d3_L, 5yr_L, 10yr_L, L_wavg, L2_wavg, 5yr_L_wavg

### Institutions and elections
- Revised Combined Polity Score (single regime score, runs from 1 (full democracy) to -1 (full autocracy)) — Source: Center for Systemic Peace.
  - Permutations: L, L2, fd_L, fd_L2, 5yr_L, 10yr_L, wavg; Selected: L, L2, 5yr_L, 10yr_L
- Checks and balances index — Source: Data base on Political Institutions (DPI) 2015.
  - Permutations: L, 5yr_L, 10yr_L, wavg; Selected: L
- Bureaucracy Quality — Source: PRS Group.
  - Permutations: L, 5yr_L, 10yr_L; Selected: L, 5yr_L, 10yr_L
- Corruption — Source: PRS Group.
  - Permutations: L, 5yr_L, 10yr_L; Selected: L, 5yr_L, 10yr_L
- Years remaining in current chief executive's term — Source: Data base on Political Institutions (DPI) 2015.
  - Permutation: L
- Legislative election held dummy — Source: Cruz, Keefer and Scartascini (2018).
  - Permutation: L
- Executive election held dummy variable — Source: Cruz, Keefer and Scartascini (2018).
  - Permutation: L
- Political Stability and Absence of Violence/Terrorism: Estimate — Source: World Bank, World Development Indicators.
  - Permutations: L, 5yr_L, 10yr_L; Selected: L, 5yr_L, 10yr_L
- Regulatory Quality: Estimate — Source: World Bank, World Development Indicators.
  - Permutations: L, 5yr_L, 10yr_L; Selected: L, 5yr_L, 10yr_L

### Demographics
- Population ages 15-64, total — Source: World Bank, World Development Indicators.
  - Permutations: L, 5yr_L, 10yr_L, wavg; Selected: L, 5yr_L, 10yr_L, 5yr_L_wavg, 10_yr_L_wavg
- Percent change of population ages 15-64, total — Source: World Bank, World Development Indicators.
  - Permutations: L, wavg; Selected: L
- Age Dependency Ratio, % of working-age population — Source: World Bank, World Development Indicators.
  - Permutations: L, 5yr_L, 10yr_L, wavg; Selected: 5yr_L, 10yr_L, 5yr_L_wavg
- Population density (people per sq. km of land area) — Source: World Bank, World Development Indicators.
  - Permutations: L, wavg; Selected: L
- Log of population relative to US — Source: IMF, World Economic Outlook.
  - Permutations: L, L2, fd_L, d3_L, 5yr_L, 10yr_L, wavg; Selected: L, L2, fd_L, d3_L, 5yr_L, 10yr_L, L_wavg, L2_wavg

### Natural resources and sector shares
- Dummy: Fuel exporter — Source: IMF, World Economic Outlook.
- Dummy: Fuel exporter or VELIC commodity exporter — Source: IMF.
- Value of oil export, % of GDP in USD — Source: IMF, World Economic Outlook.
  - Permutations: L, L2, fd_L, fd_L2, wavg; Selected: L, L2, fd_L, fd_L2
- Mineral rents (% of GDP) — Source: World Bank, World Development Indicators.
  - Permutations: L, L2, 5yr_L, 10yr_L, mean_L, wavg; Selected: L, L2, 5yr_L, 10yr_L, mean_L
- Oil rent (% of GDP) — Source: World Bank, World Development Indicators.
  - Permutations: L, L2, 5yr_L, 10yr_L, mean_L, wavg; Selected: L, L2, 5yr_L, 10yr_L, mean_L
- Total natural resources rent (% of GDP) — Source: World Bank, World Development Indicators.
  - Permutations: L, L2, 5yr_L, 10yr_L, mean_L, wavg; Selected: L, L2, 5yr_L, 10yr_L, mean_L
- Agriculture, forestry, and fishing, value added (% of GDP) — Source: World Bank, World Development Indicators.
  - Permutations: L, wavg; Selected: L

### Country category dummies
- Dummy: Monetary Union — Source: IMF, World Economic Outlook. Permutation: L
- Dummy: Island country — Source: Wikipedia. Permutation: Dummy: Island country
- Dummy: Landlocked country — Source: CIA, World Factbook.
- Dummy: Small state — Authors' calculations based on IMF, World Economic Outlook and World Bank, World Development Indicators. Permutation: Dummy: Small state

### Notes on notation and debt definitions
- Notation explained:
  - L = lag; L2 = second lag; fd_L = lag of first difference; fd_L2 = second lag of first difference; 5yr_L = lag of the trailing 5 year difference; 10yr_L = lag of the trailing 10 year difference; d3_L = lag of the trailing 3 year difference; pc3_L = lag of percentage change over trailing 3 year; mean_L = lag of trailing 10 year moving average; wavg = cross sectional weighted average for all the permutations
- Debt definitions (footnotes):
  1. Public debt includes total debt liabilities of the government with domestic and foreign creditors. In compiling public debt series for each country, different perimeters of government (non-financial public sector, general government, and central government) reported in the Global Debt Database are compared, choosing the debt category for which the time series is the longest. This often results in a narrow definition of debt (central government) but ensures consistency across time.
  2. Public external debt is defined in terms of the residency of holder. It includes general government debt and debt guaranteed by the government and, as such, it may have a wider sectoral coverage than our measure of total public debt. Attempts to construct alternative measures based on currency-denomination were limited by time series availability and were excluded.
  3. External debt includes total debt liabilities of a country (both for the government and private sector) with foreign creditors.

*Source: wpiea2020001-print-pdf - APPENDIX 4. DATA: DEFINITION, SOURCES, AND PREDICTOR GROUPINGS*

---

### APPENDIX 5. METHODOLOGICAL DETAILS

### Empirical model and prediction window
- Prediction window: two years.
- Observations: only consider observations in which a country is not in a crisis in year t; drop all crisis years after the start of a crisis episode.
- Model: Random Forest (RF), an ensemble method based on decision trees.
  - Trees trained with two random perturbation mechanisms:
    1. each tree is trained on a bootstrap sample;
    2. at each split optimal variables are chosen from a random subset m_try of the m predictors (i.e. m_try < m).
  - Prediction for each leaf is the mean outcome for observations in that leaf; trees are fit to minimize mean squared errors.
  - Overall RF prediction is the average prediction across trees.
- Training approach: pool all countries.
- Tuning parameter m_try chosen from a grid through cross-validation.
- Number of trees: 2000.
- Trees grown exhaustively (no other restrictions on tree growth).

### Sample splitting and rolling cutoff
- Sample split into training (model estimation) and test (evaluation) using a rolling cutoff year beginning in 2000.
- Procedure:
  - Start by estimating a model with data for 1980–2000.
  - Roll forward estimation and testing periods, adding one year at a time.
  - Total estimated models: 15 (each based on a larger training sample than the previous, with hyperparameter retuned each round).

### Hyperparameter tuning procedure
- For each of the 15 training samples:
  - Use k-fold cross validation where k is the number of years of that training sample.
  - Choose m_try to minimize out-of-sample log-likelihood loss.
  - Tuning length parameter set to 10 (i.e. blind search 10 times to search for the optimal m_try).
- Tuning steps:
  1. Partition training sample into k equal sized subsamples. Fit model to k-1 subsamples and predict for the kth; repeat for each subsample to obtain out-of-sample predictions for each observation. Choose best m_try in terms of log-likelihood loss.
  2. Using selected tuning parameter values, fit model to the entire training set.
  3. Using fitted model, produce predicted probability of a crisis for the corresponding testing set.
- Note: Same hyperparameter tuning procedure applied for variable selection algorithms.

### Evaluation measures
- Models evaluated on out-of-sample predictive performance.
- Primary metrics:
  - AUROC (AUC): standard accuracy measure that does not require specification of a probability threshold.
  - Mean squared error (MSE): calculated as
    1/N [ sum_{i in I_crisis} (1 − p_i)^2  +  sum_{i in I_non-crisis} p_i^2 ],
    where p_i denotes the predicted probability of crisis for observation i, I_crisis and I_non-crisis denote sets of crisis starts and non-crisis observations in the test sample, and N is the number of observations in the test sample.
  - Log-likelihood: calculated as
    1/N [ sum_{i in I_crisis} log(p_i)  +  sum_{i in I_non-crisis} log(1 − p_i) ].
- Notes on metrics:
  - MSE will have a value of 1 if the model misses entirely all crises in the test sample.
  - Log-likelihood will be -∞ if the model assigns zero probability to an observed crisis.
- All evaluation measures computed as the average of each metric across the 15 test samples.

### Comparison of out-of-sample performance across algorithms
- Algorithms compared: Random Forest (RF, using original 748 variables), LASSO, Support Vector Machine (SVM).
  - LASSO: shrinkage and selection method for logistic regression; penalizes large coefficients and forces most coefficients to zero.
  - SVM: classification algorithm that, after non-linear feature transformation, estimates a separating hyperplane.
- Empirical findings:
  - RF performs as well or better than LASSO and SVM across the three evaluation metrics.
  - In terms of AUROC, RF is clearly superior to both LASSO and SVM for advanced and emerging market economies.
  - For other metrics and income groups, RF is broadly better though differences are not statistically significant.

### Appendix Table A5.1. Out-of-Sample Performance (selected figures)
- Advanced and emerging market economies:
  - AUROC: RF 0.806; Lasso 0.763**; SVM 0.767*** (bootstrapped standard deviations in parentheses: (0.026), (0.03), (0.025))
  - Log(likelihood): RF -0.299; Lasso -0.299; SVM -0.295 (SDs: (0.026), (0.036), (0.032))
  - MSE: RF 0.087; Lasso 0.089; SVM 0.087 (SDs: (0.01), (0.013), (0.012))
- Low-income countries:
  - AUROC: RF 0.706; Lasso 0.676; SVM 0.682 (SDs: (0.043), (0.043), (0.04))
  - Log(likelihood): RF -0.556; Lasso -0.581; SVM -0.587 (SDs: (0.039), (0.052), (0.056))
  - MSE: RF 0.189; Lasso 0.2*; SVM 0.201 (SDs: (0.017), (0.021), (0.022))
- Note: Bootstrapped standard deviations based on 100 random resamples of the test sample with replacement. Stars indicate degree of confidence that a model outperforms the Random Forest estimated with the full set of 748 variables: * 90%, **=95% , and ***=99%. Models predict probability of crisis start occurring in year t+1 or t+2. Out-of-sample performance obtained from 15 rolling regressions.

### Partial Dependence Plots (PDPs)
- PDPs show the relationship between a predictor (or set of predictors) and the model-predicted outcome by marginalizing over the distribution of other features.
- Definitions and estimation:
  - Predictor set X = {x1, x2, x3, ..., xn}. Construct subset X_S containing the predictor(s) of interest (e.g., {x1} or {x1, x2}). Let X_C be the complementary set.
  - PDP of a predicted response variable for X_S:
    f_S(X_S) = E_C[ f(X_S, X_C) ] = ∫ f(X_S, X_C) p_C(X_C) dX_C
    where p_C(X_C) is the marginal probability of X_C.
  - Estimate of partial dependence using observed predictor data:
    f_S(X_S) ≈ (1/N) ∑_{i=1}^N f( X_S, X_i^C )
    where N is the number of observations and X_i = (X_i^S, X_i^C) for the ith observation.
- Interaction note:
  - If two variables X_j and X_k do not interact, the partial dependence function decomposes into the sum of individual PDPs:
    PD_jk(X_j, X_k) = PD_j(X_j) + PD_k(X_k)
  - If they interact, bivariate PDPs cannot be expressed as the sum of univariate PDPs.

*Source: wpiea2020001-print-pdf - APPENDIX 5. METHODOLOGICAL DETAILS*

---


_Source: https://www.imf.org/-/media/files/publications/wp/2020/english/wpiea2020001-print-pdf.pdf_
