## Predicting the demand for IMF arrangements: A machine learning approach

## Source details

**Canonical URL:** [Predicting the demand for IMF arrangements: A machine learning approach](https://www.imf.org/-/media/files/publications/wp/2024/english/wpiea2024054-print-pdf.pdf)

## Other formats

- [Markdown version](/-/media/files/publications/wp/2024/english/wpiea2024054-print-pdf.pdf.md)
- [Structured JSON version](/-/media/files/publications/wp/2024/english/wpiea2024054-print-pdf.pdf.json)

---

### Motivation and objectives
- IMF functions as a lender of last resort within the global financial safety net (GFSN); anticipating member countries’ financing needs is crucial to ensure the IMF remains adequately funded.
- Research questions:
  - Can machine learning (ML) improve forecasts of new IMF-supported arrangements versus traditional econometric methods, and which ML methods perform best?
  - Which factors indicate future use of IMF resources and how model-sensitive are they?
  - How should international institutions train and use ML models to optimize performance and maintain relevance over time?

### Data, sample, and preprocessing
- Sample period and coverage:
  - 1982-2021 for 189 countries
  - 710 arrangements total: 409 GRA arrangements and 301 PRGT arrangements
  - Programs requested by 129 different member countries; number of programs observed per country ranges from 0 to 13
- Exclusions and rationale:
  - Emergency financing instruments excluded: 90 Rapid Credit Facilities (RCF) and 46 Rapid Financing Instruments (RFI)
    - Total approved RCFs and RFIs ~ SDR 24 billion, compared to SDR 615 billion for approved arrangements included in the sample
  - Undrawn arrangements excluded (112 of 130 undrawn arrangements were officially classified as precautionary arrangements)
- Dependent variable (baseline):
  - Binary = 1 for country in each of the two years prior to approval of a new IMF-supported arrangement (t-1 and t-2)
  - Exclude years t, t+1, and t+2
  - Average duration of programs in the sample is 2.4
- Predictors and feature engineering:
  - Predictors span external, fiscal, financial, real, global, and structural/institutional variables (including VIX index, U.S. treasury yields, federal funds rate, corruption and polity scores, costs of natural disaster hazards, population growth, indicators of access to RFAs and central bank swap lines, political closeness to the US)
  - Transformations: growth rates, accelerations, and lags
  - Full feature set comprises 1014 features; six feature sets defined (Full Set and Sets 1–6) with exact feature counts: 1014, 480, 277, 109, 76, 70, and 40 as specified
- Missing-data imputation:
  - Methods: K-nearest neighbors’ imputation; mean imputation by country income-level group; median imputation by country income-level group
  - Around one third of predictors have share of missing values larger than 50 percent
  - KNN imputation preserved distributional shape more closely than mean/median imputation (Kolmogorov-Smirnov tests rejected equality of distributions for most variables when comparing KNN vs mean/median)
- Data unit transformations:
  - Variables in national currency units converted to ratios (proportion of GDP, exports, imports, fiscal revenues, or IMF quotas) or percentage changes

### Model set and experimental design
- Models considered:
  - Logistic regression; regularized logistic regression (L1/Lasso, L2/Ridge, Elastic Net); kernelized SVM; KNN; ensemble: random forest, extra tree; boosting: XGBoost, ADABoost, RUSBoost; deep learning: recurrent neural network (RNN)
- Training/test split:
  - Training: 1982-2017
  - Out-of-sample hold-out test: 2018-2021
- Cross-validation:
  - Expanding-window gap k-fold time-split with 4 folds (folds and gap years specified exactly)
    - Fold 1: Training 1982-1988; Validation 1990-1996; Gap 1989
    - Fold 2: Training 1982-1995; Validation 1997-2003; Gap 1996
    - Fold 3: Training 1982-2002; Validation 2004-2010; Gap 2003
    - Fold 4: Training 1982-2009; Validation 2011-2017; Gap 2010
- Class imbalance handling for training (minority = pre-arrangement periods):
  - Minority observations: 16 percent; Majority observations: 84 percent
  - Sampling methods evaluated: no sampling; random under-sampling; SMOTE; ADASYN
- Hyperparameter tuning:
  - Best hyperparameters selected by highest average ROC-AUC across four validation folds; if tied, select lower standard deviation of ROC-AUC across folds
  - Sampling techniques and imputation applied only to training sample after splitting to preserve out-of-sample validity

### Evaluation metrics and thresholding
- Primary metric: ROC-AUC
- Additional evaluation: precision and recall trade-offs via precision-recall curves across classification thresholds
- Predicted probabilities converted to binary predictions by classification thresholds; lower thresholds reduce missed arrangements at cost of more false alarms
- Models produce widely varying predicted-probability distributions; algorithm-tailored thresholding and adaptive thresholds by income groups or time-varying predicted probability recommended

### Key empirical findings — classification performance
- ML-based techniques outperform traditional logistic regression in out-of-sample prediction of new IMF-supported arrangements
- Representative models for deeper analysis: regularized logistic regression, random forest, recurrent neural network
- Top-performing algorithms on hold-out ROC-AUC:
  - Random forest: Test ROC-AUC 0.86
  - Extra Tree: Test ROC-AUC 0.86
  - RNN: Test ROC-AUC ~ 0.84
  - XGBoost: Test ROC-AUC ~ 0.84–0.85 depending on specification
- Logistic-based models:
  - Logistic and regularized variants achieve ROC-AUC roughly 0.81 on the test set; logistic variants exhibit identical ROC-AUC scores on test set at three-digit level
- Model-specific data-processing preferences:
  - Linear models prefer mean imputation and random under-sampling
  - Tree-based models prefer KNN imputation and over-sampling methods
  - RNN and SVM perform best with median imputation
  - All models tend to perform better with smaller, model-selected feature sets; model-selected sets (sets 4-6) outperform pre-defined sets (sets 1-3)
- Summary of best-performing models (selected entries with exact settings):
  - Random Forest — Imputation: KNN; Variable Set: Set 5; Sampling: SMOTE; Test ROC-AUC: 0.86
  - Extra Tree — Imputation: KNN; Variable Set: Set 6; Sampling: SMOTE; Test ROC-AUC: 0.86
  - XGBoost — Imputation: None; Variable Set: Set 5; Sampling: No Sampling; Test ROC-AUC: 0.85
  - RNN — Imputation: Median; Variable Set: Set 4; Sampling: Under-Sampling; Test ROC-AUC: 0.84
  - KNN — Imputation: KNN; Variable Set: Set 5; Sampling: Under-sampling; Test ROC-AUC: 0.84
  - SVM — Imputation: Median; Variable Set: Set 6; Sampling: ADASYN; Test ROC-AUC: 0.83
  - Logistic, Logistic Lasso, Logistic Ridge, Logistic Elastic Net — Imputation: Mean; Variable Set: Set 4; Sampling: Under-sampling; Test ROC-AUC: 0.81

### Precision–recall and thresholding insights
- Precision and recall definitions:
  - Precision = proportion of predicted pre-arrangement periods that were true pre-arrangement periods
  - Recall = proportion of true pre-arrangement periods correctly predicted
- Precision–recall findings (test period 2018-2021):
  - Extra Tree outperforms peers in precision when recall is between 0.3 and 0.7
  - RNN and XGBoost outperform Random Forest and Extra Tree at higher recall levels (above 0.7)
  - KNN achieves highest precision when recall is 0.93
- Classification thresholds that maximize F1 vary across models (examples given range from 0.35 in Random Forest to 0.72 in RNN)
- Distribution of predicted probabilities (test set 2018-2021) examples:
  - Logistic lasso: true negatives 0 to 0.96; true positives 0.16 to 0.97
  - Random forest: true negatives 0 to 0.84; true positives 0.09 to 0.92
  - RNN: true negatives 0 to 0.99; true positives 0.01 to 0.98
- Practical thresholding caution:
  - Percentile-based thresholds can misrepresent low absolute probabilities for models with concentrated near-zero outputs (e.g., RNN)

### Predictors and interpretability (Shapley analysis)
- Shapley values used to decompose feature contributions and capture non-linear interactions; do not imply causation
- Consistently important predictors across algorithms include:
  - Fiscal variables: public debt to revenue; general government revenue
  - External variables: current account balance; external debt; foreign share of general government gross debt
  - Real variables: PPP income per capita relative to the US
  - Financial variables: private credit
  - Structural variables: membership in a regional financing arrangement (RFA); access to central bank swap lines; IMF credit outstanding
- Model-specific differences in top Shapley contributors:
  - Random forest: membership in a RFA has the greatest absolute Shapley value, followed by debt and current account features
  - Logistic lasso: assigns highest Shapley value to gross international reserves in percent of imports
- Nonlinearities and interactions:
  - Shapley value thresholds and magnitudes differ by income group (AEs, EMs, LIDCs)
  - Example — reserves-to-imports:
    - EMs and LIDCs: reserves below 50 percent of annual imports amplify likelihood of IMF-supported program
    - AEs: threshold at 25 percent
  - Public debt to revenue thresholds where Shapley values escalate sharply:
    - around 150% for AEs
    - 20% for both EMs and LIDCs
  - Current account deficits:
    - In AEs, likelihood rises when deficits fall below 5% of GDP
    - In EMs and LIDCs, even small deficits linked to higher probability of program commencement
  - Membership in an RFA reduces likelihood across all groups with Shapley ranges:
    - AEs: -0.11 to -0.03
    - EMs: -0.35 to -0.04
    - LIDCs: -0.35 to -0.1

### Robustness, temporal stability, and limitations
- Robustness exercise 1 (adding literature-identified variables individually to best-performing sets):
  - 21 additional variables tested (e.g., reserves percent of ARA metric; bank capital adequacy ratios; net non-FDI liability inflows; access to central bank swap line dummy; federal funds rate; 10-year treasury yields; IMF credit outstanding)
  - Non-linear models (e.g., Random Forest) maintain prediction performance; linear models (logistic lasso) display significant performance variation
  - Largest increase in logistic lasso test ROC-AUC when IMF credit outstanding is added; RF reports a decline with this addition
  - Recommendation for operations: include credit outstanding in prediction exercise to account for debt rollover and repeated use of fund resources
- Robustness exercise 2 (temporal stability):
  - Time-split evaluations with five training/test configurations (training expanded progressively and tests defined as specified)
  - Observations:
    - Higher ROC-AUC scores for models trained on more recent data
    - Average out-of-sample ROC-AUC scores tend to decline over time; noticeable dips during global crises (Asian Financial Crisis, Global Financial Crisis, European debt crisis, COVID-19 pandemic)
    - Models failed to predict the surge in IMF-supported programs following Global Financial Crisis, particularly among AEs
    - RF outperforms competitors in near-term predictions in each fold but deteriorates more rapidly as training data becomes outdated; deterioration generally apparent when training data is outdated by at least 8-10 years
- Limitations:
  - Prediction difficulty increases around periods of global financial distress
  - Reliance on timely and reliable data; publication lags and noise in national accounts pose challenges
  - Models may miss non-economic influences (e.g., stigma, shifts in risk perception)
  - Nonlinearities and potential model misspecification, multicollinearity, and omitted variable biases noted for size prediction regressions

### Practical model-design implications and recommended practices
- Algorithm-tailored approach:
  - Data processing choices (imputation, sampling, feature selection) materially affect performance and differ by algorithm
  - Use model-selected feature sets (smaller sets) rather than full pre-defined sets when possible
- Thresholding:
  - Use algorithm-specific and adaptive thresholds (by income group or time-varying changes in predicted probability) to reflect cross-country heterogeneity and varying probability distributions
- Ensembles:
  - Ensemble averaging across representative models (regularized logistic regression, random forest, RNN) mitigates individual-model outliers and provides dispersion (standard deviation) across models
  - Ensemble performance example (test set 2018-2021):
    - Average predicted probability > 0.5 for 46 out of the 69 approved UCT arrangements between 2018-2021
    - Average predicted probability > 0.5 for 33 out of the 65 approved emergency financing instruments between 2018-2021
- Model maintenance:
  - Regular model updates recommended to maintain predictive accuracy; if frequent updates infeasible, prefer temporally more stable models even if less optimal for near-term forecasts

### Predicting the size of IMF arrangements (two-step illustrative approach)
- Approved arrangement sizes in sample range from 14.5 to 3211.8 percent of quota
- Two-step procedure:
  - Use binary prediction results from Random Forest to identify likely arrangement starts
  - Regress log approved size of IMF arrangements on lagged explanatory variables (examples use RF-selected 40-variable set and RF-40 Plus)
- Sample and estimation details:
  - Period 1982-2021; training 1982-2017; hold-out 2018-2021
  - Only drawn GRA arrangements included for illustrative size analysis
  - 585 data points in training set after dropping no-arrangement observations
- Key findings:
  - Model using set 6 underestimates size of approved arrangements in every year of 2018-2021 testing sample
  - Model using set 6 plus additional variables predicts larger use and overestimates sum of approved arrangements in some years (2019 and 2020)
  - Inclusion of credit outstanding to the Fund improves ability to predict arrangements with large access levels significantly
  - Best performing model (out-of-sample root MSE and adjusted r-square) uses RF-40 Plus variable set and includes observations from 1990-2017
- Selected numeric model-fit metrics (four specifications: 1982-2017 RF; 1990-2017 RF; 1982-2017 RF PLUS; 1990-2017 RF PLUS):
  - Observations: 585; 388; 585; 388
  - R-squared: 0.3; 0.401; 0.372; 0.468
  - Adjusted R-squared: .25; .33; .30; .38
  - Root MSE: .76; .81; .73; .78
  - Root MSE OOS: 1.01; 0.96; 1.04; 0.90
  - Constant estimates: 4.6***; 4.538***; 4.569***; 4.450*** (standard errors reported in parentheses in source)
- Selected statistically significant coefficient examples (significance codes: *** p<0.01, ** p<0.05, * p<0.1):
  - Dollar Appreciation (REER, CPI), 1990-2017 RF: -0.335*** (0.123)
  - PPP income per capita relative to the U.S., 1982-2017 RF: 1.0*** (0.4)
  - Total credit outstanding in percent of quota (RF PLUS): 0.0942*** (0.0305) and 0.0934** (0.0382) for alternate samples
  - Access to swap line (RF PLUS): 1.066*** (0.240) and 0.707** (0.286)
  - Federal Funds Rate (RF PLUS): 0.247** (0.102) and 0.469*** (0.144)
  - External bank liabilities, percent of GDP (RF PLUS): 0.627* (0.320)
- Cautions:
  - Non-linearities and outliers complicate size prediction; regression results should be interpreted cautiously due to potential misspecification and omitted variable bias

### Agreement, dispersion, and country-level illustrations
- Representative models (regularized logistic Lasso, Random Forest, RNN) show substantial agreement on many country-level risks:
  - All models identify most African economies, several Middle East economies, and several South American economies as likely future IMF resource users
  - All models ascribe medium to high likelihood to the 18 countries which approved arrangements in 2023; each model correctly assigns high probability to five of those countries
  - RNN assigns medium likelihood to the United States and Finland in 2021 predictions
  - Logistic Lasso tends to assign fewer countries above the 50th percentile medium-likelihood threshold and is the only model assigning probability below 50th percentile to Argentina and Ukraine in the 2021 illustration
- Ensemble approach (mean and standard deviation across models) provides:
  - Mitigation of individual algorithm outliers
  - Dispersion measure: high mean and low standard deviation = strong model agreement
  - Ensemble performance on test set 2018-2021 summarized above

### Main conclusions and avenues for future work
- ML models consistently outperform econometric methods in out-of-sample prediction of new IMF arrangements; tree-based ensembles and RNNs among top classifiers
- Important predictors span external, fiscal, real, financial, and institutional factors (e.g., IMF credit outstanding; RFA membership)
- Data processing decisions (imputation, sampling, feature selection) substantially affect results and should be algorithm-tailored
- Temporal stability is limited: performance dips during global crises and when training data become outdated by 8-10 years; regular model updates recommended
- Promising extensions:
  - Direct ML prediction of arrangement size
  - Expanded analysis of predictor interactions, alternative feature sets, and long-term stability under evolving global financing conditions
  - Use of natural language processing to address data lags and capture non-economic drivers

*IMF Working Paper: Predicting the demand for IMF arrangements: A machine learning approach.*

### Bibliography _______________________________________________________________________ 35

### Predicting the demand for IMF arrangements: A machine learning approach

### Motivation and objectives
- IMF functions as a lender of last resort within the global financial safety net (GFSN); anticipating member countries’ financing needs is crucial to ensure the IMF remains adequately funded.
- Study aims to assess whether machine learning (ML) techniques can improve forecasts of new IMF-supported arrangements versus traditional econometric methods and to identify best-performing methods.
- Three central questions: (1) Can ML improve forecasts and which methods perform best? (2) Which factors indicate future use of IMF resources and how model-sensitive are they? (3) How should international institutions train and use ML models to optimize performance and maintain relevance over time?

### Data, sample, and preprocessing
- Sample period: 1982-2021 for 189 countries.
- Coverage: essentially all IMF-supported arrangements (GRA and PRGT), excluding emergency financing instruments and undrawn arrangements.
- Predictors span external, fiscal, financial, real, and structural/institutional variables.
- Missing-data imputation methods used include K-nearest neighbors’ imputation and imputation based on mean and median values by income-level country categorization.
- To preserve out-of-sample validity, sampling techniques are applied only to the training sample after splitting into training and validation/test sets.

### Model set and experimental design
- Models considered: logistic regression, regularized logistic regression, kernelized support vector machine (SVM), K-nearest neighbors (KNN), ensemble models (random forest (RF), extra tree), boosting techniques (XGBoost, RUSBoost, ADABoost), and a deep learning model (recurrent neural network (RNN)).
- Training/test split: training sample covers 1982-2017; out-of-sample hold-out test set covers 2018-2021.
- Class imbalance handling: Synthetic Minority Oversampling Technique (SMOTE), Adaptive Synthetic (ADASYN) technique, and random under-sampling produce an even split between observations with and without IMF-supported arrangements for training.
- Cross-validation: training set split into four folds, each composed of an in-sample training and out-of-sample validation set based on an expanding time window.
- Hyperparameter tuning: select best hyperparameter setting via highest average ROC-AUC across the four validation sets; best models refitted on entire training sample and evaluated on hold-out test set.

### Evaluation metrics and thresholding
- Primary performance metric: ROC-AUC (receiver operating curve – area under the curve).
- Additional evaluation: precision and recall trade-off examined via precision-recall curves across classification thresholds.
- Predicted probabilities are converted to binary predictions by comparing to classification thresholds; lower thresholds reduce missed arrangements at the cost of more false alarms.
- Models produce widely varying distributions of predicted probabilities, motivating algorithm-tailored thresholding and adaptive thresholds based on income groups or time-varying changes in predicted probability.

### Key empirical findings
- ML-based techniques outperform traditional logistic regression in out-of-sample prediction of new IMF-supported arrangements.
- Top-performing algorithms on hold-out ROC-AUC: random forest, extra tree, and recurrent neural network.
- Extra tree model achieves the best combined precision and recall at most classification thresholds, followed by random forest; RNN and XGBoost outperform RF at some combinations of higher recall and lower precision.
- Representative models selected for deeper analysis: regularized logistic regression, random forest, and recurrent neural network.
- Ensemble approaches that combine representative models can mitigate individual-algorithm outliers and provide insight into dispersion across models.

### Predictors and interpretability
- Shapley values used to analyze feature importance and non-linear interactions.
- Consistently important predictors across algorithms include:
  - Fiscal variables: public debt to revenue, general government revenue.
  - External variables: current account balance, external debt.
  - Real variables: per capita income relative to the US.
  - Financial variables: private credit.
  - Structural variables: access to regional financing arrangements (RFAs).
- Selected features show minimal variation across models; non-linear relationships and interactions between predictors are documented.

### Robustness, temporal stability, and limitations
- Robustness exercise 1: adding variables previously identified in the literature but not selected by baseline models shows ensemble non-linear models (e.g., random forest) maintain prediction performance; linear models (regularized logistic regression) display significant performance variations.
- Robustness exercise 2: temporal stability tested by training RF, RNN, and regularized logistic regression on sub-samples of different time periods and evaluating out-of-sample performance in subsequent years.
  - Finding: general decline in out-of-sample performance over time, with notable performance dips during global crises (Asian Financial Crisis, Global Financial Crisis, COVID-19 pandemic).
  - Implication: changing nature of IMF-supported program use (supply changes like IMF policy adjustments; demand shifts like alternative funding sources) affects predictive stability.
- Models that are most accurate for near and medium-term (5-7 years) forecasts can experience reduced performance in the long term (beyond eight years), emphasizing the need for regular model updates or selection of more temporally stable models when updates are impractical.

### Practical model-design implications and recommended practices
- Data processing and modeling choices materially affect performance: best-performing models select different feature sets, sampling methods, and missing-data imputation techniques, suggesting an algorithm-tailored and flexible approach.
- Threshold setting should be algorithm-specific and can benefit from adaptive thresholds by country income group or changes in predicted probability over time to reflect cross-country heterogeneity.
- Ensemble methods offer a way to reduce sensitivity to individual-model outliers.
- Regular model updates are important to maintain predictive accuracy over time; if frequent updates are infeasible, prefer more stable models even if they are less optimal for near-term forecasts.

### Extensions and future work
- The study uses binary prediction outputs as input to a linear regression to forecast the size of IMF-supported arrangements; results provide suggestive evidence but warrant further exploration.
- Future work could examine ML-based techniques directly for predicting arrangement size and expand analyses of predictor interactions, alternative feature sets, and long-term stability under evolving global financing conditions.

*IMF Working Paper: Predicting the demand for IMF arrangements: A machine learning approach.*

### 2. Data Processing

### 2. Data Processing

### Empirical strategy and preprocessing steps
- Empirical strategy includes four main steps: (i) data processing, feature selection, and imputation, (ii) model selection, data splitting, and sampling, (iii) model training and evaluation, and (iv) model analysis.
- Variables in national currency units are converted to ratios (proportion of GDP, exports, imports, fiscal revenues, or IMF quotas) or percentage changes to avoid inconsistencies in currency units and the impact of changing price levels.
- Transformations generated for each variable include growth rates, accelerations, and lags.
- Missing values are imputed using three methods:
  - K-nearest neighbors’ imputation
  - Mean imputation by country income-level group
  - Median imputation by country income-level group

### Sample selection and dependent variable
- Sample covers IMF-supported arrangements approved between 1982 and 2021.
- Sample composition:
  - 189 countries
  - 710 arrangements total
  - 409 GRA arrangements
  - 301 PRGT arrangements
  - Programs requested by 129 different member countries; number of programs observed per country ranges from 0 to 13
- Exclusions:
  - Emergency financing instruments excluded: 90 Rapid Credit Facilities (RCF) and 46 Rapid Financing Instruments (RFI)
    - Rationale: emergency financing differs in design/purpose, is typically one-time disbursement without ex-post program-based conditionality, does not align with dependent variable definition excluding two periods after commencement, and is highly concentrated during the Covid-19 pandemic which could distort test-set predictions
    - Footnote figures: total amount of approved RCFs and RFIs was about SDR 24 billion, compared to SDR 615 billion for the approved arrangements included in the sample
  - Undrawn arrangements (arrangements that ended without an actual balance of payments need) are excluded
    - Note: 112 of the 130 undrawn arrangements were officially classified as precautionary arrangements
- Dependent variable (baseline specification):
  - Binary variable equal to one for a country in each of the two years prior to the approval of a new IMF-supported arrangement (i.e., t-1 and t-2)
  - Exclude the year of approval and the following two years (periods t, t+1, and t+2)
  - The definition focuses on predicting materialized IMF arrangements; program requests not resulting in approval are treated as non-program observations
  - Average duration of programs in the sample is 2.4

### Feature selection and missing value imputation
- Predictor coverage spans external, fiscal, financial, and real sectors plus global variables (VIX index, U.S. treasury yields, federal funds rate), structural variables (corruption and polity scores from ICRG), costs of natural disaster hazards, population growth, and indicators of access to alternative safety-net layers (central bank currency swap lines, regional financing arrangements) and political closeness to the US (based on UN assembly votes).
- Robustness analysis finds including credit outstanding with the IMF as an explanatory variable can significantly improve prediction performance.
- Missing-data context and approach:
  - Around one third of predictors have a share of missing values larger than 50 percent
  - Mean and median imputation were performed by three country groups: advanced economies (AEs), low-income developing economies (LIDCs), and EMDEs that are non-LIDCs
  - The KNN imputed series preserved distributional shape more closely than mean/median imputation; Kolmogorov-Smirnov tests rejected equality of distributions for most variables when comparing KNN vs mean/median imputation
- Feature engineering:
  - Full feature set (including all transformations) comprises 1014 features
  - Six feature sets defined for model comparison (feature selection methods also used):
    - Full Set: 1014 features; transformations included: 1-year growth rate, 2-year growth rate, 3-year growth rate, 5-year growth rate of levels and ratios, acceleration, first and second lags; Missing Value Threshold: None; Other criteria: Drop variables denoted in local currency units and levels; Time Period: 1982-2021
    - Set 1: 480 features; transformations: 1-year and 5-year growth rates of levels and ratios, acceleration, first lags; Missing Value Threshold: 85% (eliminates 5 base variables*); Other criteria: Drop highly correlated variables; Time Period: 1982-2021
    - Set 2: 277 features; transformations: 1-year and 5-year growth rates of levels, first lags; Missing Value Threshold: 70% (eliminates 9 base variables*); Other criteria: None; Time Period: 1982-2021
    - Set 3: 109 features; transformations: 1-year growth rate of levels; Missing Value Threshold: 50% (eliminates 9 base variables*); Other criteria: None; Time Period: 1982-2021
    - Set 4: 76 features; transformations: Not applicable; Missing Value Threshold: Not applicable; Other criteria: Lasso; Time Period: 1982-2021
    - Set 5: 70 features; transformations: Not applicable; Missing Value Threshold: Not applicable; Other criteria: Random Forest – Recursive Feature Elimination; Time Period: 1982-2021
    - Set 6: 40 features; transformations: Not applicable; Missing Value Threshold: Not applicable; Other criteria: Random Forest – Recursive Feature Elimination; Time Period: 1982-2021
  - *Notes: Base variables are those used to calculate the transformations included in the full dataset. Set 5 and 6 both rely on random forest recursive feature elimination but select different numbers of features (70 vs. 40).
  - Selected feature elimination steps do not address multicollinearity; multicollinearity likely present but not problematic for prediction-focused models

### Models considered (overview)
- Traditional and ML classification models compared:
  - Logistic regression (baseline)
  - Regularized logistic regression (L1/Lasso, L2/Ridge, Elastic Net)
  - K-nearest neighbors (KNN)
  - Kernelized support vector machine (SVM)
  - Ensemble models: random forest, extra tree
  - Boosting techniques: XGBoost, ADABoost, RUSBoost
  - Deep learning model: recurrent neural network (RNN)
- Model characteristics summarized:
  - Regularization: Lasso (L1) can set coefficients to zero; Ridge (L2) shrinks coefficients toward zero but not exactly to zero; Elastic Net mixes L1 and L2
  - Random forest: ensemble of decision trees trained on random subsets of data/features; prediction by majority vote
  - Boosting (e.g., XGBoost): sequential tree building where each tree corrects previous errors

### Data splits, cross-validation, and hyperparameter tuning
- Training/test split:
  - Training set: years 1982-2017
  - Out-of-sample test set (hold-out): years 2018-2021
- Cross-validation procedure:
  - Expanding-window gap k-fold time-split approach (4 folds) used to account for time dependence and avoid data leakage
  - Training and validation periods for each fold:
    - Fold 1: Training 1982-1988; Validation 1990-1996; Gap year: 1989
    - Fold 2: Training 1982-1995; Validation 1997-2003; Gap year: 1996
    - Fold 3: Training 1982-2002; Validation 2004-2010; Gap year: 2003
    - Fold 4: Training 1982-2009; Validation 2011-2017; Gap year: 2010
  - Candidate hyperparameter settings for each model type are evaluated on each fold; the setting with best average out-of-sample performance across the four validation sets is selected
- Hyperparameter selection metric:
  - ROC-AUC (area under the Receiver Operating Characteristic curve) averaged across the four validation sets is used as the principal evaluation metric in cross-validation
  - If two hyperparameter settings yield similar average ROC-AUC, the setting with lower standard deviation of ROC-AUC across folds is selected
  - Precision-recall curves and trade-offs between false positive and false negative rates are analyzed later in Section 4

### Class imbalance and sampling methods
- Class imbalance in dependent variable:
  - Minority observations (pre-arrangement periods): 16 percent
  - Majority observations (non-pre-arrangement periods): 84 percent
  - A naive classifier predicting all observations as non-pre-arrangement would achieve 84 percent accuracy
- Sampling strategies evaluated to address imbalance (applied during model training to equalize class counts):
  - (i) no sampling
  - (ii) random under-sampling (reduce majority class)
  - (iii) SMOTE (Synthetic Minority Over-sampling Technique)
  - (iv) ADASYN
- Model performance depends crucially on sampling method; best-performing sampling techniques differ across algorithms

*Source: wpiea2024054-print-pdf — 2. Data Processing*

### 4. Prediction Results

### 4. Prediction Results

### Horse Race and Performance Comparison
- Models evaluated by out-of-sample performance measured by ROC-AUC scores on a hold-out testing set.
- Logistic-based models achieve an ROC-AUC score of roughly 0.81 on the test set.
- Other models achieve ROC-AUC values ranging from 0.82 to 0.86 on the test set.
- Logistic regression and its regularized variants (Lasso, Ridge, Elastic Net) show very similar predictive performance and each exhibited their best performance with the Lasso feature set; ROC-AUC scores on the test set are identical at a three-digit level for these logistic variants.
- Tree-based models are most successful in out-of-sample prediction:
  - Random forest and extra tree achieve the highest test set ROC-AUC score of 0.86.
  - RNN and XGBoost achieve scores of approximately 0.84.
- Contextual benchmarks:
  - Within ML literature, models that achieve an ROC-AUC of 0.8 or higher are typically considered to be ‘good’ classifiers.
  - Existing crisis-prediction studies achieved ROC-AUC scores in the range of 0.69 to 0.77.
- Temporal variation in performance:
  - Model performance is best across all model types in 1997-2003.
  - ROC-AUC scores decline significantly in the validation set 2004-2010.
  - No observed decline in ROC-AUC scores during the COVID-19 pandemic (emergency financing loans excluded from sample).
- Model-specific notes:
  - KNN is the worst performing model class across all validation sets but scores above several competitors in the hold-out test set; improvement may be due to greater value of a large training set for KNN.
  - Random forest (and boosting methods) display most robust performance and rank in the top-three across all validation and the test set. Random forest's ensemble approach mitigates overfitting.
  - Boosting methods (e.g., XGBoost) improve performance by sequentially training weak models and emphasizing misclassified samples.
  - RNN scores below logistic-based models in first and second validation sets but ranking improves consistently as training set size expands; deep learning models tend to improve with more observations.
- Data processing and feature selection matter:
  - Linear models consistently prefer mean imputation; tree-based models perform best with KNN imputation; RNN and SVM perform best with median imputation.
  - Linear models and RNN perform best using random under-sampling; tree-based models prefer over-sampling methods.
  - All models tend to perform better with smaller feature sets; model-selected feature sets consistently outperform pre-defined feature sets.
  - Random forest prediction performance is most robust to different data processing decisions; data processing decisions are particularly important for logistic-based models.

### Summary of Best Performing Models (Table 3)
- Random Forest — Imputation Method: KNN; Variable Set: Set 5; Sampling: SMOTE; Test ROC-AUC Score: 0.86
- Extra Tree — Imputation Method: KNN; Variable Set: Set 6; Sampling: SMOTE; Test ROC-AUC Score: 0.86
- XG Boost — Imputation Method: None; Variable Set: Set 5; Sampling: No Sampling; Test ROC-AUC Score: 0.85
- XG Boost — Imputation Method: KNN; Variable Set: Set 5; Sampling: SMOTE; Test ROC-AUC Score: 0.84
- Ada Boost — Imputation Method: Mean; Variable Set: Set 6; Sampling: SMOTE; Test ROC-AUC Score: 0.82
- RUS Boost — Imputation Method: Mean; Variable Set: Set 5; Sampling: ADASYN; Test ROC-AUC Score: 0.81
- RNN — Imputation Method: Median; Variable Set: Set 4; Sampling: Under-Sampling; Test ROC-AUC Score: 0.84
- KNN — Imputation Method: KNN; Variable Set: Set 5; Sampling: Under-sampling; Test ROC-AUC Score: 0.84
- SVM — Imputation Method: Median; Variable Set: Set 6; Sampling: ADASYN; Test ROC-AUC Score: 0.83
- Logistic — Imputation Method: Mean; Variable Set: Set 4; Sampling: Under-sampling; Test ROC-AUC Score: 0.81
- Logistic Lasso — Imputation Method: Mean; Variable Set: Set 4; Sampling: Under-sampling; Test ROC-AUC Score: 0.81
- Logistic Ridge — Imputation Method: Mean; Variable Set: Set 4; Sampling: Under-sampling; Test ROC-AUC Score: 0.81
- Logistic Elastic Net — Imputation Method: Mean; Variable Set: Set 4; Sampling: Under-sampling; Test ROC-AUC Score: 0.81

### Trade-off Between False Negatives and False Positives
- Complement ROC-AUC with precision-recall analysis across classification thresholds for test period 2018-2021.
- Definitions:
  - Precision = proportion of predicted pre-arrangement periods that were true pre-arrangement periods.
  - Recall = proportion of true pre-arrangement periods correctly predicted.
- Precision-recall curve findings (Figure 8):
  - Random forest and extra trees show strong performance.
  - Extra tree outperforms others in precision when recall is between 0.3 and 0.7, indicating a lower false positive rate in this interval.
  - RNN and XGBoost outperform random forest and extra tree at higher recall levels (above 0.7).
  - KNN achieves the highest precision score when recall is at 0.93.
- Classification thresholds that maximize F1 scores vary significantly across models:
  - Example thresholds vary from 0.35 in the random forest to 0.72 in the RNN.
- Distribution of predicted probabilities (test set 2018-2021):
  - RNN outputs very low predicted probabilities for most countries without pre-arrangement periods, while random forest and logistic lasso produce more evenly distributed probabilities.
  - Predicted probability ranges for true negatives (true positives):
    - Logistic lasso: true negatives 0 to 0.96; true positives 0.16 to 0.97
    - Random forest: true negatives 0 to 0.84; true positives 0.09 to 0.92
    - RNN: true negatives 0 to 0.99; true positives 0.01 to 0.98
  - A correct classification of all true negative events by the random forest requires a classification threshold of at least 85 percent, implying a significant share of false negative alarms under that thresholding choice.
- Observed sources of prediction difficulty:
  - True negatives with high predicted probabilities are sometimes associated with emergency financing instruments (treated as no program cases).
  - True positives with low probabilities often include countries with limited data availability or countries with strong fundamentals that made drawings under precautionary facilities.
- Threshold-selection caveat:
  - Using percentiles of predicted probabilities to compare country-level predictions across models may not account for different shapes of predicted probability distributions (e.g., RNN placing many predictions near zero and few in between). Percentile-based thresholds can place low absolute probabilities into high percentile ranges; adjusting percentiles to account for distribution shapes and recall-precision trade-offs may optimize accuracy.

### Agreement and Dispersion of Prediction Results across Models
- Representative models for country-level illustration: regularized logistic regression (Lasso), random forest, recurrent neural network.
- Maps (Figure 10) for 2021 predictions and actual active arrangements as of end-October 2023:
  - All models consistently identify most African economies, several Middle East economies, and several South American economies as likely future IMF resource users.
  - All models ascribe medium to high likelihood to the 18 countries which approved arrangements in 2023.
  - Each model correctly assigns a high probability to five countries that had a program approved in 2023; there is consensus for Pakistan, Côte d'Ivoire, Niger, Senegal, and Sri Lanka between different pairs of models.
  - RNN indicates a medium likelihood for the United States and Finland seeking IMF assistance.
  - Logistic lasso suggests a smaller number of countries surpassing the medium likelihood threshold (50th percentile) and is the only model assigning a probability below the 50th percentile threshold to Argentina and Ukraine.
- Ensemble approach:
  - Ensemble constructed by calculating the average and standard deviation of predicted probabilities across the three representative models.
  - Ensemble insights:
    - Provides prediction dispersion (agreement/disagreement across models).
    - Mitigates outliers of individual algorithms.
  - Scatter plot (Figure 11) interpretation:
    - Each dot = country-year observation in test set (2018-2021), plotted by mean predicted probability and standard deviation across models.
    - Blue dots = non-pre-arrangement periods; orange dots = observed arrangements; gray dots = emergency financing instruments.
    - A dot with mean probability close to 1 and standard deviation close to 0 indicates high model agreement that country will request an IMF-supported arrangement.
  - Ensemble performance on test set:
    - Average predicted probability > 0.5 for 46 out of the 69 approved UCT arrangements between 2018-2021.
    - Average predicted probability > 0.5 for 33 out of the 65 approved emergency financing instruments between 2018-2021.
    - Mean predicted probabilities for countries with an actual UCT arrangement or actual emergency financing approved during the test period are reported in the analysis.

*IMF Working Paper — Predicting the demand for IMF arrangements: A machine learning approach.*

### 0.59 and 0.51, respectively, compared to an overall mean predicted probability of 0.28 across all observations.

### IMF WORKING PAPERS Predicting the demand for IMF arrangements: A machine learning approach.

### Model Analysis
- Final modeling stage actions:
  - Examine feature importance and contribution of different features to model predictions using Shapley values.
  - Perform two robustness exercises evaluating sensitivity of prediction performance to different feature sets and time periods.

### Feature Importance and Predictors of IMF Arrangements
- Shapley values:
  - Used to allocate the contribution of each feature to model predictions by averaging marginal contributions across all possible feature combinations.
  - Quantify how much each feature contributes to the difference between the actual prediction and the average prediction.
  - Sum of Shapley values for all features of a single observation explains that prediction’s deviation from the average prediction.
  - Do not imply causation.
- Models compared:
  - Traditional linear logistic lasso model.
  - Best performing ML-based algorithm: the random forest.31
- Top predictors and cross-model agreement:
  - Several influential features correspond with economic theory and span real, fiscal, financial, external, and structural sectors.
  - Features influential in both models include:
    - foreign share of general government gross debt
    - PPP income per capita
    - private credit in percent of GDP
  - Features influential in only one model:
    - Random forest: membership in a RFA has the greatest absolute Shapley value, followed by features related to debt and the current account.
    - Logistic lasso: assigns the highest Shapley value to gross international reserve in percent of imports.
  - Additional features among the ten most influential include:
    - trading partner growth
    - the financial inclusion index (access)
    - TED spreads

### Case Studies — Out-of-Sample Performance and Feature Importance
- Method:
  - Two countries selected (one single program, one repeated use) and excluded from the training set to ensure truly out-of-sample results.
- Four key observations:
  - Both models assign a consistently higher probability to Country B, a frequent user of IMF resources, reflecting ongoing balance of payments needs and higher likelihood of seeking IMF support.
  - Models successfully predict program starts, showing an increase in predicted probabilities before program start dates, demonstrating capability to capture use of IMF resources across diverse members.
  - Drivers differ between Country A and Country B:
    - Country A: reserves, trading partner growth, and capital market variables positively affect likelihood of an IMF-supported program; institutional and debt-related indicators reduce likelihood.
    - Country B: bank, credit, money, debt, and current account variables increase probability of IMF-supported programs; fiscal revenue and reserves-related factors are most important stabilizing factors.
  - Probabilities assigned by the logit model fluctuate less and tend to be higher during non-program years compared to probabilities assigned by the random forest.

### Nonlinearity and Predictor Interactions
- ML advantage:
  - Machine learning techniques (e.g., random forest) can handle nonlinear relationships and hidden interactions among predictors without explicit specification.
  - Use of IMF resources is challenging to forecast due to nonlinearities and complex interactions; no theoretical consensus exists on factor combinations that trigger or delay program commencement.
- Evidence from random forest interactions:
  - Shapley value distributions and magnitudes differ across income groups: Advanced Economies (AEs), Emerging Market Economies (EMs), and Low-Income Developing Countries (LIDCs).
  - Identical variable values can have disparate impacts on likelihood of an IMF-supported program depending on income group.
  - Example — reserves to imports:
    - Higher reserves, expressed as a percent of annual imports, tend to decrease probability of seeking IMF support.
    - Reserve levels at which Shapley values turn negative differ among income groups:
      - EMs and LIDCs: reserves below 50 percent of annual imports amplify likelihood of an IMF-supported program.
      - AEs: threshold is at 25 percent.
    - For LIDCs, higher reserve-to-import ratios reduce likelihood of a new program less compared to AEs (steepest) and EMs (flatter Shapley curve).
    - Countries with low reserves to imports and high external public debt ratios are less likely to start a program than countries with low reserves to imports but low public external debt, potentially pointing to binding financial constraints when debt levels are already high.
- Data representation note:
  - Values of variables presented as a percentage of different metrics are expressed in decimals; a value of 0.5 indicates the variable is 50% of the referenced metric.32

### Selected numeric model performance snapshot
- Example predicted probabilities referenced in the content:
  - 0.59 and 0.51, respectively, compared to an overall mean predicted probability of 0.28 across all observations.

*IMF Working Paper: Predicting the demand for IMF arrangements: A machine learning approach.*

### 0.5 of exports, this means reserves are equal to 50% of the export value.

### wpiea2024054-print-pdf - 0.5 of exports, this means reserves are equal to 50% of the export value.

### Reserves to GDP
- Relationship shown between gross reserves in percent of GDP and the FX share of general government gross debt.
- Advanced Economies (AEs) with high reserves and high FX share of government debt tend to have lower probability of entering IMF-supported programs than countries with similar reserves but lower FX share of public debt.
- For low levels of reserves the relationship reverses: a higher FX share of public debt increases likelihood of entering IMF-supported programs.
- For Emerging Markets (EMs) and Low-Income Developing Countries (LIDCs), low reserves significantly increase probability of engaging in a fund-supported program; no clear interaction with FX share of government debt documented.

### Private credit to GDP
- Low levels of private credit are associated with a higher likelihood of program commencement across all groups; effect particularly pronounced for EMs and LIDCs.
- Observed that in EMs, higher private credit tends to coincide with a higher stock of reserves; trend partially extends to LIDCs.
- Interpretation: low private credit often signals underdeveloped private financial systems and lower income levels; higher private credit may indicate more prosperous conditions and reduced need for IMF support.
- Literature link noted: international reserves can provide insurance against sudden capital outflows in presence of high outstanding private external debts (see Lutz and Zessner 2023).

### Public external debt to exports
- High levels of public external debt generally raise probability of seeking IMF financing across all income groups.
- Magnitudes and thresholds where Shapley values turn positive vary by group; most relevant for EMs.
- For EMs: low debt levels associated with most negative Shapley values; higher debt levels associated with most positive Shapley values.
- Once public external debt to exports surpasses a certain small threshold, likelihood of program commencement abruptly increases and remains elevated thereafter.
- AEs with substantial public external debt tend to have low FX share of government debt; EMs and LIDCs with high public external debt often have high FX share of public external debt.

### FX share of general government debt
- Positive relationship between use of IMF financing and FX share of public debt.
- Inverted U-shaped relationship for EMs and LIDCs.
- Within EMs, higher per-capita income tends to coincide with low FX share of debt and lower likelihood of seeking IMF support.
- Most LIDCs exhibit comparably high FX shares of public debt irrespective of income.

### Public debt to revenue
- Higher public debt-to-revenue increases likelihood of using IMF financing across all income groups.
- Notable thresholds where Shapley values escalate sharply:
  - around 150% for AEs
  - 20% for both EMs and LIDCs
- Beyond thresholds, EMs and LIDCs exhibit a more pronounced response than AEs (higher absolute magnitude of Shapley values).
- Observed: AEs with low public debt-to-revenue tend to have high FX share of general government debt; FX exposure elevated across all EMs.

### PPP income per capita, relative to the US
- Within each income group, low PPP income per capita (relative to the US) associated with higher likelihood of fund-supported program commencements.
- FX share of government debt highly correlated with income: lower income levels associated with higher FX share, even within income groups.
- For LIDCs, less variability in income per capita and FX share of government debt across country-year observations.

### General government revenue to GDP
- Higher government revenue to GDP associated with lower likelihood of IMF financing requests across all income groups.
- Relationship particularly pronounced in EMs, with Shapley values ranging from approximately -0.13 to 0.25.
- AEs with high government revenues and high external public debt tend to have lower probability of requesting IMF financing than countries with lower public external debt—likely due to enhanced access to global financial markets.
- For EMs and LIDCs, at equivalent levels of government revenue to GDP, a higher ratio of external public debt to GDP is associated with decreased likelihood of an IMF-supported program.

### Current account to GDP
- In AEs, likelihood of observing a fund-supported program does not significantly increase until current account deficits fall below 5% of GDP (Shapley values remain negative above this point).
- In EMs and LIDCs, even small current account deficits are linked to higher probability of program commencement.
- Maximum share of government debt in foreign currency:
  - LIDCs: nearing 100%
  - EMs: around 90%
  - AEs: just over 50%
- AEs with current account deficits and higher FX share of government debt show stronger likelihood of observing a fund-supported program.
- Evidence found for a reversed relationship for countries with current account surpluses.

### Membership in an RFA (Regional Financial Arrangement)
- Membership in an RFA reduces likelihood of observing IMF-supported programs across all income groups.
- Shapley value ranges:
  - AEs: -0.11 to -0.03
  - EMs: -0.35 to -0.04
  - LIDCs: -0.35 to -0.1
- For EMs and LIDCs, RFA membership mitigates use of IMF financing somewhat less if level of international reserves is high.

### Robustness analysis — Evaluating the impact of adding additional variables
- Best-performing variable sets maximize ROC-AUC scores during cross-validation; model-selected sets (set 4-6) outperform pre-defined sets (sets 1-3).
- Model-selected sets exclude several variables identified as key drivers in earlier literature.
- Additional 21 variables added individually and ex post to best-performing sets for RF, RNN, and logistic lasso:
  - reserves in percent of the ARA metric
  - bank capital adequacy ratios
  - net non-FDI liability inflows
  - access to a central bank swap line (dummy)
  - federal funds rate
  - 10-year treasury yields
  - terms of trade
  - debt service to revenue
  - foreign liabilities to domestic credit
  - short term deposit rates
  - bank external liabilities as a percent of GDP
  - food price inflation
  - oil price inflation
  - international risk aversion index
  - fiscal interest expenses to revenue
  - consumer prices
  - country classification dummies (PRGT, AE, non-LIDC EMDE, LIDC)
  - IMF credit outstanding
- Key robustness findings:
  - Non-linear models (e.g., random forest) maintain prediction performance; linear models (logistic lasso) show notable variability.
  - Largest increase in logistic lasso test ROC-AUC when IMF credit outstanding is added; RF reports a decline with this addition.
  - Fiscal interest expenses to revenue increases performance for logistic lasso and RNN (already included in RF-selected set).
  - Biggest RNN improvement when dummy for advanced economies is added; RF and logistic lasso ROC-AUC scores decline marginally.
  - Non-linear ensemble models like RF appear more robust to omitted important variables than simpler linear models.
- Recommendation for operational purposes: include credit outstanding in prediction exercise to account for debt rollover and repeated use of fund resources.

### Evaluating model sensitivity to training data variability and performance over time
- Time series split method for cross-validation: training set increased progressively by seven years per iteration; next seven years assigned as out-of-sample validation.
- Five training and out-of-sample test periods used:
  1. training 1982-1988; test 1989-2021
  2. training 1982-1995; test 1996-2021
  3. training 1982-2002; test 2003-2021
  4. training 1982-2009; test 2010-2021
  5. training 1982-2017; test 2018-2021
- Observations:
  - Higher ROC-AUC scores for models trained on more recent data.
  - Differences in prediction performance across folds most visible in years immediately following training set updates — models best at near-term predictions.
  - Average out-of-sample ROC-AUC scores tend to decline over time, potentially due to structural breaks and non-stationarities (e.g., changes in program supply or demand, rise of alternative funding sources).
  - Noticeable dips in performance during global financial distress episodes (Asian financial crisis, global financial crisis, European debt crisis, Covid-19 pandemic); decline in prediction performance observed approximately two years prior to crisis onset, consistent with dependent variable definition (equal to one in periods t-2 and t-1).
  - Models failed to predict surge in IMF-supported programs following Global Financial Crisis, particularly among AEs.
  - Decline in performance modest during Covid-19 pandemic, likely driven by predominant use of emergency financing instruments (treated as no-program cases).
  - RF outperforms competitors in near-term predictions in each fold but tends to deteriorate more rapidly as training data becomes outdated; deterioration generally apparent when training data is outdated by at least 8-10 years.

### Evaluating the size of IMF arrangements
- Approved IMF-supported arrangement sizes in sample range from 14.5 to 3211.8 percent of quota.
- Two-step procedure proposed: use binary prediction results from random forest to assess implications for IMF resources; then regress log approved size of IMF arrangements on lagged explanatory variables.
- Sample for size analysis:
  - period 1982-2021
  - training set 1982-2017; hold-out test set 2018-2021
  - only drawn GRA arrangements included for illustrative section; results extensible to full set of IMF-supported arrangements (including PRGT and RST accounts).
  - 585 data points in training set after dropping no-arrangement observations.
- Explanatory variables include the 40 variables selected by RF recursive feature elimination (set 6); robustness check adds literature-identified additional variables.
- Model performance evaluated via out-of-sample root mean-squared error.
- Findings from Figure 17 summary:
  - Model using set 6 underestimates size of approved arrangements in every year (2018-2021 testing sample).
  - Model using set 6 plus additional variables predicts larger use of IMF resources and overestimates sum of approved arrangements in some years (2019 and 2020).
  - Inclusion of credit outstanding to the Fund improves ability to predict arrangements with large access levels significantly.
  - Best performing model (out-of-sample root MSE and adjusted r-square) uses RF-40 Plus variable set and includes observations from 1990-2017.
- Regression observations (cautious interpretation due to potential misspecification):
  - Only a small set of predictors from set 6 significant at 10 percent: dollar appreciation, FX share of public debt, private credit, PPP income.
  - Several additional variables in set 6 plus significant: total fund credit outstanding, reserves in percent of the ARA metric, swap line access, the federal funds rate, the VIX index.
- Limitations noted:
  - Non-linearities important given wide range of approved access levels and presence of outliers.
  - Potential issues with model misspecification, multicollinearity, omitted variable biases.
  - ML-based algorithms promising for predicting size of IMF arrangements; further research recommended.

*IMF WORKING PAPERS Predicting the demand for IMF arrangements: A machine learning approach. INTERNATIONAL MONETARY FUND*

### 6. Conclusion

### 6. Conclusion

### Main findings on predictability and model performance
- The study uses a large data set spanning "thirty-five years and nearly all IMF member countries" and a broad set of ML-based algorithms to analyze predictability of countries’ demand for IMF arrangements.
- ML models consistently outperform econometric methods in out-of-sample prediction of new IMF arrangements.
- Ensemble models and recurrent neural networks are identified as among the most successful classifiers.
- There is considerable agreement across algorithms in the set of selected predictors, and prediction performances of ML-based models are largely robust to alternative feature sets.
- ML-based techniques are appealing because of their ability to:
  - recognize nonlinearities, and
  - capture sequential and temporal patterns in the data.

### Influential predictors and modeling insights
- Important predictors include variables related to multiple sectors:
  - external, fiscal, real, and financial sectors, and
  - institutional factors such as membership in a regional financing arrangement and existing fund credit outstanding.
- Data processing decisions (imputation and sampling methods) can have significant implications for model performance.
- Different models can perform best under different preprocessing and sampling settings, motivating a flexible, algorithm-tailored approach.

### Limitations and data challenges
- Noticeable dips in model performances occur during periods of global financial distress, indicating such events remain inherently difficult to predict.
- Prediction performance depends on reliable and near real-time data; publication lags and noise in national accounts present challenges, especially in vulnerable economies that are often most reliant on IMF support.
- The models primarily rely on economic fundamentals and institutional factors and may miss non-economic influences that could deter countries from requesting IMF-supported programs, such as changes in risk perception or stigma.
- Natural language processing techniques are suggested as one way to help address data challenges and publication lags.

### Scope and avenues for future research
- The paper concentrates on classification models and provides only a brief discussion of the overall size of approved arrangements.
- The use of ML-based algorithms to predict:
  - the size of approved arrangements, and
  - repeated use of IMF-supported arrangements,
  is identified as a promising area for future research.

*Source: wpiea2024054-print-pdf — 6. Conclusion*

### Annex IV. Predicting the Size of IMF

### Annex IV. Predicting the Size of IMF Arrangements

### Key predictors and estimated coefficients (four specifications)
- EMDE Dummy:
  - 0.2 (1982-2017 RF)
  - 0.236 (1990-2017 RF)
  - 0.167 (1982-2017 RF PLUS)
  - 0.247 (1990-2017 RF PLUS)
  - Standard errors: (0.2), (0.327), (0.241), (0.326)
- Dollar Appreciation (Real Effective Exchange Rate, based on CPI):
  - -0.1** (1982-2017 RF) — (0.0)
  - -0.335*** (1990-2017 RF) — (0.123)
  - -0.152** (1982-2017 RF PLUS) — (0.0701)
  - -0.448*** (1990-2017 RF PLUS) — (0.145)
- Membership in RFA:
  - 0.1; 0.179; 0.0356; 0.0748
  - Standard errors: (0.1), (0.144), (0.102), (0.148)
- (Ideal Point Estimate US – ideal point estimate country x)*(-1):
  - 0.1; -0.0583; 0.00565; -0.0878
  - Standard errors: (0.1), (0.147), (0.115), (0.147)
- (Ideal Point Estimate US – ideal point estimate country x)*(-1), lag:
  - -0.2; 0.0844; -0.0392; 0.133
  - Standard errors: (0.1), (0.152), (0.114), (0.155)
- Current Account, in USD billion, percent of GDP (in USD):
  - -0.1; -0.0874; -0.0956; -0.0404
  - Standard errors: (0.1), (0.0925), (0.0679), (0.0929)
- Current Account, in USD billion, percent of GDP (in USD), lag:
  - 0.0; 0.00712; 0.0624; 0.0213
  - Standard errors: (0.1), (0.0750), (0.0596), (0.0751)
- Exchange Rate, national currency units per U.S. dollar, period average, 1-year growth rate:
  - -0.1; 0.0353; -0.0524; 0.00930
  - Standard errors: (0.1), (0.248), (0.0749), (0.240)
- External debt, in USD billion, percent of exports:
  - 0.2; 0.157; 0.0984; 0.0539
  - Standard errors: (0.1), (0.141), (0.102), (0.140)
- External debt, in USD billion, percent of exports, lag:
  - 0.1; 0.0680; 0.0410; 0.0488
  - Standard errors: (0.1), (0.101), (0.0740), (0.0996)
- Gross Reserves, in USD billions, percent of GDP (USD):
  - 0.1; 0.00719; 0.134; 0.00881
  - Standard errors: (0.1), (0.175), (0.139), (0.179)
- Gross Reserves, in USD billions, percent of imports:
  - 0.0; 0.0892; 0.0399; 0.0942
  - Standard errors: (0.1), (0.120), (0.0922), (0.121)
- FX share of general government debt:
  - -0.1*; -0.163*; -0.132*; -0.146
  - Standard errors: (0.1), (0.0925), (0.0707), (0.0921)
- FX share of general government debt, lag:
  - -0.1; -0.141**; -0.0629; -0.108
  - Standard errors: (0.1), (0.0702), (0.0564), (0.0686)
- Hard Peg Indicator (=1 if fine exchange rate regime =1 or =2; =0 otherwise):
  - 0.1; 0.00497; 0.0596; -0.00596
  - Standard errors: (0.1), (0.108), (0.0752), (0.107)
- Capital Outflow restrictions:
  - 0.0; 0.0988; 0.0208; 0.0860
  - Standard errors: (0.0), (0.0636), (0.0461), (0.0630)
- Capital Inflows restrictions, lag:
  - 0.1; 0.0365; 0.0672; 0.0389
  - Standard errors: (0.0), (0.0674), (0.0487), (0.0661)
- Trading Partner Growth:
  - -0.0; 0.0298; -0.00642; -0.00323
  - Standard errors: (0.0), (0.0520), (0.0386), (0.0576)
- Overall balance, in USD billion, percent of GDP (in USD):
  - -0.0; -0.00118; -0.00798; -0.00364
  - Standard errors: (0.0), (0.0190), (0.0165), (0.0191)
- General government revenue, in national currency, percent of GDP:
  - -0.0; 0.0404; -0.0508; -0.0362
  - Standard errors: (0.1), (0.139), (0.104), (0.145)
- General government revenue, in national currency, percent of GDP, lag:
  - -0.1; -0.0903; -0.0162; 0.0115
  - Standard errors: (0.1), (0.0858), (0.0764), (0.0977)
- Fiscal interest expenses / Fiscal revenue (%):
  - 0.1*; 0.0690; 0.00452; -0.00672
  - Standard errors: (0.0), (0.0507), (0.0511), (0.0714)
- Public debt to revenue:
  - 0.1; 0.104; 0.0646; 0.0414
  - Standard errors: (0.1), (0.200), (0.0599), (0.205)
- Public debt to revenue, lag:
  - 0.0; 0.00458; 0.0191; 0.0187
  - Standard errors: (0.0), (0.0626), (0.0237), (0.0646)
- Public external debt, in USD billion, percent of exports, lag:
  - -0.1; -0.0467; -0.0791*; -0.0624
  - Standard errors: (0.0), (0.0586), (0.0423), (0.0589)
- External debt, in USD billion, percent of exports:
  - -0.1; -0.181; -0.0705; -0.0337
  - Standard errors: (0.1), (0.192), (0.134), (0.192)
- External debt, in USD billion, percent of exports, lag:
  - -0.1; -0.00195; -0.0146; 0.0280
  - Standard errors: (0.1), (0.152), (0.112), (0.151)
- Loan to deposit ratio:
  - 0.1; 0.108*; 0.0733*; 0.0959*
  - Standard errors: (0.0), (0.0577), (0.0433), (0.0570)
- Financial Inclusion - Access:
  - 0.0; 0.00824; 0.0230; -0.0553
  - Standard errors: (0.2), (0.295), (0.239), (0.291)
- Financial Inclusion – Access, lag:
  - 0.3; 0.425; 0.250; 0.371
  - Standard errors: (0.2), (0.298), (0.244), (0.295)
- Total credit, in USD billion, percent of GDP (in USD):
  - 0.1; 0.0616; 0.0773; 0.0752
  - Standard errors: (0.1), (0.0784), (0.0722), (0.0855)
- Private credit, in USD billion, 5-year growth ratee:
  - 0.0*; 0.0316*; 0.0322**; 0.0301*
  - Standard errors: (0.0), (0.0174), (0.0154), (0.0170)
- Private credit, in USD billion, percent of GDP (in USD):
  - -0.0; -0.179; -0.0979; -0.224
  - Standard errors: (0.2), (0.252), (0.155), (0.249)
- Private credit, in USD billion, percent of GDP (in USD), lag:
  - 0.1; 0.284; 0.0975; 0.172
  - Standard errors: (0.1), (0.234), (0.148), (0.233)
- PPP income per capita, relative to the U.S.:
  - 1.0***; 0.780*; 0.870**; 0.637
  - Standard errors: (0.4), (0.407), (0.355), (0.402)
- PPP income per capita, relative to the U.S., lag:
  - -0.9**; -0.726*; -0.864**; -0.600
  - Standard errors: (0.4), (0.422), (0.356), (0.420)
- Consumer Prices, period average, 1-year growth rate:
  - 0.0; -0.0904; 0.0415; -0.0387
  - Standard errors: (0.1), (0.144), (0.0623), (0.143)
- Bureaucracy quality (ICRG):
  - -0.0; 0.0293; 0.0181; 0.0820
  - Standard errors: (0.0), (0.0730), (0.0491), (0.0718)
- Institutional Quality (ICRG):
  - -0.1; 0.0285; -0.0459; 0.00689
  - Standard errors: (0.0), (0.0643), (0.0430), (0.0630)
- Dollar Appreciation (Real Effective Exchange Rate, based on CPI, Index), 5-year growth rate, lag:
  - -0.0; 0.118; -0.0127; 0.0990
  - Standard errors: (0.1), (0.116), (0.0576), (0.127)
- Remittances, in USD billion, 1-year growth rate:
  - -0.0; -0.0200; -0.0131; -0.0430
  - Standard errors: (0.1), (0.0563), (0.0500), (0.0555)

### RF PLUS specific additional predictors (RF PLUS specifications)
- Total credit outstanding in percent of quota:
  - 0.0942*** (1982-2017 RF PLUS)
  - 0.0934** (1990-2017 RF PLUS)
  - Standard errors: (0.0305), (0.0382)
- Reserves as percent of the ARA metric:
  - -0.100* (1982-2017 RF PLUS)
  - -0.0524 (1990-2017 RF PLUS)
  - Standard errors: (0.0514), (0.0640)
- Capital Adequacy Ratio:
  - 0.00872 (1982-2017 RF PLUS)
  - -0.0754 (1990-2017 RF PLUS)
  - Standard errors: (0.0435), (0.0593)
- Access to swap line:
  - 1.066*** (1982-2017 RF PLUS)
  - 0.707** (1990-2017 RF PLUS)
  - Standard errors: (0.240), (0.286)
- Net non-FDI liability inflows, lag:
  - 0.160* (1982-2017 RF PLUS)
  - 0.286** (1990-2017 RF PLUS)
  - Standard errors: (0.0947), (0.138)
- Federal Funds Rate, Percent, Monthly, Not Seasonally Adjusted:
  - 0.247** (1982-2017 RF PLUS)
  - 0.469*** (1990-2017 RF PLUS)
  - Standard errors: (0.102), (0.144)
- Market Yield on U.S. Treasury Securities at 10-Year Constant Maturity:
  - -0.188 (1982-2017 RF PLUS)
  - -0.558*** (1990-2017 RF PLUS)
  - Standard errors: (0.117), (0.180)
- Terms of Trade:
  - -0.0270; -0.0587
  - Standard errors: (0.0595), (0.0807)
- Public debt service to revenue:
  - 0.0584; 0.110
  - Standard errors: (0.0727), (0.102)
- Foreign liability to domestic credit, lag:
  - 0.106; 0.0899
  - Standard errors: (0.0809), (0.100)
- Short-term deposit rate:
  - -0.0158; -0.0299
  - Standard errors: (0.0220), (0.0436)
- External bank liabilities, in USD, percent of GDP (in USD):
  - 0.627* (1982-2017 RF PLUS)
  - 0.599 (1990-2017 RF PLUS)
  - Standard errors: (0.320), (0.460)
- Food price inflation:
  - -0.0169; 0.00263
  - Standard errors: (0.0402), (0.0505)
- Oil price inflation:
  - -0.0322; -0.0639
  - Standard errors: (0.0406), (0.0619)
- CBOE Volatility Index: VOX, Index, Daily, Not Seasonally Adjusted, lag:
  - 0.0369; 0.0756
  - Standard errors: (0.0401), (0.0535)

### Constants, sample, and model fit statistics
- Constant:
  - 4.6*** (1982-2017 RF) — (0.2)
  - 4.538*** (1990-2017 RF) — (0.284)
  - 4.569*** (1982-2017 RF PLUS) — (0.224)
  - 4.450*** (1990-2017 RF PLUS) — (0.287)
- Observations:
  - 585 (1982-2017 RF and RF PLUS)
  - 388 (1990-2017 RF and RF PLUS)
- R-squared:
  - 0.3 (1982-2017 RF)
  - 0.401 (1990-2017 RF)
  - 0.372 (1982-2017 RF PLUS)
  - 0.468 (1990-2017 RF PLUS)
- Adjusted R-squared:
  - .25; .33; .30; .38
- Root MSE:
  - .76; .81; .73; .78
- Root MSE OOS:
  - 1.01; 0.96; 1.04; 0.90
- Sample Period:
  - 1982 - 2017 (RF and RF PLUS)
  - 1990 - 2017 (RF and RF PLUS)
- Standard errors presented in parentheses
- Significance codes:
  - *** p<0.01, ** p<0.05, * p<0.1

*IMF WORKING PAPERS: Predicting the demand for IMF arrangements: A machine learning approach. Working Paper No. WP/2024/054*

---


_Source: https://www.imf.org/-/media/files/publications/wp/2024/english/wpiea2024054-print-pdf.pdf_
