## _wp0452

## Source details

**Canonical URL:** [_wp0452](https://www.imf.org/-/media/websites/imf/imported-full-text-pdf/external/pubs/ft/wp/2004/_wp0452.pdf)

## Other formats

- [Markdown version](/-/media/websites/imf/imported-full-text-pdf/external/pubs/ft/wp/2004/_wp0452.pdf.md)
- [Structured JSON version](/-/media/websites/imf/imported-full-text-pdf/external/pubs/ft/wp/2004/_wp0452.pdf.json)

---

### I. Introduction — purpose and scope
- Research focus: performance of early-warning-system (EWS) models of currency crisis, with emphasis on out-of-sample (real-time) prediction.
- Time horizon of interest: forecasts made since 1999; reexamination of run-up to the Asia crisis (1997/1998).
- Comparative objective: contrast EWS model forecasts with non-model-based indicators (bond spreads, agency ratings, Economist Intelligence Unit risk scores).
- Key prior references cited in text: Berg and others (1999); Kaminsky, Lizondo, and Reinhart (1998); IMF (2002); Berg and Pattillo (1999a, 1999b); Mulder, Perrelli, and Rocha (2002); Abiad (2003); Manasse, Roubini, and Schimmelpfennig (2003).

### II. Models tracked and methodological features
- Models tracked by the IMF staff since 1999:
  - DCSD (Developing Country Studies Division)
  - KLR (Kaminsky, Lizondo, and Reinhart)
  - GS—WATCH (Goldman Sachs)
  - EMRI / CSFB (Credit Suisse First Boston)
  - DB Alarm Clock (Deutsche Bank)
- Methods and implementation highlights:
  - DCSD: Probit regression with rhs variables measured in (country specific) percentile terms; main horizon 24 months.
  - KLR: Indicator (0/1) thresholds combined as a weighted average of indicators; horizon 24 months.
  - GS—WATCH: Logit regression with rhs variables as 0/1 indicators; horizon 3 months.
  - EMRI/CSFB: Logit regression with logged and standardized rhs variables; horizon 1 month.
  - DBAC: Logit two-equation simultaneous system on exchange-rate and interest-rate “events”; horizon 1 month.
- Variable taxonomy (terminology preserved): Overvaluation, Current account, Reserve Losses, Export Growth, ST Debt/Reserves, Reserves/M2 (level), Reserves/M2 (growth), Domestic Credit Growth, Financing Requirement, Debt/Exports, Growth of credit to private sector, Money Multiplier Change, Real Interest Rate, Excess M1 Balances, Reserves/Imports (level), Oil prices, Industrial Production, Stock Market, Stock price growth, GDP growth, Political Event, Global Liquidity Contagion, Regional Contagion, Devaluation Contagion, Market Pressure Contagion, Regional Dummies, Interest Rate “event”.
- Data shortcoming noted: omission of market-determined interest rates in many EWS definitions due to historical data deficiencies; DBAC attempts to incorporate interest rate events using IFS money market rates.
- Forecasting horizons: in-house models adopt longer horizons to allow time for policy changes (24 months); private-sector models use shorter horizons (1–3 months) and sometimes different evaluation criteria (e.g., trading-rule performance).

### III. Crisis definitions — differences across models (preserved text)
- DCSD / KLR crisis definition:
  - "Weighted average of one-month changes in exchange rate and reserves more than 3 (country-specific) standard deviations above country average"
  - Horizon: 2 years
- GS—WATCH crisis definition:
  - "Weighted average of three-month changes in exchange rate and reserves above country-specific threshold"
  - Horizon: 3 months
- EMRI / CSFB crisis definition:
  - "Depreciation > 5% and at least double of preceding month’s"
  - Horizon: 1 month
- DB Alarm Clock crisis definition:
  - "Depreciation > 10% and Interest rate increase > 25% typical" (various “trigger points”; separates exchange rate and interest rate events but jointly estimates probabilities)
  - Horizon: 1 month

### IV. Crisis-dating, incidence, and model behavior — key empirical findings
- Crisis counts (common sample of 16 countries, 1996 through mid-2001):
  - CSFB crisis definition produces 34 crises
  - DCSD/KLR definition produces 34 crises
  - GS definition produces 150 crises, grouped into 47 distinct episodes
- Character of crises by definition:
  - DCSD/KLR crises are "almost always isolated"; GS crises are more frequent and often consecutive.
- Practical caveat: simple formulae may miss "close calls" or country-specific events (example: Sri Lanka in 2000–January 2001 — reserve loss of about 40 percent and currency depreciation of nearly 15 percent registered as a close call by KLR/DCSD).

### V. Value added of EWS models — evaluation approach and major comparative findings
- Evaluation principle: focus on out-of-sample (real-time) performance rather than in-sample fit.
- Benchmarks used for comparison:
  - Bond spreads on dollar-denominated sovereign debt
  - Sovereign credit ratings
  - Economist Intelligence Unit (EIU) country expert assessments of currency crisis risk
- Historical Asia-crisis result:
  - KLR implemented as it might have been in early 1997 "decisively outperformed all the comparators" for that period (Berg and others (1999)).
- Comparative numeric summaries (Table 2 averages, preserved):
  - Crisis Countries: KLR 20, DCSD 34, Spread 90, Rating 34, EIU Currency Risk 35
  - Non-Crisis Countries: KLR 15, DCSD 17, Spread 201, Rating 57, EIU Currency Risk 46
- Rank Correlation (indicator ranking vs crisis severity, preserved):
  - KLR: 0.52
  - DCSD: 0.53
  - Spread: -0.31
  - Rating: -0.49
  - EIU: -0.33
- Comparative performance since January 1999 and July 1999 forecasts (selected points):
  - In July 1999 forecasts: both countries with DCSD probabilities above 50 percent subsequently had crises; no crisis country had a DCSD probability below 26 percent.
  - Using the models’ own definition, there have been only three crises in the roughly two years since July 1999 (as reported in the text).

### VI. Goodness-of-fit measures, cut-off choice, and statistical evaluation
- Alarm conversion: predicted probability above a cut-off threshold signals a crisis within the model horizon (24 months for KLR/DCSD).
- Loss function used: sum of false alarms as a share of total tranquil periods and missed crises as a share of total pre-crisis periods; equal weight placed on false alarms and missed crises.
- ROC-style/Cut-off independent insights:
  - Cutoff = 0 produces 100 percent crises correctly called and 100 percent false alarms.
  - Cutoff = 100 produces 0 percent crises correctly called and 0 percent false alarms.
  - In-sample: DCSD dominates KLR (DCSD curve lies to the right and below KLR curve).
- Regression evaluation (regression: c24_it = α + β PredProb_it + ε_it; OLS with HAC standard errors):
  - In-sample: estimated β is always statistically different from 0 and a value of 1 cannot be rejected for both KLR and DCSD.
  - Out-of-sample:
    - KLR forecasts are highly significant; hypothesis that β = 1 cannot be rejected.
    - DCSD forecasts not significant at traditional levels: p-value for β = 0 is 12 percent; p-value for β = 1 is 31 percent.
- Cut-off guidance from DCSD out-of-sample behavior:
  - A much higher cutoff (around 50 percent) would have reduced false alarms without substantially increasing missed crises; out-of-sample curve approaches the in-sample curve near a cutoff probability of 47 percent.

### VII. Out-of-sample performance summary — model-level results (preserved figures and statistics)
- Forecast periods evaluated:
  - DCSD and KLR: January 1999 to December 2000 (24-month horizon).
  - GS: January 1999 to April 2001.
  - CSFB: April 2000 to June 2001 (out-of-sample assessed through August 2001 for CSFB).
- KLR model (out-of-sample):
  - Calls 58 percent of pre-crisis months correctly.
  - When the crisis probability was below the cut-off, a crisis ensued 9 percent of the time.
  - When the crisis probability was above the cut-off, a crisis ensued 35 percent of the time.
  - KLR performed better out-of-sample than in-sample and the forecasts were highly informative.
- DCSD model (out-of-sample):
  - Calls 31 percent of pre-crisis months correctly.
  - Crises followed above-cut-off signals 22 percent of the time and followed below-cut-off signals 14 percent of the time.
  - DCSD performance deteriorated substantially out-of-sample relative to in-sample and produced more false alarms out-of-sample for any cut-off level examined.
  - Out-of-sample period contains 528 observations but only eight crises (limiting statistical power).
- Short-horizon private-sector models (GS and CSFB) — Table 5 key exact figures preserved:
  - Sample periods:
    - In Sample: GS 1996:1 to 1998:12; CSFB 1994:1 to 2000:7.
    - Out of Sample: GS 1999:1 to 2001:4; CSFB 2000:8 to 2001:8.
  - Cut-off values used in Table 5 headings: 10 (GS), 35 (CSFB).
  - Value of loss function 2/: In Sample: 66 (GS), 58 (CSFB). Out of Sample: 97 (GS), 88 (CSFB).
  - Percent of observations correctly called: In Sample: 66 (GS), 76 (CSFB). Out of Sample: 50 (GS), 83 (CSFB).
  - Percent of crisis correctly called 3/: In Sample: 66 (GS), 65 (CSFB). Out of Sample: 54 (GS), 27 (CSFB).
  - Percent of tranquil correctly called 4/: In Sample: 66 (GS), 76 (CSFB). Out of Sample: 50 (GS), 85 (CSFB).
  - False alarms as percent of total alarms 5/: In Sample: 74 (GS), 92 (CSFB). Out of Sample: 87 (GS), 96 (CSFB).
  - Probability of crisis given signal 6/: In Sample: 26 (GS), 8 (CSFB). Out of Sample: 14 (GS), 4 (CSFB).
  - Statistical regression coefficients (actual on predicted, HAC standard errors):
    - Coefficient: In Sample: 1.41 (GS), 0.17 (CSFB). Out of Sample: 0.56 (GS), 0.06 (CSFB).
    - Standard error: In Sample: 0.42 (GS), 0.03 (CSFB). Out of Sample: 0.41 (GS), 0.06 (CSFB).
    - p-value (coefficient <= 0): In Sample: 0.00 (GS), 0.00 (CSFB). Out of Sample: 0.17 (GS), 0.29 (CSFB).
- Interpretation: short-horizon private-sector models performed poorly out of sample despite strong in-sample performance; predictions were largely uninformative out-of-sample.

### VIII. Statistical power, false alarms, and interpretation caveats
- Limited number of crises and serial correlation reduce effective information and increase standard errors; small changes in sample can materially change goodness-of-fit indicators.
- Simulation for DCSD: simulating the implied data-generating process 500 times showed that for ideal forecasts the hypothesis that β = 0 would not be rejected 28 percent of the time — illustrating lack of power.
- False alarms may still be valuable:
  - Crisis definitions may miss events an analyst would want a model to warn about.
  - Warnings may lead to policy adjustments or luck that avoid a crisis; a warning could therefore be useful even if classified mechanically as a false alarm.
  - Empirical note: of 90 DCSD false-alarm observations for 1999:1 to 2000:12, half (46) can be classified as either missed crisis classifications or useful warnings (countries include Argentina, Chile (before July 1999), Pakistan, Venezuela, Turkey, Uruguay).

### IX. Structural context, limitations, and research directions
- Emerging trends affect currency-crisis prediction:
  - Resurgence of crises where sovereign and domestic debt dynamics play a central role; debt and currency crises are related but distinct.
  - Increased prevalence of de jure floating regimes with capital account openness and substantial de facto flexibility (examples listed in source: Brazil, Chile, Colombia, Korea, Mexico, Poland, Thailand, and historically South Africa).
- Implications:
  - Floating regimes do not eliminate sharp depreciations or crises; no firm evidence that floating rates have been more resistant.
  - Research direction: recent work has turned to predicting debt crises and fiscal vulnerabilities.
- Specification and implementation cautions:
  - Predictand ambiguity (exchange rate, reserves, interest rate events) complicates model design.
  - Danger of overfitting through data-mining; emphasis on parsimonious models capturing robust symptoms across episodes.
  - Data availability and frequency constraints for financial-sector and political variables limit inclusion of potentially important predictors.

### X. Policy-relevant conclusions and recommended role of EWS models
- Strengths of EWS models:
  - Objective, systematic processing of data.
  - Potential to identify vulnerabilities that other analyses miss and to avoid biases rooted in past experience.
  - Historical example: models identified vulnerability signals (example: Korea in 1996–97) not reflected in market/analyst signals prior to crises.
- Limitations and practical use:
  - EWS models are not accurate enough to be sole methods for anticipating crises.
  - Models should be used as one component in broader surveillance and vulnerability assessment frameworks rather than as standalone predictors.
- Expectation: painful currency crises are likely to remain a feature of emerging markets for the foreseeable future.

*Source: _wp0452 - References (author’s calculations and analysis in the provided content unit).*

### References .............................................................................................................

### _wp0452 - References .............................................................................................................

### I. Introduction — purpose and scope
- Research focus: performance of early-warning-system (EWS) models of currency crisis, with emphasis on out-of-sample (real-time) prediction.
- Time horizon of interest: forecasts made since 1999; reexamination of run-up to the Asia crisis (1997/1998).
- Comparative objective: contrast EWS model forecasts with non-model-based indicators (bond spreads, agency ratings, Economist Intelligence Unit risk scores).
- Key prior references cited in text: Berg and others (1999); Kaminsky, Lizondo, and Reinhart (1998); IMF (2002); Berg and Pattillo (1999a, 1999b); Mulder, Perrelli, and Rocha (2002); Abiad (2003); Manasse, Roubini, and Schimmelpfennig (2003).

### II. Early-Warning-System (EWS) models — characterization and implementation
- Models tracked by the IMF staff since 1999:
  - DCSD (Developing Country Studies Division)
  - KLR (Kaminsky, Lizondo, and Reinhart)
  - GS—WATCH (Goldman Sachs)
  - EMRI / CSFB (Credit Suisse First Boston)
  - DB Alarm Clock (Deutsche Bank)
- Table 1 documents: crisis definition, prediction horizon, method, and predictor variables for each model.
- Appendix I contains fuller model descriptions; Table 6 (Appendix I) lists all crisis dates for these models since January 1999.
- Specification choices noted as empirical and judgmental; predictive variables inspired by balance-of-payments crisis theory but constrained by data availability.
- The in-house IMF models adopt longer horizons to allow time for policy changes; private sector models use shorter horizons and sometimes different evaluation criteria (e.g., trading-rule performance).

### III. Crisis definitions — differences across models (Box 1 and Table 1 excerpts)
- DCSD / KLR crisis definition:
  - "Weighted average of one-month changes in exchange rate and reserves more than 3 (country-specific) standard deviations above country average"
  - Horizon: 2 years
- GS—WATCH crisis definition:
  - "Weighted average of three-month changes in exchange rate and reserves above country-specific threshold"
  - Horizon: 3 months
- EMRI / CSFB crisis definition:
  - "Depreciation > 5% and at least double of preceding month’s"
  - Horizon: 1 month
- DB Alarm Clock crisis definition:
  - "Depreciation > 10% and Interest rate increase > 25% typical" (various “trigger points”; separates exchange rate and interest rate events but jointly estimates probabilities)
  - Horizon: 1 month
- Methods summary:
  - DCSD: Probit regression with rhs variables measured in (country specific) percentile terms
  - KLR: Weighted (by frequency of correct predictions) average of indicators; variables measured as 0/1 indicators according to thresholds chosen to minimize noise/signal ratio
  - GS: Logit regression with (most) rhs variables measured as 0/1 indicators based on thresholds found in autoregression with dummy (SETAR)
  - EMRI/CSFB: Logit regression with rhs variables measured in logs, then deviation from mean and standardized
  - DBAC: Logit two-equation simultaneous systems on ex rate and i rate “events”
- Variable examples included across models (as listed in Table 1), preserving terminology exactly:
  - Overvaluation, Current account, Reserve Losses, Export Growth, ST Debt/Reserves, Reserves/M2 (level), Reserves/M2 (growth), Domestic Credit Growth, Financing Requirement, Debt/Exports, Growth of credit to private sector, Money Multiplier Change, Real Interest Rate, Excess M1 Balances, Reserves/Imports (level), Oil prices, Industrial Production, Stock Market, Stock price growth, GDP growth, Political Event, Global Liquidity Contagion, Regional Contagion, Devaluation Contagion, Market Pressure Contagion, Regional Dummies, Interest Rate “event”
- Noted data shortcoming: omission of market-determined interest rates in many EWS definitions due to historical data deficiencies; DBAC attempts to incorporate interest rate events using IFS money market rates.

### IV. Empirical features and model behavior
- Crisis-dating and incidence differences:
  - Across a common sample of 16 countries from 1996 through mid-2001:
    - CSFB crisis definition produces 34 crises
    - DCSD/KLR definition produces 34 crises
    - GS definition produces 150 crises, grouped into 47 distinct episodes of one or more consecutive crisis months
  - DCSD/KLR crises are "almost always isolated"; GS crises are more frequent and often consecutive.
- Practical caveat: simple formulae may miss "close calls" or country-specific events; analysts may adjust dating based on country knowledge (example: Sri Lanka in 2000–January 2001 — reserve loss of about 40 percent and currency depreciation of nearly 15 percent that the KLR/DCSD formula registered as a close call but not a crisis).
- Models’ intended targets:
  - Private-sector models may focus on predicting successful speculative attacks measured by large short-run exchange rate moves.
  - In-house models and GS aim to predict both successful and unsuccessful attacks via an exchange market pressure index.

### V. Value added of EWS models — evaluation approach and findings
- Main evaluation principle: focus on out-of-sample (real-time) performance rather than in-sample fit.
- Benchmarks used for comparison:
  - Bond spreads on dollar-denominated sovereign debt
  - Sovereign credit ratings
  - Economist Intelligence Unit (EIU) country expert assessments of currency crisis risk
- Historical finding for the Asia crises: KLR implemented as it might have been in early 1997 produced forecasts that "decisively outperformed all the comparators" for that period.
- General conclusion emphasized in text:
  - EWS models were statistically significant predictors of crises in earlier studies, but economic significance (usefulness to an already-informed observer) is less clear.
  - Comprehensive, country-specific, holistic analyst assessments plausibly outperform simple EWS models, but no definitive studies proving this superiority are available; paper gains perspective by direct comparison to market and analyst indicators.
- Ongoing monitoring since early 1999: the Fund produced forecasts from KLR and DCSD and monitored GS and CSFB (and later DBAC) to assess usefulness in providing early warnings of crises.
- Testing approaches described:
  - Reexamination of KLR predictions for Asia crises using a pre-crisis implementation benchmark
  - Detailed analysis of first forecasts produced within the Fund in May 1999 and subsequent out-of-sample tests (text references to Figures and Tables: Text Tables 2–5; Figures 1–6; Appendix Tables 6–7)

*Source: _wp0452 - References .............................................................................................................*

### 1999. These are compared to the alternative indicators, such as bond spreads, ratings and

### _wp0452 - 1999. These are compared to the alternative indicators, such as bond spreads, ratings and

### A. EWS Models and Alternative Indicators in Asia Crisis
- Main conclusion from Berg and others (1999): the original KLR model, designed prior to 1997, had substantial predictive power over the Asian episodes.
- Example country rankings in early 1997 (KLR): Korea and Thailand were among the top third of countries in terms of vulnerability.
- Country rank in the predicted vulnerability list is a statistically significant predictor of actual crisis incidence.
- Shortcomings identified in earlier models:
  - Important potential indicators had not been tested, e.g., current account deficit as a share of GDP and the ratio of short-term external debt to GDP.
  - Regression-based estimation techniques were promising compared to the KLR “indicator” method.
- Model improvements:
  - A revamped KLR-based model and the DCSD model were developed; these perform substantially better in predicting the Asia crises (Column 2 of Table 2 shows DCSD results).
- Comparative performance vs non-model indicators:
  - Sovereign spreads (Q1 1997) averaged 90 basis points in the countries that subsequently suffered a crisis, while they averaged 201 in the other countries.
  - Sovereign ratings (Q1 1997), based on a quantitative conversion of Moody’s and S&P’s ratings where higher numbers correspond to a better rating, were on average substantially better in the ten most affected countries than in the ten least affected countries.
  - Economist Intelligence Unit (EIU) currency crisis risk scores (risk of a 20 percent real depreciation over two years) gave generally positive assessments to the Asian economies that subsequently suffered severe episodes; countries with higher EIU risk scores in the second quarter of 1997 were systematically less likely to have a crisis during the 1997–1998 period.
  - Surveys of market participants (e.g., Financial Times Currency Forecaster) provided no useful early warning in a large sample of emerging markets or key cases such as Mexico 1994 or Thailand 1997.
- Overall summary: little evidence that “market views” (spreads, ratings, surveys) are reliable crisis predictors, although they remain important for market access and sentiment.

### Table 2: Key numeric comparisons (as presented)
- Averages for crisis vs non-crisis countries (as shown in Table 2):
  - Crisis Countries: KLR 20, DCSD 34, Spread 90, Rating 34, EIU Currency Risk 35
  - Non-Crisis Countries: KLR 15, DCSD 17, Spread 201, Rating 57, EIU Currency Risk 46
- Rank Correlation (indicator ranking vs crisis severity):
  - KLR: 0.52
  - DCSD: 0.53
  - Spread: -0.31
  - Rating: -0.49
  - EIU: -0.33
- Notes from Table 2 footnotes:
  - Probabilities of currency crisis over 24 months horizon reported for KLR and DCSD models.
  - The spread is expressed in basis points (difference between yield on US dollar denominated Foreign government eurobond and equivalent maturity US treasury bond).
  - Rating: average of S&P and Moody's ratings converted to numerical rating ranging from 100 (S&P SD) to 0 (S&P AAA or Moody's Aaa); a lower number means a better rating (unlike Ferri, Liu, and Stiglitz).
  - EIU Currency risk: "Scores and ratings assess the risk of a devaluation against the dollar of 20% or more in real terms over the two year forecast period".

### B. Korea — Time-series indicators (Figure 1)
- Indicators plotted for Korea, 1996–2000 (through November 20, 2000):
  - Forward premium defined as log of the ratio of the 12-month forward to the spot exchange rate.
  - KDB Bond Spread: spread of KDB Eurobond over comparable US Treasuries.
  - Estimated Probability of Crisis by DCSD model: probability of a crisis over the next 24 months estimated using the DCSD model.
- Sources for Figure 1: Bloomberg, J.P. Morgan; and IMF staff estimates.

### Continued implementation, hindsight, and limitations
- The in-sample and out-of-sample results of EWS testing were sufficiently promising to support continued implementation and further research.
- Caveats:
  - DCSD model benefited from hindsight in its formulation, though only pre-Asia-crisis data was used to produce the forecasts.
  - Main benefit from hindsight: inclusion of short-term-debt/reserves as a predictive variable (the original KLR model had focused on M2/reserves instead).
  - It is possible that good Asia-crisis performance was due to luck or that subsequent crisis episodes differ enough that model performance may degrade.

### C. Performance since January 1999 and July 1999 forecasts
- Two approaches to evaluation since implementation in January 1999:
  1. Results of the models the first time forecasts were produced “officially” in July 1999.
  2. Systemic measures of model “goodness of fit,” including comparison of in-sample and out-of-sample performance and the tradeoff between missing crises and false alarms.
- July 1999 forecasts (Table 3 summary in text):
  - The KLR and particularly the DCSD model did fairly well in these first official forecasts.
  - Both countries with DCSD probabilities above 50 percent subsequently had crises.
  - No crisis country had a DCSD probability below 26 percent.
  - Using the models’ own definition, there have been only three crises in the roughly two years since July 1999.

*Sources: Kaminsky, Lizondo, and Reinhart (1998); Berg and others (1999); Economist Intelligence Unit; Standard and Poor’s; and Moody’s.*

### Appendix II explains how in-sample and out-of-sample periods are determined for each of

### _wp0452 - Appendix II explains how in-sample and out-of-sample periods are determined for each of

### Determination of in-sample and out-of-sample periods
- For the KLR and DCSD models the out-of-sample period extends from January 1999 through December 2000.
- For the private sector GS and CSFB models, which have three-month and one-month horizons respectively, goodness-of-fit can be examined through April 2001 (GS) and August 2001 (CSFB).
- The threshold probability for an alarm is chosen by minimizing a weighted sum of the share of false alarms and the share of missed crises; throughout the chapter equal weight is placed on false alarms and missed crises.

### Crisis-probability snapshot (as of July 1999) — comparative indicators
- Table 3 (as of July 1999) reports crisis probabilities/scores from different models and indicators (KLR, DCSD, Spread in basis points, Ratings, Economist Intelligence Unit).
- Selected country examples (as shown in Table 3):
  - Colombia (Aug. 99): KLR 42, DCSD 61, Spread 544, Ratings 45, Economist Intelligence Unit 42
  - Turkey (Feb. 01): KLR 45, DCSD 50, Spread 554, Ratings 68, Economist Intelligence Unit 58
  - Zimbabwe (Aug. 00): KLR 24, DCSD 26, Spread na, Ratings na, Economist Intelligence Unit 77
  - India: KLR 11, DCSD 9, Spread na, Ratings na, Economist Intelligence Unit 38
  - Indonesia: KLR 32, DCSD 1, Spread 872, Ratings 78, Economist Intelligence Unit 50
- Averages reported in Table 3:
  - Crisis Countries: 37, 46, 549, 56, 59
  - Non-Crisis Countries: 22, 16, 462, 55, 42
- Notes on interpretation:
  - The spread is expressed in basis points (difference between yield on US dollar denominated foreign government eurobond and equivalent maturity US treasury bond).
  - Ratings are an average of S&P and Moody’s converted to numerical ratings from 100 (S&P SD) to 0 (S&P AAA or Moody’s AAA); a lower number means a better rating.
  - Economist Intelligence Unit currency risk scores assess the risk of a devaluation against the dollar of 20 percent or more in real term over the two year forecast period.

### Performance of non-model indicators vs. model-based predictors
- The alternative (non-model-based) predictors improved relative performance compared to their performance before the Asia crises, but still did not perform well overall.
- Spreads: average 549 for the three crisis countries versus 462 for the others (as of July 1999).
- Average sovereign rating: 56 for the three crisis countries and 55 for the rest.
- Economist Intelligence Unit risk scores: average 59 for the three crisis countries versus 42 for the others.
- Observations on distinction between sovereign/default risk and currency crisis risk:
  - Colombia’s August 1999 crisis involved a drop in the exchange rate with little concern about sovereign default; ratings and spreads unsurprisingly did not predict this incident.
  - Pakistan experienced a debt crisis but no currency crisis over the period; high spreads (widening since 1998 sanctions) raised the non-crisis-country average. Excluding Pakistan, the average spread for non-crisis countries declines to 312 from 462.

### Overall goodness-of-fit since January 1999 — model-level findings
- General findings:
  - In-sample goodness-of-fit of models has been reasonably stable as new data arrived; models’ coefficients and statistical significance have not changed much for some models (e.g., DCSD).
  - The out-of-sample period studied was comparatively calm: two crises per year during the recent period versus almost three crises per year in the estimation period.
  - The lower incidence of crises during the out-of-sample period increased the challenge for models to avoid false alarms rather than to predict unforeseen crises.
- Model-specific out-of-sample performance:
  - KLR model:
    - Calls 58 percent of pre-crisis months correctly (out-of-sample).
    - When the crisis probability was below the cut-off, a crisis ensued 9 percent of the time.
    - When the crisis probability was above the cut-off, a crisis ensued 35 percent of the time.
    - Overall, KLR performed better out-of-sample than in-sample and the forecasts were highly informative.
  - DCSD model:
    - Calls 31 percent of pre-crisis months correctly (out-of-sample).
    - Crises followed above-cut-off signals 22 percent of the time and followed below-cut-off signals 14 percent of the time.
    - DCSD performance deteriorated substantially out-of-sample relative to in-sample.
    - The DCSD model produced more false alarms out-of-sample than in-sample for any cut-off level in the examined figures.
    - A much higher cutoff (around 50 percent) would have reduced false alarms without substantially increasing missed crises; this is reflected by an out-of-sample curve that approaches the in-sample curve near a cutoff probability of 47 percent.
- Implication:
  - While it is impossible to know the optimal cut-off until after the fact, the successful goodness-of-fit performance for some cut-off values implies the models could rank observations reasonably well by crisis probability (higher probabilities concentrated on pre-crisis months).

### Goodness-of-fit measures and evaluation approach
- The loss function used equals the sum of:
  - false alarms as a share of total tranquil periods, and
  - missed crises as a share of total pre-crisis periods.
- Key goodness-of-fit concepts (as used in Tables and notes):
  - Cut-off: the probability above which a forecast is deemed to signal a crisis.
  - Percent of pre-crisis periods correctly called: share of pre-crisis periods for which estimated probability is above the cut-off and a crisis ensues within 24 months.
  - Percent of tranquil periods correctly called: share of tranquil periods for which estimated probability is below the cut-off and no crisis ensues within 24 months.
  - False alarm: observation with estimated probability above the cut-off not followed by a crisis within 24 months.
  - Probability of crisis given signal: pre-crisis periods correctly called as a share of total predicted pre-crisis periods.
  - Probability of crisis given no signal: periods where tranquility is predicted and a crisis actually ensues as a share of total predicted tranquil periods.
  - Statistical evaluation: crisis dummy (1 if pre-crisis month, 0 otherwise) regressed on forecast probabilities with HAC standard errors; reported regression coefficient, standard error, and p-value.

*Source: _wp0452 - Appendix II explains how in-sample and out-of-sample periods are determined for each of*

### Box 2. Goodness-of-Fit Measures and Trade-Offs

### Box 2. Goodness-of-Fit Measures and Trade-Offs

### Overview
- Early warning systems produce a predicted probability of crisis. Goodness-of-fit calculations measure how well predicted probabilities compare to the subsequent incidence of crisis.
- To evaluate performance, predicted probabilities are converted into alarms: an alarm signals that a crisis will ensue within 24 months (the model’s horizon).
- An alarm is defined as a predicted probability of crisis above some threshold level (the cutoff threshold). Each observation (a particular country in a particular month) is categorized as:
  - whether it is an alarm (predicted probability above the cutoff threshold), and
  - whether it is an actual pre-crisis month.

### Thresholds, Loss Function, and Trade-Offs
- The threshold probability for an alarm can be chosen to minimize a “loss function” equal to the weighted sum of:
  - false alarms (as a share of total tranquil periods), and
  - missed crises (as a share of total crisis periods).
- In this paper, equal weight is placed on the share of alarms that are false and the share of crises that are missed.
  - A higher weight on missed crises implies a lower cut-off threshold, resulting in fewer missed crises and more false alarms.
- Only in-sample information can be used to calculate a threshold for actual forecasting purposes. Out-of-sample, a different threshold could provide better goodness-of-fit.

### Goodness-of-Fit Tables and Measures
- Columns of Table 4 and Table 5 show how model signals compare to actual outcomes over various periods; each number is the number of observations satisfying the listed criteria.
  - Example: For the DCSD model (Table 4, column 7) over the 1999:1 to 2000:12 period, there were a total of 443 tranquil months, and for 90 of them the probability was above the cutoff threshold.
- From these contingency counts various accuracy measures can be calculated; e.g.:
  - Percent of crises correctly called = (number of observations with alarm and subsequent crisis) / (total number of actual crises).
- Footnotes to Table 4 define the various measures.

### Cut-off Independent Comparison (ROC-style presentation)
- Models may perform differently depending on whether emphasis is on avoiding false alarms (high cutoffs) or avoiding missed crises (low cutoffs), but for the models examined this is generally not the case.
- For a given model and sample, each candidate cutoff threshold produces a percentage of crises correctly called and a percentage of false alarms.
  - Cutoff = 0 produces 100 percent crises correctly called and 100 percent false alarms.
  - Cutoff = 100 produces 0 percent crises correctly called and 0 percent false alarms.
- The upper panel of Figure 2 traces points for each cut-off between 1 and 100 for the DCSD and KLR models over the in-sample period.
  - Points further toward the lower right (higher percent crises correctly called, lower percent false alarms) are unambiguously preferred for any loss function.
  - The DCSD model dominates the KLR model for all cut-off frequencies in the in-sample period: the DCSD curve lies to the right and below the KLR curve. For any given percent of crises correctly called, the DCSD model calls fewer false alarms.
- Text Figures 3 through 6 show similar results for individual models’ in-sample and out-of-sample results.

### Loss Function Graph and Interpretation
- The lower panel of Figure 2 shows the loss function (number of false alarms in percent of tranquil periods plus number of missed crises in percent of pre-crisis periods) for various cut-off thresholds and models.
  - Example interpretation: a loss function value of 50 implies 20 percentage points fewer false alarms and/or missed crises than a loss function of 70.
- Notes on Figure 2:
  - Point A corresponds to the cut-off that minimizes the loss function in-sample for the KLR model.
  - Point B indicates the same point for the DCSD model.
  - DCSD: Berg and others (1999). DCSD stands for Developing Country Studies Division.
  - KLR: Kaminsky, Lizondo, and Reinhart (1998).

### Regression Evaluation and Statistical Significance
- Table 4 also reports regressions of the actual crisis indicator on the model’s predicted probability of crisis:
  - Regression form: c24_it = α + β PredProb_it + ε_it
    - c24_it = 1 if there is a crisis in the 24 months after period t (for country i) and 0 otherwise.
    - PredProb_it is the predicted crisis probability for period t and country i.
  - For informative forecasts, β should be significant; a coefficient of 1 implies that forecasts are unbiased.
- Estimation details:
  - Regression is estimated using OLS with HAC standard errors to correct for serial correlation and heteroskedasticity concerns.
- Empirical findings:
  - In-sample: the estimated β is always statistically different from 0 and a value of 1 cannot be rejected for both KLR and DCSD (strong in-sample results).
  - Recent out-of-sample period:
    - KLR model’s forecasts are highly significant; the hypothesis that the true β is 1 cannot be rejected.
    - DCSD forecasts are not significant at traditional confidence levels:
      - p-value for β = 0 is 12 percent.
      - p-value for β = 1 is 31 percent.
  - Tests lack power in the out-of-sample period due to the small amount of information.

### Simulation Exercise and Power Illustration
- To illustrate lack of power, a simulation supposes that the data-generating process for the out-of-sample period is exactly the DCSD model as estimated in-sample, implying β = 1. The paper uses this exercise to show how sampling variability and limited out-of-sample information affect the ability to reject β = 0 or confirm β = 1.

*Source: Author’s calculations, Box 2 text in the provided document.*

### conclusions that these procedures are unsatisfactory and suggest that OLS with HAC

### _wp0452 - conclusions that these procedures are unsatisfactory and suggest that OLS with HAC

### Performance of monitored EWS models
- Models monitored: two long-horizon, in-house models (DCSD and KLR) and two short-horizon, private sector models (GS and CSFB).
- Forecast periods evaluated:
  - DCSD and KLR: January 1999 to December 2000 (24-month horizon).
  - GS: January 1999 to April 2001.
  - CSFB: April 2000 to June 2001.
- Nature of forecasts: “Pure” out-of-sample forecasts — no information about actual outcomes used in forecasts or model development/estimation.

### Main empirical findings
- KLR model:
  - Forecasts are statistically and economically significant predictors of actual crises.
  - Out-of-sample accuracy is only slightly inferior to in-sample accuracy.
- DCSD model:
  - Performs substantially worse out of sample than in sample.
  - Forecasts remain somewhat informative.
  - Statistical evidence: hypothesis that forecasts are unbiased and informative is more likely (p value = 0.31) than the hypothesis that the forecasts were useless (p value = 0.12).
  - Out-of-sample period contains 528 observations but only eight crises.
- Short-horizon private-sector models (GS and CSFB):
  - Performed poorly out of sample despite strong in-sample performance.
  - Both sets of crisis predictions were largely uninformative; probability of a crisis was about the same whether forecast probability was above a cutoff threshold or not.

### Table 5 (Goodness of Fit: Short-Horizon Models) — key exact figures
- Sample periods reported:
  - In Sample: GS 1996:1 to 1998:12; CSFB 1994:1 to 2000:7.
  - Out of Sample: GS 1999:1 to 2001:4; CSFB 2000:8 to 2001:8.
- Cut-off values used: 1/ = 10 (GS), 35 (CSFB) in the table headings.
- Value of loss function 2/: In Sample: 66 (GS), 58 (CSFB). Out of Sample: 97 (GS), 88 (CSFB).
- Percent of observations correctly called: In Sample: 66 (GS), 76 (CSFB). Out of Sample: 50 (GS), 83 (CSFB).
- Percent of crisis correctly called 3/: In Sample: 66 (GS), 65 (CSFB). Out of Sample: 54 (GS), 27 (CSFB).
- Percent of tranquil correctly called 4/: In Sample: 66 (GS), 76 (CSFB). Out of Sample: 50 (GS), 85 (CSFB).
- False alarms as percent of total alarms 5/: In Sample: 74 (GS), 92 (CSFB). Out of Sample: 87 (GS), 96 (CSFB).
- Probability of crisis given signal 6/: In Sample: 26 (GS), 8 (CSFB). Out of Sample: 14 (GS), 4 (CSFB).
- Statistical tests of forecasts 8/ (regression of actual on predicted with HAC standard errors):
  - Coefficient in regression of actual on predicted: In Sample: 1.41 (GS), 0.17 (CSFB). Out of Sample: 0.56 (GS), 0.06 (CSFB).
  - Standard error: In Sample: 0.42 (GS), 0.03 (CSFB). Out of Sample: 0.41 (GS), 0.06 (CSFB).
  - p-value (coefficient <= 0): In Sample: 0.00 (GS), 0.00 (CSFB). Out of Sample: 0.17 (GS), 0.29 (CSFB).

### Interpretation and caveats regarding statistical power and sample size
- Limited number of crises observed yields a small effective sample size; small changes in sample can materially change goodness-of-fit indicators.
- Example: out-of-sample period for DCSD contains 528 observations but only eight crises — low number of crises constrains information.
- Serial correlation in pre-crisis observations and predicted probabilities reduces effective information and greatly increases standard errors.
- Simulations for the DCSD model: simulating the implied data-generating process 500 times showed that for these ideal forecasts the hypothesis that β = 0 would not be rejected 28 percent of the time.

### False alarms and missed crises — nuanced interpretation
- “False alarm” defined mechanically: an alarm not followed by a crisis (GS: within 3 months; CSFB: within 1 month).
- Warnings classified as false may still be valuable because:
  - Crisis definitions may miss events one would want a model to warn about.
  - Warnings may lead to policy adjustments or luck that avoid a crisis; warning could still have been useful.
- Empirical note: examination of 90 observations that generated false alarms in DCSD model for 1999:1 to 2000:12 suggests that half (46) can be classified as either missed crisis classifications or useful warnings (countries: Argentina, Chile (before July 1999), Pakistan, Venezuela, Turkey, Uruguay).

### Comparative performance versus non-model indicators
- During the Asian crisis, DCSD and KLR EWS models performed much better than non-model predictors (bond spreads, ratings, analysts’ assessments).
- During more recent periods, some alternative indicators improved, narrowing the models’ relative superiority.
- Overall conclusion: EWS models are not accurate enough to be sole methods for anticipating crises but can contribute to vulnerability analysis alongside traditional surveillance and other indicators.

### Structural and contextual considerations for future EWS models
- Emerging trends affecting currency-crisis prediction:
  - Resurgence of crises in which sovereign and domestic debt dynamics play a central role; debt and currency crises are related but distinct.
  - Increased importance of floating exchange rates in emerging markets and more prevalent de jure floating regimes with capital account openness and substantial de facto flexibility (examples listed in source: Brazil, Chile, Colombia, Korea, Mexico, Poland, Thailand, and historically South Africa).
- Implications:
  - Floating regimes do not eliminate the possibility of sharp depreciations or crises; floating regimes might still evolve toward de facto pegs under speculative pressure.
  - No firm conclusion on whether floating rates are more resistant to currency crises; IMF de jure classification shows no evidence that floating rates have been more resistant.
- Research direction noted in source: recent work has turned to predicting debt crises and fiscal vulnerabilities.

### Overall conclusions and policy-relevant points
- EWS models’ strength: objective, systematic processing of data; potential to identify vulnerabilities that other analyses miss and to avoid biases rooted in past experiences.
- Practical use: EWS models should be used as one component in broader surveillance and vulnerability assessment frameworks rather than as standalone predictors.
- Historical lesson: models identified vulnerability signals (example: Korea in 1996–97) that were not reflected in market/analyst signals prior to crises.
- Expectation: painful currency crises are likely to remain a feature of emerging markets for the foreseeable future.

*Source: Author’s calculations and analysis in the provided content unit.*

### 1. Models Implemented at the IMF

### 1. Models Implemented at the IMF

### Kaminsky-Lizondo-Reinhart (KLR) model
- Purpose: Predicts the probability of a crisis within the next 24 months.
- Crisis definition: Extreme changes in a weighted average of the monthly exchange rate depreciation and reserve loss.
- Methodology:
  - Monitors a large set of monthly indicators that signal a crisis whenever they cross a certain threshold.
  - Uses a variable-by-variable (indicator) approach to identify which variables are “out-of-line”.
  - Information from separate variables is combined using each variable’s forecasting track record to produce a composite measure of the probability of crisis.
- Typical predictors included:
  - Overvaluation, the current account, reserve losses, export growth.
  - Reserves to broad money as a measure of reserve adequacy.
  - Monetary variables: domestic credit growth, real interest rates, excess M1 balances.
- Implementation note: Staff has implemented a version of the KLR model supplemented with several additional variables.

### Developing Country Studies Division (DCSD) model
- Origins and evolution:
  - Influenced by out-of-sample testing of KLR and other models; later work tested discrete thresholds versus linear effects (Berg and Pattillo, 1999a,b).
  - Embedded the KLR approach into a multivariate probit regression and found the probability of crisis tends to go up linearly with changes in predictive variables.
- Measurement approach: Variables measured in percentiled form (relative to their own history).
- Current model composition:
  - Initially a “linear” probit composed of five variables: real exchange rate deviations from trend, the current account to GDP ratio, export growth reserve growth, and the level of M2 to reserves.
  - After adding short-term debt to reserves (in response to the Asian crises), the ratio of M2 to reserves lost significance and was dropped, resulting in the current five-variable DCSD model.
- Data frequency: Uses mainly monthly data, but some quarterly or annual series are interpolated or extrapolated to generate monthly crisis predictions.

### Policy Development and Review (PDR) model
- Additions to DCSD: Incorporates balance sheet variables and proxies for standards.
- Important predictors identified:
  - Corporate level: leveraged financing and a high ratio of short-term debt to working capital.
  - Balance sheet indicators: bank and corporate debt to foreign banks as a share of exports.
  - Institutional/legal: a legal regime variable proxying shareholders rights.
- Data constraints:
  - Corporate sector data are available only on an annual basis and with a significant lag.
  - Despite slow-moving nature of these variables, they can improve forecasting accuracy.
  - More up-to-date data and a larger, more stable corporate sample would increase usefulness.

### Private Sector Models (overview)
- Motivation: Investment banks developed EWSs to advise clients on FX trading, risk assessment, and to supplement economic forecasts.
- Evolution: Many banks created models after the Asian crisis; some later abandoned or simplified them. Examples cited include Lehman Brothers, Citicorp, JP Morgan, Deutsche Bank, and Morgan Stanley Dean Witter.
- Typical horizons: Private sector models usually target shorter horizons (one to three months) and often provide weekly updates.
- Common features across private models:
  - Inclusion of contagion measures and market sentiment indicators (e.g., stock prices).
  - Use of models for trading strategies; claims of substantial profits over particular periods.
  - Some models propose “action triggers” to convert probability outputs into trading decisions.

### Goldman Sachs’s GS-WATCH
- Horizon: Predicts likelihood of a crisis in three-month time.
- Crisis definition: Weighted average of three-month exchange rate and reserve changes.
- Methodology:
  - Logit regression where most explanatory variables are converted into zero/one signals.
  - Predictor variables include credit booms, real exchange rate misalignment, export growth, reserve growth, external financing requirements, changes in stock prices, political risk, contagion and global liquidity.
  - Political risk is represented as a simple zero/one signal (e.g., around elections or major unrest).
  - Contagion is a weighted average of changes in the exchange rate and reserve change index for other countries, with weights based on historical relationships across countries.
- Frequency: Model estimated with monthly data; predictions updated weekly for analysts’ reports.
- Behavior: Week-to-week movements in crisis probabilities are often driven by changes in the contagion variable.

### Credit Suisse First Boston’s Emerging Markets Risk Indicator (CSFB)
- Re-specification: Model changed in September 2000, reducing number of predictor variables.
- Horizon: Predicts the one-month ahead probability of depreciation greater than 5 percent and at least double the preceding month’s depreciation.
- Measurement: Variables standardized (measured relative to the country-specific mean and variability).
- Typical predictors: Real exchange rate deviations from trend, the ratio of debt to exports, growth in credit to the private sector, output changes, reserves to imports, changes in stock prices, oil prices, and a regional contagion dummy (number of countries in the region recently experiencing a crisis).

### Deutsche Bank Alarm Clock (DBAC)
- Event definitions:
  - Exchange rate “events”: depreciations greater than sizes estimated separately for levels ranging from 5 to 25 percent.
  - Interest rate “events”: increases in the money market interest rates of more than 25 percent in a month.
- Methodology:
  - Jointly estimates probabilities of exchange rate and interest rate events, allowing interaction effects between the two event types.
  - Exchange rate event predictors: changes in stock prices, domestic credit, industrial production, real exchange rate deviations, and a contagion variable.
- Trading application: Proposes an “action trigger” cutoff probability level to maximize profits from a strategy switching long/short local currency positions based on the trigger.

### Specification Issues (methodological considerations)
- Predictand ambiguity:
  - Currency crisis definitions vary: sudden large exchange rate depreciation, large reserve losses, or hikes in domestic interest rates can all qualify.
  - Forecasting may aim to capture both successful (resulting in depreciation) and unsuccessful attacks.
- Variable selection guided by theory:
  - Theoretical models suggest three classes of variables:
    1. Macroeconomic fundamentals: real exchange rate overvaluation, fiscal deficit, excess money growth, terms of trade, domestic credit, current account deficit, output growth.
    2. Vulnerability measures: adequacy of international reserves relative to short-run liabilities, external financing needs, financial sector soundness.
    3. Market expectations/sentiment: interest rate differentials, bond spreads, forward exchange rate, number of crises elsewhere/contagion channels, investors’ “risk appetite”.
- Overfitting and parsimony:
  - Danger of overfitting through data-mining; models that explain a particular historical episode may perform poorly forecasting future crises.
  - Emphasis on parsimonious models capturing robust “symptoms” common across crisis episodes.
- Interrelated indicators:
  - Many indicators overlap; including all is unnecessary if some variables capture manifestations of others (e.g., fiscal deficit and inflation effects may be captured by real exchange rate overvaluation and current account).
- Data availability and frequency constraints:
  - Consistent, high-frequency cross-country data are often unavailable for financial sector health (e.g., non-performing loans) and political risk.
  - Contagion is difficult to incorporate for longer-horizon prediction windows because it can operate very rapidly.
- Aggregation and modeling choices:
  - Signaling (indicator) approach: issues include jumpy probabilities as variables cross thresholds and the inability to capture inter-relationships among indicators.
  - Probit/logit regression: allows joint consideration of variable correlations and statistical testing, but is nonlinear—marginal contribution of a variable depends on magnitudes of others, reducing transparency.
  - Alternative aggregation: weighted-sum composite probabilities, where indicators are weighted by reliability; BIS uses aggregate scores with judgmental weights.
  - Other explored methods include linear discriminant analysis, switching regime models, and non-parametric safe-zone identification.
- Forecasting horizon implications:
  - KLR and DCSD target a 24-month horizon (probability that a crisis will occur sometime in the following 24 months).
  - Research on the DCSD model found relatively little difference using horizons between nine months and two years.
  - Private sector models focus on 1–3 month horizons to suit trading and short-term portfolio decisions; variables such as contagion, stock prices, and domestic credit growth are especially relevant at short horizons.

*Source: _wp0452 - 1. Models Implemented at the IMF*

### appendix defines these terms and explains their implementation in this paper. The designer of

### APPENDIX II

### In-sample testing: purpose and cautions
- The designer of an EWS chooses variables and estimates parameters to best fit observations in an estimation sample.
- In-sample testing measures how well models fit the crises in this sample, i.e. in-sample.
- Good in-sample testing is a sign of a useful model but must be interpreted cautiously:
  - Good in-sample performance may be a coincidence, perhaps resulting from a search through a large number of specifications until a good fit occurs by chance.
  - The determinants of crises may vary through time.

### Out-of-sample testing: definition and practical constraints
- In out-of-sample testing, predictions of an existing model are compared to a new set of observations not belonging to the estimation sample.
- An unavoidable difficulty: a forecast can only be properly judged after the entire forecast window has closed.
- Example timing constraints stated:
  - This paper examines the forecasts through June 1999 for the KLR and DCSD models.
  - A prediction of risks as of August 1999 cannot be fully judged until September 2001 because it is not yet known whether a crisis followed within 24 months.
  - Given the two-year model horizon, these forecasts apply to realizations during a two-and-a-half year period, through July 2001.
  - For private sector models:
    - GS model horizon: three-months — goodness-of-fit observable through April 2001.
    - CSFB model horizon: one-month — goodness-of-fit observable through August 2001.

### Best practice for out-of-sample testing and model creation
- Out-of-sample testing should mimic the process by which a forecasting model would be used in practice.
- Strictest form: the modeler has no knowledge of the out-of-sample observations when generating forecasts to be tested.
- Less strict approach: the modeler withholds recent observations from the estimation sample and uses them for subsequent out-of-sample testing but may still use information from those observations to create the model.
  - Example: The DCSD model was estimated over the pre-Asia-crisis period and used to predict the Asia crises out-of-sample in OP 186. However, the authors created the model in 1998, after the Asia crises, and added the short-term debt/reserves variable because they knew it was likely important in explaining the Asia crises.

### Estimation period timing constraint for KLR and DCSD
- In-sample estimation periods for KLR and DCSD must end some 24 months before the time the model is being estimated, because it is not yet known whether later observations are pre-crisis or tranquil periods.
- Example: the in-sample period for the DCSD model in OP 186 ended in May 1995 so that the estimation did not reflect knowledge of the Asia crises that began in July 1997.

### Model samples (as reported in Table 7)
- Asia crisis
  - In-sample
    - KLR/DCSD: 1985:12 to 1995:4
  - Out-of-sample
    - KLR/DCSD: 1995:5 to 1996:12
- Recent experience
  - In-sample
    - KLR/DCSD: 1985:12 to 1997:5
    - GS: 1996:1 to 1998:12
    - CSFB: 1994:1 to 2000:7
  - Out-of-sample
    - KLR/DCSD: 1999:1 to 2000:12
    - GS: 1999:1 to 2001:4
    - CSFB: 2000:8 to 2001:8

### Notes on the choice of out-of-sample start dates and forecast vintages
- Start dates for out-of-sample periods are chosen to be after the date at which they could have informed estimation.
- KLR and DCSD forecasts examined for the 1999:1 to 2000:12 period correspond to versions used for the “official” internal July 1999 forecasts and subsequent internal forecasts; model specifications were finalized in late-1998.
- GS out-of-sample forecasts come directly from contemporary monthly publications over 1999:1 to 2001:4, so they could not have reflected out-of-sample information.
- CSFB out-of-sample estimates for the April 2000 to June 2001 period were produced in August 2001 using the model as it had been estimated a year earlier, so in principal they should not have been influenced by out-of-sample events.

*Source: Author’s calculations.*

---


_Source: https://www.imf.org/-/media/websites/imf/imported-full-text-pdf/external/pubs/ft/wp/2004/_wp0452.pdf_
