## 5. How does the accuracy of WEO forecasts compare to those of CE forecasts?

## Source details

**Canonical URL:** [5. How does the accuracy of WEO forecasts compare to those of CE forecasts?](https://www.imf.org/-/media/files/publications/wp/2021/english/wpiea2021216-print-pdf.pdf)

## Other formats

- [Markdown version](/-/media/files/publications/wp/2021/english/wpiea2021216-print-pdf.pdf.md)
- [Structured JSON version](/-/media/files/publications/wp/2021/english/wpiea2021216-print-pdf.pdf.json)

---

### Structure and scope
- Paper analyzes WEO GDP growth forecast accuracy, bias, sources of errors, and compares WEO forecasts to Consensus Economics (CE) forecasts.
- Relevant sections summarized:
  - Section 2: data description.
  - Section 3: predictive accuracy across horizons, countries, and time; comparison to a naïve historical-average forecast.
  - Section 4: bias in growth forecasts and its evolution.
  - Section 5: regressions on sources of errors and use of external information (terms of trade, output gap, and growth forecasts for China, the euro area, and the United States).
  - Section 6: comparison to Consensus Economics forecasts.
  - Section 7: summary and steps to enhance WEO forecast performance.

### Data, timing conventions, and sample coverage
- Forecast horizons: h = 0 to h = 5 (h = 0 current-year, h = 1 next year), two forecast rounds per year (Spring and Fall) → 12 forecast horizons in total.
  - Current-year Fall (h = 0, F) and current-year Spring (h = 0, S) are hybrids of nowcasts and forecasts.
- Sample vintages and coverage:
  - Full dataset: back to 1990 for current-year forecasts and 1995 for five-year-ahead forecasts, ending in 2017.
  - Sample size: 28 outturn observations for current-year horizon; 23 outturn observations for five-year horizon.
  - Most calculations use 2004 till 2017 sample; comparisons calculated for 1990–2003.
- Outturn vintage choice: actual value for year t taken as reported in the following year's (t + 1) Fall WEO.
- Country groups (12): World; Advanced Economies (AE); Emerging Market Economies (EME); Low-Income countries (LIC); EMDE (EME ∪ LIC); EMDE_Fuel Exporters (EMDE_FE); EMDE_Fuel Importers (EMDE_FI); Emerging and Developing Europe (EEUR); Emerging and Developing Asia (DASIA); Latin America and the Caribbean (LAC); MENAP_CIS; Sub-Saharan Africa (SSA); IMF-program observations (Program).

### Outliers handling (ANNEX 1 summary)
- Filters applied to drop observations associated with large, unforecastable shocks and data anomalies:
  - Filter one (Conflict): exclude years with 100+ conflict deaths per 1 million population and adjacent years for listed countries.
  - Filter two (Disasters): drop country-year observations where total disaster damage > 3 percent of GDP and growth is negative; Ebola-affected observations dropped for Guinea, Sierra Leone, Liberia (2014–2015).
  - Filter three (Ad-hoc): e.g., Ireland 2015, Equatorial Guinea volatility, drop year 2009 from subsample 2004–2016.
  - Filter four: drop horizons with precisely zero forecast errors.
  - Filter five: drop horizons with fewer than eight observations in a subsample (affects 1990–2003 more).
- Dropping outliers increases mean forecast error (closer to median) and reduces extreme values.

### Predictive accuracy: measures, horizon patterns, and key RMSE findings (2004–17)
- Accuracy metric: root mean squared forecast error (RMSE); presentations include inter-quartile ranges, medians, and GDP-weighted means by group and horizon.
- Typical horizon ordering: (h = 0, F), (h = 0, S), (h = 1, F), (h = 1, S), (h = 2, F) ... (h = 5, F).
- Representative RMSE findings (2004–17):
  - Advanced Economies (AE):
    - Median RMSE ≈ 0.7 percentage point for current-year Fall; rises to ≈ 1.8 percentage points by next-year Spring; remains near that level for two- to five-year horizons.
  - Emerging Market Economies (EME) and Low-Income Countries (LIC):
    - Median RMSE ≈ 1.2–1.5 percentage points for current-year Fall; rises slightly above 2 percentage points by next-year Fall and stays around that level over longer horizons.
    - LICs have a slightly wider interquartile range than EMEs.
  - EMDE Fuel Exporters (EMDE_FE):
    - Median RMSE ≈ 2 percentage points for current-year Fall; ≈ 3.7 percentage points for three- to five-year forecasts.
  - Regional pattern:
    - Forecasts least accurate in MENAP_CIS (high share of fuel exporters); other EMDE regional groups and Program sample broadly comparable.
- Horizon effects:
  - Current-year forecasts (h = 0, F and h = 0, S) have notably lower RMSEs due to partial-year indicators and preliminary GDP values.
  - Accuracy deteriorates with horizon but nonlinearly: largest losses moving from current-year Spring (h = 0, S) to current-year Fall (h = 0, F) and to next-year Fall (h = 1, F); changes much less beyond about 1.5 years.

### Changes in forecasting performance over time (1990–2003 vs 2004–17)
- For more than half of countries, forecasts for most horizons were more accurate in 2004–17 than in 1990–2003.
- Median declines in RMSEs:
  - ≈ 0.7 percentage point for the shortest, current-year Fall projection.
  - ≈ 0.3 percentage point for the four- and five-year ahead horizons.
- Group-specific notes:
  - AE and EMDE_FE: only about half of countries saw improvements at four- to five-year horizons.
  - MENAP_CIS and SSA: only about half the countries saw improvements at the five-year horizon.
  - EEUR: large improvements reflecting large early-1990s output declines.
  - Program episodes: RMSE changes similar to EME and LIC; close to 1 percentage point through four-year horizon and < 0.5 percentage point in five-year horizon.
- Statistical significance:
  - Increases in accuracy tend to be mostly statistically significant (especially within LAC, MENAP_CIS, and SSA); worsenings often statistically insignificant.
  - Virtually no instances where increases in RMSE-values for current- and next-year forecasts between subsamples are statistically significant.
- Effect of including 2009:
  - For World sample, accuracy for horizons up to four years improves for most countries in 2004–17 even if 2009 errors are kept.
  - AE exception: median difference negative for one-year ahead Spring and longer horizons when 2009 is kept.

### Adjusting for underlying variability: Theil U-statistic and relative accuracy
- Theil U-statistic computed by scaling mean squared error by recursively-updated historical variance (prevailing mean).
  - U < 1: WEO more accurate than historical average.
  - U > 1: historical average more accurate than WEO.
- Findings:
  - At shortest two horizons (h = 0, F) and (h = 0, S), U-statistic values small—about 0.4 and 0.6 respectively, for three quarters of economies globally.
  - As horizon expands, WEO’s advantage over historical average deteriorates.
  - Much of worsening in U-statistic occurs as horizon lengthens from three months (h = 0, F) to about two years (h = 1, S).
  - For about a quarter of economies the U-statistic climbs above unity as of next-year Spring forecast (h = 1, S).
  - Median U-statistic for World sample remains only modestly below unity for three- to five-year forecasts—close to half of economies have longer-term WEO forecasts less accurate than recursive mean.
- Regional/group patterns in long-horizon underperformance:
  - AEs: median reaches unity by next-year Spring forecast.
  - LAC and SSA: median U ≈ 1 from three-year ahead forecasts onwards.
  - DASIA: median U very close to unity from next-year Spring forecast onwards.
- Fuel-exporter dominated groups:
  - EMDE_FE and MENAP_CIS: although absolute RMSE worst, they score strongest on relative accuracy (lower median U-statistics).

### Biases in WEO real GDP growth forecasts: patterns and magnitudes
- Bias definition: bias_g,t0::t1 = (1/(t1 − t0 + 1)) ∑ e_g,t|t−h; positive (negative) values correspond to underpredictions (overpredictions).
- World sample (2004–17) patterns:
  - No significant tendency for upward or downward bias at same- and next-year horizons.
  - Tendency for overprediction at two-year and longer horizons.
  - Median forecast errors:
    - Current-year: small and positive (about 0.1).
    - Next-year: close to zero.
    - Two-year ahead: median forecast error ≈ -0.3 percentage point.
    - Three- to five-year ahead: median forecast error ≈ -0.5 percentage point.
  - "The growth forecasts of one quarter of all countries are biased upward by more than 1 percentage point at the three- to five-year forecast horizons."
- Bias across income and regional groups:
  - AE, EME, LIC: median bias similar for two-year and longer horizons.
  - LIC: optimism at shorter horizons; median bias ≈ 0.35 percentage point for next-year Spring and Fall forecasts; a quarter of LICs have biases > ≈ 0.75–1 percentage point at those horizons.
  - Biases more diverse in EME and LIC (wider interquartile ranges) than AEs.
  - GDP-weighted mean of bias consistently larger than median → greater tendency for overprediction in countries with lower GDP levels.
  - EMDE Fuel Exporters: median bias generally smaller than Fuel Importers, but wider range of biases.
  - Regional: EEUR and SSA mostly optimistic at all but shortest horizons; DASIA often pessimistic; LAC and MENAP_CIS slight underprediction at short horizons turning optimistic at two-years+.
  - Program group: median bias similar to EME and LIC at four- to five-year horizons; slightly more optimistic at one- and two-year horizons.
- Statistical significance of biases (country-level):
  - Current year: growth systematically underpredicted for close to a fifth of World economies; overpredicted for about 3-6 percent.
  - Three- to five-years ahead: underpredicted for about 6 percent; overpredicted for more than a quarter of economies.
  - EEUR: ≈ two thirds of countries have statistically significant upward biases for three- to five-year forecasts.
  - Fuel Exporters: ≈ 10-20 percent have systematic biases in shorter horizons only.
  - Fuel Importers: 25-30 percent have statistically significant upward biases for three- to five-year ahead growth.
- Summary judgment:
  - Country-level WEO forecasts (2004–17): modest downward bias at current- and next-year horizons (except Fuel Exporter, LIC, SSA groups); stronger tendency for overprediction for two-year and longer horizons for all groups except EMDE Fuel Exporters.

### Shifts in bias over time (1990–2003 vs 2004–17)
- Change in bias computed as mean forecast error in 2004–17 minus that in 1990–2003.
- Findings:
  - For majority of subgroups and horizons, median of the difference is negative → lessened tendency for overprediction in 2004–17.
  - EME and LIC: decline in median bias broadly in range 0.4–1.0 percentage point across horizons.
  - AE: reduced overprediction in one- to two-year ahead forecasts; little change for same-year forecasts; more upward bias at three-year+ horizons for slight majority.
  - AE, MENAP_CIS, SSA: median error for four- to five-year forecasts largely unchanged or larger than in 1990–2003.
  - Program forecasts: optimism declined for 75 percent of cases for same year, next-year, and two-year horizons; for more than half for three- to five-year forecasts.
- Sensitivity to including 2009:
  - Including 2009 makes decline in bias smaller, largest effect in AE group where decline becomes very small and confined to current- and next-year forecasts.

### Bootstrap permutation test of bias changes and serial correlation
- Bootstrap permutation test rejects equality of absolute biases in many cases:
  - Aggregated results: for most horizons many countries have significantly reduced biases; World sample shows weighted average of significant declines in overprediction outweighs increases for all horizons except current-year ones.
  - Exceptions: AE (two- to five-year horizons) and DASIA (zero- to two-year horizons).
- Serial correlation in forecast errors:
  - Regression estimator computes serial correlation coefficient 휌̂g over 2004:2017.
  - Presence of serial correlation in Fall current-year and Fall next-year vintages in 15–30 percent of countries except AE group; higher than 5 percent expected under null, especially LIC group.
  - Local serial correlation: world GDP local serial correlation plunged to negative after Global Financial Crisis and returned to positive toward end of sample → overpredictions during crisis followed by underpredictions during recovery.

### Sources of forecast errors and predictability
- Main exercises:
  1. Trace WEO growth forecast errors to contemporaneous errors in US, China, and Euro Area forecasts.
  2. Test informational efficiency using indicators observed at forecasting time (projected TOT, large-country growth forecasts).
  3. Examine predictability using output gap estimates.
- Contemporaneous spillovers (Table 3 highlights):
  - Median pass-through of US same-year forecast error to other countries’ same-year Fall forecasts ≈ 13 percent.
  - Share of countries with positive and statistically significant US spillover typically 5–16 percent; negative and significant 2–8 percent.
  - For significant cases, US forecast errors can explain ≈ one third of variation in country forecast errors.
  - China: median pass-through positive for all groups, 27–46 percent globally across horizons; positive and significant coefficients for ≈ 30–40 percent of economies; R-squared for positive-coefficient regressions ≈ 40–50 percent.
  - Euro Area: median pass-through sizable—50 and 37 percent globally for same- and next-year Fall horizons; strong pass-through (≈ 100 percent) for Eastern European countries.
- Terms of trade (TOT) forecast errors (Table 4):
  - Median impact of TOT forecast error on growth forecast errors typically positive; shares of positive and significant correlations much larger than negative ones.
  - EMDE Fuel Exporter, MENAP_CIS, and LAC show high shares of positive and significant correlations.
  - Median correlations between TOT and growth forecast errors always < 0.05 but TOT errors can explain a meaningful fraction (30–50 percent) of growth forecast errors in key groups.
  - Up to 56 percent of EMDE Fuel Exporter economies have TOT errors as significant predictors; typically 5–15 percent for other EMDE groups.
  - TOT forecast errors more significant at longer horizons (two and five years).
- Are WEO growth forecast errors predictable (Table 5 and 6)?
  - Informational inefficiencies:
    - For 10–25 percent of world economies, a higher US growth forecast predicts a higher growth forecast error.
    - For China, higher projected growth often coincides with overpredictions: almost one third of economies globally show this at five-year horizon.
    - Euro Area forecasts show under-reaction at two- and five-year horizons for ≈ 28 and 33 percent of economies.
  - Terms-of-trade forecasts significantly predict growth forecast errors for many groups and horizons; better accounting for TOT outlook can improve forecasts of 10–30 percent of world economies.
- Output gap and forecast errors (Advanced Economies; equation (18)):
  - Strong negative correlation of -0.6 (significant at 1 percent) for AE group for Spring next-year forecasts (h = 1, S).
  - About one third of AE countries have a significant negative correlation individually.
  - Interpretation: for ~one third of AE countries, growth systematically overpredicted when economy estimated to have spare capacity—possibly due to assumption gap will be closed via above-potential growth.

### Comparison with Consensus Economics (CE) forecasts
- CE vintages monthly (m = 1..12); pairings:
  - CE current-year March (m = 3) ↔ WEO current-year Spring (h = 0, S).
  - CE current-year September (m = 9) ↔ WEO current-year Fall (h = 0, F).
  - Analogous pairings for next-year forecasts.
- Performance measure: Uh = RMSE_WEO(h)/RMSE_CE(h) (equation (19)); formal Diebold-Mariano tests applied.
- Empirical results (Table 7):
  - Comparing WEO Spring/Fall with CE September/March:
    - For a majority of countries, WEO forecasts have lower RMSE than CE for Fall and current-year Spring forecasts.
    - Proportion where WEO RMSE is lower ranges between 56 percent and 66 percent for these horizons.
    - Diebold-Mariano test statistics significant in only a few cases—close to 5 percent.
    - For Spring next-year forecasts, WEO more accurate than CE only about 40 percent of the time.
  - Comparing WEO Spring/Fall with CE April/October:
    - CE forecasts more accurate in the majority of cases → small timing/informational advantages of a few additional weeks of data.

### Main takeaways and quantitative summaries
- Outliers:
  - Upward bias partly because unpredicted growth booms rarer than severe collapses (natural disasters, wars, systemic financial crises).
- Accuracy:
  - In 2004–17, highest accuracy for Advanced Economies; lower and comparable accuracy for EME and LIC.
  - Within EMDEs, forecasts more accurate for fuel importers than fuel-exporters.
  - Predictive accuracy drops with horizon nonlinearly: largest reductions moving from three-month horizon (Fall current-year) to two-year horizon (Spring next year); reductions beyond two years relatively small.
  - Short-term forecasts more accurate in 2004–17 than 1990–2003 for well more than half of countries.
  - Growth forecast errors broadly proportional to growth volatility; after accounting for volatility, accuracy similar across groups in 2004–17.
  - WEO forecasts generally more accurate than naïve recursively-computed historical average for horizons up to 2–3 years in more than half of countries; for about half of countries in AEs, LAC, SSA, and DASIA the naïve forecast is more accurate at horizons > 2–3 years.
- Bias:
  - No discernible bias for current-year and next-year (except LIC optimism).
  - Two-year and longer horizons tend to be upward biased:
    - Median overprediction ≈ 0.3 percentage point for two-year horizon; ≈ 0.5 percentage point for three- to five-year horizons.
  - Optimism declined: bias lower in 2004–2017 than in 1990–2003 for at least half of countries at all horizons; larger declines among EMDEs.
- Sources and predictability:
  - Forecast errors for US, Euro Area, and especially China correlated with many other economies’ forecast errors.
  - TOT errors contribute to growth forecast errors, especially for commodity exporters.
  - For ≈ 8 percent of AEs and 20–30 percent of EMDEs, current-year growth errors tend to persist—near-term forecasts could be improved by learning from recent errors for ≈ 1 in 10 countries globally.
  - Forecasts for up to a quarter of economies can be improved by more efficiently incorporating US growth forecasts; up to 30 percent (especially commodity exporters) by better using own TOT forecasts.
  - Growth systematically overpredicted when output estimated below potential, but less so than in previous evaluation.
- Comparison with CE:
  - Spring and Fall current- and next-year WEO forecasts broadly comparable to CE March and September forecasts.
  - Timing differences explain many small accuracy differences rather than systematic performance gaps.

### Policy recommendations and suggested improvements
- Reduce optimism bias on long-term, multi-year forecasts and some Spring one-year-ahead forecasts:
  - Closely scrutinize two- to five-year ahead growth forecasts that are significantly higher than recent historical averages.
- Modify incorporation of output gap information:
  - Avoid imposing assumption that gap will be closed through above-potential growth over forecast horizon.
- Improve accuracy of TOT forecasts and of growth forecasts for China, Euro Area, and United States to indirectly improve many countries’ forecasts:
  - Better incorporation of information on US growth forecasts can improve forecasts for up to a quarter of economies.
  - Assess uncertainty surrounding TOT and major-economy growth forecasts to evaluate GDP forecast uncertainty.
- Inspect cases where CE forecasts dominate WEO for major economies to identify reasons for weaker WEO performance.
- Reduce serially correlated forecast errors by reacting more strongly to recently observed errors, while avoiding systematic underprediction following large downturns to compensate for past optimism.

_Italic: Source: wpiea2021216-print-pdf - 0.4 and 0.6 respectively, for three quarters of economies globally. WEO forecasts thus appear_

### 5. How does the accuracy of WEO forecasts compare to those of CE forecasts?

### 5. How does the accuracy of WEO forecasts compare to those of CE forecasts?

### Structure and scope
- The paper analyzes WEO GDP growth forecast accuracy, bias, and sources of errors, and compares WEO forecasts to Consensus Economics (CE) forecasts.
- Sections relevant here:
  - Section 2: data description.
  - Section 3: predictive accuracy of WEO GDP growth forecasts across horizons, countries, and over time; comparison to a naïve historical-average forecast.
  - Section 4: bias in growth forecasts and its evolution.
  - Section 5: regressions on sources of errors and use of external information (terms of trade, output gap, and growth forecasts for China, the euro area, and the United States).
  - Section 6: comparison of WEO GDP growth forecast performance to Consensus Economics forecasts.
  - Section 7: summary and steps to enhance WEO forecast performance.

### Data and timing conventions
- Forecast horizons: h = 0 to h = 5 (h = 0 current-year, h = 1 next year, etc.), with two forecast rounds per year (Spring and Fall), yielding 12 forecast horizons in total.
  - Current-year Fall (h = 0, F) and current-year Spring (h = 0, S) are hybrids of nowcasts and forecasts because part of the year’s outcome is observed when forecasts are produced.
- Sample vintages and coverage:
  - Full dataset goes back to 1990 for current-year forecasts and 1995 for five-year-ahead forecasts, ending in 2017.
  - Gives a sample of 28 outturn observations for the current-year forecast horizon and 23 outturn observations for the five-year horizon.
  - Most calculations use the 2004 till 2017 sample, with comparisons calculated for 1994-2003.
- Outturn vintage choice: actual value for year t is taken as reported in the following year's (t + 1) Fall issue of the WEO (to avoid mixing forecast errors with later data revisions).

### Country groups and sample partitions
- Twelve groups analyzed: World; Advanced Economies (AE); Emerging Market Economies (EME); Low-Income countries (LIC); EMDE (union of EME and LIC); EMDE split into Fuel Exporters (EMDE_FE) and Fuel Importers (EMDE_FI); regional groupings: Emerging and Developing Europe (EEUR); Emerging and Developing Asia (DASIA); Latin America and the Caribbean (LAC); Middle East, North Africa, Afghanistan, and Pakistan and Commonwealth of Independent States (MENAP_CIS); Sub-Saharan Africa (SSA); and IMF-program observations (Program).

### Outliers handling
- Growth distributions are left-skewed; mean often falls short of the median due to occasional severe negative shocks.
- The authors drop outliers associated with years of major natural disasters, cross-border conflict, apparent data entry errors, and all forecast errors for 2009 for the baseline analysis (but also report results when 2009 is kept).
- Dropping these observations increases the mean forecast error (bringing it closer to the median) and meaningfully reduces outliers.

### Forecasting instruments used in error-source analysis
- WEO forecasts of the output gap (for advanced economies), terms of trade, and commodity terms of trade for all countries.
- Terms of trade forecasts are based on WEO projections of import and export prices.
- Separate Spring and Fall forecasts for these instruments, covering the same horizons h = 0 through h = 5 (12 forecast horizons).

### Predictive accuracy: measures and patterns
- Accuracy measure: root mean squared forecast error (RMSE), computed as the square root of the average squared h-step-ahead forecast error over the sample.
- Presentation includes inter-quartile ranges, medians, and GDP-weighted means of RMSE by group and horizon.
- Typical horizon ordering shown: (h = 0, F), (h = 0, S), (h = 1, F), (h = 1, S), (h = 2, F) ... (h = 5, F).

Key RMSE findings (2004–17):
- Advanced Economies (AE):
  - Median RMSE starts at about 0.7 percentage point for the current-year Fall forecast and rises to about 1.8 percentage points by the next-year Spring forecast, remaining near that level for the two- to five-year horizons.
- Emerging Market Economies (EME) and Low-Income Countries (LIC):
  - Median RMSE starts at about 1.2–1.5 percentage points for the current-year Fall horizon, rising slightly above 2 percentage points by the next-year Fall forecast and staying around that level over longer horizons.
  - LICs have a slightly wider interquartile range than EMEs, indicating more diverse forecast accuracy.
- EMDE Fuel Exporters (EMDE_FE):
  - Median RMSE close to 2 percentage points for the current-year Fall forecasts and around 3.7 percentage points for the three- to five-year forecasts.
- Regional patterns:
  - Forecasts tend to be least accurate in MENAP_CIS, reflecting a high share of fuel exporters.
  - Median RMSE values broadly comparable across other EMDE regional groups and the Program sample.

Horizon effects:
- Current-year forecasts (h = 0, F and h = 0, S) have notably lower RMSEs than longer-horizon forecasts due to availability of partial-year outcome indicators and preliminary GDP values.
- Accuracy deteriorates as horizon lengthens, but nonlinearly:
  - Largest losses in accuracy occur moving from current-year Spring (h = 0, S) to current-year Fall (h = 0, F) and to next-year Fall (h = 1, F).
  - Accuracy changes much less beyond about one and a half years (beyond next-year Spring forecasts).

### Changes in forecasting performance over time (1990–2003 vs 2004–17)
- Comparison acknowledges that underlying growth volatility/predictability can differ across subsamples; absolute RMSE changes do not alone imply improved forecasting skill.
- For more than half of countries in the overall sample, forecasts for most horizons were more accurate in 2004–17 than in 1990–2003.
- Median declines in RMSEs:
  - About 0.7 percentage point for the shortest, current-year Fall projection.
  - About 0.3 percentage point for the four- and five-year ahead horizons.
- Group-specific notes:
  - In AE and EMDE_FE groups, only about half of countries saw improvements at four- to five-year horizons.
  - In MENAP_CIS and SSA groups, only about half the countries saw improvements at the five-year horizon.
  - EEUR shows large improvements reflecting large output declines and forecasting errors in the early 1990s after the Soviet Union collapse.
  - Program episodes: RMSE changes similar to EME and LIC groups; close to 1 percentage point through the four-year horizon and less than half a percentage point in the five-year horizon.
- Statistical significance:
  - Increases in accuracy tend to be mostly statistically significant (especially within LAC, MENAP_CIS, and SSA).
  - Worsenings in accuracy are often statistically insignificant.
  - Virtually no instances where increases in RMSE-values for current- and next-year forecasts between subsamples are statistically significant.
- Effect of including 2009:
  - For the World sample, accuracy for horizons up to four years improves for most countries in 2004–17 even if 2009 errors are kept.
  - AE is an exception: median difference is negative for the one-year ahead Spring forecast and longer horizons when 2009 is kept, reflecting large AE errors in 2009.

Overall conclusion on time changes:
- Small but notable improvements in predictive accuracy of short-term WEO GDP growth forecasts for a clear majority of countries in 2004–17 relative to 1990–2003.
- Slightly more than half of countries show improvements for longer horizons.
- Evidence points to improved ability to forecast growth over one and a half years, with mixed results for longer horizons.

### Adjusting for underlying variability: Theil U-statistic
- To account for how difficult outcomes were to predict, authors compute a Theil U-statistic by scaling the mean squared error of WEO forecasts by the variance of the predicted variable, where the variance uses a recursively-updated historical average (prevailing mean) known in real time.
- U-statistic interpretation:
  - Values below unity: WEO forecasts more accurate than the historical average (naïve benchmark).
  - Values above unity: historical average is more accurate than WEO forecasts.
- The U-statistic provides a fairer cross-country comparison by accounting for differing underlying growth variability.
- Figure 4 (referenced) shows U-statistic values over 2004–17: at the shortest two horizons (h = 0, F) and (h = 0, S), U-statistic values are small—below about [value cutoff referenced in source].

*Source: wpiea2021216-print-pdf - 5. How does the accuracy of WEO forecasts compare to those of CE forecasts?*

### 0.4 and 0.6 respectively, for three quarters of economies globally. WEO forecasts thus appear

### wpiea2021216-print-pdf - 0.4 and 0.6 respectively, for three quarters of economies globally. WEO forecasts thus appear

### WEO forecasts versus historical average; short- versus long-horizon performance
- WEO forecasts incorporate valuable information during the current year that facilitates substantially more accurate forecasting than simply using the historical average of growth outturns.
- Exact short-horizon performance indicators cited:
  - "0.4 and 0.6 respectively, for three quarters of economies globally."
- As the forecast horizon expands, the ability of the WEO forecasts to dominate the historical average deteriorates.
- Much of the worsening in the Theil U-statistic generally occurs as the horizon lengthens from three months (h=0, F) to about two years (h=1, S).
- For about a quarter of economies the U-statistic climbs above unity as of the next-year Spring forecast (h=1, S).
- The median U-Statistic for the World sample remains only modestly below unity for the three to five year ahead forecasts, meaning that for close to half of the economies longer term WEO forecasts are less accurate than a recursively-computed mean at these horizons.

### Regional and group patterns in long-horizon underperformance
- Underperformance by horizon and region/group:
  - Advanced Economies (AEs): median reaches unity by the next year Spring forecast.
  - LAC and SSA: median U-statistic is around unity from the three-year ahead forecasts onwards.
  - DASIA: median U-statistic is very close to unity from the next-year Spring forecast onwards.

### Accounting for growth volatility and relative accuracy
- Accounting for growth volatility significantly changes the rankings of forecast accuracy across country groups—suggesting more uniform performance across groups.
- Fuel-exporter dominated groups:
  - EMDE_FE and MENAP_CIS: although absolute forecast accuracy (by RMSE) was the lowest, they score strongest on relative accuracy as reflected by their lower median U-statistics.
  - Interpretation: forecasts of fuel-exporting economies are more accurate than would be predicted by their high degree of GDP growth volatility.

### Changes in relative accuracy over time (1990–2003 vs 2004–16)
- General pattern:
  - Relative accuracy has improved for most countries for horizons of up to two years, with noticeable gains for many countries in the LIC group.
  - Exception: AE group—relative accuracy has declined in most countries in the AE group (except for the current year forecasts, for which about half of the countries have seen improvements and the other half declines).
  - Relative accuracy has declined for the next year and longer-horizon forecasts for most economies in the DASIA group.
  - Relative accuracy declined for the four- and five-year ahead forecasts for the SSA and Program groups.
- Summary: improvements in shorter horizons (with the notable exception of the AE countries) and a mixed record for the longer ones generally holds for relative accuracy.

### Biases in WEO real GDP growth forecasts (definitions and overall patterns)
- Definition:
  - bias_g,t0::t1 = (1/(t1 − t0 + 1)) ∑ e_g,t|t−h (equation (4)).
  - Based on equation (1), positive (negative) values of the bias correspond to under predictions (overpredictions) of growth.
- World sample, 2004–17:
  - No significant tendency for upward or downward bias at the same- and next-year horizons.
  - Tendency for overprediction at the two-year and longer horizons.
  - Median forecast errors:
    - Current-year: small and positive (about 0.1).
    - Next-year: close to zero.
    - Two-year ahead: median forecast error declines to about -0.3 percentage point.
    - Three-to-five year ahead: median forecast error declines to about -0.5 percentage point.
  - "The growth forecasts of one quarter of all countries are biased upward by more than 1 percentage point at the three- to five-year forecast horizons."

### Bias patterns across income and regional groups
- Income groups:
  - AE, EME, LIC: median bias similar for the two-year and longer horizons.
  - LIC group: tendency for optimism also at the shorter horizons:
    - Median bias at about 0.35 percentage point for the next-year spring and fall forecasts.
    - A quarter of LICs having biases in excess of about 0.75-1 percentage point at those two horizons, respectively.
  - Biases more diverse in EME and LIC (wider inter-quartile ranges) than in AEs.
  - GDP-weighted mean of bias is consistently larger than the median, indicating greater tendency for overpredictions for countries with lower GDP levels.
- Fuel exporters vs importers:
  - Median bias among EMDE Fuel Exporters is generally smaller than for EMDE Fuel Importers.
  - Range of biases is wider for Fuel Exporters—especially at short and medium horizons—consistent with larger terms-of-trade shocks and output volatility.
- Regional differences:
  - EEUR and SSA: growth forecasts mostly optimistic at all but the shortest horizons.
  - DASIA: forecasts have often been pessimistic.
  - LAC and MENAP_CIS: slight tendency for underprediction at shorter horizons, turning optimistic at two-years and beyond.
- Program group:
  - Median bias similar to EME and LIC at four- to five-year horizons.
  - Slightly more optimistic at the one- and two-year ahead horizons.

### Statistical significance of biases (country-level)
- Share of countries with statistically significant bias (Table 2 referenced):
  - Current year: growth systematically underpredicted for close to a fifth of World economies; overpredicted for about 3-6 percent of economies.
  - Three- to five-years ahead: underpredicted for about 6 percent of economies; overpredicted for more than a quarter of economies.
  - EEUR countries: about two thirds of countries have statistically significant upward biases for the three- to five-year ahead forecasts.
  - Program countries: shares of statistically significant biases similar to broader EM and LIC groups.
  - Fuel Exporters: relatively high shares (about 10-20 percent) of systematically upward or downward biases in shorter horizons only.
  - Fuel Importers: high shares of upward biased forecasts for the longer horizons (25-30 percent of Fuel Importer EMDEs have statistically significant upward biases for three- to five-year ahead growth).

### Summary judgment on bias across groups and horizons
- Country-level WEO growth forecasts in 2004–17:
  - Modestly downward biased at the current- and next-year horizons (except in the Fuel Exporter and to some extent the LIC and SSA groups).
  - Comparatively stronger tendency for overprediction for the two-year and longer horizons for all groups except the EMDE Fuel Exporter group.

### Formal tests of shifts in the bias (1990–2003 vs 2004–17)
- Method: change in bias computed by subtracting mean forecast error in 2004–17 from that in 1990–2003; interquartile ranges shown in Figure 7.
- Findings:
  - For the clear majority of subgroups and horizons, the median of the difference is negative → lessened tendency for overpredictions in 2004–17.
  - EME and LIC countries: decline in median bias relatively uniform across horizons, broadly in the range of 0.4-1.0 percentage point.
  - AE group: many countries saw reduced tendency for overprediction in one- to two-years ahead forecasts, little change for same-year forecasts, but more upward bias in forecasts at the three-year and longer horizons for a slight majority of countries.
  - For AE, MENAP_CIS, and SSA: median error in forecasting growth four-to-five years out is largely unchanged or larger than in 1990–2003—biases increased for slightly more than half of countries for the four- and five-year ahead forecasts.
  - Program forecasts: optimism declined for 75 percent of cases for same year, next-year, and two-year-ahead horizons, and for more than half of cases for three to five-year forecasts.

### Sensitivity to including 2009 forecast errors (Global Financial Crisis)
- Including 2009 forecast errors:
  - Decline in bias is smaller when 2009 is included.
  - Greatest impact is for Advanced Economies (epicenter of the crisis): once 2009 errors included, median decline in overprediction in AE group becomes very small and confined to current- and next-year forecasts.
  - Including 2009 reduces the decline in bias for EMDE Fuel Exporters.
  - For other groups, change in bias remains broadly similar whether 2009 errors are included or excluded.

_Italic: Source: wpiea2021216-print-pdf - 0.4 and 0.6 respectively, for three quarters of economies globally. WEO forecasts thus appear_

### conclusions  of  improved  accuracy  and  reduced  bias  are  not  driven  by  dropping  outturns  for

### wpiea2021216-print-pdf - conclusions  of  improved  accuracy  and  reduced  bias  are  not  driven  by  dropping  outturns  for

### Bootstrap permutation test of bias changes
- Null hypothesis tested: equality of absolute biases in two subsamples (1990–2003 vs 2004–2016).
- Aggregated (weighted) results summarized in Annex Figure 4:
  - The sum of negative and positive bars corresponds to the change in the weighted mean bias in Figure 7.
  - Darker colored segments = weighted average of statistically significant changes; lighter = statistically insignificant.
  - For all groups, dominant dark red bars indicate that for most horizons many countries have significantly reduced biases.
  - For the World sample, the weighted average of significant declines in overprediction bias outweighs statistically significant increases for all horizons except the current year ones.
  - Clear exceptions: Advanced Economy group (at two-to-five year horizons) and DASIA groups for zero-to-two year horizons.
- Conclusion: Findings confirm statistically significant declines in overprediction bias in many economies, especially the larger ones.

### Serial correlation in forecast errors
- Under squared error loss, forecast errors should be serially uncorrelated when horizons are non-overlapping.
- Estimator: regression with no intercept to compute serial correlation coefficient 휌̂g (equations (6) and (7)), computed over the sample 2004:2017.
- Test: H0: 휌g = 0 versus two-sided alternative.
  - Positive 휌̂g indicates persistence of over- or under-predictions.
  - Negative 휌̂g indicates tendency to reverse errors in consecutive years.
- Empirical results (Figure 8):
  - Presence of serial correlation in Fall current-year and Fall next-year WEO vintages in 15–30 percent of countries except in the AE group.
  - These rejection rates are significantly higher than the 5 percent expected under the null, especially for the LIC group.
- Implication: Forecast errors—and biases—could be reduced if forecasts react more strongly to recently observed errors.

### Local serial correlation
- Motivation: errors may be serially correlated only during certain periods (e.g., Global Financial Crisis).
- Local estimator defined in equation (8) pools short-window covariances across countries in a group to estimate average local serial correlation.
- Findings (Figure 9, world economy):
  - Local serial correlation of errors in forecasting world GDP growth temporarily plunges from positive levels (strongest around 2005) to negative values in the aftermath of the Global Financial Crisis.
  - At all four forecast horizons, local serial correlation estimates increase to positive levels again towards the end of the sample.
  - Pattern indicates overpredictions during the Global Financial Crisis were followed by underpredictions during the recovery, suggesting a tendency to “compensate” with pessimism following large overprediction errors.

### Sources of forecast errors — overview
- WEO forecasting emphasizes integrated, globally consistent projections; analysis focuses on global environment drivers.
- Exercises:
  1. Trace WEO growth forecast errors to errors in predicting external factors—systemically-important economies and terms of trade (TOT).
  2. Test informational efficiency of forecasts using indicators observed at forecasting time (e.g., projected terms of trade, large-country growth forecasts).
  3. Examine whether forecast errors can be predicted by the output gap estimate made in the same round.

### Contemporaneous errors in forecasts of external factors
- Regression specifications (equations (9)-(11)) relate h-step-ahead forecast error in economy g to contemporaneous forecast errors for US, China, and Euro Area.
- Key empirical findings (Table 3):
  - Median pass-through of US same-year forecast error to other countries’ same-year Fall forecasts ≈ 13 percent.
  - Share of countries with positive and statistically significant US spillover: generally 5 to 16 percent; negative and significant: 2-8 percent.
  - For countries with significant coefficients, US forecast errors can explain about a third of the variation in country forecast errors (columns 13-21).
  - Chinese forecast errors: median pass-through positive for all country groups, 27-46 percent globally across horizons (columns 1-4).
    - Positive and significant coefficients for about 30-40 percent of economies globally (columns 5-8).
    - R-squared for regressions with positive coefficients generally between 40 and 50 percent.
    - Chinese forecast errors statistically significant and positive for close to 40 percent of world economies at the five-year horizon.
  - Euro Area forecast errors: median pass-through sizable—50 and 37 percent globally for same- and next-year Fall horizons.
    - Share of countries with positive and significant impact less than China but comparable to US—10-25 percent globally depending on horizon.
    - Pass-through rates around 100 percent for Eastern European countries; strong explanatory power for Advanced Economies.
- Conclusion: More accurate forecasts for Euro Area, United States, and especially China would help improve global growth forecast accuracy directly and indirectly.

### Terms of trade (TOT) forecast errors and growth errors
- TOT forecast error defined in equation (12): actual final vintage minus previous forecast.
- Regression (equation (13)) relates country growth forecast errors to TOT forecast errors.
- Findings (Table 4):
  - Median impact of TOT forecast error on growth forecast errors typically positive.
  - Shares of countries with positive and significant correlations (columns 5-8) typically much larger than those with negative correlations (columns 9-12).
  - EMDE Fuel Exporter, MENAP_CIS, and LAC groups show high shares of positive and significant correlations.
  - Median correlations between TOT and growth forecast errors generally small (always less than 0.05), typically smaller than coefficients on major-economy growth forecast errors.
  - Magnitude of TOT forecast errors larger than those of growth errors for US, China, Euro Area; overall impacts comparable.
  - TOT forecast errors can explain a meaningful fraction (30-50 percent) of growth forecast errors.
  - Up to 56 percent of economies in EMDE Fuel Exporter group have TOT errors as significant predictors; typically 5-15 percent for other EMDE groups.
  - TOT forecast errors more significant at longer horizons (two and five years).

### Are WEO growth forecast errors predictable?
- Tests of informational efficiency use contemporaneous forecasts (equations (14)-(16)) as predictors.
- Results (Table 5):
  - Some degree of inefficiency: for 10-25 percent of world economies a higher US growth forecast predicts a higher growth forecast error (columns 5-8).
  - For China, forecasts are more often overly sensitive: higher shares of significant and negative coefficients in columns 9-12; for almost one third of economies globally, a higher five-year ahead projected growth rate for China has typically coincided with growth overpredictions (column 12).
  - Euro Area forecasts: coefficients more often positive than negative; elevated share of under-reaction at two- and five-year horizons (about 28 and 33 percent of all economies correlated with Euro Area forecast at two- and five-year horizons).
- Terms-of-trade forecasts as predictors (equation (17), Table 6):
  - Terms of trade forecasts significantly predict growth forecast errors for many groups and horizons.
  - Evidence that in many cases forecasts fail to respond strongly enough to TOT outlook (columns 5-8) and in other cases overreact.
  - Accounting for TOT outlook more carefully can improve forecasts of 10-30 percent of world economies (summing specific shares reported in Table 6).

### Output gap and forecast errors
- Regression (equation (18)) relates output gap predicted at time t−h to forecast errors.
- Analysis limited to Advanced Economies due to data sparseness.
- Findings:
  - Strong negative correlation of -0.6 (significant at the one percent level) for the AE group as a whole for Spring next year forecasts (h=1, S).
  - Correlations negative but not significant for same-year forecasts (h=0, F, S) and next-year Fall forecasts (h=1, F).
  - About one third of AE countries have a significant negative correlation between output gap and forecast error individually.
  - Interpretation: for ~one third of AE countries, growth is systematically overpredicted when the economy is estimated to have spare capacity, possibly reflecting assumption that the output gap will close through above-potential growth.

### Comparison with Consensus Economics (CE) forecasts
- CE provides monthly updates for current- and next-year forecasts; vintages labeled m=1..12.
- Pairings used for comparison:
  - CE current-year March (h = 0, m = 3) with WEO current-year Spring (h = 0, S).
  - CE current-year September (h = 0, m = 9) with WEO current-year Fall (h = 0, F).
  - Analogous pairings for next-year forecasts.
- First performance measure: ratio of RMSE values (equation (19)): Uh = RMSE_WEO(h)/RMSE_CE(h).
  - Uh > 0 indicates WEO has higher RMSE than CE; Uh < 0 indicates opposite.
- Formal test for equal predictive accuracy follows Diebold and Mariano (1995) using loss differential (equations (20)-(21)).
- Empirical results (Table 7):
  - Comparing WEO Spring/Fall with CE September/March:
    - For a majority of countries, WEO forecasts generate lower RMSE values than CE for Fall and current-year Spring forecasts.
    - Proportion of countries where WEO RMSE is lower ranges between 56 percent and 66 percent for these horizons.
    - Diebold-Mariano test statistics significant only in a few cases—close to 5 percent.
    - For Spring next-year forecasts, WEO more accurate than CE only about 40 percent of the time.
  - Comparing WEO Spring/Fall with CE April/October:
    - CE forecasts more accurate in the majority of cases, showing small timing/informational advantages of even a few additional weeks of data.

### Main takeaways (from Conclusion)
- Outliers:
  - WEO growth forecasts biased upward partly because unpredicted growth booms are rarer than severe growth collapses (natural disasters, wars, systemic financial crises such as the Global Financial Crisis of 2008-09).
  - Removing unforecastable growth collapses helps assess forecasting performance under normal circumstances.
- Accuracy:
  - In 2004-17, forecast accuracy highest for advanced economies; lower for, and comparable between, emerging-market and low-income economies.
  - Among emerging-market and low-income groups, forecasts more accurate for fuel importers than for fuel-exporters.
  - Predictive accuracy drops as horizon lengthens, but not linearly:
    - Largest reductions moving from three-month horizon (Fall forecasts for current year) to two-year horizon (Spring forecasts for the year ahead).
    - Reductions beyond two years are relatively small; accuracy of two-year and five-year-ahead forecasts not very different.
  - Short-term forecasts more accurate in 2004-17 than in 1990-2003 for well more than half of countries.
  - Longer horizons: accuracy increased for most fuel-importing EMDEs; about half of AEs and fuel-exporter EMDEs saw improved accuracy.
  - Growth forecast errors broadly proportional to volatility of growth; once volatility is accounted for, accuracy similar across country groups in 2004-17.
  - WEO forecasts generally more accurate than a naïve recursively-computed historical average for horizons up to 2-3 years in more than half of countries.
  - For about half of countries in AEs, LAC, SSA, and Developing Asia, naïve forecasts more accurate at horizons longer than 2-3 years.
- Bias:
  - No discernible bias for current-year and next-year (except some optimism in LIC group).
  - Two-year and longer horizons tend to be upward biased.
    - Median overprediction across all countries ≈ 0.3 percentage point for two-year horizon and ≈ 0.5 percentage point for three- to five-year horizons.
  - Median bias in Program group similar to broader EMDE and LIC groups at four- to five-year horizons; slightly higher at one- and two-year horizons.
  - Optimism (overprediction) generally declined: bias lower in 2004-2017 than in 1990-2003 for at least half of countries at all horizons, with larger declines among EMDEs.
- Sources and predictability:
  - Forecast errors for US, Euro Area, and especially China correlated with many other economies’ forecast errors.
  - Errors in predicting TOT contribute to growth forecast errors, especially for commodity exporters.
  - For about 8 percent of AEs and 20-30 percent of EMDEs, current-year growth over- or underpredictions tend to persist—near-term forecasts can be improved by learning from recent errors for about one in ten countries globally.
  - Forecasts for up to a quarter of economies can be improved by more efficiently incorporating information on US growth forecasts; up to 30 percent (especially commodity exporters) by better using country’s own TOT forecasts.
  - Growth systematically overpredicted when output estimated below potential, but less so than in previous evaluation.
- Comparison with Consensus Economics:
  - Spring and Fall current-year and next-year WEO forecasts broadly comparable to CE March and September forecasts.
  - Timing differences explain many small accuracy differences rather than systematic performance gaps.

### Policy recommendations and improvements
- Focus efforts to reduce optimism bias on long-term, multi-year forecasts and in some cases Spring one-year-ahead forecasts.
  - Two- to five-year ahead growth forecasts significantly higher than recent historical averages should be closely scrutinized.
- Modify how output gap information is incorporated to avoid imposing an assumption that the gap will be closed through above-potential growth over the forecast horizon.
- Improve accuracy of TOT forecasts and of growth forecasts for China, Euro Area, and United States to indirectly improve many countries’ forecasts.
  - Better incorporation of information on US growth forecasts can improve forecasts for up to a quarter of economies.
  - Assess uncertainty surrounding TOT and major-economy growth forecasts to evaluate GDP forecast uncertainty.
- Inspect cases where Consensus Economics forecasts dominate WEO forecasts for major economies to identify reasons for weaker WEO performance.
- Reduce serially correlated forecast errors by reacting more strongly to recently observed errors, while avoiding systematic underprediction following large downturns to compensate for past optimism.

*Source: https://www.imf.org/-/media/files/publications/wp/2021/english/wpiea2021216-print-pdf.pdf*

### ANNEX 1. Limiting Outliers in the Forecast Evaluation Exercise

### ANNEX 1. Limiting Outliers in the Forecast Evaluation Exercise

### Overview and starting point
- GDP forecasts from past WEO vintages are used to construct a database with 12 series corresponding to 6 horizons and two bi-annual forecasting exercises (Spring and Fall) for the period 1990-2016.

### Filter one: Conflict
- Conflict indicator constructed using the UCDP Georeferenced Event Dataset:
  - Indicator = 1 in years when 100 or more people are killed due to conflict per 1 million population.
- Exclude years when conflict=1 and adjacent years for countries where economic performance was severely affected:
  - Angola (1998)
  - Georgia (1994)
  - Iraq (1992-1993)
  - Libya (2012-2013)
  - Moldova (1993-1994)
  - Ukraine (2015)
  - Tajikistan (1994)
- Additional exclusion:
  - Sudan: errors for 2011 are excluded given the separation between Sudan and South Sudan.

### Filter two: Disasters
- Using EM-DAT the International Disaster Database:
  - Drop country-year observations where total disaster damage exceeds 3 percent of GDP and growth is negative.
- Ebola exclusions:
  - Drop country-year observations affected by the Ebola epidemic: 2014 and 2015 for Guinea, Sierra Leone, and Liberia.

### Filter three: Ad-Hoc
- Specific ad-hoc exclusions:
  - Ireland: errors for 2015 are dropped due to a jump in GDP that year from the relocation of intellectual property by large multinational companies.
  - Equatorial Guinea: dropped due to large volatility surrounding oil discovery and extraction.
  - Drop year 2009 from the subsample of 2004-2016 because of the Global Financial Crisis.
- Note: the previous filters drop all forecasts made for a given country-year based on the outturn being unforecastable.

### Filter four: Zero forecast errors
- Drop forecast horizons with precisely zero forecast errors as it is highly unlikely that forecasts match outturns exactly.
- Note on outturns: since growth estimates for year t recorded in the Fall WEO of year t+1 are used as outturns, it is possible those outturns are carried-over forecasts rather than actual statistics, resulting in zero forecast errors.

### Filter five: Minimum number of observations in subsamples
- Analysis divided into two subsamples: years 1990-2003 and 2004-2016.
- Drop forecast horizons with observations less than eight in a given period.
- This filter affects the 1990-2003 subsample more than the 2004-2016 subsample.

*Source: ANNEX 1. Limiting Outliers in the Forecast Evaluation Exercise (wpiea2021216-print-pdf).*

---


_Source: https://www.imf.org/-/media/files/publications/wp/2021/english/wpiea2021216-print-pdf.pdf_
