## wpiea2021216app-print-pdf

## Source details

**Canonical URL:** [wpiea2021216app-print-pdf](https://www.imf.org/-/media/files/publications/wp/2021/english/wpiea2021216app-print-pdf.pdf)

## Other formats

- [Markdown version](/-/media/files/publications/wp/2021/english/wpiea2021216app-print-pdf.pdf.md)
- [Structured JSON version](/-/media/files/publications/wp/2021/english/wpiea2021216app-print-pdf.pdf.json)

---

### Forecast data, coverage, and groups
- Forecast frequency and horizons:
  - WEO forecasts reported twice a year in April (Spring) and October (Fall).
  - Forecast horizons: six yearly horizons h = 0 to h = 5 (h = 0 corresponds to current-year outcomes).
  - Within each year there are two issues (Fall, Spring) yielding 12 forecast horizons in total.
  - Current-year fall (h = 0;F) and spring (h = 0;S) forecasts are hybrids of nowcasts and forecasts.
- Sample and timing:
  - Dataset span: 1990 through 2016.
  - Sample sizes: 27 yearly observations for the shortest current-year horizon and 22 observations for the five-year horizon (earliest five-year-ahead forecast reported for 1995).
  - Forecast start dates: earliest observations for current-year forecasts start in 1994; longer horizons begin in 1994 + h.
  - Terms of trade / commodity terms of trade forecasts dataset runs from 2003 to 2016.
- Groups analyzed (G): world; advanced economies (ae); G7; emerging market and developing economies (emde) with subgroups including lics, eeur, dasia, lac, menap, cis, ssa; fuel exporters; Program countries.
- Outcomes and instruments:
  - Outcomes: real GDP growth and consumer price inflation (CPI).
  - Forecast instruments include WEO forecasts of the output gap, terms of trade, and commodity terms of trade (basket of 40 commodities).
- Actual-value convention: use actual value for year t reported in the following year’s (t + 1) Fall WEO issue.

### Measures of accuracy and RMSE definitions
- Main metric: root mean squared forecast error (RMSE).
- Definitions preserved verbatim:
  - h-step-ahead forecast error for country i: e_it|t-h = y_it − ŷ_it|t-h.
  - Country RMSE_i,h = (T − h + 1)^(−1) Σ_{t=h}^T e_it|t-h^2.
  - Group GDP-weighted outcome y_G,t = Σ_{i∈G} ω_i,G,t y_it.
  - Group GDP-weighted forecast ŷ_G,t|t-h = Σ_{i∈G} ω_i,G,t ŷ_it|t-h.
  - Group MSE_G,h = (T − h)^(−1) Σ_{t=h}^T e_G,t|t-h^2.
  - Country weights ω_i,G,t computed from three-year trailing average dollar-denominated GDP: GDP_i,t = 1/3 [GDP_i,t-1 + GDP_i,t-2 + GDP_i,t-3], ω_i,G,t = GDP_i,t / Σ_{j∈G} GDP_j,t.
  - For comparability across horizons Table 1 RMSEs use common sample 1995-2016: gMSE_G,h = 1/22 Σ_{t=1995}^{2016} e_G,t|t-h^2.

### Term structure of GDP growth forecast errors — key statistics and patterns
- World RMSE pattern (top row, Table 1):
  - current-year horizons: 0.30 and 0.54
  - one-year horizons (Fall and Spring): 1.24 and 1.51
  - 2-5 year horizons: stabilize around 1.60
- Cross-group accuracy ordering (higher RMSE = less accurate):
  - Least accurate: Commonwealth of Independent States, Latin America and the Caribbean, emerging and developing Asia and Europe, fuel exporting economies, program countries.
  - More accurate: advanced economies, G7 economies, Sub-Sahara Africa, low income developing countries.
- Specific conclusions for World GDP growth:
  - Current-year forecasts (h = 0;S and h = 0;F) are notably more accurate than longer-horizon forecasts.
  - One-year-ahead world forecasts reported in the Fall of the previous year are more accurate (on the order of a 10-20% reduction in the RMSE) than 2-5 year horizon forecasts.
  - Little evidence of improvements in predictive accuracy between spring one-year-ahead forecasts (h = 1;S) and forecasts up to five years previously.
- Group-specific horizon patterns:
  - Advanced economies and G7: greater precision for current-year Spring and Fall forecasts; term structure flattens at 2-5 year horizons.
  - Emerging market and developing economies, dasia, lac: RMSE increases steadily with forecast horizon.
  - Declines in RMSE at longest horizons (h = 4;F or h = 5;F) observed for lics, cis, ssa in some cases.

### RMSE-ratio (improvement as horizon shrinks) and monotonicity
- Four RMSE ratios computed for adjacent horizon comparisons (definitions preserved verbatim).
- Histograms (Figure 1) over 1990-2016 show:
  - Large improvements, sometimes over 100%, moving from current-year Spring to current-year Fall (h = 0;S → h = 0;F) and from prior-year Fall to current-year Spring (h = 1;F → h = 0;S).
  - Smaller gains and some large losses for one-year-ahead comparisons, though many countries still show substantial gains at longer horizons.
- Monotonicity test (Patton and Timmermann (2012)):
  - Null: E[e^2_{0;F}] ≤ E[e^2_{0;S}] ≤ E[e^2_{1;F}] ≤ E[e^2_{1;S}] ≤ E[e^2_{2;F}] ≤ ... ≤ E[e^2_{5;F}].
  - Only one group rejects the null at the 10% level: Sub-Sahara Africa (p-value = 0.06).
  - At the country level, few countries (including Switzerland) reject monotonicity for GDP growth.

### Time variation and subsample comparisons (1990-2003 vs 2004-2016)
- Subsample RMSE summary for World current-year:
  - h = 0;F: 0.29 versus 0.31 (second subsample versus first subsample)
  - h = 0;S: 0.41 versus 0.60 (second subsample versus first subsample)
- Observations:
  - Marked improvements in current-year forecast accuracy in the second subsample for many groups (lics, eeur, lac, cis, menap, ssa, fuel exporters).
  - One-year-ahead forecasts for the world are less accurate in the second subsample, largely driven by advanced and emerging market and developing economies (impact of Global Financial Crisis).
  - Excluding 2009 materially affects results: one-year-ahead forecasts in 2004-2016 (excluding 2009) are more accurate in the second subsample than in the first subsample in all cases.
- Formal bootstrap permutation tests for h = 0 (10,000 repetitions) find:
  - No instances of significantly negative ΔRMSE (no RMSE values significantly increased in second subsample).
  - Significant improvements in many groups, notably current-year Fall forecasts for eeur, lac, cis, and fuel exporters; current-year Spring forecasts for world, G7, ae, eeur, lac, cis, menap, ssa, and fuel exporters.
  - Excluding 2009 strengthens evidence of significant RMSE reductions.

### Tracking against historical-average benchmark (CSSED)
- Historical-average benchmark defined as ȳ_G,t|t-h = (t − max(h,1))^(−1) Σ_{s=1}^{t − max(h,1)} y_G,s.
- CSSED_G,t,h = Σ_{s=1990}^t [ (e_bmk_G,s|s-h)^2 − (e_WEO_G,s|s-h)^2 ].
- World economy (Figure 2):
  - Current-year forecasts (h = 0;S and h = 0;F): WEO consistently beats historical average by a wide margin.
  - One-year horizon:
    - Spring forecasts (h = 1;S): less accurate than historical average in most years and on average (CSSED trends downward), exception 2004-2009 subsample.
    - Fall forecasts (h = 1;F): slightly more accurate on average, driven entirely by 2003–2009; post-2010 WEO Fall and historical average roughly equally accurate.
- Advanced economies (Figure 3):
  - Current-year forecasts clearly more accurate than the historical mean for almost all years; one-year-ahead performance dominated by 2009 improvements but underperforms historical average in many other years.

### Theil U-statistics — relative accuracy versus prevailing mean
- Theil U definition preserved verbatim:
  - U_G;t;h = (∑_{t=1990+h}^{2016} (y_Gt − ^y_WEO_G;tjt−h)^2) / (∑_{t=1990+h}^{2016} (y_Gt − y_G;tjt−h)^2).
  - Interpretation: U < 1 means WEO better than historical average; U > 1 means historical average better.
- Empirical findings (Table 4 summary):
  - Shortest horizon (h = 0;F): U-values very small — below 0.32 for all groups; around 0.05 for world, advanced, and G7.
  - Short current-year Spring (h = 0;S): U between 0.13 and 0.82.
  - One-year horizons world: U = 0.87 (h = 1;F) and U = 1.27 (h = 1;S).
  - Longer horizons: U-statistic averages 1.4 for the world; U exceeds unity for many groups including advanced economies, G7, dasia, lac, menap, and Program countries.
  - Exceptions where WEO remains more accurate at longer horizons: eeur, cis, ssa, and fuel exporters.
- Conclusion: WEO GDP growth forecasts fail to outperform a simple historical average at longer horizons for a majority of economies.

### Local serial correlation in forecast errors
- Local estimator ^G;t uses four adjacent cross-products and average of five squared forecast errors across NG countries to capture time-local serial correlation.
- Empirical patterns (world, Figure 8):
  - h = 0: estimates generally between −0.1 and 0.3; peak in 2005 above 0.2.
  - h = 1: estimates between 0.1 and 0.3 in first half of sample; plunge to negative in 2011.
  - Local serial correlation increases towards the end of the sample across horizons.
- Interpretation: serial correlation appears episodic; episodes such as the Global Financial Crisis induced sequences of over- or underpredictions.

### Sources of forecast errors — attribution via regressions on instruments
- Objective: regress forecast errors on monitoring instruments to attribute errors to contemporaneous errors or forecasted drivers.
- Major international spillovers (Tables 11–12):
  - Regress h-step-ahead forecast error in economy G on contemporaneous forecast errors for the US, China, and Euro Area.
  - US GDP forecast errors:
    - G7: beta estimates range from 0.49 and 0.51 at the shortest horizons to 0.69 and 0.67 at the one-year horizons.
    - Advanced economies: beta estimates range from 0.42 to 0.67.
    - World GDP: beta estimates range from 0.3 to 0.6.
    - Proportion of countries with statistically significant beta at 5% tends to be much higher than expected at the one-year horizon, often exceeding 50%.
  - Chinese GDP forecast errors:
    - Very large and significant effects for developing and emerging Asia: coefficients between 0.40 and 0.62.
    - Highly significant and large for emde (h = 1), lac (h = 1), and cis (h = 0).
  - Euro Area GDP forecast errors:
    - Emerging and Developing Europe: coefficients exceeding unity and highly significant.
    - Significant results also for world, G7, advanced, and emde.
  - Note: these regressions are attributional (contemporaneous errors), not predictive; little evidence that lagged US forecast errors predict future global forecast errors.
- Output gap (Table 13):
  - Current-year (h = 0): output gap strongly explains current-year forecast errors for G7, advanced, eeur, lac, cis, ssa, and Program countries.
  - One-year-ahead (h = 1): lagged output gap more strongly correlated; majority of groups show negative and significant coefficients (larger output gap associated with larger negative forecast error — overprediction of growth).
  - Country-level proportions: contemporaneous output gap significant for 0.33 and 0.39 of countries (h = 0;F and h = 0;S); rises to 0.50 and 0.51 for h = 1;F and h = 1;S.
  - Notable country examples where output gap predicts one-year-ahead errors: France, Italy, Korea.
- Terms of trade and commodity terms of trade (Tables 14–17):
  - Terms of trade forecast errors significantly related to GDP forecast errors, especially for menap, cis, ssa, and fuel exporters.
  - Commodity terms of trade:
    - World: proportion of significantly positive coefficients around 10% for current-year; increases from about 14% at one-year to 29% at four-year horizon.
    - Country examples for short-term commodity TOT effects: Argentina, Brazil, Saudi Arabia, United Arab Emirates, and many commodity-exporting economies.
  - Using predicted terms of trade as instruments yields fewer positive slopes and more negative slopes (especially for fuel exporters).
- Interpretation: forecast errors in terms of trade and commodity terms of trade are closely related to GDP growth forecast errors, with stronger relations at longer horizons.

### Relative performance versus Consensus Economics (CE)
- Relative RMSE measure: Uh = RMSE(WEO_h) / RMSE(CE_h) − 1.
  - Uh > 0: CE more accurate on average; Uh < 0: WEO more accurate on average.
- Cross-country RMSE ratio results (Table 18):
  - Median RMSE ratios range from −0.05 at h = 0;S to −0.00 at h = 1;S.
  - Proportion of countries where WEO RMSE < CE RMSE ranges from 52% to 69%.
  - Certain countries where WEO worse than CE for majority of horizons: Argentina, Australia, Hungary, Korea, New Zealand, Thailand, Turkmenistan.
  - Countries where WEO notably more accurate: Euro area, Ireland, Italy, Japan, South Africa, Spain, United States.
- Bias and equal-predictive-accuracy tests:
  - CE current-year bias: negative bias for the G7 and advanced economies in second subsample (2004-2016).
  - Diebold-Mariano tests show relatively few significant differences; power is low.
- Forecast encompassing tests:
  - Current-year CE forecasts encompass current-year WEO forecast errors in 18% (Fall) and 25% (Spring) of cases.
  - WEO encompasses CE for current-year in 6% (Fall) and 12% (Spring).
  - Mixed patterns: sometimes both forecasts contain separate valuable information; examples include Denmark, China, UK, US.
- CSSED evolution (Figures 9–12) shows country-specific time patterns (United States, China, Japan, Germany highlighted).

### Efficiency and sources of inflation forecast errors
- Mincer-Zarnowitz efficiency tests (Table 26):
  - Efficient forecast null (intercept = 0 and slope = 1) holds approximately for world, G7, and advanced economies (exception: current-year Fall world inflation).
  - Significant inefficiency for lics, lac, ssa, and Program countries.
  - Country-level rejection rates: null rejected for between 37% (h = 1;F) and 45% (h = 1;S) of countries; United States rejects for three of four shortest horizons.
- Serial correlation in inflation forecast errors:
  - High rejection rates of no serial correlation across many groups; proportion rejecting rises from 32% for h = 0;F to 52% for h = 1;S.
  - Local serial correlation (Figure 15) shows strong positive values at or above 0.4 in 2000 and adjacent years; negative around 2005 at short horizons.
- Instruments and drivers (Tables 28–35):
  - US inflation forecast errors strongly and positively correlated with contemporaneous inflation forecast errors for G7 and advanced economies (coefficients 0.49 to 0.77).
  - Output gap is not a strong predictor for most groups, but positive and significant for dasia and ssa (three of four shortest horizons) and for fuel exporters and Program countries (two of four horizons).
  - Terms of trade:
    - For horizons ≥ one year, proportion with significantly negative coefficients ranges 15%–31%; positive proportions 8%–15%.
    - Negative estimates more prevalent for G7 and advanced economies.
  - Commodity terms of trade:
    - For horizons ≥ one year, negative and significant coefficients for between 38% and 56% of world economies; positive in 8%–14%.
    - Strong relation: overpredicted commodity TOT coincided with underpredicted inflation.
  - Global commodity/fuel/non-fuel prices regressions (Tables 31–33):
    - For horizons ≥ one year, proportion with significantly positive slopes varies from 54% to 64% for the world; negative slopes rare.
    - Advanced, eeur, and dasia show particularly high positive proportions.
  - GDP growth vs inflation forecast errors (Tables 34–35):
    - Overpredictions of GDP growth tend to be associated with underpredictions of inflation; proportion with significantly negative slopes close to 15% on average.
    - Using WEO GDP forecasts as predictors shows positive coefficients for 11%–18% of countries at horizons ≥ one year.

### WEO versus CE for inflation forecasts
- RMSE ratio distribution (Table 36): win rates range from 53% (h = 1;S) to 72% (h = 0;S) in favor of WEO.
- Proportion where WEO significantly more accurate is low but at least as large as reverse at all four horizons.
- Countries where WEO more accurate for at least three out of four horizons: China, France, Germany, United Kingdom, United States.
- Encompassing tests reveal varied directional information shares between WEO and CE across horizons.

### Conclusions — key empirical findings (summary)
- Sample focus and comparative results:
  1. Some evidence that WEO short-term GDP growth forecasts became less biased and more accurate over time across most groups; mixed evidence for inflation forecasts.
  2. WEO current-year and next-year GDP growth forecasts notably more accurate than a naive historical average benchmark; little evidence that 2-5 year forecasts beat the historical-average benchmark.
  3. Errors in WEO forecasts of GDP growth and inflation closely related to errors in forecasts of terms of trade and commodity terms of trade; output gap sometimes helps predict one-year forecast errors.
  4. For between 15% and 30% of countries, errors in WEO terms of trade forecasts are positively correlated with WEO GDP growth forecast errors, strongest at multi-year horizons. Commodity terms of trade errors positively correlated with multi-year GDP growth forecast errors for close to 30% of economies.
  5. Errors in WEO inflation forecasts strongly negatively correlated with terms of trade forecast errors (particularly G7 and advanced economies); even stronger negative correlation with commodity terms of trade errors at horizons > one year.
  6. Accuracy of Spring and Fall current-year and next-year WEO forecasts broadly comparable to Consensus Economics forecasts reported for March and September.

### Policy implications and recommendations
- Improve incorporation and use of output gap information:
  - "Our analysis indicates that the current output gap has some predictive power over GDP growth and inflation, although the effect varies considerably across countries and forecast horizons."
- Place greater emphasis on terms of trade and commodity terms of trade:
  - "Given the close and often significant relation between errors in forecasting GDP growth or inflation on the one hand and errors in predicting individual countries’ terms of trade or their commodity terms of trade, strategies for improving the accuracy of forecasts of these terms of trade measures could have positive spillover effects on forecasts of output growth and prices."
  - "At a minimum, an assessment of the uncertainty surrounding terms of trade and commodity terms of trade forecasts will be helpful in evaluating the uncertainty of GDP growth and inflation forecasts."
- Address country-specific weaknesses versus Consensus Economics:
  - "While the accuracy of the WEO and Consensus Economics forecasts of GDP growth are broadly similar, there are some major economies for which the Consensus Economics forecasts appear to dominate the WEO, including China."
  - "Inspecting the reasons for the slightly worse performance of the WEO GDP growth forecasts for China seems sensible."
- Additional methodological recommendations (Section 10):
  1. Enhance predictive accuracy for long-term, multi-year forecasts (and in some cases the one-year-ahead Spring WEO forecasts).
  2. Explore methods to improve multi-year forecasts, including evaluating direct versus iterative forecast methods and applying a variety of forecast combination methods.

*Source: wpiea2021216app-print-pdf — excerpts from sections 2.1, 3.4, 4.5, 5, 6.1, 7.4, 8, 9, 10, and related appendix descriptions.*

### 2.1  Data and Forecasts of GDP growth and ináation

### 2.1  Data and Forecasts of GDP growth and ináation

### Forecast coverage and horizons
- WEO forecasts reported twice a year in April (Spring) and October (Fall).
- Forecast horizons: six yearly horizons h = 0 to h = 5 (h = 0 corresponds to current-year outcomes).
- Within each year there are two issues (Fall, Spring) yielding 12 forecast horizons in total.
- Current-year fall (h = 0;F) and spring (h = 0;S) forecasts are hybrids of nowcasts and forecasts because part of the target year has been observed (Fall current-year benefits from preliminary data for at least half of the current year).
- Dataset span: 1990 through 2016, giving a sample of 27 yearly observations for the shortest current-year horizon and 22 observations for the five-year horizon (earliest five-year-ahead forecast reported for 1995).
- Forecast start dates: earliest observations for current-year forecasts start in 1994; longer horizons begin in 1994 + h (e.g., five-year-ahead forecasts begin in 1999).
- Terms of trade / commodity terms of trade forecasts dataset runs from 2003 to 2016.

### Groups of economies analyzed
- Main groups (G) used in the study:
  - world
  - advanced economies (ae)
  - G7 economies (G7)
  - Emerging market and developing economics (emde)
  - Emerging market and developing economies: low income developing countries (lics)
  - Emerging and developing Europe (eeur)
  - Emerging and Developing Asia (dasia)
  - Latin America and the Caribbean (lac)
  - Middle East, North Africa, Afghanistan, and Pakistan (menap)
  - Commonwealth of Independent States (cis)
  - Sub-Sahara Africa (ssa)
  - Fuel exporters
  - Program countries
- Forecasting performance is also reported for individual countries in appendices.

### Forecasting instruments and outcomes
- Outcomes examined: real GDP growth and consumer price inflation (CPI).
- Actual value convention: use actual value for year t reported in the following year’s (t + 1) Fall WEO issue.
- Forecast instruments include WEO forecasts of the output gap, terms of trade (based on import/export price projections), and commodity terms of trade (basket of 40 commodities).

### Measures of predictive accuracy
- Main metric: root mean squared forecast error (RMSE).
- Definitions:
  - h-step-ahead forecast error for country i: e_it|t-h = y_it − ŷ_it|t-h.
  - Country RMSE_i,h = (T − h + 1)^(−1) Σ_{t=h}^T e_it|t-h^2.
  - Group GDP-weighted outcome y_G,t = Σ_{i∈G} ω_i,G,t y_it.
  - Group GDP-weighted forecast ŷ_G,t|t-h = Σ_{i∈G} ω_i,G,t ŷ_it|t-h.
  - Group MSE_G,h = (T − h)^(−1) Σ_{t=h}^T e_G,t|t-h^2.
  - Country weights ω_i,G,t computed from three-year trailing average dollar-denominated GDP: GDP_i,t = 1/3 [GDP_i,t-1 + GDP_i,t-2 + GDP_i,t-3], ω_i,G,t = GDP_i,t / Σ_{j∈G} GDP_j,t.
- For comparability across horizons Table 1 RMSEs use common sample 1995-2016: gMSE_G,h = 1/22 Σ_{t=1995}^{2016} e_G,t|t-h^2.

### Term structure of GDP growth forecast errors — key statistics and patterns
- World RMSE pattern (top row, Table 1):
  - current-year horizons: 0.30 and 0.54
  - one-year horizons (Fall and Spring): 1.24 and 1.51
  - 2-5 year horizons: stabilize around 1.60
- Cross-group accuracy ordering (higher RMSE = less accurate):
  - Least accurate: Commonwealth of Independent States, Latin America and the Caribbean, emerging and developing Asia and Europe, fuel exporting economies, program countries.
  - More accurate: advanced economies, G7 economies, Sub-Sahara Africa, low income developing countries.
- Specific conclusions for World GDP growth:
  1. Current-year forecasts (h = 0;S and h = 0;F) are notably more accurate than longer-horizon forecasts (explained by concurrent indicators and partial-year data).
  2. One-year-ahead world forecasts reported in the Fall of the previous year are more accurate (on the order of a 10-20% reduction in the RMSE) than 2-5 year horizon forecasts.
  3. Little evidence of improvements in predictive accuracy between spring one-year-ahead forecasts (h = 1;S) and forecasts up to five years previously, suggesting a limit to improved accuracy beyond the long-run average.
- Group-specific horizon patterns:
  - Advanced economies and G7: notably greater precision for current-year Spring and Fall forecasts; term structure flattens at 2-5 year horizons.
  - Emerging market and developing economies, dasia, lac: RMSE increases steadily with forecast horizon.
  - Declines in RMSE at longest horizons (h = 4;F or h = 5;F) observed for lics, cis, ssa in some cases.

### RMSE ratio analysis (improvement as horizon shrinks)
- Four RMSE ratios computed for adjacent shorter-to-longer horizon comparisons:
  - ΔRMSE_{h=0;F→h=0;S} = (RMSE(h = 0;S) − RMSE(h = 0;F)) / RMSE(h = 0;F)
  - ΔRMSE_{h=0;S→h=1;F} = (RMSE(h = 1;F) − RMSE(h = 0;S)) / RMSE(h = 0;S)
  - ΔRMSE_{h=1;F→h=1;S} = (RMSE(h = 1;S) − RMSE(h = 1;F)) / RMSE(h = 1;F)
  - ΔRMSE_{h=1;S→h=2;F} = (RMSE(h = 2;F) − RMSE(h = 1;S)) / RMSE(h = 1;S)
- Histograms (Figure 1) over 1990-2016 show:
  - Large improvements, sometimes over 100%, moving from current-year Spring to current-year Fall (h = 0;S → h = 0;F) and from prior-year Fall to current-year Spring (h = 1;F → h = 0;S).
  - Smaller gains and some large losses for one-year-ahead comparisons (h = 1;S vs h = 1;F and h = 1;S vs h = 2;F), though many countries still show substantial gains at longer horizons.

### Monotonicity test of term structure (Patton and Timmermann (2012))
- Null hypothesis: E[e^2_{0;F}] ≤ E[e^2_{0;S}] ≤ E[e^2_{1;F}] ≤ E[e^2_{1;S}] ≤ E[e^2_{2;F}] ≤ ... ≤ E[e^2_{5;F}].
- p-values reported in Table 1; low p-values indicate rejection of monotonic increase.
- Only one group rejects the null at the 10% level: Sub-Sahara Africa (p-value = 0.06). For this region RMSE increases up to h = 2;F then declines monotonically at longer horizons (five-year forecasts more accurate than one-year spring forecasts).
- At the country level, only a handful of countries (including one developed economy: Switzerland) reject monotonicity for GDP growth.

### Time variation in forecasting performance — subsample comparisons
- Full sample RMSEs reported for 1995-2016 (Table 1). Subsample analysis uses:
  - First subsample: 1990-2003
  - Second subsample: 2004-2016
- Current-year RMSE comparisons for World:
  - h = 0;F: 0.29 versus 0.31 (second subsample versus first subsample)
  - h = 0;S: 0.41 versus 0.60 (second subsample versus first subsample)
- Observations:
  - Marked improvements in current-year forecast accuracy in the second subsample for many groups, including lics, eeur, lac, cis, menap, ssa, and fuel exporters — notable given that the Global Financial Crisis is included in 2004-2016.
  - One-year-ahead forecasts (h = 1;F and h = 1;S) for the world are less accurate in the second subsample, largely driven by advanced and emerging market and developing economies (impact of Global Financial Crisis).
  - For many other groups (lics, eeur, dasia, lac, cis, ssa) RMSEs are smaller in the second subsample across the four shortest horizons.
- Excluding 2009 (Table 3) materially affects results:
  - One-year-ahead forecasts in 2004-2016 (excluding 2009) are more accurate in the second subsample than in the first subsample in all cases. With few exceptions the same holds at horizons h = 2;F and h = 5;F.

### Formal bootstrap permutation tests of subsample RMSE changes
- Bootstrap procedure for h = 0:
  - Randomly sample without replacement 14 forecast errors from 27 years (1990–2016) to compute RMSE_b^I; remaining 13 errors compute RMSE_b^II.
  - Compute ΔRMSE_b = RMSE_b^I − RMSE_b^II. Repeat 10,000 times; proportion of bootstraps with ΔRMSE_b as large as observed ΔRMSE_data gives test outcome.
- Panel II in Table 2 reports ΔRMSE_data and significance:
  - No instances of significantly negative ΔRMSE (i.e., no RMSE values significantly increased in second subsample).
  - Significant improvements found in many groups, notably current-year Fall forecasts for eeur, lac, cis, and fuel exporters; current-year Spring forecasts for world, G7, ae, eeur, lac, cis, menap, ssa, and fuel exporters.
  - Significant improvements at the one-year horizon for lics and ssa (Spring).
- Excluding 2009 (Panel II, Table 3) strengthens evidence of significant RMSE reductions between subsamples for many groups (e.g., eeur and cis improve by more than two and three percent per year at many horizons).

### Tracking performance against a historical-average benchmark
- Historical-average benchmark forecast for group G at horizon h: ȳ_G,t|t-h = (t − max(h,1))^(−1) Σ_{s=1}^{t − max(h,1)} y_G,s.
- Cumulative sum of squared error differential (CSSED) up to time t:
  - CSSED_G,t,h = Σ_{s=1990}^t [ (e_bmk_G,s|s-h)^2 − (e_WEO_G,s|s-h)^2 ].
  - Positive rising CSSED indicates WEO forecasts outperform the historical-average benchmark; negative declining indicates the benchmark outperforms WEO.
- World economy (Figure 2):
  - Current-year forecasts (h = 0;S and h = 0;F): WEO consistently beats historical average by a wide margin on both cumulative basis and nearly every year (informational advantage).
  - One-year horizon:
    - Spring forecasts (h = 1;S): less accurate than historical average in most years and on average (CSSED trends downward), exception 2004-2009 subsample.
    - Fall forecasts (h = 1;F): slightly more accurate on average, driven entirely by 2003–2009; otherwise, WEO less accurate than historical mean in early sample years; post-2010 WEO Fall and historical average roughly equally accurate.
- Advanced economies (Figure 3):
  - Current-year forecasts clearly more accurate than historical mean for almost all years; especially pronounced in 2009.
  - One-year-ahead Fall forecasts trail historical mean up to 2008, but are far more accurate in 2009.
  - One-year-ahead Spring forecasts also notably more accurate in 2009 but underperform historical average in 2010 and most remaining years; overall Spring one-year-ahead WEO underperforms the historical average in MSE over the full sample.
  - Conclusion: one-year-ahead WEO forecasts for advanced economies fared very well in 2009 but were less accurate than the historical average forecast in most years.

*Source: wpiea2021216app-print-pdf - 2.1  Data and Forecasts of GDP growth and ináation*

### 3.4  TheilU-statistics

### 3.4  TheilU-statistics

### Theil U definition and purpose
- RMSE-values estimate absolute forecast accuracy but do not account for difficulty of prediction across economies with differing variability.
- The variance of the predicted variable is computed relative to a recursively updated historical average (prevailing mean):
  - y_G;tjt−h = (t−max(h;1))^−1 ∑_{ξ=1}^{t−max(h;1)} y_Gξ.
- The TheilU-statistic used is the ratio of MSE of the WEO forecasts to the MSE of the historical mean:
  - U_G;t;h = (∑_{t=1990+h}^{2016} (y_Gt − ^y_WEO_G;tjt−h)^2) / (∑_{t=1990+h}^{2016} (y_Gt − y_G;tjt−h)^2).
- Interpretation:
  - Values of U_G;t;h show the proportion of the variance of the outcome that was predicted at a given horizon.
  - Smaller values indicate WEO forecasts are relatively more accurate than the historical average.
  - Values below unity: WEO forecasts better; values above unity: historical average more accurate.

### Empirical findings on relative accuracy (summary of Table 4 results)
- Sample used: 1990-2016 or the longest subsample available at a given forecast horizon.
- Shortest forecast horizon (h = 0;F):
  - U-values are very small — below 0.32 for all groups of economies.
  - Around 0.05 for the world as a whole, and for the advanced and G7 economies.
- Short current-year Spring horizon (h = 0;S):
  - U-statistic lies between 0.13 and 0.82.
  - Indicates WEO forecasts incorporate valuable current-year information improving on historical average.
- One-year horizons:
  - World: U = 0.87 (h = 1;F) and U = 1.27 (h = 1;S).
    - Suggests prevailing mean dominates one-year-ahead Spring forecast (h = 1;S) but not the Fall issue (h = 1;F).
  - Similar ranking holds for most individual economic groups.
- Longer horizons (averages and exceedances):
  - At longer forecast horizons, U-statistic averages 1.4 for the world.
  - U exceeds unity for advanced economies, G7 countries, emerging and developing Asia, Latin America and the Caribbean, Middle East, North Africa, Afghanistan, Pakistan, and for Program countries.
- Exceptions where WEO remains more accurate at longer horizons:
  - Developing and emerging Europe, Commonwealth of Independent States, Sub-Sahara Africa, and fuel exporters.
- Conclusion:
  - WEO GDP growth forecasts fail to be more accurate than a simple historical average at longer horizons for a majority of economies.

### Connection to subsequent analysis
- Section 4 proceeds to conduct classical tests for bias and efficiency to better understand performance results from Tables 1–4.

*Source: wpiea2021216app-print-pdf - 3.4  TheilU-statistics*

### 4.5  Periods with signiÖcant serial correlation

### 4.5  Periods with signiÖcant serial correlation

### Local estimator of serial correlation
- Definition: the local estimator ^G;t uses four adjacent cross-products and the average of five squared forecast errors across NG countries as in equation (16).
- Purpose: captures serial correlation that is “local in time” by using covariance between e i;t and past (e i;t 1) and future (e i;t+1) forecast errors, scaled by neighboring squared forecast errors to produce a correlation-type measure.
- Rationale: single-country annual outcomes are sparse; pooling across NG countries smooths noisy local estimates derived from only four adjacent observations.

### Empirical patterns (world economy, Figure 8)
- Shortest current-year forecast horizon (h= 0):
  - Local serial correlation estimates generally hover between -0.1 and 0.3.
  - Peak in 2005 at a level above 0.2.
- One-year forecast horizon (h= 1):
  - Local serial correlation estimates between 0.1 and 0.3 in the first half of the sample.
  - Plunge to a negative value in 2011 as overpredictions during the early Global Financial Crisis were followed by underpredictions during the recovery.
- All four forecast horizons:
  - Local serial correlation increases towards the end of the sample.
- Overall conclusion:
  - Evidence of some positive local serial correlation in forecast errors, strongest around 2005, declining to negative values in the aftermath of the Global Financial Crisis.

### Interpretation
- Significant serial correlation appears episodic rather than uniformly persistent across the full sample.
- Episodes such as the Global Financial Crisis induced sequences of over- or underpredictions when forecasting models did not adapt quickly to shocks.

---

### 5  Sources of forecast errors

### Overview
- Objective: identify instruments that explain WEO forecast errors and point to potential sources of suboptimal forecasting performance.
- Approach: regress forecast errors on monitoring instruments (Timmermann and Zhu (2017) framework) to attribute errors to contemporaneous errors or to forecasted drivers (e.g., major-country GDP forecast errors, output gap, terms of trade).

### 5.1  Contemporaneous errors in forecasts of US, Chinese, and EU GDP growth
- Method: regress h-step-ahead forecast error in economy G, e G;t;h, on contemporaneous forecast errors for the US, China, and Euro Area (equations (17)-(19)). Results reported in Panel I-III of Table 11.
- US GDP forecast errors (equation (17)):
  - G7: beta (-estimates) range from 0.49 and 0.51 at the shortest horizons to 0.69 and 0.67 at the one-year horizons.
  - Advanced economies: -estimates range from 0.42 to 0.67.
  - World GDP: -estimates range from 0.3 to 0.6.
  - Proportion of countries with statistically significant beta at 5% level tends to be much higher than expected at the one-year horizon for both advanced economies and G7 countries, often exceeding 50%.
  - Country-level significant relations (Appendix Table A9): notable countries with significant relations include Australia, Canada, France, Germany, Mexico, and the United Kingdom.
- Chinese GDP forecast errors (equation (18), Panel II):
  - Very large and significant effects for developing and emerging Asia: coe¢ cients between 0.40 and 0.62.
  - Highly significant and large for emerging market and developing economies (h= 1), Latin America and the Caribbean (h= 1), and Commonwealth of Independent States (h= 0).
- Euro Area GDP forecast errors (equation (19), Panel III):
  - Emerging and Developing Europe: coe¢ cients exceeding unity and highly significant.
  - Significant results also for World, G7, advanced, and emerging market and developing economies.
  - Significant correlations also found for developing Asia, Latin America and the Caribbean, Commonwealth of Independent States (h= 1), and fuel exporting economies.
- Attribution vs prediction:
  - These regressions use contemporaneous major-economy forecast errors for attribution (not predictive regressions).
  - Little to no evidence that past one-year-lagged US GDP forecast errors predict future forecast errors across economies (only one test statistic significant at 10% with wrong sign).

### 5.1 (continued) — Using forecasts (predictor rather than error)
- Regressions using forecasts (equations (20)-(22), Table 12) — weaker evidence overall:
  - WEO next-year Fall forecasts of US GDP growth (h= 1;F): slope significant for seven economic groups, positive coefficients indicating higher US forecasts associated with larger forecast errors (underpredictions) elsewhere.
  - Forecasts of Chinese GDP growth: significant correlations at one-year horizon outside advanced economies (negative slope coefficients in many cases, indicating higher China forecasts associated with smaller forecast errors — i.e., over-predictions elsewhere).
  - Forecasts of Euro-Area GDP growth: significantly negative correlations with forecast errors for G7 and advanced economies, Commonwealth of Independent States, and Sub-Sahara Africa (small slopes), suggesting high Euro Area forecasts translate into overpredictions for these economies; positive coefficients for Middle East, North Africa, Afghanistan, and Pakistan (h= 1), Commonwealth of Independent States, and fuel exporters.

### 5.2  Output Gap
- Instrument: predicted output gap GAP G;tjt h (equation (23)).
- Current-year (h= 0) results (Panel I, Table 13):
  - Output gap has strong explanatory power over current-year forecast errors for: G7, advanced economies, emerging and developing Europe, Latin America and the Caribbean, Commonwealth of Independent States, Sub-Sahara Africa, and Program countries.
- One-year-ahead (h= 1) results:
  - Lagged output gap more strongly correlated with one-year-ahead forecast errors.
  - For a clear majority of groups, the coefficient on the output gap is negative and statistically significant, indicating a larger output gap is associated with a larger negative forecast error (stronger tendency to overpredict GDP growth).
  - Rejection rates at 5% level:
    - Current-year regressions (h= 0): more than 30% of countries reject null that output gap is not correlated with one-year-ahead forecast error.
    - One-year-ahead predictive regressions (h= 1): rejection rates close to 50%.
  - Country-level proportions (Table A10):
    - Proportion of countries with contemporaneous output gap significantly correlated with GDP growth: 0.33 and 0.39 for h= 0;F and h= 0;S.
    - Proportion rises to 0.50 and 0.51 for h= 1;F and h= 1;S.
    - Notable countries where output gap predicts one-year-ahead errors: France, Italy, Korea.
- Two-year-ahead Fall forecasts (h= 2;F, Panel II):
  - Output gap is significantly negatively correlated with subsequent two-year-ahead GDP growth for most groups.
- Horizons longer than two years:
  - Regressions feasible only for World, G7, advanced economies, and program countries.
  - No evidence of significant relationship for World, G7, advanced groups.
  - For program countries, some evidence that output gap predicts forecast errors at long horizons.

### 5.3  Terms of trade and commodity terms of trade forecasts
- Notation:
  - tot it: term of trade for country i in year t; c tot itjt h: h-year-ahead forecast.
  - ctot it: commodity terms of trade; d ctot itjt h: h-year-ahead forecast.
  - Forecast errors defined in final vintage as e tot itjt h = tot it   c tot itjt h and e ctot itjt h = ctot it   d ctot itjt h (equations (24)-(25)).
- Regressions of GDP growth forecast errors on terms of trade forecast errors (equation (26)) and commodity terms of trade errors (equation (27)).
- Terms of trade results (Table 14):
  - For horizons beyond current year and for the world as a whole, proportion of countries with significantly positive i;h varies between 0.14 and 0.31.
  - Proportion with significantly negative i;h varies between 0.07 and 0.14.
  - Proportion of significant i;h tends to be smaller for current-year forecasts (takes time for terms of trade to affect output).
  - Groups where terms of trade most important: Middle East, North Africa, Afghanistan, and Pakistan; Commonwealth of Independent States; Sub-Sahara Africa; fuel exporters.
  - Advanced economies and G7: generally less impacted by terms of trade forecast errors.
  - Country-level examples (Table A11): short-term terms-of-trade forecast errors significantly correlated with GDP forecast errors for Australia, Brazil, Chile, United Kingdom, and many African economies.
- Commodity terms of trade results (Table 15):
  - World: proportion of significantly positive i;h around 10% for current-year forecasts; increases from about 14% at one-year horizon to 29% at four-year horizon.
  - For longer horizons, positive proportions are two to three times larger than proportions with significantly negative i;h.
  - Proportion of significant positive i;h varies across horizons and groups but smaller for horizons up to one year for G7 and advanced economies.
  - Country-level examples (Table A11): short-term commodity terms-of-trade forecast errors significantly correlated with GDP forecast errors for Argentina, Brazil, Saudi Arabia, United Arab Emirates, and many commodity-exporting economies.
- Using predicted (forecasted) terms of trade as instruments (equations (28)-(29), Tables 16-17):
  - Fewer cases with significantly positive slopes compared to using errors in predicting terms of trade.
  - Larger overall proportion of countries for which slope on predicted terms of trade or commodity terms of trade is significantly negative compared to using forecast errors, particularly strong among fuel exporting countries.
- Interpretation:
  - Forecast errors in predicting terms of trade and commodity terms of trade are closely related to GDP growth forecast errors, with stronger relations at longer horizons.

---

*Source: wpiea2021216app-print-pdf - 4.5  Periods with signiÖcant serial correlation (excerpts).*

### 6.1  Relative forecasting performance

### 6.1  Relative forecasting performance

### Relative RMSE measure and interpretation
- The Örst performance measure is the ratio of RMSE values: Uh = RMSE(WEO_h) / RMSE(CE_h) - 1 (equation (30)).
- Uh > 0: RMSE of WEO forecasts exceeds RMSE of CE forecasts (CE more accurate on average).
- Uh < 0: RMSE of WEO forecasts is lower (WEO more accurate on average).
- The magnitude of Uh quantiÖes relative performance.

### Cross-country RMSE ratio results (Table 18)
- Percentile summary: 25th, 50th (median), and 75th percentiles of cross-sectional distribution of RMSE ratios computed across countries covered by both CE and WEO.
- Median RMSE ratios range from -0.05 at the longest current-year horizon (h= 0;S) to -0.00 at the one-year horizon (h= 1;S), indicating a small majority of countries where WEO forecasts generate lower RMSE than CE.
- The proportion of countries for which WEO RMSE < CE RMSE ranges from 52% to 69%.
- For three of four horizons the 25th percentiles are equally or more negative than corresponding 75th percentiles are positive, suggesting percentage improvements in WEO precision are slightly larger in magnitude than percentage reductions where CE is more accurate.

### Country-level highlights
- WEO forecasts worse than CE for majority of horizons: Argentina, Australia, Hungary, Korea, New Zealand, Thailand, Turkmenistan.
- WEO forecasts notably more accurate than CE for majority of horizons: Euro area, Ireland, Italy, Japan, South Africa, Spain, United States.

### Bias evidence for CE forecasts (Table 19)
- Biases in CE GDP growth forecasts assessed using same groupings and GDP-weighting as WEO analysis.
- Current-year forecasts (Panel I): few instances of signiÖcant biases; negative bias for the G7 and advanced economies in the second subsample (2004-2016) indicating CE forecasts were too high on average during this period.
- One-year horizon (Panel II): larger biases; signiÖcant evidence for World (Örst subsample and full sample), G7 and advanced economies, Sub-Saharan economies, and Program countries.

### Tests of equal predictive accuracy (Diebold-Mariano)
- Null hypothesis: equal MSE in expectation (equation (31)).
- Loss di§erential dif_i;tjt-h = (e_CE_i;tjt-h)^2 - (e_WEO_i;tjt-h)^2; regression dif = β_ih + ε (equation (32)); t-test on β_ih.
- Positive and signiÖcant β_ih: CE generates higher MSE than WEO. Negative and signiÖcant β_ih: WEO less accurate than CE.
- Table 18 annotates DM test outcomes with stars. Relatively few significant cases, consistent with weak power given short samples.
- Slightly more countries with signiÖcantly negative β estimates than positive ones; overall proportion of significant cases is very small and close to test size (5%).

### Forecast encompassing tests
- Purpose: test if one forecast "encompasses" the other (i.e., whether one forecast adds information once the other is available).
- Regressions:
  - e_WEO = β_i + γ_i * ŷ_CE + ε (equation (33)). Null γ_i = 0: CE forecasts add no value to WEO.
  - e_CE = β_i + γ_i * ŷ_WEO + ε (equation (34)). SigniÖcant γ_i: CE fails to encompass WEO.
- Regressions estimated for four shortest common horizons: pairs (e_WEO_0;S vs ŷ_CE_0;M3), (e_WEO_0;F vs ŷ_CE_0;M9), (e_WEO_1;S vs ŷ_CE_1;M3), (e_WEO_1;F vs ŷ_CE_1;M9) and vice versa.
- Power is low; null often fails to be rejected.
- Proportions where current-year CE forecasts encompass current-year WEO forecast errors:
  - Fall forecasts: 18%
  - Spring forecasts: 25%
  - Versus WEO encompassing CE for current-year: 6% and 12% (Fall and Spring respectively).
- Next-year forecasts: WEO predicts CE errors for slightly higher proportions (10% and 17%) versus CE predicting WEO errors (8% and 8%).
- In many countries both encompassing regressions reject, indicating each forecast contains separate valuable information (example: Denmark current-year GDP growth).
- Examples of one-way encompassing:
  - China: one-year-ahead WEO forecast errors predicted by CE forecasts; CE errors not predicted by WEO.
  - UK: WEO forecasts appear to encompass CE forecasts for two horizons.

### Cumulative sum of squared error differences (CSSED) and evolution over time (equation (35))
- CSSED measures cumulative (e_CE^2 - e_WEO^2) from 1990 to t for four horizons: CSSED_0;S;t, CSSED_0;F;t, CSSED_1;S;t, CSSED_1;F;t.
- Interpretation: positive and rising CSSED indicates WEO more accurate in squared-error sense; negative and declining indicates CE more accurate.
- Economy-speciÖc patterns (Figures 9-12):
  - United States: modestly better current-year WEO performance is stable; one-year improvement due mainly to 2009, with smaller reversal in 2010.
  - China: current-year WEO forecasts generally less accurate than CE except 2004-2008 Spring; one-year horizon shows mixed switching over time.
  - Japan: WEO systematically better for current-year Fall (h= 0;F); WEO underperformed at one-year horizon up to 1999 then improved thereafter.
  - Germany: current-year WEO and CE broadly similar; one-year horizon CE tends to be more accurate most years except 2008.

*Source: wpiea2021216app-print-pdf - 6.1  Relative forecasting performance*

### 7.4  E¢ ciency tests for ináation forecasts

### 7.4  E¢ ciency tests for ináation forecasts

### Efficiency test outcomes (Mincer-Zarnowitz regressions)
- Table 26: Mincer-Zarnowitz regressions of actual ináation on an intercept and ináation forecasts at the four shortest horizons.
- Efficient forecast null hypothesis: intercept = 0 and slope = 1 (per regression in (13)).
- Findings:
  - Holds, to a close approximation, for:
    - the world as a whole,
    - the G7 countries,
    - advanced economies.
  - Notable exception: current-year Fall forecasts of world ináation.
  - Significant evidence of ine¢ ciency at many horizons for:
    - low income developing countries,
    - Latin America and Caribbean,
    - sub-Saharan African economies,
    - Program countries.
  - For these economies, actual ináation is far less correlated with short-term forecasts than the predicted one-to-one relation would imply.

### Country-level rejection rates and US case
- Appendix Table A22: null of e¢ ciency rejected for between 37% (h= 1;F) and 45% (h= 1;S) of countries.
- United States: null rejected for three of the four shortest horizons.

### Serial correlation in forecast errors
- Groups with significant biases largely coincide with groups showing first-order serial correlation in ináation forecast errors.
- Table 27: strong evidence of serial correlation in individual countries' ináation forecast errors for large fractions of:
  - low income developing countries,
  - emerging market and developing economies,
  - Latin America and Caribbean,
  - Commonwealth of Independent States,
  - Sub-Sahara Africa,
  - Program countries.
- Appendix Table A23: proportion of countries rejecting null of no serial correlation across four shortest horizons rises from 32% for h= 0;F to 52% for h= 1;S.
- Figure 15 (local serial correlation using equation (16)):
  - For current-year and next-year forecasts, strong positive serial correlations in forecast errors at or above 0.4 in 2000 and some adjacent years.
  - Serial correlation generally smaller prior to 2000 and turns negative around 2005 at short horizons.

### Sources of ináation forecast errors (overview of regressions in Tables 28–35)
- Table 28 regressions: individual economies' ináation forecast errors on contemporaneous US ináation forecast error (Panel I) and the output gap (Panel II), per equations (17) and (23).

### 8.1 US ináation forecast errors
- Errors in forecasting current-year and next-year US ináation are strongly and significantly positively correlated with contemporaneous ináation forecast errors for:
  - G7 countries and advanced economies, with coefficients ranging from 0.49 to 0.77.
- Correlations also strong for next-year forecasts in emerging and developing Asia and Middle East, North Africa, Afghanistan and Pakistan.
- Appendix Table A24: proportion of cases where US ináation forecast error significantly correlates with another country's ináation forecast error rises from 7% at the shortest horizon (h= 0;F) to 29% at one-year horizons.
- Significant correlations include developed economies closely resembling the US: Canada, France, Germany, Italy, Japan, UK.
- Note: very weak evidence that lagged US ináation forecast errors help predict future forecast errors for other economies.

### 8.2 Output gap
- Panel II of Table 28: output gap is not a particularly strong predictor of short-term ináation forecast errors for most groups.
- Exceptions where output gap is a strong predictor (positive regression slope) at three of the four shortest horizons:
  - emerging and developing Asia,
  - Sub Saharan Africa.
- Output gap also positive and significant for two of four horizons for:
  - fuel exporting economies,
  - Program countries.
- Implication: higher output gaps associated with larger ináation forecast errors (tendency to underpredict ináation) for these economies.
- Heterogeneous direction across countries can cancel out effects at aggregate level.
- World as a whole: output gap is a significant predictor of one-year-ahead ináation forecast errors for around 25% of individual countries.
- Larger proportions for one-year-ahead prediction in:
  - emerging market and developing economies,
  - emerging and developing Europe,
  - Latin America and the Caribbean.
- Appendix Table A25: null that output gap does not matter rejected for around one-third of countries, including China (for three of four horizons) and France (for two of four horizons).

### 8.3 Terms of trade
- Table 29: proportion of countries with significantly positive (Panel I) and significantly negative (Panel II) estimates of ih from regression (26) of ináation forecast errors on terms of trade forecast errors.
- Aggregate (world) and at horizons of one year or longer:
  - proportion with significantly negative ih estimates ranges from 15% to 31%.
  - proportion with significantly positive ih estimates ranges from 8% to 15%.
  - Negative estimates roughly twice as large as positive ones; rejection rates for negative estimates close to four times what chance would suggest at close to 20% on average.
- Particularly high proportion of significantly negative ih estimates for:
  - G7 countries,
  - advanced economies,
  - higher-than-average rejection rates for Latin America and Caribbean.
- Proportion with significantly positive slope somewhat higher for:
  - Middle East, North Africa, Afghanistan, and Pakistan,
  - fuel exporting economies.
- Appendix Table A26 (country level): at one-year horizon rejection rates:
  - negative ih estimates: 0.15 for h= 1;F and 0.20 for h= 1;S,
  - positive ih estimates: 0.08 for h= 1;F and 0.09 for h= 1;S.

### 8.4 Commodity terms of trade
- Regression per equation (27); Table 30 shows far higher rejection rates than for GDP growth or total terms of trade.
- At horizons of one year or longer:
  - negative and statistically significant ih estimates for between 38% and 56% of world economies,
  - compared to only 8-14% with significantly positive ih estimates.
- Interpretation: WEO overpredicted commodity terms of trade tended to coincide with underpredicted ináation; commodity terms of trade coming in higher than expected generally associated with realized ináation being lower than expected.
- Relation particularly strong for:
  - advanced economies,
  - G7 countries,
  - emerging and developing Europe,
  - Latin America and the Caribbean.
- Country-level evidence at one-year horizon:
  - proportions with significantly negative ih coefficients equal 0.38 (h= 1;F) and 0.46 (h= 1;S),
  - vs. 0.08 for countries with significantly positive ih coefficients.

### 8.5 Global commodity, fuel and non-fuel prices
- Table 31: regressions of country-level ináation forecast errors on contemporaneous global commodity price forecast errors per equation (36).
- For horizons of one year or longer, world as a whole:
  - proportion with significantly positive slope estimates varies from 54% to 64%.
  - proportion with significantly negative slope estimates varies between 0% and 11%.
- High proportions of significantly positive coe¢ cients in:
  - advanced economies,
  - emerging and developing Europe,
  - emerging and developing Asia.
- Conclusion: over- and underpredictions of global commodity prices generally map into similar over- and underpredictions of ináation in the same countries.
- Table 32 (global fuel prices) and Table 33 (global non-fuel prices): similar findings; proportion of countries with significantly positive coe¢ cients similar to Table 31 and few with significant negative coe¢ cients.

### 8.6 GDP growth and its relation to ináation forecast errors
- Table 34: regressions of ináation forecast errors on WEO GDP growth forecast errors per equation (37).
- World averages:
  - proportion with significantly positive slope coe¢ cients (Panel I) around the test size (5%).
  - proportion with significantly negative slope coe¢ cients (Panel II) substantially higher, close to 15% on average across horizons.
- Proportion with significantly negative coe¢ cients particularly high for:
  - low income countries,
  - developing Asia,
  - markedly lower among G7 economies.
- Interpretation: tendency for overpredictions of GDP growth to be associated with underpredictions of ináation and vice versa.
- Regressions using WEO forecasts of GDP growth per equation (38) (Table 35):
  - At horizons one year or longer, significantly positive coe¢ cients for between 11% and 18% of countries (Panel I), strong for emerging Europe and Middle East, North Africa, Afghanistan, Pakistan.
  - Proportion with significantly negative coe¢ cients mostly close to 10% (Panel II), notably lower among G7 and advanced economies.

### 9 Comparison with Consensus Economics ináation forecasts
- Table 36: distribution of RMSE ratios U_h = RMSE(WEO_h)/RMSE(CE_h) - 1 per equation (39).
- Negative U_h: WEO more accurate on average than CE.
- Win rates (proportion of countries where WEO more accurate, even if not significantly so):
  - range from 53% (h= 1;S) to 72% (h= 0;S).
- Proportion of countries where WEO ináation forecasts are significantly more accurate than CE is low but ≥ proportion where the reverse holds at all four horizons.
- Countries where WEO more accurate than CE for at least three out of four horizons include: China, France, Germany, United Kingdom, United States.
- Forecast encompassing tests (Appendix Table A30):
  - Proportion of countries where CE forecasts can significantly predict WEO forecast errors but not vice versa: 25% (h= 0;F) and 32% (h= 0;S).
  - Proportion where WEO forecasts can significantly predict CE forecast errors but not vice versa: 8% (h= 0;F) and 4% (h= 0;S).
  - At one-year-ahead horizon:
    - CE predicts WEO errors but not vice versa: 11% (h= 1;F) and 19% (h= 1;S).
    - WEO predicts CE errors but not vice versa: 18% (h= 1;F) and 10% (h= 1;S).
  - For the two shortest horizons (h= 0;F and h= 0;S):
    - CE predicts WEO errors but not vice versa: 26% and 22%, respectively.
    - WEO predicts CE errors but not vice versa: 44% and 50%, respectively.

### 10 Conclusion — key findings
- Sample focused on 2004-2016 with comparison to 1990-2003; several findings:
  1. Some evidence that WEO short-term GDP growth forecasts became less biased and more accurate over time across most groups; mixed evidence for ináation forecasts.
  2. WEO current-year and next-year GDP growth forecasts notably more accurate than a naive historical average benchmark; little evidence that 2-5 year forecasts beat the historical average benchmark.
  3. Errors in WEO forecasts of GDP growth and ináation closely related to errors in forecasts of terms of trade and commodity terms of trade; output gap sometimes helps predict one-year forecast errors.
  4. For between 15% and 30% of countries, errors in WEO terms of trade forecasts are positively correlated with WEO GDP growth forecast errors, strongest at multi-year horizons. Commodity terms of trade errors positively correlated with multi-year GDP growth forecast errors for close to 30% of economies.
  5. Errors in WEO ináation forecasts strongly negatively correlated with terms of trade forecast errors (particularly G7 and advanced economies); even stronger negative correlation with commodity terms of trade errors at horizons > one year.
  6. Accuracy of Spring and Fall current-year and next-year WEO forecasts broadly comparable to Consensus Economics forecasts reported for March and September.

### Policy implications and recommendations
- Points to consider to improve future WEO forecast performance:
  1. Enhance predictive accuracy for long-term, multi-year forecasts (and in some cases the one-year-ahead Spring WEO forecasts). Evidence of weak gains in predictive accuracy beyond next-year Fall WEO forecasts suggests further efforts.
  2. Explore methods to improve multi-year forecasts, including:
     - evaluating direct versus iterative forecast methods,
     - applying a variety of forecast combination methods.

*Source: wpiea2021216app-print-pdf - 7.4  E¢ ciency tests for ináation forecasts*

### 2. Changes to the way information on the output gap is incorporated into

### 2. Changes to the way information on the output gap is incorporated into

### Predictive power of the current output gap
- "Our analysis indicates that the current output gap has some predictive power over GDP growth and ináation, although the e§ect varies considerably across countries and forecast horizons."

### Terms of trade and forecast spillovers
- "Given the close and often signiÖcant relation between errors in forecast-ing GDP growth or ináation on the one hand and errors in predicting individual countriesí terms of trade or their commodity terms of trade, strategies for improving the accuracy of forecasts of these terms of trade 42 measures could have positive spillover e§ects on forecasts of output growth and prices."
- "At a minimum, an assessment of the uncertainty surrounding terms of trade and commodity terms of trade forecasts will be helpful in evaluating the uncertainty of GDP growth and ináation forecasts."

### Comparing WEO and Consensus Economics forecasts
- "While the accuracy of the WEO and Consensus Economics forecasts of GDP growth are broadly similar, there are some major economies for which the Consensus Economics forecasts appear to dominate the WEO, including China."
- "Inspecting the reasons for the slightly worse performance of the WEO GDP growth forecasts for China seems sensible."

*Source: wpiea2021216app-print-pdf - 2. Changes to the way information on the output gap is incorporated into*

### References

### References

### Bibliography
- [1] Artis, Michael J., 1988, How Accurate Is the World Economic Outlook? A Post Mortem on Short-Term Forecasting at the International Monetary Fund. Sta§ Studies for the World Economic Outlook (Washington, International Monetary Fund), 1ñ49.
- [2] Artis, Michael J., 1997, How Accurate Are the WEOís Short-Term Forecasts? An Examination of the World Economic Outlook. Sta§ Studies for the World Economic Outlook (Washington, International Monetary Fund).
- [3] Barrionuevo, J.M., 1993, How Accurate Are the World Economic Outlook Projections? Sta§ Studies for the World Economic Outlook (Washington, International Monetary Fund).
- [4] Chong, Y. Y., and D. F. Hendry. 1986. Econometric evaluation of linear macro-economic models. Review of Economic Studies 53:671ñ90
- [5] Diebold, F. X., and R. S. Mariano. 1995. Comparing predictive accuracy. Journal of Business & Economic Statistics 13:253ñ63.
- [6] Elliott, Graham, and Allan Timmermann, 2016. Economic Forecasting. Princeton University Press.
- [7] Patton, Andrew J., and Allan Timmermann, 2010, Monotonicity in Asset Returns: New Tests with Applications to the Term Structure, the CAPM, and Portfolio Sorts. Journal of Financial Economics 98, 605-625.
- [8] Rossi, Barbara, 2013, Advances in Forecasting under Instability. Chapter 21, pages 1203-1324 in G. Elliott and A. Timmermann (eds.) Handbook of Economic Forecasting vol 2B. North-Holland.
- [9] Timmermann, Allan, 2007, An Evaluation of the World Economic Outlook Forecasts. IMF Sta§ Papers vol 54, 1, 1-33.
- [10] Timmermann, Allan, and Yinchu Zhu, 2017, Monitoring Forecasting Performance. Unpublished manuscript, UCSD and University of Oregon.

### Appendix: Results for Individual Countries — Inventory of Tables A1–A30
- General note: Tables A1-A30 present country-level statistics on the performance of the WEO and/or Consensus Economics forecasts of real GDP growth and inflation in the sample. The appendix organizes results by forecast horizon, test type, and comparisons (full sample 1990-2016 and sub-samples 1990-2003 and 2004-2016).
- GDP Growth Forecasts (Tables A1–A15; descriptions preserved verbatim):
  - Table A1: full-sample (1995-2016) RMSE values for the four shortest forecast horizons (h= 0;F, h= 0;S, h= 1;F, h= 1;S) and the Fall forecasts for the two through Öve-year horizons (h= 2;F, h= 3;F, h= 4;F, and h= 5;F) along with p-values from the Patton-Timmermann (2010) test for a monotonically (weakly) rising term structure of MSE values against the alternative of a non-monotonic term structure of mean squared GDP growth forecast errors.
  - Table A2: sub-sample comparison of RMSE values across the four shortest forecast horizons, using the 1990-2003 and 2004-2016 sample split adopted in the main analysis.
  - Table A3: Theil U-statistics comparing WEO forecasts to the historical average GDP growth forecasts for the four shortest forecast horizons and for the Fall forecasts of outcomes two through Öve years ahead in time.
  - Table A4: estimates of biases in individual countriesí current-year (h= 0;F and h= 0;S) forecast errors for the two sub-samples (1990-2003 and 2004-2016) as well as for the full sample (1990-2016). The top line ("Proportion Sign.") shows the proportion of individual countries for which the null of a zero bias in the forecast error is rejected at the 5% level, using a two-sided test.
  - Table A5: structured the same way as Table A4, but adopts a one-year forecast horizon (h= 1;F and h= 1;S).
  - Table A6: presents full-sample bias estimates for forecast horizons ranging from two years (h= 2;S and h= 2;F) through Öve years (h= 5;S and h= 5;F).
  - Table A7: results of efficiency tests using Mincer-Zarnowitz regressions for the four shortest current-year and next-year forecasts of GDP growth in the individual countries.
  - Table A8: estimates of Örst-order serial correlation in the errors in WEO forecasts of individual countriesí GDP growth for the four shortest forecast horizons.
  - Tables A9–A12: regression slopes from projections of individual countriesí forecast errors on (i) the contemporaneous error in the US GDP growth forecast (Table A9); (ii) the output gap (Table A10); (iv) Errors in WEO forecasts of terms of trade (first four columns in Table A11) or their commodity terms of trade (columns 5-8 in Table A11).
  - Tables A12–A14: tests applied to the Consensus Economics forecasts of GDP growth in individual countries: Table A12 reports the estimated sample bias in the Consensus Economics forecasts of GDP growth in individual countries forecasts for both the second sub-sample (2004-2016) and for the full sample (1990-2016). Table A13 reports estimates of first-order serial correlation in the errors associated with the Consensus Economicsí GDP growth forecasts, again computed for the individual countries. Table A14 shows Mincer-Zarnowitz tests of forecast efficiency for Consensus Economics forecasts of GDP growth in individual countries.
  - Table A15: estimates from forecast encompassing regressions that compare the WEO forecasts to the Consensus Economics forecasts of GDP growth.
- Inflation Rate Forecasts (Tables A16–A30; descriptions preserved verbatim):
  - Table A16: full-sample (1990-2016) RMSE values for the four shortest forecast horizons (h= 0;F, h= 0;S, h= 1;F, h= 1;S) and the Fall forecasts for the two through Öve-year horizons (h= 2;F, h= 3;F, h= 4;F, and h= 5;F) along with p-values from the Patton-Timmermann (2010) test for a monotonically rising term structure of MSE values.
  - Table A17: sub-sample comparison of RMSE values across the four shortest forecast horizons, using the 1990-2003 and 2004-2016 sample split adopted in the main analysis.
  - Table A18: Theil U-statistics comparing WEO inflation forecasts to random walk inflation forecasts for the four shortest forecast horizons and for the Fall forecasts of outcomes two- through Öve years ahead in time.
  - Table A19: estimates of biases in individual countriesí current-year (h= 0;F and h= 0;S) forecast errors for the two sub-samples (1990-2003 and 2004-2016) as well as for the full sample (1990-2016).
  - Table A20: structured the same way as Table A19, but for next-year forecasts (h= 1;F and h= 1;S).
  - Table A21: full-sample bias estimates for forecast horizons ranging from two years (h= 2;F) through Öve years (h= 5;F).
  - Table A22: efficiency tests for the WEO inflation forecasts using Mincer-Zarnowitz regressions for the four shortest current-year and next-year forecast horizons.
  - Table A23: estimates of first-order serial correlation in individual countriesí inflation forecast errors for the four shortest forecast horizons.
  - Tables A24–A26: regression slopes from projections of individual countriesí inflation forecast errors on (i) the contemporaneous error in the US inflation forecast (Table A24); (ii) individual countriesí contemporaneous output gap (Table A25); (iii) Errors in WEO forecasts of the countriesí terms of trade (first four columns in Table A26) or their commodity terms of trade (columns 5-8 in Table A26).
  - Tables A27–A29: tests applied to the Consensus Economics forecasts of inflation in individual countries. Specifically, Table A27 reports the bias in the Consensus Economics forecasts of inflation in individual countries in the second sub-sample (2004-2016) and for the full sample (1990-2016). Table A28 reports estimates of first-order serial correlation in Consensus Economicsí inflation forecast errors, again computed for the individual countries. Table A29 shows Mincer-Zarnowitz tests of forecast efficiency for Consensus Economics forecasts of inflation in individual countries.
  - Table A30: estimates from forecast encompassing regressions that compare the WEO inflation forecasts to the forecasts produced by Consensus Economics.

*Source: References and Appendix (Tables A1–A30) as provided in the content unit.*

---


_Source: https://www.imf.org/-/media/files/publications/wp/2021/english/wpiea2021216app-print-pdf.pdf_
