## _wp11254

## Source details

**Canonical URL:** [_wp11254](https://www.imf.org/-/media/websites/imf/imported-full-text-pdf/external/pubs/ft/wp/2011/_wp11254.pdf)

## Other formats

- [Markdown version](/-/media/websites/imf/imported-full-text-pdf/external/pubs/ft/wp/2011/_wp11254.pdf.md)
- [Structured JSON version](/-/media/websites/imf/imported-full-text-pdf/external/pubs/ft/wp/2011/_wp11254.pdf.json)

---

### Introduction and motivation
- Futures prices performed poorly as forecasters during the recent commodity price cycle (Figure 1; Bloomberg L.P. data).
- Three stated reasons for revisiting forecasting by futures:
  - Futures did a poor job during the recent cycle and it is natural to ask whether we can do better.
  - The issue has not been decisively settled; enhancing measurement of futures prices, updating the sample period, and broadening coverage may bring us closer to a definitive answer.
  - Forecasting commodity prices is important and often costly for policymakers in countries where commodity prices significantly affect the terms of trade, inflation, and poverty levels; fitting structural or reduced-form models or applying informed judgment for a wide range of commodities can be costly and may add no more value than extrapolating from the current price.

### Four contributions of the paper
- Precise measurement:
  - Provide a careful measure of futures prices that exactly matches the horizon of the subsequent change in spot prices and addresses problems posed by illiquid long-dated contracts.
- Horizon coverage:
  - Compare futures and other candidate models for useful forecasts at horizons stretching out two years; note policymakers’ relevant horizons are typically longer than the standard 3 to 12 months.
- Updated and broadened assessment:
  - Update and broaden assessment including years since 2002 when commodity market liquidity increased; include a broad range of commodities.
- Conditional performance:
  - Assess forecasting ability conditional on futures curve shape (backwardation and contango) and spot price trends (bull and bear markets) to test market efficiency and the contention that financial investors impeded futures price discovery.

### Structure of the paper
- Section 2: models and intuition.
- Section 3: data.
- Section 4: empirical results and discussion.
- Section 5: concluding remarks.

### Key methodological points (model and test specification)
- Futures pricing and forecast tests:
  - Common unbiasedness/efficiency regression: s_{t+k} − s_t = α + β (f_{t,t+k} − s_t) + ε_{t+k}; unbiased forecast requires α = 0, β = 1.
  - Three in-sample notions of rationality tested: (1) weak (no persistent in-sample prediction errors; cointegration), (2) weak-form efficiency (current futures incorporate useful prediction information; ECM tests), (3) unbiasedness (weak-form efficiency plus no risk premium).
  - For I(1) series, Engle-Granger cointegration tests applied; Newey-West HAC standard errors with bandwidth = days to maturity − 1 to control overlapping observations.
- Out-of-sample evaluation:
  - Benchmark: random walk without drift, s_{t+1} = s_t + ε_{t+1}.
  - Forecast metric: mean squared forecast error (MSFE); Diebold-Mariano (DM) test used for equality of MSFEs; Clarke and West (2005) adjusted MSFE test reported for nested models.
  - Conditional tests: regress loss-differential d_t = ε_{s,t+k}^2 − ε_{f,t+k}^2 on dummies for backwardation and bull markets to test whether futures perform relatively better in backwardation or worse in bull markets.
- Candidate forecasting models compared (structure preserved): random walk, ARIMA (1,1,1), ARMA (1,1), W-ARIMA (weekly), Holt-Winters exponential smoother, futures price (and variants with risk premium, basis, ECM, levels with lags).

### Data (sample and construction)
- Sample period: January 1990 to June 2011.
- Sampling frequency: weekly.
- Futures price source: Bloomberg.
- Futures contracts: first 24 contracts ordered by days to delivery (for many commodities curve stretches two years; some commodities shorter or longer depending on contract spacing).
- Commodities covered (spot and futures contract specifications preserved in source): Aluminum, Copper, Corn, Cotton, Crude Oil (WTI), Gasoline, Gold, Natural Gas, Soy (Soybeans), Wheat.
- Fitted futures curve method:
  - Weekly third-order polynomial regression of futures prices on days to delivery to produce a continuous curve and extract implied futures prices at exact horizons t1 = 0, 91, 182, 364, and 728 days (implied spot = γ0).
  - Average R-squared across all commodities for polynomial fits: 0.8.
  - Fit weaker (average R-squared < 0.8) in cases with seasonal curve shapes (corn, wheat) or illiquid long-dated quotes (gold).

### Empirical results — in-sample rationality
- Average R-squared statistics for ECM/cointegration equations (average r-square for all equations per commodity):
  - Aluminum 0.96
  - Copper 0.92
  - Corn 0.80
  - Cotton 0.89
  - Gasoline 0.75
  - Gold 0.72
  - Natural gas 0.54
  - Soybeans 0.72
  - Wheat 0.78
  - WTI crude oil 0.93
- Unit root and cointegration:
  - Most spot and futures prices are nonstationary (I(1)) over Jan-1990 to Jun-2011; Engle-Granger tests reject φ = 0 in equation (13) in all I(1) cases—futures and realized spot are cointegrated.
- Efficiency and unbiasedness tests (Wald test p-value summary):
  - Efficient market hypothesis rejected for most commodities and most horizons.
  - At the 91 and 182 day horizons, less than half of commodities are efficient at the 5 percent significance level.
  - Efficiency can be ruled out for all commodities at the one- and two-year horizons except for cotton.
- Lagged explanatory power:
  - Lagged spot and futures often provide significant in-sample explanatory power, especially at one- and two-year horizons; Bayesian selection often chose long lag lengths (often five or six periods).

### Empirical results — out-of-sample forecasting
- Forecasting difficulty and RMSEs:
  - RMSEs are large across approaches, reflecting high volatility.
  - Example magnitudes: RMSE for crude oil, copper, and corn futures range between 15 to 19 percent at the three month (91 day) horizon.
  - For the same three contracts RMSE range was 31 to 45 percent at the two year (728 day) horizon.
- Relative RMSE performance (91 to 364 day horizons):
  - Futures price and random walk models outperform most other models.
  - At the 91 day horizon, futures price and random walk RMSEs averaged across all 10 commodities are 16.1 percent and 16.4 percent, respectively.
  - Time-series models’ RMSEs are about 9 percentage points higher than benchmark at short- to medium-term horizons.
  - Models producing k-step ahead forecasts (including exponential smoother and weekly ARIMA) are about 2 percentage points higher than benchmarks.
  - Gaps narrow at the 2 year horizon; Holt-Winters exponential smoother performs well at longer horizons, suggesting trend persistence.
- Crude oil extended results (up to 3 years / 1,092 days):
  - Futures forecasting performance deteriorates with horizon.
  - Futures forecast better than a random walk up to 2 years; statistical significance at the 10 percent level only at the 91 day horizon.
  - DM test: exponential smoother is a better forecaster than the random walk beyond 2 years (2 year and 3 year horizons).
- Statistical tests:
  - DM tests: random walk is a better forecaster than most naïve reduced-form models for most commodities.
  - Clarke and West adjusted MSFE tests: less conclusive; ARIMA specifications sometimes outperform random walk, especially at longer horizons for some agricultural commodities.
- Summary on futures performance:
  - In almost all cases futures price outperformed the random walk, though statistical significance at 90 percent and 95 percent levels was less common.
  - Futures did better at shorter horizons and became less accurate relative to a random walk at the two-year horizon.
  - Time-series models generally performed worse than the random walk.

### Empirical results — conditioning on curve shape (contango vs backwardation)
- Conditional test using loss differential d_t and backwardation dummy:
  - Theory predicts negative coefficient on backwardation dummy (futures better in backwardation).
  - Empirical findings:
    - In most cases, cannot reject the null that both prices are equally good in contango and backwardation.
    - In some cases (notably crude oil) futures are worse forecasters in backwardated markets.
  - Extreme backwardation (spot > futures by more than 5 percent; 2 percent threshold used for corn) analysis:
    - Aluminum and gold dropped due to few occurrences.
    - Futures somewhat better relative to random walk in extreme backwardation, but results inconsistent across commodities and horizons.
- Interpretation:
  - Results are puzzling relative to conventional cost-of-carry predictions; market segmentation is a possible explanation but would require strict limits to arbitrage.

### Empirical results — conditioning on market phase (bull vs bear)
- Test regressing d_t on bull-market dummy (bull/bear turning points via Bry-Boschan algorithm):
  - Hypothesis: futures perform worse in bull markets if financial speculation drives futures away from fundamentals.
- Findings:
  - Very little evidence that futures forecasting ability depends on bull/bear phase for most commodities.
  - Exceptions where relative forecasting ability depends on market phase:
    - 1-year crude oil contracts
    - 1-year gasoline contracts
    - Near-dated wheat contracts
  - Overall: limited support that index investing or financial speculation in bull markets systematically distorts futures price discovery; alternative explanations include speculation affecting both spot and futures or lack of evidence for sustained inventory buildups.

### Conclusion — main findings and implications
- Main findings on forecasting performance:
  - Futures price-based forecasts are hard to beat.
  - Futures prices perform at least as well as a random walk for most commodities and horizons, and in some cases significantly better.
  - The failure of futures to clearly and statistically significantly outperform the random walk in almost all cases is a puzzle; absent arbitrage constraints, cost-of-carry theory would predict futures should outperform.
  - Many naïve time-series models perform much worse than a random walk; parameter instability undermines many models.
  - Relative forecasting ability of futures deteriorates with forecast horizon, likely reflecting lower liquidity at the back end of futures curves.
- Performance during different market conditions:
  - Theory predicts better futures performance when market is backwardated; empirically no significant difference is found between backwardation and contango, even in strong backwardation.
  - Market segmentation could explain results but would require unrealistic limits to arbitrage; authors identify this as a puzzle for further research.
- Performance across market trends:
  - No significant difference in futures forecasting ability during bull and bear markets generally, suggesting recent financialization has not distorted futures price discovery.

*Source: _wp11254*

### 1. Introduction

### _wp11254 - 1. Introduction

### Motivation for the paper
- Futures prices performed poorly as forecasters during the recent commodity price cycle (see Figure 1).
- Three stated reasons for revisiting the forecasting ability of futures:
  - Futures did a poor job during the recent cycle and it is natural to ask whether we can do better.
  - The issue has not been decisively settled; enhancing measurement of futures prices, updating the sample period, and broadening coverage may bring us closer to a definitive answer.
  - Forecasting commodity prices is important and often costly for policymakers in countries where commodity prices significantly affect the terms of trade, inflation, and poverty levels; fitting structural or reduced-form models or applying informed judgment for a wide range of commodities can be costly and may add no more value than extrapolating from the current price.
- Figure 1 illustrates futures price curves and spot price developments, 2005-10, for:
  - Crude Oil: Spot and Futures Prices (USD per barrel)
  - Corn: Spot and Futures Prices (USD cents per bushel)
  - Copper: Spot and Futures Prices (USD cents per pound)
  - Natural Gas: Spot and Futures Prices (USD per mmBtu)
- Source for figure data: Bloomberg L.P.

### Four contributions of this paper
- Precise measurement:
  - Provide a careful measure of futures prices that exactly matches the horizon of the subsequent change in spot prices and addresses problems posed by illiquid long-dated contracts.
- Horizon coverage:
  - Compare the ability of futures prices (and other candidate models) to provide useful forecasts of spot prices at various horizons stretching out two years.
  - Note: For policymakers, the relevant forecast horizon for commodity prices is typically longer than the standard 3 to 12 months tested in much of the literature.
- Updated and broadened assessment:
  - Update and broaden the assessment of futures prices as forecasters, including the years since 2002 when commodity market liquidity has greatly increased.
  - Until now, there have been few studies of forecasting by futures markets for a broad range of commodities.
- Conditional performance:
  - Assess forecasting ability of futures prices during different market conditions, defined by the futures curve shape (backwardation and contango) and spot price trends (bull and bear markets).
  - This innovation enables a stricter test of commodity market efficiency and assessment of the contention that financial investors—attracted by rising commodity prices—have impeded the price discovery process in futures markets.

### Structure of the paper
- Section 2: outlines the models to be estimated and the intuition behind the specifications.
- Section 3: describes the data.
- Section 4: outlines the main empirical results and discusses their significance.
- Section 5: provides brief concluding remarks.

*Source: _wp11254 - 1. Introduction*

### 2. Model and Empirical Test Specification

### 2. Model and Empirical Test Specification

### Model specification
- Futures pricing model: price of a futures contract equals discounted expected spot price:
  - F(t,T) = E_f(t)[S(T)] e^{-ρ(T−t)}  (equation (1) in text)
- Log-linearized and written for forecast horizon k = T − t:
  - f_{t,t+k} = E_{t}^{f} s_{t+k} − ρ k + ½ var(s)  (equation (2))
  - Jensen’s term ½ var(s) is ignored (or absorbed in ρk under homoscedasticity).
- Basis relation (current log futures minus current log spot):
  - f_{t,t+k} − s_t = E_{t}^{f} s_{t+k} − s_t − ρ k  (equation (3))
- Common in-sample test regression (unbiasedness/efficiency test):
  - s_{t+k} − s_t = α + β (f_{t,t+k} − s_t) + ε_{t+k}  (equation (4))
  - Unbiased forecast: α = 0, β = 1, E[ε_{t+k}|Ω_t] = 0
- Under rational expectations with white-noise prediction error ν:
  - E_{t}^{f} s_{t+k} − s_{t+k} = ν_{t,t+k}, with E[ν_{t,t+k}] = 0 and covariances zero for j = 0,1,2,3,...,k,k+k (equation (6))
- Implied linear relationship between futures and realized spot (cointegration context):
  - f_{t,t+k} = ρ k − β ω_{t+k} + v_{t+k}  (equation (7) notation preserved)

### In-sample rationality tests
- Three notions of rationality tested:
  1. Weak notion: futures market does not make persistent in-sample prediction errors; long-run expected errors are zero. Assessed via cointegration tests when f and s are I(1).
     - Engle-Granger applied to specification (6) with realized spot as dependent variable (equation (13)).
     - Lag length for residual test determined by Bayesian information criterion.
     - Newey-West HAC standard errors with bandwidth = futures contract horizon (days to maturity) − 1 to control for overlapping observations.
  2. Weak-form efficiency: current futures price incorporates all information useful for in-sample prediction.
     - Rearrangement of (7) yields:
       - (φ_0) + φ_1 s_{t} + v_{t} + ...  (equation (8) mapping: φ_0 = ρ k / β, φ_1 = 1/β)
     - Error-correction specification (ECM) allowing lagged price changes (equation (9)):
       - Δ s_{t+k} = α + δ (f_{t} − μ s_{t−k}) + Σ φ_j Δ s_{t−j} + Σ γ_l Δ f_{t−l} + v_{t+k}  (structure preserved per (9))
     - Weak-form short-run market efficiency requires conditions (10):
       - 1) δ = 1
       - 2) 1 − δ φ_1 μ ≠ 0
       - 3) φ_j + γ_l = 0 for j,l conditions, and sign restrictions 0, 0_{jl} ∀ etc. (conditions presented verbatim as in text)
     - Rejection of (10) implies lagged levels of spot and futures help predict spot, violating efficiency.
  3. Unbiasedness (strictest): incorporates weak-form efficiency and no risk premium.
     - Restrictions analogous to (10) but with the coefficient on futures in cointegrating equation equal to unity (conditions (12) preserved).
     - Tests of efficiency (10) and unbiasedness (12) implemented via Wald tests on ECM (9) using robust covariance estimates.
- For commodities where both spot and futures are I(0):
  - Use equation in log levels (no ECM) (equation (14)):
    - s_{t+k} = α + μ f_{t,t+k} + Σ φ_j s_{t−j} + Σ γ_l f_{t−l} + v_{t+k}  (structure per (14))
  - Efficiency restrictions: φ_j = γ_l = 0 for j ≥ 1 and l ≥ 0; unbiasedness adds μ = 1.
- Lag specification K and M for each commodity/horizon chosen by Bayesian information criterion.
- Robust standard errors: Newey-West HAC to adjust for overlapping observations.

### Out-of-sample forecasting ability test specifications
- Forecasting metric: mean squared forecast error (MSFE) under symmetric loss; minimizes MSFE.
- Benchmark: random walk without drift:
  - s_{t+1} = s_t + ε_{t+1}
- Cost-of-carry and commodity-specific adjustments:
  - Financial asset relation (no frictions): f_{t,t+k} = s_t + r k  (equation (15))
  - Commodities add storage costs m (proportion of spot) and marginal convenience yield ψ(N):
    - f_{t,t+k} = s_t + (r + m − ψ) k  (equation (16) structure preserved)
  - Spot pricing asset equation with same risk premium:
    - E[s_{t+k}] = s_t + (−r + m − ψ) k − ρ k  (equation (17) notation preserved)
- Forecast errors definitions:
  - ε_{s,t+k} = s_{t+k} − E_{t}^{s} s_{t+k}
  - ε_{f,t+k} = s_{t+k} − f_{t,t+k}  (equation (18) preserved)
- Difference in squared forecast errors d_t = ε_{s,t+k}^2 − ε_{f,t+k}^2 expressed as (19) (preserve full expression):
  - d_t = (r + m − ψ)^2 + cross terms that cancel under arbitrage linking spot and futures expectations, leading to inequality (20):
    - ε_{f,t+k}^2 ≤ ε_{s,t+k}^2  (equation (20))
- Conditioning the loss function on slope/backwardation:
  - Estimate regression with dependent variable d_t and independent variables: constant and dummy (1 if futures curve is backwardated at forecast time, 0 otherwise).
  - Prediction: coefficient on backwardation dummy is negative (better relative performance of futures in backwardation).
  - Alternative dummy: 1 when spot price is more than 5 percent higher than futures price (periods with very high ψ).
- Candidate models compared against random walk benchmark (each model produces one-step-ahead forecasts; notation t+1 used). Candidate models (Table 1) with their equations preserved:
  - Random walk benchmark:
    - s_{t+1} = s_t + ε_{t+1}
  - ARIMA (1,1,1):
    - Δ s_t = a_{10} + a_{11} Δ s_{t−1} + a_{12} ε_{t−1} + ε_t  (notation per Table 1)
  - ARMA (1,1):
    - s_t = b_{10} + b_{11} s_{t−1} + b_{12} ε_{t−1} + ε_t
  - W-ARIMA (1,1,1):
    - Δ s_t = a_{02} + a_{1} Δ s_{t−k} + a_{2} ε_{t−k} + ε_t  (weekly variant producing k-step ahead forecasts)
  - Holt-Winters:
    - s_{t+k} = a_t + b_k + c  (structure noted per Table 1)
  - Futures price:
    - s_{t+1} = f_{t,t+1} + ε_{t+1}
  - Futures with risk premium:
    - s_{t+1} = α + β f_{t,t+1} + ε_{t+1}
  - Basis:
    - Δ s_{t+1} = α + β (f_{t,t+1} − s_t) + ε_{t+1}
  - Error correction:
    - Δ s_{t+1} = α + δ ECM_term + Σ φ_j Δ s_{t−j} + Σ γ_l Δ f_{t−l} + ε_{t+1}  (structure per Table 1)
  - Levels with lags:
    - s_{t+1} = α + μ f_{t,t+1} + Σ φ_j s_{t−j} + Σ γ_l f_{t−l} + ε_{t+1}
- Forecast-horizon-specific lagging: Lags in time-series models match forecast horizon (e.g., 91 day ARIMA uses previous 91-day change as lag). Exceptions: W-ARIMA and Holt-Winters produce k-step ahead forecasts with t−j lags representing variables from j weeks previous.
- Forecast comparison tests:
  - Diebold-Mariano (DM) test for equality of MSFEs:
    - DM = (d̄) / ω̂  with d̄ = (1/T) Σ d_t  and ω̂ a consistent long-run variance estimate accounting for overlapping observations (equation (21))
    - Under null, DM ~ N(0,1).
  - For nested models, also report Clarke and West (2005) adjusted MSFE test to account for noise due to estimating additional parameters that are zero under the null.
- Test for financialization / conditional performance in bull vs bear markets:
  - Hypothesis: forecasting performance of futures relative to spot deteriorates during bull markets if financial investors with less information or momentum trading dominate.
  - Procedure: regress d_t on constant and bull-market dummy (bull/bear turning points identified by Bry-Boschan algorithm). Null: coefficient on dummy = 0 (no difference in relative forecasting ability between bull and bear markets).

*Source: _wp11254 — 2. Model and Empirical Test Specification*

### 3. Data

### _wp11254 - 3. Data

### Overview of data
- Sample period: January 1990 to June 2011.
- Sampling frequency: weekly.
- Futures price source: Bloomberg.
- Futures contracts used: the set of the first 24 contracts, ordered by days to delivery.
  - For commodities with equally-spaced monthly delivery dates, the futures curve stretches out for two years.
  - For some commodities the number of traded contracts is less than 24, but the futures curve may stretch out further than two years (typically agricultural commodities with no delivery date each month).
  - Insufficient data to undertake analysis at the two year horizon for aluminum and gasoline.
- Table of spot price and futures contract specifications (as provided in the source):
  - Aluminum — Spot price: London Metal Exchange-Aluminum 99.7% Cash; Futures contract: London Metal Exchange
  - Copper — Spot price: London Metal Exchange-Copper, Grade A Cash; Futures contract: London Metal Exchange
  - Corn — Spot price: Corn Number 2 Yellow; Futures contract: Chicago Mercantile Exchange
  - Cotton — Spot price: Cotton, 1 1/16Str Low - Middling, Memphis; Futures contract: ICE: Intercontinental Exchange
  - Crude Oil — Spot price: West Texas Intermediate Spot Cushing; Futures contract: Chicago Mercantile Exchange
  - Gasoline — Spot price: Unleaded Regular Oxygenated New York; Futures contract: New York Mercantile Exchange
  - Gold — Spot price: Gold Bullion London Bullion Market; Futures contract: Chicago Mercantile Exchange
  - Natural Gas — Spot price: Natural Gas-Henry Hub; Futures contract: Chicago Mercantile Exchange
  - Soy — Spot price: Soyabeans, Number 1 Yellow; Futures contract: Chicago Mercantile Exchange
  - Wheat — Spot price: Wheat, Number 2 Hard (Kansas); Futures contract: Chicago Mercantile Exchange

### Estimating fitted futures prices
- Objective: provide a smoothed futures curve each week so that the horizon of the futures price matches the corresponding spot price change exactly (e.g., obtain an implied futures price at a delivery date exactly 28 days forward when sampling weekly).
- Method: estimate a smoothed futures curve for each week with a third order polynomial regression to produce a continuous curve and extract implied futures prices for exactly specified future dates.
- Weekly cross-sectional regression (as stated in the source):
  - 23 0123 t t t t F t t v
    γ  γ  γ  γ =  +  +  +  +
  - where F is the vector of futures prices for a particular commodity at the close of Friday trading with increasing maturity, γ0 is a constant, t represents the number of days until delivery for each contract, and γ1, γ2, and γ3 are coefficients estimated for each commodity each week.
- Implied futures price extraction:
  - To obtain the implied futures price for a notional contract with delivery exactly t1 days forward, substitute t1 for t in the estimated polynomial.
  - The paper uses t1 = 0, 91, 182, 364, and 728.
  - Result: implied spot price (t = 0) equals γ0.
  - Benefit: provides an implied price for illiquid long-dated contracts based on the liquid part of the curve rather than non-tradable price quotes; useful where contract specification differences complicate direct comparison of futures to actual spot prices.

### Measures of fit and diagnostics
- Average R-squared results:
  - The average R-squared for all commodities stands at 0.8.
  - A third-order polynomial provides an acceptable fit for most commodities.
  - Cases with average R-squared below 0.8 reflect two main causes:
    - Seasonal factors causing multiple local maxima or several inflection points in the curve (example: corn and wheat).
    - Quoted futures prices at the back end of the 2-year curve being illiquid and not fully representative of market conditions (example: gold, where interest-arbitrage considerations should suggest a very good fit but illiquidity at long horizons can reduce fit).
  - The least-squares method tends to put more weight on the liquid and smooth part of the curve where arbitrage considerations and actual market conditions are more important.

*Source: _wp11254 - 3. Data*

### 4. Empirical results

### 4. Empirical results and Discussion

### 4.1. In-sample rationality
- R-squared statistics (average r-square for all equations per commodity):
  - Aluminum 0.96
  - Copper 0.92
  - Corn 0.80
  - Cotton 0.89
  - Gasoline 0.75
  - Gold 0.72
  - Natural gas 0.54
  - Soybeans 0.72
  - Wheat 0.78
  - WTI crude oil 0.93
- Unit root testing (Jan-1990 to Jun-2011):
  - Most spot and futures prices were nonstationary (I(1)) based on augmented Dickey-Fuller tests with lag length selected by Bayesian information criteria; Newey-West HAC standard errors with a bandwidth equal to days to maturity minus one used to control residual correlation.
  - For commodities with I(1) spot and futures prices, the null hypothesis that the coefficient φ = 0 in equation (13) was rejected in all cases (Engle-Granger tests). Interpretation: the commodity futures price and the realized spot price are cointegrated and futures markets do not make persistent forecasting errors in-sample.
- Efficiency and unbiasedness (Wald test p-values; summary):
  - The efficient market hypothesis was rejected for most commodities and at most horizons.
  - At the 91 and 182 day horizons, less than half of the commodities are efficient at the 5 percent significance level.
  - Efficiency can be ruled out for all commodities at the one and two year horizon except for cotton.
- Explanatory power of lags:
  - Lagged values of spot and futures prices often provide significant in-sample explanatory power for realized spot prices, especially at one- and two-year horizons.
  - Bayesian criteria typically selected long lag lengths (often five or six periods) for test equations, particularly for horizons beyond 3 months.
  - The signs and sizes of lag coefficients varied widely across commodities, complicating interpretation and ruling out a simple common mechanism such as long-run mean reversion.

### 4.2. Out-of-sample forecasting
- Root mean squared errors (RMSEs) and volatility of forecasting:
  - RMSEs are large across approaches, confirming commodity prices are volatile and difficult to forecast.
  - Examples: RMSE for crude oil, copper, and corn futures range between 15 to 19 percent at the three month horizon.
  - For the same three contracts the RMSE range was 31 to 45 percent at the two year horizon.
- Relative RMSE performance (horizons 91 to 364 days):
  - Futures price and random walk models outperform most other models.
  - At the 91 day horizon, the futures price and random walk RMSEs averaged across all 10 commodities are 16.1 percent and 16.4 percent, respectively.
  - RMSEs for the time series models are about 9 percentage points higher than benchmark at short- to medium-term horizons.
  - Models producing k-step ahead forecasts (including the exponential smoother and the weekly ARIMA) are about 2 percentage points higher than benchmarks.
  - Gaps narrow significantly at the 2 year horizon.
  - Holt-Winters exponential smoother performs well at longer horizons, suggesting some trend persistence not captured by other models.
- Detailed results for crude oil (extended to 3 years / 1092 days):
  - Forecasting performance of futures deteriorates with forecast horizon.
  - Futures prices forecast better than a random walk at all horizons up to 2 years; statistical significance (10 percent level) only at the 91 day horizon.
  - Based on the DM test, the exponential smoother is a better forecaster than the random walk beyond 2 years (2 year and 3 year horizon rejection of null).
  - Suggested mechanism: medium-term reversion to persistent trends may help explain why the exponential smoother outperforms at long horizons.
- Statistical hypothesis tests of out-of-sample accuracy:
  - Diebold-Mariano (DM) tests: the random walk is a better forecaster than most naïve reduced form models for the majority of commodities.
  - Clarke and West adjusted mean squared error tests (accounting for parameter uncertainty): tests less conclusive; in most cases it was not possible to reject the null and ARIMA specifications appear to outperform the random walk, especially at longer horizons for some agricultural commodities.
- Performance for futures prices specifically:
  - In almost all cases, the futures price outperformed the random walk, although statistical significance at the 90 percent and 95 percent levels was less common.
  - Futures did better at shorter horizons and became less accurate relative to a random walk at the two-year horizon.
  - Time series models generally performed worse than the random walk.

### 4.3. Conditioning on curve shape: contango and backwardation
- Conditional test (equation (20)) of whether futures are better forecasters when the market is backwardated:
  - Regression uses dummy = 1 when market is backwardated; dependent variable is loss differential d; theory predicts negative coefficient.
  - In most cases, the null that both prices are equally good forecasters in contango and backwardation cannot be rejected.
  - In some cases, notably crude oil, futures prices are worse forecasters in backwardated markets.
  - Implication: forecasters should not disregard spot prices when making forecasts in a backwardated (and tight) market.
- Extreme backwardation analysis:
  - Extreme backwardation defined as spot price more than 5 percent higher than corresponding futures price (2 percent used for corn due to insufficient observations).
  - Aluminum and gold were dropped due to very few occurrences of extreme backwardation.
  - Results: futures are somewhat better forecasters relative to a random walk in extreme backwardation, but this is not consistent across commodities or horizons.
  - Contrast: results are not consistently supportive of findings in other studies (e.g., Reeve and Vigfusson (2011) as referenced in source).
- Interpretation and puzzle:
  - One possible explanation is market segmentation: spot market participants may make better forecasts than futures market participants.
  - However, segmentation does not fully explain why arbitrage does not eliminate differences (e.g., spot participants could sell spot and buy futures in backwardation if they expected rises).
  - The empirical inability to reject the null when conditioning on curve shape is puzzling relative to conventional commodity pricing models.

### 4.4. Conditioning on market phase: bull and bear markets
- Conditional test (Table 11) of whether futures forecasting performance relative to a random walk depends on whether spot prices are in a bull market (dummy = 1 in bull market):
  - Proponents of the view that financial speculation during bull markets drives futures away from fundamentals would anticipate positive coefficients (futures do worse in bull markets).
- Findings:
  - Very little evidence that futures forecasting ability depends on bull/bear phase for most commodities.
  - Exceptions where the relative forecasting ability does depend on market phase:
    - 1-year crude oil contracts
    - 1-year gasoline contracts
    - Near-dated wheat contracts
  - Interpretation: limited support for the hypothesis that index investing or financial speculation in bull markets systematically distorts the price discovery process in futures markets.
  - Alternative explanation offered: speculation could affect both spot and futures markets simultaneously; additionally, persistent inventory buildup required by the speculation story is not supported by evidence from much of the last decade (inventories often declined as prices rose).

*Source: IMF Working Paper — 4. Empirical results and Discussion*

### 5. Conclusion

### 5. Conclusion

### Main findings on forecasting performance
- Futures price-based forecasts are hard to beat.
- Futures prices perform at least as well as a random walk for most commodities and at most horizons and, in some cases, do significantly better.
- The failure of futures prices to clearly (and statistically significantly) outperform the random walk in almost all cases is a puzzle. The spot price reflects the cost of carry and is more influenced by current physical market conditions (and less by expectations of the future) than is the futures price. In the absence of constraints on arbitrage, this should mean that futures prices outperform the random walk, on average.
- Many other naïve time series models, including some that maximize in-sample fit, tend to do much worse than a random walk. Parameter instability renders many time-series models as useless, at best.
- The relative forecasting ability of futures prices deteriorates the longer the forecast horizon, which likely reflects lower liquidity at the back end of futures curves.

### Performance during different market conditions
- Theory predicts that futures prices should do much better than the random walk when the market is in backwardation because the influence of current market conditions on spot prices is particularly strong during these periods.
- Empirically, no significant difference in forecasting ability is found between periods of backwardation and contango. This result holds even when spot prices are significantly above futures prices, in strong backwardation.
- Potential explanations considered:
  - Over small sample periods, permanent shocks that increase prices could lead to better “forecasts” by spot prices, but over long periods and assuming a symmetric distribution of shocks, this cannot be the explanation.
  - Segmented markets could explain the result, with backwardated markets reflecting different and better information in the spot market about future spot prices than futures markets. However, this would require strict and unrealistic limits to arbitrage (for example, preventing spot market participants from buying cheaper futures contracts).
- The authors identify this as an apparent puzzle that would benefit from further research.

### Performance across market trends (bull and bear markets)
- No significant difference is found in the forecasting ability of futures markets during bull and bear markets, defined as when spot prices are trending higher or lower.
- This new result suggests that the recent period of financialization has not distorted the futures price discovery process.

*From: 5. Conclusion, _wp11254 - 5. Conclusion.*

---


_Source: https://www.imf.org/-/media/websites/imf/imported-full-text-pdf/external/pubs/ft/wp/2011/_wp11254.pdf_
