## wpiea2025252-source-pdf

## Source details

**Canonical URL:** [wpiea2025252-source-pdf](https://www.imf.org/-/media/files/publications/wp/2025/english/wpiea2025252-source-pdf.pdf)

## Other formats

- [Markdown version](/-/media/files/publications/wp/2025/english/wpiea2025252-source-pdf.pdf.md)
- [Structured JSON version](/-/media/files/publications/wp/2025/english/wpiea2025252-source-pdf.pdf.json)

---

### Introduction — nowcasting context and purpose
- Nowcasting exploits high-frequency information to monitor economic activity and has become essential in academia, finance, and government.
- The paper categorizes nowcasting methods into two classes: (1) traditional econometric models and (2) machine learning algorithms (ML).
- Performance metric: out-of-sample root-mean-square errors (OOS RMSE) reported relative to a benchmark over an evaluation period (test set).
  - Relative OOS RMSE = OOS RMSE of a model, f (over i in Test: (y_i − f(X_i))) / OOS RMSE of a benchmark model (over i in Test: (y_i − ŷ_benchmark_i))
- Empirical scope:
  - Countries analyzed in the pseudo-real-time empirical section: China, France, India, Mauritius, Türkiye, and Vietnam.
  - In simulations the benchmark model is AR(1); in empirical evaluation the benchmark is the automatic ARIMA model.
- High-frequency indicator treatment:
  - Forecasting unreleased indicators is required in practice; the paper assumes quarter-end evaluation with all high-frequency indicators fully released (it does not analyze indicator-forecasting errors).
  - Frequency mismatch typically mitigated via temporal aggregation (averaging or blocking); the paper reports the lower RMSE from these two choices.

### Major evaluated models
- Traditional Econometric Models:
  - Bridge; MIDAS; U-MIDAS; MIDAS-PC; U-MIDAS-PC; Three Pass Regression Filter (TPRF); Dynamic Factor Model (DFM)
- Machine Learning Algorithms:
  - LASSO; Elastic Net; Decision Tree (DT); Random Forest (RF); Light GBM; XGBoost; Support Vector Regression – Radial basis (SVRR); Support Vector Regression – Linear basis (SVRL); Support Vector Regression – Polynomial basis (SVRP); Neutral Network (NNET); Long Short-Term Memory Neutral Network (LSTM)
  - Ensembles: LASSO-DT; LASSO-DT-SVRP; LASSO-SVRP-NNET; LASSO-RF-SVRR

### Key high-level findings (simulation and empirical)
- Simulation findings:
  - TPRF often ranks first in both linear and non-linear Data Generating Processes (DGPs).
  - Linear ML algorithms such as Elastic Net and Lasso frequently follow TPRF in performance.
  - Linear ML algorithms tend to outperform more complex ML algorithms when:
    - the number of observations is sufficiently large, i.e. exceeding 100,
    - and the DGP is linear.
  - When the sample size is small, TPRF performs best.
- Empirical findings (six-country pseudo-real-time nowcasting):
  - In most of the six countries, traditional econometric models, particularly Bridge or DFM, are top performers.
  - Bridge model delivers strong nowcasting performance in all six countries.
  - In Türkiye and India, linear ML algorithms—Lasso or Elastic Net—are the best performers.
  - Across countries, linear ML algorithms outperform complex, non-linear algorithms.
  - Complex non-linear models can overfit short GDP series, producing poorer nowcasting performance relative to simple models.
- Overall recommendation from analyses:
  - Use traditional econometric models (Bridge or DFM) and linear ML algorithms for single-country GDP nowcasting given short GDP series and overfitting risk in complex models.
- Caveat: conclusions may depend on simulation design, high-frequency indicators, and choice of countries.

### Relation to existing literature
- Early nowcasting methods summarized: Bridge (Baffigi et al. (2004)), MIDAS (Ghysels et al. (2004, 2007)), U-MIDAS (Forini et al. (2011)), MIDAS-PC and U-MIDAS-PC, DFM (Giannone et al. (2008); Doz et al. (2011)), and TPRF (Kelly and Pruitt (2015)).
- Mixed evidence in literature: no single model is universally superior; performance varies by data, country, and period.

---

### 3.1 Traditional econometric models — model descriptions and specifications

- Bridge
  - Purpose: Convert high-frequency indicators into low-frequency (quarterly) aggregates so OLS can be applied.
  - Practical note: Unreleased high-frequency indicators must be forecasted to aggregate into the low frequency.
  - Specification:
    - y_q = β_0 + sum_{l=1}^m γ_l y_{q−l} + sum_{i=1}^n β_i x_{i.q} + ε_q

- MIDAS
  - Purpose: Mixed-frequency regression splitting monthly indicator into three quarterly variables; Almon polynomial restrictions of degree 2 reduce coefficients to two per indicator.
  - Estimation: Non-linear least squares; AR terms included.
  - Specification:
    - y_q = β_0 + sum_{l=1}^m γ_l y_{q−l} + sum_{i=1}^n sum_{h=1}^{L(i)} δ_{i,h} x_{i.3q−h} + ε_q
    - δ_{i.h}(w_{1,i},w_{2,i}) = exp(w_{1,i} h + w_{2,i} h^2) / sum_{h=1}^{L(i)} exp(w_{1,i} h + w_{2,i} h^2)

- U-MIDAS
  - Purpose: Unrestricted MIDAS estimated by OLS without coefficient restrictions.
  - Specification:
    - y_q = β_0 + sum_{l=1}^m γ_l y_{q−l} + sum_{i=1}^n sum_{h=1}^{L(i)} β_{i,h} x_{i.3q−h} + ε_q

- MIDAS-PC and U-MIDAS-PC
  - Purpose: Use principal components of high-frequency indicators to curb parameter proliferation; U-MIDAS-PC uses OLS when number of factors is small.

- Three-pass Regression Filter (TPRF)
  - Purpose: Extracts factors that maximally correlate with the forecast target.
  - Estimation steps:
    - Step 1: For each indicator, run y_q = β_0 + β_i x_{i,q} + ε_{i,q}
    - Step 2: For each t, run x_{i,t} = β_0 + TPRF_t β_i + ε_i to yield TPRF_t
    - Step 3: Use TPRF in nowcasting regression:
      - y_q = β_0 + sum_{l=1}^m γ_l y_{q−l} + θ TPRF_q + ε_q
  - Practical note: No look-ahead bias because Steps 1 and 2 use lagged GDP; nowcast uses current-quarter TPRF_q.

- Dynamic Factor Model (DFM)
  - Purpose: Address frequency mismatch and ragged-edge problems; allows immediate nowcast updates when indicators are released asynchronously.
  - Estimation: Kalman Filtering.
  - System:
    - z_{i,q} = μ + Λ F_q + ε_{i.q}, z_{i,q} ∈ Z_q
    - F_q = A F_{q−1} + B u_q, u_q ∼ N(0, I)
    - Assumptions: ε'_{i.q} are cross-sectionally white-noise and orthogonal to u_q.
  - Nowcast: regress y_q on estimated factor ŜF_q with controls; use current ŜF_q to compute nowcast.

---

### 3.2 Machine Learning Algorithms — tuning, architectures, and ensembles

- Cross-validation and tuning
  - Five-fold cross-validation across hyperparameter grids per algorithm; Table 2 in source summarizes parameter grids and total combinations.

- Lasso
  - Variable selection via L1 shrinkage.
  - Objective:
    - β = argmin ( sum_{q=1}^T (y_q − β_0 − sum_{j=1}^k x_{j,q} β_j)^2 + λ sum_{j=1}^k |β_j| )

- Elastic Net
  - Combines L1 and L2 penalties with mixing parameter α.
  - Objective:
    - β = argmin ( sum_{q=1}^T (y_q − β_0 − sum_{j=1}^k x_{j,q} β_j)^2 + λ [ α sum_{j=1}^k |β_j| + (1−α) sum_{j=1}^k β_j^2 ] )
  - Tuning: 3,150 combinations of λ and α using 5-fold cross-validation.

- Decision Tree (DT)
  - Complexity controlled by cp (complexity parameter).

- Random Forest
  - Tuned via "mtry" controlling number of variables sampled at each split.

- LightGBM and XGBoost
  - Each tuned across 18 combinations:
    - n_estimators: 50, 100
    - max_depth: 2, 4, 8
    - learning_rate: 0.01, 0.05, 0.1

- Support Vector Regression (SVR)
  - SVRL (Linear Kernel): C ∈ [10^{−3}, 10^2] (10 values)
  - SVRR (Radial Kernel): σ ∈ {0.001, 0.01, 0.05, 0.1, 0.2}; C ∈ {0.1, 1, 10, 100, 1000}
  - SVRP (Polynomial Kernel): C ∈ [10^{−2}, 10^2] (5 values); Degree ∈ {2, 4, 8}; Scale ∈ {0.1, 1, 10}
  - Objective:
    - L(ω,b) = 1/2 ||ω||^2 + C sum_{i=1}^n L_ε(y_i, f(x_i))

- Neural Networks
  - NNET: Feedforward with single hidden layer; hidden sizes {3, 5, 10}; weight decay ∈ {0.0001, 0.001, 0.01, 0.1} — 12 combinations.
  - LSTM: Keras Tuner RandomSearch (max 5 trials).
    - LSTM Layers: 1, 2, 3
    - Units per Layer: 16 to 128 (step 16)
    - Activation: tanh
    - Learning Rate: 0.0001 to 0.01 (log scale)
    - Epochs: 30; Batch Size: 8; Validation Split: 0.1

- Ensemble learning
  - Simple averaging of constituent model predictions; ensembles constructed:
    - LASSO_DT; LASSO_DT_SVRP; LASSO_RF_SVRBF; LASSO_SVRP_NN

---

### 4.1–4.2 Simulation setup and Monte Carlo experiments — DGPs, parameterization, and procedures

- Factor structure and notation
  - Monthly observed data and quarterly nowcast targets with q = 3m.
  - Unobservable factor F(m) drives both X(m) and y(q); additional unobservable factors G(m) drive observed high-frequency data only.
  - F and G follow autoregressive structures with heterogeneous loadings; λ ∼ Uniform(λ_{Lo}, λ_{Hi}).
  - High-frequency data X(m) = X^U(m) (unobservable components) + X^O(m) (observed components).
  - Only unobservable components enter DGP for nowcast target Y; observed components correlate with Y via common factors F.

- Simulation parameter values
  - K = 5 and N = 105 (five unobservable components generate GDP; total predictors N = 105).
  - Components X_1, X_2, and X_3 are included in all DGPs.
  - Loadings: X_1 homogeneous across months in a quarter; X_2 and X_3 heterogeneous across months.

- True DGP specifications
  - DGP 1: Linear
    - y_t = 1.2 + 0.5 x^{(0)}_1 + 0.5 x^{(−1)}_1 + 0.5 x^{(−2)}_1 + 0.3 x^{(0)}_2 + 0.6 x^{(−1)}_2 + 0.4 x^{(−2)}_2 − 0.25 x^{(0)}_3 − 0.2 x^{(−1)}_3 − 0.3 x^{(−2)}_3
  - DGP 2: Non-Linear Quadratic
    - y_t = 1.2 + 0.5 x^{(0)}_1 + 0.5 x^{(−1)}_1 + 0.5 x^{(−2)}_1 + 0.3 x^{(0)}_2 + 0.6 x^{(−1)}_2 + 0.4 x^{(−2)}_2 − 0.25 x^{(0)}_3 − 0.2 x^{(−1)}_3 − 0.3 x^{(−2)}_3 + 0.05 X^2_{2,(0)} + 0.02 X^2_{2,(−1)} + 0.03 X^2_{2,(−2)}
  - DGP 3: Quadratic and Interaction Terms
    - y_t = 1.2 + 0.5 x^{(0)}_1 + 0.5 x^{(−1)}_1 + 0.5 x^{(−2)}_1 + 0.3 x^{(0)}_2 + 0.6 x^{(−1)}_2 + 0.4 x^{(−2)}_2 − 0.25 x^{(0)}_3 − 0.2 x^{(−1)}_3 − 0.3 x^{(−2)}_3 + 0.05 X^2_{2,(0)} + 0.02 X^2_{2,(−1)} + 0.03 X^2_{2,(−2)} + 0.15 X_{3,(0)} X_{4,(0)} + 0.15 X_{3,(−1)} X_{4,(−1)} + 0.15 X_{3,(−2)} X_{4,(−2)} − 0.2 X_{4,(0)} X_{5,(0)} − 0.3 X_{4,(−1)} * X_{5,(−1)} − 0.1 X_{4,(−2)} X_{5,(−2)}

- Monte Carlo experiments
  - Number of Monte Carlo data draws: 500.
  - Number of factors: F and G set to 5 and 3, respectively.
  - Number of observed high-frequency indicators: 120 (source also states 120 and elsewhere 120 / 120 monthly indicators leading to 360/361 low-frequency variables after splitting).
  - Parameter draws per iteration: ηf, ηg, β, and λ generated to create synthetic data (500 iterations).
  - Parameterization:
    - λ ∼ iid Uniform(0.4,0.6)
    - β ∼ iid Uniform(0,0.9)
    - ηf, ηg, ε ∼ iid N(0,1)
    - Observed correlations between simulated X and nowcast targets range from 0.2 to 0.6
  - Time-dimensions considered:
    - Quarterly Tq = 40,50,60,70,80,100,120,150,200
    - Monthly Tm = 120,150,180,210,240,300,360,450,600 respectively
  - Training/test split: 70% training, 30% test (no rolling updates in theoretical simulation).
    - Example: when T = 80, training and test sets comprise 56 and 34 observations of XU and Y, respectively.

- Estimation and variable selection in simulations
  - Relative OOS RMSE on the test set is the evaluation metric.
  - Recursive Feature Elimination (RFE) restricts maximum number of variables based on time observations; number selected ranges from 2 to quarterly time periods divided by 5.
  - The same selected indicators are used across traditional econometric models and neural networks to isolate model differences.
  - ML algorithms employ built-in feature selection and 5-fold cross-validation within the training set.

- Simulation results summary
  - TPRF frequently ranks first across DGPs, followed by Elastic Net and Lasso.
  - Elastic Net outperforms all models when number of observations exceeds 100 and DGP is linear.
  - TPRF delivers superior performance in DGPs 2 and 3 despite nonlinearity.
  - TPRF, MIDAS-PC, U-MIDAS-PC, Elastic Net, and Lasso perform well in short samples (e.g., T = 40).
  - When the DGP is highly nonlinear with interactions, non-linear SVR with Radial basis occasionally achieves the best results by a small margin.

---

### 5.2 Performance Evaluation — pseudo-real-time empirical results and diagnostics

- Evaluation setup
  - Pseudo-real-time out-of-sample evaluation at quarter end using an expanding rolling window.
  - Evaluation window: 2018Q1 to 2024Q1 (test period described elsewhere as 2018Q1–2024Q2 for OOS RMSE computation).
  - Each model re-estimated every quarter upon release of new GDP data; coefficients vary quarterly.
  - Simplification: high-frequency indicators assumed available at quarter end prior to GDP release; GDP revisions not considered.
  - No look-ahead bias: coefficients estimated using previous-quarter data; current-quarter data then inserted to derive nowcasts.
  - Benchmark: univariate automatic ARIMA (“autoarima” package in R).

- Aggregate nowcast results (summary)
  - Traditional econometric models tend to outperform ML algorithms across simulations and country applications.
  - Best-performing models across six countries: Bridge and DFM each rank first in two countries.
    - When DFM outperforms Bridge, the margin is substantial.
    - When Bridge outperforms DFM, the margin is relatively modest.
  - Bridge models consistently rank within the top four models in five of the six countries.
  - Linear ML algorithms outperform nonlinear ML methods across all six countries.
    - Elastic Net and Lasso perform best among ML methods, except in India where SVRL slightly outperforms Lasso by a narrow margin.
  - In Türkiye and Mauritius, linear ML algorithms are the best performers.
  - Simpler models with fewer parameters tend to perform well, with the notable exception of DFM.

- COVID-19 episode (2020Q1–2022Q4)
  - Nowcasting during COVID-19 is particularly challenging.
  - In China, France, and Türkiye, traditional econometric models or linear ML algorithms outperform the autoARMA benchmark.
  - In India, Mauritius, and Vietnam, both classes of models fail to beat the benchmark during COVID-19.
  - Commonalities among successful cases during COVID-19: a large number of high-quality high-frequency indicators and a sufficiently long analysis sample.
  - Best performers during COVID-19 in specific countries:
    - China: TPRF
    - France: MIDAS
    - Türkiye: Elastic Net

- Ensemble forecasts and ML comparisons
  - Ensembles evaluated: LASSO-SVRP-NNET; LASSO-RF-SVRR; LASSO-DT
  - Ensembles do not outperform the best traditional econometric models or the linear ML algorithms.
  - Nonlinear ML algorithms risk overfitting given limited GDP observations, reducing nowcasting performance.
  - Simpler linear methods (Lasso, Elastic Net) tend to outperform more complex counterparts.

- Key raw OOS RMSE statistics (selected values preserved exactly as reported)
  - Panel A: All periods (average OOS RMSEs over 2018Q1–2024Q2)
    - Traditional econometric models (selected countries):
      - China: AutoARMA 0.06; Bridge 0.02; UMIDAS 0.04; MIDAS 0.05; UMIDAS_PC 0.04; MIDAS_PC 0.04; TPRF 0.03; DFM 0.01
      - France: AutoARMA 0.06; Bridge 0.02; UMIDAS 0.02; MIDAS 0.04; UMIDAS_PC 0.05; MIDAS_PC 0.04; TPRF 0.05; DFM 0.02
      - India: AutoARMA 0.08; Bridge 0.03; UMIDAS 0.05; MIDAS 0.06; UMIDAS_PC 0.06; MIDAS_PC 0.06; TPRF 0.04; DFM 0.02
      - Mauritius: AutoARMA 0.08; Bridge 0.06; UMIDAS 0.13; MIDAS 0.23; UMIDAS_PC 0.27; MIDAS_PC 0.14; TPRF 0.06; DFM 0.08
      - Türkiye: AutoARMA 0.05; Bridge 0.02; UMIDAS 0.02; MIDAS 0.03; UMIDAS_PC 0.03; MIDAS_PC 0.02; TPRF 0.05; DFM 0.02
      - Vietnam: AutoARMA 0.03; Bridge 0.01; UMIDAS 0.03; MIDAS 0.02; UMIDAS_PC 0.03; MIDAS_PC 0.03; TPRF 0.02; DFM 0.02
    - Machine learning algorithms (selected values):
      - China: DT 0.03; Elastic Net 0.02; LASSO 0.03; NNET 0.03; RF 0.03; SVRL 0.03; SVRP 0.04; SVRR 0.03; LightGBM 0.03; XGBoost 0.03; LSTM 0.03
      - France: DT 0.04; Elastic Net 0.02; LASSO 0.03; NNET 0.04; RF 0.04; SVRL 0.03; SVRP 0.04
      - India: DT 0.07; Elastic Net 0.04; LASSO 0.04; NNET 0.06; RF 0.07; SVRL 0.04
      - Mauritius: DT 0.08; Elastic Net 0.06; LASSO 0.06; NNET 0.08; RF 0.09
      - Türkiye: DT 0.04; Elastic Net 0.01; LASSO 0.02; NNET 0.05; RF 0.04
      - Vietnam: DT 0.02; Elastic Net 0.01; LASSO 0.01; NNET 0.02; RF 0.02
  - Panel B: COVID-19 period (average OOS RMSEs over 2020Q1–2022Q4)
    - Traditional econometric models (selected countries):
      - China: AutoARMA 0.05; Bridge 0.03; MIDAS 0.03; UMIDAS 0.03; MIDAS-PC 0.02; UMIDAS-PC 0.02; TPRF 0.02; DFM 0.04
      - France: AutoARMA 0.05; Bridge 0.04; MIDAS 0.02; UMIDAS 0.05; MIDAS-PC 0.03; UMIDAS-PC 0.04; TPRF 0.04; DFM 0.03
      - India: AutoARMA 0.04; Bridge 0.10; MIDAS 0.10; UMIDAS 0.06; MIDAS-PC 0.10; UMIDAS-PC 0.06; TPRF 0.10; DFM 0.10
      - Mauritius: AutoARMA 0.01; Bridge 0.13; MIDAS 0.30; UMIDAS 0.17; MIDAS-PC 0.12; UMIDAS-PC 0.31; TPRF 0.10; DFM 0.07
      - Türkiye: AutoARMA 0.02; Bridge 0.08; MIDAS 0.05; UMIDAS 0.06; MIDAS-PC 0.07; UMIDAS-PC 0.05; TPRF 0.02; DFM 0.06
      - Vietnam: AutoARMA 0.01; Bridge 0.04; MIDAS 0.02; UMIDAS 0.03; MIDAS-PC 0.02; UMIDAS-PC 0.03; TPRF 0.03; DFM 0.03
    - Machine learning algorithms (selected values):
      - China: DT 0.05; Elastic Net 0.04; LASSO 0.04; NNET 0.04; RF 0.05; SVRL 0.04; SVRP 0.05
      - France: DT 0.06; Elastic Net 0.04; LASSO 0.04; NNET 0.06; RF 0.06
      - India: DT 0.10; Elastic Net 0.06; LASSO 0.06; NNET 0.09; RF 0.10; SVRL 0.05
      - Mauritius: DT 0.10; Elastic Net 0.06; LASSO 0.06; NNET 0.09; RF 0.10
      - Türkiye: DT 0.06; Elastic Net 0.02; LASSO 0.02; NNET 0.06; RF 0.05
      - Vietnam: DT 0.03; Elastic Net 0.02; LASSO 0.02; NNET 0.02; RF 0.03

- Conclusion and implications
  - Simulations and six country applications show traditional econometric models often deliver superior nowcasting performance to ML algorithms.
  - In simulations, TPRF mostly outperforms others, losing only to Elastic Net under the linear DGP and long data.
  - In the six country applications, DFM or Bridge outperform ML algorithms in four out of six countries.
  - Among ML algorithms, Lasso and Elastic Net tend to outperform more complex ML algorithms and even surpass traditional econometric models in two out of six countries.
  - Practical considerations:
    - Number of GDP observations is relatively short by ML standards; small sample sizes hamper complex models and increase overfitting risk.
    - Relationship between GDP growth and high-frequency indicators may be adequately captured by linear models.
  - Recommendation:
    - Use simpler models like Bridge or linear ML algorithms for GDP nowcasting due to consistent superior performance and lower overfitting risk; DFM is a notable complex model that performs well in some cases.
  - Caveats:
    - No universally superior model; performance may vary for other countries or data sets.
    - Analysis primarily relies on traditional macroeconomic data; using non-traditional data (e.g., satellite data) might yield different results.
    - Study focuses on single-country time-series applications and does not extend to panel data setups.

---

### Appendix I — notation and country indicator lists
- Notation:
  - ( dl ) represents the log difference.
  - ( d ) denotes the first difference.
- Data source:
  - All indicators are sourced from Haver Analytics.
- Country-specific indicator lists (variable names preserved as in source):
  - China: dl_cpi; dl_export_value; dl_export_volume; dl_import_value; dl_import_volume; dl_p_steel; dl_invest_real_estate; dl_fiscal_rev; dl_fiscal_exp; dl_yuan_usd; dl_m2; dl_m1; dl_ip_crude_oil; dl_ip_natural_gas; dl_ip_iron_ore; dl_ip_salt; dl_ip_yarn; dl_ip_cotton; dl_ip_gasoline; dl_ip_kerosene; dl_ip_diesel; dl_ip_coke; dl_output_sulfuric; dl_output_sodiumhy; dl_output_sodiumcar; dl_output_ethylene; dl_output_fertilizer; dl_output_chemical; dl_output_plastic; dl_output_synthetic; dl_ip_plastic; dl_output_cement; dl_output_plate; dl_output_pigiron; dl_output_steel; dl_output_steelprod; dl_output_nonferrous; dl_ip_aluminumo; dl_output_copper; dl_ip_aluminumprod; dl_output_boiler; dl_output_icecar; dl_output_metalcut; dl_output_electric_tool; dl_ip_packaging; dl_ip_cement_equip; dl_ip_smelt; dl_ip_feeding; dl_ip_tractor; dl_ip_airclean; dl_ip_railway; dl_ip_motor; dl_ip_ship; dl_ip_powergen; dl_ip_acmotor; dl_ip_electricity_prod; dl_output_fridge; dl_output_freezer; dl_output_ac; dl_output_wastemac; dl_ip_mobilehandset; dl_output_microcomputer; dl_output_semiconductor; dl_output_electricmeasure; dl_output_copymac; dl_freight_traffic; dl_freight_turnover; dl_passenger_traffic; dl_passenger_turnover; dl_volume_port; dl_volume_foreigntradegoods; dl_no_expressmail; dl_no_cellphoneuser; dl_epu; d_fiscal_balance; d_neer; d_reer; d_pmi_manu; d_pmi_output; d_pmi_neworder; d_pmi_majorinput; d_pmi_emplotment; d_pmi_deliverytime; d_pmi_backlog; d_pmi_inputp; d_pmi_inputbuy; d_pmi_finishedgoods; d_pmi_newexport; d_pim_import; d_consumer_confidence.
  - France: dl_us_eur; dl_m1; dl_m2; dl_m3; dl_private_credit; dl_m1_dom; dl_m2_dom; dl_cb_asset; dl_saving; dl_unreg_saving; dl_credit_to_private; dl_loan_to_gov; dl_business_creation; dl_gov_rev; dl_gov_exp; dl_cpi; dl_core_cpi; dl_ppi; dl_imp_price; dl_brent_oil; dl_house_start; dl_house_permit; dl_no_job_seeker; dl_job_vacan; dl_job_vacan_dom; dl_job_vacan_skill; dl_job_vacan_unski; dl_hh_consumption; dl_export; dl_import; dl_num_dealth; dl_construction_creation; d_fis_bal; d_income_tax; d_corp_tax; d_vat; d_trade_bal; d_busness_climate_indicator; d_epu; d_manufacture_survey; d_service_survey; d_construction_survey; d_retail_survey; d_wholesale_survey; d_employment_survey; d_hh_confidence_indicator; d_retail_sale_survey; d_industry_sentiment; d_service_sentiment; d_cap_utilize_survey; r_tbill_1m; r_tbill_3m; r_tnote_1yr; r_bond_10yr; r_saving; r_industrial; r_housing; r_newloan; turnover_indus; turnover_service; retail_sales; turning_point_survey; in- dustry_turn_survey; service_turn_survey; construction_turn_survey; wholesale_turn_survey; surprise_indicator; unemply- oment_t12_survey; live_standard_p12; current_order_book; d_foreign_order_survey; d_total_order_survey; d_employment_survey_m; emp_forecast_survey; d_production_survey; production_forecast_survey; inventory_survey.
  - India: dl_ex_sugar; dl_ex_spices; dl_ex_rice; dl_cpi; dl_ecma; dl_agri; dl_consu_petro_coke; dl_consu_gas; dl_consu_motor; dl_consu_diesel; dl_consu_petro; d_pmi_new; d_ip; reverse_repo; bank_rate.
  - Mauritius: dl_cpi; dl_stock_semdex; dl_stock_demex; dl_stock_semtri; dl_stock_demtri; dl_stoc_sem10; dl_gross_reserve; dl_fx_euro; dl_fx_gbp; dl_fx_usd; dl_fx_rand; repo; d_tourist_arr.
  - Türkiye: dl_ppi_int; dl_ppi_dur; dl_ppi_ene; dl_ppi_cap; dl_ppi_mine; dl_ppi_manu; dl_ppi_f; dl_tot_veh_prod; dl_tot_auto_prod; dl_gross_electricity; dl_export_oil; dl_export_gem; dl_export_capital; dl_export_interm; dl_export_consu; dl_export_auto; dl_cpi; dl_cpi_f; dl_cpi_h; dl_cpi_g; dl_cpi_e; dl_cpi_s; dl_ppi; dl_p_gold; dl_ipi; dl_ipi_mine; dl_ipi_manu; dl_ipi_gas; dl_ipi_food; dl_ipi_gar; dl_ipi_electronic; dl_ipi_machine; dl_ipi_car; dl_motor_sale; dl_commercial_vehicle_sale; dl_tourism_arr; dl_cli; dl_consu_confi; dl_export_fish; dl_export_mine; dl_export_manu; dl_neer; dl_reer_cpi_jpm; dl_reep_ppi_jpm; dl_reer_cpi_cb; dl_reer_cpi_developed; dl_reer_cpi_developing; dl_try_us_avg; dl_stock_index_avg; dl_int_reserve; dl_cb_reserve_asset; dl_fx_reserve_commerical; dl_m3; i_consu; i_deposit.
  - Vietnam: dl_retail_sales; dl_im; dl_cpi; dl_powder_milk; dl_stock_index; dl_int_visitor; dl_reer; dl_liquid_gas; dl_been; dl_seafood; dl_cig; dl_paint; dl_tv; dl_cement; dl_motorbike; dl_garment; dl_coal; dl_polyester; dl_electricity; dl_water; dl_cotton; dl_retail_service; d_pmi_manu_cap; d_pmi_manu; d_pmi_manu_new; d_pmi_manu_new_ex.

*Source: IMF Working Paper — Introduction, 3.1, 3.2, 4.1–4.2, 5.2, Appendix I sections of the provided chapter excerpt*

### Introduction ...........................................................................................................

### Introduction

### Nowcasting context and purpose
- Nowcasting exploits high-frequency information to monitor economic activity and has become essential in academia, finance, and government.
- Stock and Watson (2017) regarded nowcasting as one of the top ten developments in twenty years of time series econometrics research.
- The paper categorizes nowcasting methods into two classes: (1) traditional econometric models and (2) machine learning algorithms (ML).

### Evaluated models (Table 1)
- Traditional Econometric Models:
  - Bridge
  - MIDAS
  - U-MIDAS
  - MIDAS-PC
  - U-MIDAS-PC
  - Three Pass Regression Filter (TPRF)
  - Dynamic Factor Model (DFM)
- Machine Learning Algorithms:
  - LASSO
  - Elastic Net
  - Decision Tree (DT)
  - Random Forest (RF)
  - Light GBM
  - XGBoost
  - Support Vector Regression – Radial basis (SVRR)
  - Support Vector Regression – Linear basis (SVRL)
  - Support Vector Regression – Polynomial basis (SVRP)
  - Neutral Network (NNET)
  - Long Short-Term Memory Neutral Network (LSTM)
  - Ensemble methods: LASSO-DT, LASSO-DT-SVRP, LASSO-SVRP-NNET, and LASSO-RF-SVRR

### Key empirical scope and approach
- The paper thoroughly evaluates real-time GDP nowcasting performance of an extensive set of models in a single-country time series context, covering:
  - multiple country cases,
  - both theoretical (simulation) and empirical settings,
  - a broad range of models.
- Countries analyzed in the pseudo-real-time empirical section: China, France, India, Mauritius, Türkiye, and Vietnam.
- In simulations the benchmark model is AR(1); in empirical evaluation the benchmark is the automatic ARIMA model.
- Performance metric: out-of-sample root-mean-square errors (OOS RMSE) reported relative to a benchmark over an evaluation period (test set).
  - Relative OOS RMSE is computed by:
    Relative OOS RMSE = OOS RMSE of a model, f (over i in Test: (y_i − f(X_i))) / OOS RMSE of a benchmark model (over i in Test: (y_i − ŷ_benchmark_i))

### Simulation findings (linear and nonlinear DGPs)
- TPRF often ranks first in both linear and non-linear Data Generating Processes (DGPs).
- Linear ML algorithms such as Elastic Net and Lasso frequently follow TPRF in performance.
- Among ML algorithms, linear algorithms (Elastic Net, Lasso) tend to outperform more complex ones including Random Forest, SVRL, SVRP, SVRR, and neural networks (NNET).
- Linear ML algorithms emerge as best performers when:
  - the number of observations is sufficiently large, i.e. exceeding 100,
  - and the DGP is linear.
- When the sample size is small, TPRF performs best.

### Empirical findings (six-country pseudo-real-time nowcasting)
- In most of the six countries, traditional econometric models, particularly Bridge or DFM, are top performers.
- Bridge model delivers strong nowcasting performance in all six countries.
- In Türkiye and India, linear ML algorithms—Lasso or Elastic Net—are the best performers.
- Across countries, linear ML algorithms outperform complex, non-linear algorithms, suggesting limited nonlinearity or complex interactions in the underlying DGP.
- Both simulation and empirical analyses support using traditional econometric models (Bridge or DFM) and linear ML algorithms for single-country GDP nowcasting.
- Complex non-linear models can overfit short GDP series, producing poorer nowcasting performance relative to simple models.
- Caveat: conclusions may depend on simulation design, high-frequency indicators, and choice of countries.

### Data constraints and practical considerations
- A key challenge: short span of GDP data for many countries.
  - Over half of the countries in the world have fewer than 100 GDP observations, with maximum around 300 observations.
- Short sample sizes increase overfitting risk for complex/highly parameterized models.
- Forecasting high-frequency indicators is integral to nowcasting; the paper evaluates models at quarter-end assuming all high-frequency indicators are fully released (i.e., it does not analyze errors from forecasting indicators).
- Frequency mismatch between high-frequency indicators and low-frequency targets is typically mitigated via temporal aggregation (averaging or blocking). The paper reports the lower RMSE from these two choices.

### Relation to existing literature
- Early nowcasting methods: Bridge (Baffigi et al. (2004)), MIDAS (Ghysels et al. (2004, 2007)), U-MIDAS (Forini et al. (2011)), MIDAS-PC and U-MIDAS-PC, Dynamic Factor Model (DFM) (Giannone et al. (2008); Doz et al. (2011)), and Three-Pass Regression Filter (TPRF) (Kelly and Pruitt (2015)).
- Mixed evidence in the literature: no single model is universally superior; performance varies by data, country, and period (Dauphin et al. (2022); Forini and Marcellino (2013, 2014)).
- Recent ML applications have shown mixed results: some studies find ML (including tree-based and boosting methods) superior in specific contexts, while others find linear ML (Ridge, Lasso, Elastic Net) or traditional models superior depending on country and sample.
- The paper positions itself as providing a rigorous valuation comparing an exhaustive list of ML algorithms and traditional econometric models across multiple countries and simulations.

*Source: IMF Working Paper — Introduction section*

### 3.1  Traditional econometric models

### 3.1  Traditional econometric models

### Bridge
- Purpose: Convert high-frequency indicators into low-frequency (quarterly) aggregates (period averages, sums, or end-period values) so ordinary least squares (OLS) can be applied.
- Key practical note: Unreleased high-frequency indicators must be forecasted to aggregate into the low frequency.
- Model specification (as given):
  - y_q = β_0 + sum_{l=1}^m γ_l y_{q−l} + sum_{i=1}^n β_i x_{i.q} + ε_q

### MIDAS
- Purpose: Mixed-frequency regression that splits high-frequency indicators into low-frequency variables relative to quarter end.
- Implementation detail: Splits monthly indicator into three quarterly variables (last month, second-to-last, first month of a quarter); each split variable adds a coefficient unless restrictions imposed.
- Coefficient reduction: Almon polynomial restrictions of degree 2 (δ_{i,h}) used to reduce parameters to two per indicator (w_{1,i}, w_{2,i}).
- Estimation: Non-linear least squares; AR terms included (methodology of Clements and Galvo (2008, 2009)).
- Model specification (as given):
  - y_q = β_0 + sum_{l=1}^m γ_l y_{q−l} + sum_{i=1}^n sum_{h=1}^{L(i)} δ_{i,h} x_{i.3q−h} + ε_q
  - δ_{i.h}(w_{1,i},w_{2,i}) = exp(w_{1,i} h + w_{2,i} h^2) / sum_{h=1}^{L(i)} exp(w_{1,i} h + w_{2,i} h^2)

### U-MIDAS (Unrestricted MIDAS)
- Purpose: Split high-frequency indicators into multiple low-frequency variables and estimate without coefficient restrictions using OLS.
- Best-suited when frequency mismatch is minimal.
- Model specification (as given):
  - y_q = β_0 + sum_{l=1}^m γ_l y_{q−l} + sum_{i=1}^n sum_{h=1}^{L(i)} β_{i,h} x_{i.3q−h} + ε_q

### MIDAS-PC
- Purpose: Mitigate parameter proliferation when many high-frequency indicators exist by using principal components of high-frequency indicators as regressors in a MIDAS framework.
- Estimation: Similar to MIDAS but with principal components replacing original indicators.

### U-MIDAS-PC
- Purpose: Unrestricted MIDAS using principal components; when the number of factors is small, coefficient restrictions are unnecessary and OLS can be applied in low frequency.

### Three-pass Regression Filter (TPRF)
- Purpose: Extracts factors that maximally correlate with the forecast target; assigns loadings to indicators based on correlation with real GDP growth.
- Estimation steps (as given):
  - Step 1: For each high-frequency indicator, run time series regression of the nowcast target on the indicator:
    - y_q = β_0 + β_i x_{i,q} + ε_{i,q}
  - Step 2: For each time period t, run a cross-sectional regression of x_{i,t} on the beta coefficients from Step 1 to yield TPRF_t:
    - x_{i,t} = β_0 + TPRF_t β_i + ε_i
  - Step 3: Use the TPRF factor in the nowcasting regression in low frequency:
    - y_q = β_0 + sum_{l=1}^m γ_l y_{q−l} + θ TPRF_q + ε_q
- Practical note: No look-ahead bias because Steps 1 and 2 use lagged GDP; nowcast uses current-quarter TPRF_q.

### Dynamic Factor Model (DFM)
- Purpose: Address frequency mismatch and ragged-edge problems; allows immediate nowcast updates when high-frequency indicators are released asynchronously.
- Estimation: Kalman Filtering (Doz et al. (2011)).
- Notation and system (as given):
  - Let Z_q = [y_q, x_{1,q}, ..., x_{N,q}] ( (N+1)×q matrix ), F = [f_1, ..., f_r] ( r×q matrix ).
  - Measurement equation: z_{i,q} = μ + Λ F_q + ε_{i.q}, z_{i,q} ∈ Z_q
  - Transition equation: F_q = A F_{q−1} + B u_q, u_q ∼ N(0, I)
  - Assumptions: ε'_{i.q} are cross-sectionally white-noise and orthogonal to u_q.
- Nowcast procedure: Regress y_q on estimated factor ŜF_q with controls, then compute nowcast using current value of ŜF_q.

---

### 3.2  Machine Learning Algorithms

### Cross-validation and tuning
- Approach: Five-fold cross-validation across various hyperparameter configurations to select parameters.
- Implementation note: Table 2 summarizes parameter settings and total number of parameter combinations used for tuning each ML algorithm.

### Lasso
- Purpose: Variable selection via L1 shrinkage; yields parsimonious models by shrinking coefficients toward or equal to zero.
- Objective function (as given):
  - β = argmin ( sum_{q=1}^T (y_q − β_0 − sum_{j=1}^k x_{j,q} β_j)^2 + λ sum_{j=1}^k |β_j| )

### Elastic Net
- Purpose: Combines L1 (Lasso) and L2 (Ridge) penalties with mixing parameter α.
- Objective function (as given):
  - β = argmin ( sum_{q=1}^T (y_q − β_0 − sum_{j=1}^k x_{j,q} β_j)^2 + λ [ α sum_{j=1}^k |β_j| + (1−α) sum_{j=1}^k β_j^2 ] )
- Tuning: 3,150 different combinations of λ and α using 5-fold cross-validation (as described in Table 2).

### Decision Tree (DT)
- Structure: Root node, internal nodes, branches, leaf nodes.
- Complexity control: cp (complexity parameter) controls pruning; smaller cp allows more splits.

### Random Forest
- Principle: Ensemble of decision trees; tuned via "mtry" parameter controlling number of variables randomly sampled at each split.

### LightGBM
- Purpose: Efficient gradient boosting framework; tested 18 different parameter combinations via grid search in implementation.
- Tuned parameters summarized in Table 2:
  - n_estimators: 50, 100 (2 values)
  - max_depth: 2, 4, 8 (3 values)
  - learning_rate: 0.01, 0.05, 0.1 (3 values)

### XGBoost
- Purpose: Gradient boosting with regularization and sparse-aware learning; tuned across 18 combinations in implementation.
- Tuned parameters summarized in Table 2:
  - n_estimators: 50, 100 (2 values)
  - max_depth: 2, 4, 8 (3 values)
  - learning_rate: 0.01, 0.05, 0.1 (3 values)

### Support Vector Regression (SVR)
- Purpose: Fit a hyperplane in continuous space using kernel functions; evaluated with three kernels:
  - SVRL (Linear Kernel)
    - C ∈ [10^{−3}, 10^2] (log scale), total of 10 values
  - SVRR (Radial Kernel)
    - σ ∈ {0.001, 0.01, 0.05, 0.1, 0.2} (5 values)
    - C ∈ {0.1, 1, 10, 100, 1000} (5 values)
  - SVRP (Polynomial Kernel)
    - C ∈ [10^{−2}, 10^2] (log scale), 5 values
    - Degree ∈ {2, 4, 8} (3 values)
    - Scale ∈ {0.1, 1, 10} (3 values)
- SVR objective (as given):
  - L(ω,b) = 1/2 ||ω||^2 + C sum_{i=1}^n L_ε(y_i, f(x_i))

### Deep Neural Network (NNET)
- Architecture: Feedforward neural network with a single hidden layer.
- Tuning grid: Hidden layer sizes {3, 5, 10} (3 values); Regularization (weight decay) ∈ {0.0001, 0.001, 0.01, 0.1} (4 values) — 12 total combinations tuned.

### Long Short-Term Memory (LSTM)
- Purpose: Capture long-term dependencies for sequence prediction.
- Tuning via Keras Tuner (RandomSearch, 5 max trials).
- Tuning grid summarized in Table 2:
  - LSTM Layers: 1, 2, 3 (3 values)
  - Units per Layer: 16 to 128 (step 16, tuned per layer) (8 values)
  - Activation: tanh
  - Learning Rate: 0.0001 to 0.01 (log scale)
  - Epochs: 30
  - Batch Size: 8
  - Validation Split: 0.1

### Ensemble learning
- Approach: Simple averaging ensembles of selected models; no extra parameters beyond constituent models.
- Ensembles constructed:
  - LASSO_DT: average of LASSO and Decision Tree predictions
  - LASSO_DT_SVRP: average of LASSO, Decision Tree, and SVRP predictions
  - LASSO_RF_SVRBF: average of LASSO, Random Forest, and SVRR predictions
  - LASSO_SVRP_NN: average of LASSO, SVRP, and Neural Network predictions

---

### 4.1  Simulation setup and data-generating processes (DGPs)

### Factor structure and notation
- Frequency mismatch addressed: monthly observed data and quarterly nowcast targets, with q = 3m.
- Unobservable factor F(m) drives both high-frequency data X(m) and low-frequency nowcast target y(q).
- Additional unobservable factors G(m) drive observed high-frequency data only.
- Both F and G follow autoregressive structures with heterogeneous loadings and independent errors.
- Process equations (as given):
  - F = (f_1, f_2, ..., f_k)
  - G = (f_1, f_2, ..., g_l)
  - F_t = λ_i F_{t−1} + η_{f t}, η_{f t} ∼ iid N(0,1)
  - G_t = λ_j G_{t−1} + η_{g t}, η_{g t} ∼ iid N(0,1)
  - λ ∼ Uniform(λ_{Lo}, λ_{Hi})
- High-frequency data X(m) composed of unobservable components X^U(m) and observed components X^O(m):
  - X^U(i,t) = β F + ε^U, ε^U ∼ iid N(0,1), j = 1 to k
  - X^O(i,t) = β_U F + β_O G + ε^O, ε^O ∼ iid N(0,1), j = k+1 to N
  - β, β_U, β_O ∼ Uniform(β_{Lo}, β_{Hi})
- Only unobservable components enter DGP for nowcast target Y; observed components are correlated with Y via common factors F.
- Monthly partitioning: each month of the quarter impacts the nowcast target and is denoted X^U(I, 3m−0), X^U(I, 3m−1), X^U(I, 3m−2).

### Simulation parameter values (as given)
- K = 5 and N = 105 (i.e., five unobservable components generate GDP; total predictors N = 105).
- Components X_1, X_2, and X_3 are included in all DGPs.
- Loadings:
  - X_1 loadings homogeneous across each month in a quarter.
  - X_2 and X_3 loadings heterogeneous across months in a quarter.

### True DGP specifications (as given)

- DGP 1: Linear Data-Generating Process.
  - y_t = 1.2 + 0.5 x^{(0)}_1 + 0.5 x^{(−1)}_1 + 0.5 x^{(−2)}_1 + 0.3 x^{(0)}_2 + 0.6 x^{(−1)}_2 + 0.4 x^{(−2)}_2 − 0.25 x^{(0)}_3 − 0.2 x^{(−1)}_3 − 0.3 x^{(−2)}_3

- DGP 2: Non-Linear Quadratic Data-Generating Process.
  - y_t = 1.2 + 0.5 x^{(0)}_1 + 0.5 x^{(−1)}_1 + 0.5 x^{(−2)}_1 + 0.3 x^{(0)}_2 + 0.6 x^{(−1)}_2 + 0.4 x^{(−2)}_2 − 0.25 x^{(0)}_3 − 0.2 x^{(−1)}_3 − 0.3 x^{(−2)}_3 + 0.05 X^2_{2,(0)} + 0.02 X^2_{2,(−1)} + 0.03 X^2_{2,(−2)}

- DGP 3: Non-Linear Data-Generating Process with Quadratic and Interaction Terms.
  - y_t = 1.2 + 0.5 x^{(0)}_1 + 0.5 x^{(−1)}_1 + 0.5 x^{(−2)}_1 + 0.3 x^{(0)}_2 + 0.6 x^{(−1)}_2 + 0.4 x^{(−2)}_2 − 0.25 x^{(0)}_3 − 0.2 x^{(−1)}_3 − 0.3 x^{(−2)}_3 + 0.05 X^2_{2,(0)} + 0.02 X^2_{2,(−1)} + 0.03 X^2_{2,(−2)} + 0.15 X_{3,(0)} X_{4,(0)} + 0.15 X_{3,(−1)} X_{4,(−1)} + 0.15 X_{3,(−2)} X_{4,(−2)} − 0.2 X_{4,(0)} X_{5,(0)} − 0.3 X_{4,(−1)} * X_{5,(−1)} − 0.1 X_{4,(−2)} X_{5,(−2)}

### Conceptual implications highlighted by the simulation setup
- The true DGP for GDP is unknown due to the unobserved nature of X_1 to X_5.
- The true DGP could be nonlinear (DGPs 2 and 3), which linear models may not approximate well.
- High-frequency indicators are relevant but imperfect predictors of GDP: they share common factor F with GDP but also load on unrelated factors G, creating noise and potential mis-specification for nowcasting models.

*Source: wpiea2025252-source-pdf - 3.1  Traditional econometric models*

### 4.2  Monte Carlo Experiments

### 4.2 Monte Carlo Experiments

### Monte Carlo setup
- Number of Monte Carlo data draws: 500.
- Number of factors: F and G set to 5 and 3, respectively.
- Nowcast target generated solely by unobserved X1 to X5, influenced only by factors F.
- Number of observed high-frequency indicators: 120.
- High-frequency indicators derived from both factors F and G.
- Parameter draws per iteration: ηf, ηg, β, and λ generated to create synthetic data (500 iterations).

### Parameterization and stochastic components
- λ ∼ iid Uniform(0.4,0.6)
- β ∼ iid Uniform(0,0.9)
- ηf, ηg, ε ∼ iid N(0,1)
- Observed correlations between simulated X and nowcast targets range from 0.2 to 0.6.
- Interpretation: factors are independent and the data exhibits mild persistence; the β range implies some observed indicators may have significant loading on the factor driving GDP (F).

### Time-dimension variation and training/test split
- Quarterly time dimensions considered: Tq = 40,50,60,70,80,100,120,150,200.
- Corresponding monthly frequencies: Tm = 120,150,180,210,240,300,360,450,600 respectively.
- For each time dimension, 500 iterations of ηf and ηg, β, and λ are generated.
- For each iteration, data in the time dimension are randomly split: 70% training set, 30% test set.
  - Example: when T = 80, training and test sets comprise 56 and 34 observations of XU and Y, respectively.
- No rolling updates performed in theoretical simulation to calculate test-set error.

### Estimation and variable selection procedures
- Relative out-of-sample RMSE on the test set is the metric for evaluating model performance.
- Traditional econometric models face variable-selection challenges when regressors exceed time periods.
  - Recursive Feature Elimination (RFE) is applied algorithmically to add/drop indicators with a constraint on the maximum number of variables based on time observations.
  - The number of indicators selected ranges from 2 to the number of quarterly time periods divided by 5.
  - The same selected indicators are then used across traditional econometric models and neural networks to isolate model performance differences.
- Many ML algorithms include built-in feature selection; ML simulation benefits from these techniques.
- In the simulation, 361 low-frequency variables (split from 120 HF indicators) can far exceed available time observations (maximum of 150).
- Estimation for ML models uses 5-fold cross-validation within the training set to select parameters that minimize RMSE within each fold.
- Monthly-to-quarterly frequency conversion for empirical application is discussed in Section 5 but noted here: monthly indicators can be split into three quarterly variables (split-sampling/blocking) resulting in 360 quarterly variables from 120 monthly indicators, plus one extra variable for the first lag of Y.

### Key simulation findings (summary of Section 4.3 results)
- Overall model ranking tendencies:
  - TPRF frequently ranks first across DGPs, followed by linear ML algorithms such as Elastic Net and Lasso.
  - Linear ML algorithms can outperform when conditions favor linearity and sufficient observations.
- Specific insights:
  - When the number of time observations exceeds 100 and the DGP is linear, Elastic Net outperforms all models.
  - TPRF delivers superior performance in DGPs 2 and 3 despite nonlinearity; possible explanations include TPRF proxying non-linearity via factor loadings and mismatch between ML functional forms and DGP quadratic/interaction terms.
  - TPRF, MIDAS-PC, U-MIDAS-PC, Elastic Net, and Lasso perform well in short samples (e.g., T = 40) due to factor modeling or regularization that limits overfitting.
  - When the DGP is highly nonlinear with interactions, both model classes may underperform; TPRF often remains top but without economically or statistically significant outperformance. In such cases, non-linear SVR with Radial basis achieves the best results by a small margin over linear ML algorithms.
- Overall conclusion from simulations:
  - Traditional econometric models (factor-based) or linear ML algorithms (with regularization) are favored, especially given short GDP series where complex non-linear models risk overfitting.
  - Results may depend on simulation design and parameterization; empirical testing on six countries follows.

*Source: 4.2 Monte Carlo Experiments (and summary of 4.3 Results) from the provided IMF chapter excerpt.*

### 5.2  Performance Evaluation

### 5.2  Performance Evaluation

### Evaluation setup and methodology
- Pseudo-real-time out-of-sample evaluation at the end of each quarter using an expanding rolling window.
- Evaluation window: 2018Q1 to 2024Q1 (test period described elsewhere as 2018Q1–2024Q2 for OOS RMSE computation).
- Each model is re-estimated every quarter upon the release of new GDP data; model coefficients vary quarterly and are used to generate nowcasts for each month within that quarter.
- Simplification: high-frequency indicators are assumed available at the end of each quarter prior to the release of real GDP for that quarter (not using actual data release dates).
- GDP revisions are not considered (evaluation is pseudo-real-time).
- No look-ahead bias: all model coefficients estimated using previous-quarter data; current quarter data are then inserted to derive nowcasts.
- Nowcasts are compared with actual real GDP growth (in log).
- Reported performance metric: out-of-sample RMSEs (OOS RMSE) relative to a univariate automatic ARIMA model (Automatic ARIMA implemented via the “autoarima” package in R, Box-Jenkins methodology).

### Nowcast evaluation results (aggregate findings)
- Traditional econometric models tend to outperform ML algorithms in most cases across simulations and country applications.
- Best-performing models across six countries: Bridge and DFM each rank first in two countries.
  - When DFM outperforms Bridge, the margin is substantial.
  - When Bridge outperforms DFM, the margin is relatively modest.
- Bridge models consistently rank within the top four models in five of the six countries.
- Linear ML algorithms outperform nonlinear ML methods across all six countries.
  - Elastic Net and Lasso perform best among ML methods, except in India where SVRL slightly outperforms Lasso by a narrow margin.
- In Türkiye and Mauritius, linear ML algorithms are the best performers.
- Simpler models with fewer parameters tend to perform well, with the notable exception of DFM.
- Evidence suggests limited nonlinearity or complex interactions in the underlying data-generating process (DGP); even preferred complex ML algorithms may rely on linear basis functions.

### COVID-19 episode (2020Q1–2022Q4) performance
- Nowcasting during COVID-19 is particularly challenging.
- In China, France, and Türkiye, traditional econometric models or linear ML algorithms outperform the autoARMA benchmark.
- In the other three countries (India, Mauritius, Vietnam), both classes of models fail to beat the benchmark during COVID-19.
- Commonalities among successful cases during COVID-19: a large number of high-quality high-frequency indicators and a sufficiently long analysis sample.
- Best performers during COVID-19 in specific countries:
  - China: TPRF performs best.
  - France: MIDAS performs best.
  - Türkiye: Elastic Net performs best.
- Noted external example: Federal Reserve Bank of New York’s temporary suspension of external publication of the New York Fed Staff Nowcast during the COVID-19 period.

### Ensemble forecasts and ML comparisons
- Three ensemble combinations among ML algorithms were evaluated:
  1. LASSO-SVRP-NNET
  2. LASSO-RF-SVRR
  3. LASSO-DT
- Ensemble methods do not outperform the best traditional econometric models or the linear ML algorithms in this context.
- Nonlinear ML algorithms risk overfitting given the limited number of GDP observations, reducing nowcasting performance.
- Among ML algorithms, simpler linear methods (Lasso, Elastic Net) tend to outperform more complex counterparts.

### Key raw OOS RMSE statistics (selected reporting structure preserved)
- Panel A: All periods (average OOS RMSEs over 2018Q1–2024Q2)
  - Benchmark and traditional econometric model RMSEs by country (selected values shown as in source):
    - China: AutoARMA 0.06; Bridge 0.02; UMIDAS 0.04; MIDAS 0.05; UMIDAS_PC 0.04; MIDAS_PC 0.04; TPRF 0.03; DFM 0.01
    - France: AutoARMA 0.06; Bridge 0.02; UMIDAS 0.02; MIDAS 0.04; UMIDAS_PC 0.05; MIDAS_PC 0.04; TPRF 0.05; DFM 0.02
    - India: AutoARMA 0.08; Bridge 0.03; UMIDAS 0.05; MIDAS 0.06; UMIDAS_PC 0.06; MIDAS_PC 0.06; TPRF 0.04; DFM 0.02
    - Mauritius: AutoARMA 0.08; Bridge 0.06; UMIDAS 0.13; MIDAS 0.23; UMIDAS_PC 0.27; MIDAS_PC 0.14; TPRF 0.06; DFM 0.08
    - Türkiye: AutoARMA 0.05; Bridge 0.02; UMIDAS 0.02; MIDAS 0.03; UMIDAS_PC 0.03; MIDAS_PC 0.02; TPRF 0.05; DFM 0.02
    - Vietnam: AutoARMA 0.03; Bridge 0.01; UMIDAS 0.03; MIDAS 0.02; UMIDAS_PC 0.03; MIDAS_PC 0.03; TPRF 0.02; DFM 0.02
  - Machine learning algorithm RMSEs by country (selected values shown as in source):
    - China (DT 0.03; Elastic Net 0.02; LASSO 0.03; NNET 0.03; RF 0.03; SVRL 0.03; SVRP 0.04; SVRR 0.03; LightGBM 0.03; XGBoost 0.03; LSTM 0.03)
    - France (DT 0.04; Elastic Net 0.02; LASSO 0.03; NNET 0.04; RF 0.04; SVRL 0.03; SVRP 0.04)
    - India (DT 0.07; Elastic Net 0.04; LASSO 0.04; NNET 0.06; RF 0.07; SVRL 0.04)
    - Mauritius (DT 0.08; Elastic Net 0.06; LASSO 0.06; NNET 0.08; RF 0.09)
    - Türkiye (DT 0.04; Elastic Net 0.01; LASSO 0.02; NNET 0.05; RF 0.04)
    - Vietnam (DT 0.02; Elastic Net 0.01; LASSO 0.01; NNET 0.02; RF 0.02)
- Panel B: COVID-19 period (average OOS RMSEs over 2020Q1–2022Q4)
  - Benchmark and traditional econometric model RMSEs by country (selected values shown as in source):
    - China: AutoARMA 0.05; Bridge 0.03; MIDAS 0.03; UMIDAS 0.03; MIDAS-PC 0.02; UMIDAS-PC 0.02; TPRF 0.02; DFM 0.04
    - France: AutoARMA 0.05; Bridge 0.04; MIDAS 0.02; UMIDAS 0.05; MIDAS-PC 0.03; UMIDAS-PC 0.04; TPRF 0.04; DFM 0.03
    - India: AutoARMA 0.04; Bridge 0.10; MIDAS 0.10; UMIDAS 0.06; MIDAS-PC 0.10; UMIDAS-PC 0.06; TPRF 0.10; DFM 0.10
    - Mauritius: AutoARMA 0.01; Bridge 0.13; MIDAS 0.30; UMIDAS 0.17; MIDAS-PC 0.12; UMIDAS-PC 0.31; TPRF 0.10; DFM 0.07
    - Türkiye: AutoARMA 0.02; Bridge 0.08; MIDAS 0.05; UMIDAS 0.06; MIDAS-PC 0.07; UMIDAS-PC 0.05; TPRF 0.02; DFM 0.06
    - Vietnam: AutoARMA 0.01; Bridge 0.04; MIDAS 0.02; UMIDAS 0.03; MIDAS-PC 0.02; UMIDAS-PC 0.03; TPRF 0.03; DFM 0.03
  - Machine learning algorithm RMSEs by country (selected values shown as in source):
    - China (DT 0.05; Elastic Net 0.04; LASSO 0.04; NNET 0.04; RF 0.05; SVRL 0.04; SVRP 0.05)
    - France (DT 0.06; Elastic Net 0.04; LASSO 0.04; NNET 0.06; RF 0.06)
    - India (DT 0.10; Elastic Net 0.06; LASSO 0.06; NNET 0.09; RF 0.10; SVRL 0.05)
    - Mauritius (DT 0.10; Elastic Net 0.06; LASSO 0.06; NNET 0.09; RF 0.10)
    - Türkiye (DT 0.06; Elastic Net 0.02; LASSO 0.02; NNET 0.06; RF 0.05)
    - Vietnam (DT 0.03; Elastic Net 0.02; LASSO 0.02; NNET 0.02; RF 0.03)

### Conclusion and implications
- Simulations and six country applications show traditional econometric models often deliver superior nowcasting performance to ML algorithms.
- In simulations, TPRF mostly outperforms others, losing only to Elastic Net under the linear DGP and long data.
- In the six country applications, DFM or Bridge outperform ML algorithms in four out of six countries.
- Among ML algorithms, Lasso and Elastic Net tend to outperform more complex ML algorithms and even surpass traditional econometric models in two out of six countries.
- Practical considerations:
  - Number of GDP observations is relatively short by ML standards; small sample sizes hamper complex models and increase overfitting risk.
  - Relationship between GDP growth and high-frequency indicators may be adequately captured by linear models.
- Recommendation: use simpler models like Bridge or linear ML algorithms for GDP nowcasting due to consistent superior performance and lower overfitting risk; DFM is a notable complex model that performs well in some cases.
- Caveats:
  - No universally superior model; performance may vary for other countries or data sets.
  - Analysis primarily relies on traditional macroeconomic data; using non-traditional data (e.g., satellite data) might yield different results.
  - Study focuses on single-country time-series applications and does not extend to panel data setups.

*Source: IMF working paper content (5.2 Performance Evaluation and related sections).*

### Appendix I

### Appendix I

### Notation and data source
- ( dl ) represents the log difference.
- ( d ) denotes the first difference.
- All indicators are sourced from Haver Analytics.
- Variable names are abbreviated for coding efficiency but remain intuitive and interpretable.

### China
- dl_cpi; dl_export_value; dl_export_volume; dl_import_value; dl_import_volume; dl_p_steel; dl_invest_real_estate;
- dl_fiscal_rev; dl_fiscal_exp; dl_yuan_usd; dl_m2; dl_m1; dl_ip_crude_oil; dl_ip_natural_gas; dl_ip_iron_ore; dl_ip_salt;
- dl_ip_yarn; dl_ip_cotton; dl_ip_gasoline; dl_ip_kerosene; dl_ip_diesel; dl_ip_coke; dl_output_sulfuric; dl_output_sodiumhy;
- dl_output_sodiumcar; dl_output_ethylene; dl_output_fertilizer; dl_output_chemical; dl_output_plastic; dl_output_synthetic;
- dl_ip_plastic; dl_output_cement; dl_output_plate; dl_output_pigiron; dl_output_steel; dl_output_steelprod; dl_output_nonferrous;
- dl_ip_aluminumo; dl_output_copper; dl_ip_aluminumprod; dl_output_boiler; dl_output_icecar; dl_output_metalcut;
- dl_output_electric_tool; dl_ip_packaging; dl_ip_cement_equip; dl_ip_smelt; dl_ip_feeding; dl_ip_tractor; dl_ip_airclean;
- dl_ip_railway; dl_ip_motor; dl_ip_ship; dl_ip_powergen; dl_ip_acmotor; dl_ip_electricity_prod; dl_output_fridge; dl_output_freezer;
- dl_output_ac; dl_output_wastemac; dl_ip_mobilehandset; dl_output_microcomputer; dl_output_semiconductor;
- dl_output_electricmeasure; dl_output_copymac; dl_freight_traffic; dl_freight_turnover; dl_passenger_traffic; dl_passenger_turnover;
- dl_volume_port; dl_volume_foreigntradegoods; dl_no_expressmail; dl_no_cellphoneuser; dl_epu; d_fiscal_balance; d_neer;
- d_reer; d_pmi_manu; d_pmi_output; d_pmi_neworder; d_pmi_majorinput; d_pmi_emplotment; d_pmi_deliverytime;
- d_pmi_backlog; d_pmi_inputp; d_pmi_inputbuy; d_pmi_finishedgoods; d_pmi_newexport; d_pim_import; d_consumer_confidence.

### France
- dl_us_eur; dl_m1; dl_m2; dl_m3; dl_private_credit; dl_m1_dom; dl_m2_dom; dl_cb_asset; dl_saving; dl_unreg_saving;
- dl_credit_to_private; dl_loan_to_gov; dl_business_creation; dl_gov_rev; dl_gov_exp; dl_cpi; dl_core_cpi; dl_ppi; dl_imp_price;
- dl_brent_oil; dl_house_start; dl_house_permit; dl_no_job_seeker; dl_job_vacan; dl_job_vacan_dom; dl_job_vacan_skill;
- dl_job_vacan_unski; dl_hh_consumption; dl_export; dl_import; dl_num_dealth; dl_construction_creation; d_fis_bal;
- d_income_tax; d_corp_tax; d_vat; d_trade_bal; d_busness_climate_indicator; d_epu; d_manufacture_survey; d_service_survey;
- d_construction_survey; d_retail_survey; d_wholesale_survey; d_employment_survey; d_hh_confidence_indicator; d_retail_sale_survey;
- d_industry_sentiment; d_service_sentiment; d_cap_utilize_survey; r_tbill_1m; r_tbill_3m; r_tnote_1yr; r_bond_10yr;
- r_saving; r_industrial; r_housing; r_newloan; turnover_indus; turnover_service; retail_sales; turning_point_survey; in-
- dustry_turn_survey; service_turn_survey; construction_turn_survey; wholesale_turn_survey; surprise_indicator; unemply-
- oment_t12_survey; live_standard_p12; current_order_book; d_foreign_order_survey; d_total_order_survey; d_employment_survey_m;
- emp_forecast_survey; d_production_survey; production_forecast_survey; inventory_survey.

### India
- dl_ex_sugar; dl_ex_spices; dl_ex_rice; dl_cpi; dl_ecma; dl_agri; dl_consu_petro_coke; dl_consu_gas; dl_consu_motor;
- dl_consu_diesel; dl_consu_petro; d_pmi_new; d_ip; reverse_repo; bank_rate.

### Mauritius
- dl_cpi; dl_stock_semdex; dl_stock_demex; dl_stock_semtri; dl_stock_demtri; dl_stoc_sem10; dl_gross_reserve; dl_fx_euro;
- dl_fx_gbp; dl_fx_usd; dl_fx_rand; repo; d_tourist_arr.

### Türkiye
- dl_ppi_int; dl_ppi_dur; dl_ppi_ene; dl_ppi_cap; dl_ppi_mine; dl_ppi_manu; dl_ppi_f; dl_tot_veh_prod; dl_tot_auto_prod;
- dl_gross_electricity; dl_export_oil; dl_export_gem; dl_export_capital; dl_export_interm; dl_export_consu; dl_export_auto;
- dl_cpi; dl_cpi_f; dl_cpi_h; dl_cpi_g; dl_cpi_e; dl_cpi_s; dl_ppi; dl_p_gold; dl_ipi; dl_ipi_mine; dl_ipi_manu; dl_ipi_gas;
- dl_ipi_food; dl_ipi_gar; dl_ipi_electronic; dl_ipi_machine; dl_ipi_car; dl_motor_sale; dl_commercial_vehicle_sale; dl_tourism_arr;
- dl_cli; dl_consu_confi; dl_export_fish; dl_export_mine; dl_export_manu; dl_neer; dl_reer_cpi_jpm; dl_reep_ppi_jpm;
- dl_reer_cpi_cb; dl_reer_cpi_developed; dl_reer_cpi_developing; dl_try_us_avg; dl_stock_index_avg; dl_int_reserve;
- dl_cb_reserve_asset; dl_fx_reserve_commerical; dl_m3; i_consu; i_deposit.

### Vietnam
- dl_retail_sales; dl_im; dl_cpi; dl_powder_milk; dl_stock_index; dl_int_visitor; dl_reer; dl_liquid_gas; dl_been; dl_seafood;
- dl_cig; dl_paint; dl_tv; dl_cement; dl_motorbike; dl_garment; dl_coal; dl_polyester; dl_electricity; dl_water; dl_cotton;
- dl_retail_service; d_pmi_manu_cap; d_pmi_manu; d_pmi_manu_new; d_pmi_manu_new_ex.

*GDP Nowcasting Performance of Traditional Econometric Models vs Machine-Learning Algorithms: Simulation and Case Studies Working Paper No. WP/2025/252*

---


_Source: https://www.imf.org/-/media/files/publications/wp/2025/english/wpiea2025252-source-pdf.pdf_
