## _wp11230 — 1. Model Selection

## Source details

**Canonical URL:** [_wp11230 — 1. Model Selection](https://www.imf.org/-/media/websites/imf/imported-full-text-pdf/external/pubs/ft/wp/2011/_wp11230.pdf)

## Other formats

- [Markdown version](/-/media/websites/imf/imported-full-text-pdf/external/pubs/ft/wp/2011/_wp11230.pdf.md)
- [Structured JSON version](/-/media/websites/imf/imported-full-text-pdf/external/pubs/ft/wp/2011/_wp11230.pdf.json)

---

### I. Introduction: motivation and contribution
- Model uncertainty arises from weak theoretical guidance and trade-offs in regressor choice, producing many plausible specifications.
- Two Bayesian responses:
  - Bayesian model selection: choose single model with largest posterior model probability.
  - Bayesian Model Averaging (BMA): average posterior distributions across models using posterior model probabilities as weights.
- Limitations in prior literature:
  - Most BMA work focuses on static, cross-section models and typically ignores dynamic relationships and endogenous regressors.
  - Existing dynamic-panel / GMM contexts have not fully incorporated model uncertainty allowing endogenous regressors together with lagged dependent variables.
- Contribution:
  - Proposes Limited Information Bayesian Model Averaging (LIBMA) for panel data with short T that permits lagged dependent variable and endogenous regressors.
  - Constructs model likelihoods and posteriors from moment conditions only (no distributional assumptions), yielding a limited-information likelihood based on GMM asymptotics and a unit information prior with a BIC-like posterior approximation.
  - Evaluates LIBMA with Monte Carlo simulations and applies it to a dynamic gravity model for bilateral trade.

### II. Bayesian model uncertainty: key formulas and priors
- Model selection formulas:
  - p(Mj | D) = p(D | Mj) p(Mj) / Σ_{l=1}^K p(D | Ml) p(Ml)
  - p(D | Mj) = ∫ p(D | θj, Mj) p(θj | Mj) dθj
  - Posterior odds: p(Mj | D) / p(Mi | D) = [p(D | Mj) / p(D | Mi)] × [p(Mj) / p(Mi)]
- BMA identities:
  - p(ψ | D) = Σ_{j=1}^K p(ψ | D, Mj) p(Mj | D)
  - Inclusion probability p(Zi | D) = Σ_{j=1}^K I(Zi ∈ Mj) p(Mj | D)
  - E(θl | D) = Σ_{j=1}^K E(θl | D, Mj) p(Mj | D)
- Priors used:
  - Unit information prior: k_l-dimensional Normal with mean the estimate θ̂ and variance I(θ̂)^{-1}.
  - Model prior: uniform over model space, p(Mj) = 1/K.

### III. LIBMA setup: model, regressors, and assumptions
- Universe U partitions: {1} (lagged dependent variable), X (exogenous, dimension m), W (endogenous, dimension q). Total K = m + q + 1.
- Dynamic panel model:
  - y_it = β y_{i,t-1} + x'_it θ_x + w'_it θ_w + u_it
  - u_it = α_i + v_it
- Observed across N individuals and T periods, with T small relative to N.
- Assumptions:
  - E(v_it) = 0; v_it not serially correlated.
  - Exogenous x_it orthogonal to v_is for all i,t,s.
  - Endogenous w_it may be correlated with contemporaneous and past v_is but instruments exist from lags.
- Goal: perform Bayesian model selection/averaging using only moment conditions (no full parametric likelihood).

### IV. Moment conditions and system GMM instruments
- System GMM used to build instruments and moment conditions (level and first-difference equations).
- First-difference instruments / moments:
  - Δy_{i,t-1}: instruments y_{i,t-2}, ..., y_{i,0} giving E(y_{i,t-s} Δv_{it}) = 0 for t = 2,3,...,T; s = 2,3,...,t (requires T ≥ 2).
  - Exogenous x_{l,it}: instrument itself, giving E(x_{l,it} Δv_{it}) = 0 for t = 2,3,...,T; l = 1,...,m.
  - Endogenous w_{l,it}: instruments w_{l,i,t-2}, ..., w_{l,i1}, giving E(w_{l,i,t-s} Δv_{it}) = 0 for t = 3,4,...,T; s = 2,...,t-1 (requires T ≥ 3).
- Counts of first-difference moment conditions:
  - Lagged dependent variable: T(T-1)/2 moment conditions.
  - Exogenous variables: m (T-1) moment conditions.
  - Endogenous variables: q (T-2)(T-1)/2 moment conditions.
- Level-equation instruments / moments:
  - Δy_{i,t-1}: E(Δy_{i,t-1} u_it) = 0 for t = 2,3,...,T.
  - Δx_{l,it}: E(Δx_{l,it} u_it) = 0 for t = 2,3,...,T; l = 1,...,m.
  - Δw_{l,i,t-1}: E(Δw_{l,i,t-1} u_it) = 0 for t = 3,4,...,T; l = 1,...,q (requires T ≥ 3).
- Additional Ahn and Schmidt (1995) linear moments under homoskedastic v_it:
  - E(y_{i,t} u_{i,t} - y_{i,t-1} u_{i,t-1}) = 0 for t = 2,3,...,T.
- Aggregation for exogenous variables produces one moment per exogenous variable:
  - Σ_{t=2}^T E(x_{l,it} Δv_{it}) + Σ_{t=2}^T E(Δx_{l,it} u_it) = 0, for l = 1,...,m.
- Matrix form: all moments written as E[(G'_i U_i)'] = 0 (definitions of G_i and U_i in Appendix I).

### V. Limited-information likelihood and BIC-like posterior criterion
- Challenge: GMM provides moments but not a full parametric likelihood p(D | θ, Mj).
- Approach:
  - Build a limited-information likelihood using only moment conditions and model linearity.
  - Use GMM asymptotics (CLT) to obtain a limited-information likelihood suitable for Bayesian inference.
  - Combine with a unit information prior to produce a posterior with a BIC-like approximation.
- Laplace approximation and asymptotics (selected expressions preserved):
  - Under Lemma 1: N^{1/2} \bar{g}_N(θ0) d→ N(0; S(θ0)).
  - Model likelihood (up to proportionality):
    - ∝ ∫ exp(−1/2 N \bar{g}_N(θ)' S(θ)^{-1} \bar{g}_N(θ)) p(θ) dθ.
  - Laplace approximation with unit information prior yields approximate model likelihood ∝ exp(−1/2 N \bar{g}_N(bθ0) S^{-1}(bθ0) \bar{g}_N(bθ0) − K/2 log N), where K = dim(θ).
- Posterior odds (BIC-like):
  - For models M1 and M2:
    - p(M1 | data) / p(M2 | data) = p(M1)/p(M2) × exp(−1/2 N \bar{g}_N(bθ0;1)' S^{-1}(bθ0;1) \bar{g}_N(bθ0;1)
      + 1/2 N \bar{g}_N(bθ0;2)' S^{-1}(bθ0;2) \bar{g}_N(bθ0;2)
      − (k1 − k2)/2 log N).
  - Uniform model prior used: p(Mj) = 1/K.

### VI. Implementation scope and notes
- Limited-information likelihood can be used with any prior; paper focuses on unit information prior + uniform model prior.
- LIBMA permits:
  - Model selection via marginal likelihoods from moment conditions.
  - Model averaging across dynamic panel specifications with endogenous regressors and lagged dependent variables.
- Evaluated by Monte Carlo and applied to dynamic gravity model for bilateral trade.

### Theoretical assumptions and asymptotic lemma
- Assumptions (as stated):
  - 1. It is continuous on ;
  - 2. E[g(ξ_i; θ)] exists and is finite for every θ ∈ ;
  - 3. E[g(ξ_i; θ)] is continuous on θ.
  - Moment conditions E[g(ξ_i; θ)] = 0 hold for a unique unknown θ0 ∈ .
  - E[g(ξ_i; θ0) g'(ξ_i; θ0)] and S(θ0) ≡ lim_{N→∞} Var[N^{1/2} \bar{g}_N(θ0)] exist and are finite positive definite matrices.
- Lemma 1:
  - Under above assumptions, N^{1/2} \bar{g}_N(θ0) d→ N(0; S(θ0)).

### Likelihood approximation and unit information prior (analytic details)
- GMM estimate b(θ0) ≡ arg min_θ N \bar{g}_N(θ)' S(θ)^{-1} \bar{g}_N(θ).
- Approximate model likelihood ∝ exp(−1/2 N \bar{g}_N(bθ0)' S^{-1}(bθ0) \bar{g}_N(bθ0) − K/2 log N).
- Unit information prior precision set to the observed GMM precision: ∂^2/∂θ^2 [\bar{g}_N' S^{-1} \bar{g}_N] |_{θ=bθ0}.

### Practical estimation
- Iterative GMM with consistent estimate of weighting matrix used to compute approximate Bayes factors and posterior model probabilities.
- Prior over model space: uniform.

---

### Monte Carlo simulation: design
- Universe: 9 potential explanatory variables — 6 exogenous, 2 endogenous, and lagged dependent variable.
- T = 4 fixed.
- N varied: N = 200; 500; and 2000.
- Exogenous x1–x4: (x1it x2it x3it x4it)' = (0:3  0:4  0:8  0:5)' + r_t with r_t ∼ N(0; I4) for t = 0,1,...,T; i = 1,...,N.
- x5, x6 correlated with x1, x2: detailed linear relation with r_t ∼ N(0; I2).
- Endogenous (w1, w2):
  - (w1it w2it)' = 0:71 (w1_{i,t−1} w2_{i,t−1})' + 6:7 v_it (−1  1)' + r_t for t = 1,2,...,T.
  - (w1i0 w2i0)' = 6:7 v_i0 (−1  1)' + r0.
  - v_it ∼ N(0; σ^2_v) and r_t ∼ N(0; I2).
- Dependent variable:
  - t = 0: yi0 = 1/(1−ρ) (xi0 θ_x + wi0 θ_w + α_i + v_i0) with v_i0 ∼ N(0; σ^2_v) and α_i ∼ N(0; σ^2_α).
  - t = 1,...,T: y_it = ρ y_{i,t−1} + θ_x x_it + θ_w w_it + α_i + v_it.
  - θ_x = (−0:05  0  0 −0:05  0  0:05  0)' and θ_w = (0  0:13)'.
- Robustness: v_it distribution relaxed to discrete constructed via N_v points uniformly sampled on [−1;1] with exponential weights (mean zero, variance σ^2_v).
- Replications: 1000 instances of DGP.
- Reported: medians, means, variances, quartiles.
- Parameters varied:
  - N = 200; 500; 2000.
  - ρ = 0:95 and 0:50.
  - σ^2_v = 0:05; 0:10; 0:20; σ^2_α = 0:10.

### Monte Carlo results — model selection (key quantitative findings)
- Posterior probability of true model (Table 1 highlights):
  - N = 200: mean posterior probability of true model ranges from 0:031 to 0:218 depending on parameters.
  - N = 2000: mean posterior probability of true model ranges from 0:633 to 0:655.
  - Median posterior model probabilities for N = 2000 range from 0:690 to 0:705.
- Posterior-probability ratio (true model vs best alternative) (Table 2):
  - For N ≥ 500: ratio above unity in all cases considered.
  - For N = 200: ratio decreases from 1:591 and 1:039 to 0:422 and 0:249 as σ^2_v increases from 0:05 to 0:20.
  - Ratios increase with N, reaching values above 6:5 for N = 2000.
- Recovery rate of true model (Table 3):
  - N = 200: recovery rate varies from 7 percent to 59 percent, decreasing as σ^2_v increases from 0:05 to 0:20.
  - N = 500: success rate ranges from 51 percent to 83 percent.
  - N = 2000: recovery rate ranges from 91 to 94 percent.
- Theoretical R^2 of generated model varies between 0.50 and 0.60.

### Monte Carlo results — model averaging and parameter recovery
- Inclusion probabilities (Table 4):
  - Prior inclusion probability = 0:50.
  - For N ≥ 500: median inclusion probability for all relevant explanatory variables > 0:95 in all cases.
  - As N increases, posterior inclusion probabilities approach 1 for relevant variables.
  - For irrelevant variables, median posterior inclusion probability decreases with N; upper bound < 0:07 for all cases when N = 2000.
  - Even in poor true-model recovery scenarios (e.g., N = 200; ρ = 0:95; σ^2_v = 0:20 with recovery 12 percent), inclusion probabilities still differentiate relevant vs non-relevant variables.
- Parameter estimation (Table 5):
  - Median estimated parameters across 1000 replications are very close to true parameter values.
  - As N increases, variance of estimates decreases and medians converge to true values.
  - Variance over 1000 replications often < 10^{−4} in many cases.

### Robustness to non-Gaussian errors (Section 3 findings)
- Normality of v_it relaxed to discrete non-Gaussian construction; results similar to Gaussian case.
- Table 6 (posterior probability of true model with non-Gaussian errors):
  - For N = 2000: median posterior model probabilities range from 0:684 to 0:708.
- Table 7 (probability ratio):
  - For N ≥ 500: ratio above unity in all cases.
  - For N = 200: ratio decreases from 1.587 and 1.205 to 0.350 and 0.254 as σ^2_v increases from 0:05 to 0:20.
  - Averages increase with N, reaching values above 6.4 for N = 2000.
- Table 8 (recovery rates with non-Gaussian errors):
  - N = 200: recovery rate varies from 7 percent to 59 percent (decreasing with σ^2_v).
  - N = 500: success rate ranges from 51 percent to 85 percent.
  - N = 2000: recovery rate ranges from 92 to 93 percent.
- Tables 9 and 10 (inclusion probabilities and parameter estimates):
  - For N ≥ 500, median inclusion probability for relevant variables > 0:90.
  - Estimated parameter medians and variances very close to Gaussian-case Table 5; performance improves with N.
- Repeated methodological note: "The error terms are constructed using discrete distributions."

### Simulation tables: selected exact numeric entries (representative)
- Table 1 (selected rows):
  - N=200 (ρ=0.95): Mean 0.218  0.112  0.052; Median 0.213  0.078  0.027.
  - N=500 (ρ=0.95): Mean 0.448  0.419  0.275; Median 0.485  0.455  0.266.
  - N=2000 (ρ=0.95): Mean 0.633  0.652  0.655; Median 0.690  0.705  0.702.
- Table 3 (probability of retrieving true model):
  - N=200 (ρ=0.95, σ^2_v = 0.05  0.10  0.20): % Correct 59    29    12.
  - N=500 (ρ=0.95): % Correct 83    80    59.
  - N=2000 (ρ=0.95): % Correct 91    93    94.
- Table 4 (example medians for N=200, ρ=0.95, σ^2_v = 0.05):
  - yt-1: med 1.00000 var 0.00000.
  - x1 (true=1): med 0.98188 var 0.03253.
  - x4 (true=1): med 0.97537 var 0.03739.
  - w2 (true=1): med 0.98348 var 0.07007.
- Table 5 (example for N=200, ρ=0.95, σ^2_v = 0.05):
  - yt-1 (True 0.95): med 0.95174 var 0.00003.
  - x1 (True 0.05): med 0.04854 var 0.00027.
  - w2 (True 0.13): med 0.14005 var 0.00194.

---

### Application: LIBMA for a dynamic gravity model of bilateral trade
- Model (dynamic augmented gravity, equation (19)):
  - log(X_ijt) = φ_0 + β log(X_{ijt−1}) + Σ_{k=1}^K φ_k Z_{mijt} + γ_1 CU_ijt + γ_2 DirPeg_ijt + γ_3 IndirPeg_ijt + γ_4 VolSR_ijt + γ_5 VolLR_ijt + Σ_{m=1}^M ρ_cr_m FTA_mijt + Σ_{m=1}^M ρ_dv_m dFTA_mit + τ_t + u_ijt
  - u_ij ~ N(0, σ^2).
- Data:
  - Dataset from Qureshi and Tsangarides (2010) extended with WTO Regional Trade Agreements entries.
  - Eight-year intervals yield up to six panels (1960-1967, 1968-1975, 1976-1983, 1984-1991, 1992-1999, 2000-2007); maximum possible T = 5 due to lag.
  - 42 proxies and time effects identified.
  - Coverage: 159 countries over 1960-2007, yielding 9,628 country pairs and 22,875 observations; average of 3 observations per pair (max 5).
- Estimators compared:
  - OLS (no model uncertainty, no endogeneity).
  - System-GMM (SGMM) (accounts for endogeneity, not model uncertainty).
  - BMA (accounts for model uncertainty, not endogeneity).
  - LIBMA (accounts for both).
- Table 11 (selected results and comparisons):
  - Lagged trade (Trade_ijt-1):
    - OLS Mean 0.6210 (0.0074) ***, SGMM 0.3844 (0.0734) ***, BMA Mean 0.6252 (0.0048) P(incl.) 1.0000, LIBMA Mean 0.4287 (0.0192) P(incl.) 1.0000.
  - Currency union (CU_ijt):
    - OLS 0.4170 (0.0808) ***, SGMM 0.0595 (0.2200) P(incl.) 0.9960, BMA 0.3650 (0.0859) P(incl.) 0.9941, LIBMA 0.6113 (0.2730).
  - Long-run volatility (VolLR_ijt):
    - OLS −0.2540 (0.0790) ***, SGMM −0.3380 (0.3310) P(incl.) 0.9900, BMA −0.3844 (0.0917) P(incl.) 0.0261, LIBMA −0.0033 (0.0042).
  - Short-run volatility (VolSR_ijt):
    - OLS −0.1600 (0.0461) ***, SGMM −0.0848 (0.0621) P(incl.) 0.2630, BMA −0.1664 (0.0606) P(incl.) 0.8362, LIBMA −0.2828 (0.0674).
  - Core gravity (log GDP product):
    - OLS 0.3920 (0.0101) ***, SGMM 0.3785 (0.0892) ***, BMA 0.3974 (0.0086) P(incl.) 1.0000, LIBMA 0.7184 (0.0355) P(incl.) 0.9963.
  - Distance:
    - OLS −0.5150 (0.0168) ***, SGMM −0.6323 (0.1690) ***, BMA −0.5128 (0.0151) P(incl.) 1.0000, LIBMA −0.8137 (0.0361) P(incl.) 0.9975.
  - Selected FTAs (examples):
    - AFTA_ijt: OLS 0.3990 (0.1310) ***, SGMM 1.2670 (2.1740) P(incl.) 0.0010, BMA 0.4155 (0.1912) P(incl.) 0.4043, LIBMA 0.2074 (0.1227).
    - AP_ijt: OLS 0.4480 (0.1840) **, SGMM 2.0460 (0.7580) *** P(incl.) 0.0000, BMA 0.6799 (0.3212) P(incl.) 0.4527, LIBMA 0.5452 (0.1207).
  - Diagnostics:
    - AR(1) p-value 0.00; AR(2) p-value 0.14; Hansen test p-value 0.25.
- Comparative findings:
  - LIBMA identifies 25 variables out of 42 as robust, including lagged trade, several exchange-rate regime variables (currency union, indirect peg, long-run volatility), some FTAs, and a sub-Saharan dummy uniquely.
  - OLS finds 33 of 42 variables statistically significant (24 at 1 percent) but misstates some potentially endogenous variables (e.g., indirect peg and long-run volatility).
  - SGMM accounts for endogeneity and finds fewer significant variables than OLS but still differs from LIBMA when model uncertainty is considered.
  - BMA (no endogeneity control) identifies a subset of OLS-significant variables as robust but misses potentially endogenous variables that LIBMA identifies as robust.
  - Conclusion: joint treatment of endogeneity and model uncertainty affects which determinants are identified as robust and materially changes coefficient estimates.

---

### Conclusion and implications
- LIBMA: limited-information Bayesian Model Averaging for dynamic panel data with lagged dependent and endogenous regressors.
- Features:
  - Integrates GMM-based limited-information likelihood with unit information prior to approximate marginal likelihoods in a BIC-like form.
  - Accounts jointly for model uncertainty and endogeneity in short panels.
- Performance:
  - Monte Carlo evidence: asymptotically performs very well for model uncertainty in dynamic panels with endogenous regressors; inclusion probabilities and parameter recovery improve with N.
  - Robust to non-Gaussian idiosyncratic errors constructed discretely.
  - Empirical application to a dynamic gravity trade model shows LIBMA yields materially different inference from OLS, SGMM, or BMA by jointly addressing endogeneity and model uncertainty.
- Suggested future research:
  - Explore LIBMA performance in applications constrained by small sample sizes.

*Italic: Source: _wp11230 — 1. Model Selection (PDF chapter/section content provided).*

### 1. Model Selection . . . . . . . . . . . . . . . . . . . . . . . . . . . . .   15

### 1. Model Selection

### I. Introduction: model uncertainty and Bayesian responses
- Model uncertainty arises from weak theoretical guidance and trade-offs in regressor choice, producing many plausible specifications (ìopen-endednessî).
- Two common Bayesian approaches:
  - Bayesian model selection: choose the single model with largest posterior model probability.
  - Bayesian Model Averaging (BMA): average posterior distributions across models using posterior model probabilities as weights to fully incorporate model uncertainty.
- Limitations of prior literature:
  - Most BMA work focuses on static, cross-section models and typically ignores dynamic relationships and endogenous regressors.
  - Existing dynamic-panel / GMM contexts have not incorporated full model uncertainty allowing endogenous regressors together with lagged dependent variables.
- Contribution of this paper:
  - Proposes Limited Information Bayesian Model Averaging (LIBMA) for panel data with short T that permits lagged dependent variable and endogenous regressors.
  - Constructs model likelihoods and posteriors from moment conditions only (no distributional assumptions), yielding a limited-information likelihood based on GMM asymptotics and a unit information prior with a BIC-like posterior approximation.
  - Evaluates LIBMA with Monte Carlo simulations and applies it to a dynamic gravity model for bilateral trade.

### II. Model uncertainty in the Bayesian context — concepts and formulas
- Model selection:
  - For a model Mj, posterior model probability:
    - p(Mj | D) = p(D | Mj) p(Mj) / Σ_{l=1}^K p(D | Ml) p(Ml)
  - Marginal likelihood (model evidence):
    - p(D | Mj) = ∫ p(D | θj, Mj) p(θj | Mj) dθj
  - Posterior odds and Bayes factor:
    - posterior odds p(Mj | D) / p(Mi | D) = [p(D | Mj) / p(D | Mi)] × [p(Mj) / p(Mi)]
- Bayesian Model Averaging (BMA):
  - Full posterior for quantity of interest ψ:
    - p(ψ | D) = Σ_{j=1}^K p(ψ | D, Mj) p(Mj | D)
  - Inclusion probability for regressor Zi:
    - p(Zi | D) = Σ_{j=1}^K I(Zi ∈ Mj) p(Mj | D)
  - Posterior mean for parameter θl:
    - E(θl | D) = Σ_{j=1}^K E(θl | D, Mj) p(Mj | D)
- Choice of priors:
  - Common priors for Gaussian regression include Normal-Gamma for (θ, σ^2) and Zellner's g-prior for θ | σ^2.
  - Unit information prior: k_l-dimensional Normal with mean the estimate θ̂ and variance I(θ̂)^{-1} (I is expected Fisher information for one observation); provides roughly one observation's worth of information.
  - Model priors: this paper adopts a uniform prior over model space, so p(Mj) = 1/K for all models.

### III. Limited Information Bayesian Model Averaging (LIBMA): setup
- Model and universe of regressors:
  - Universe U partitions into {1} (lagged dependent variable), X (exogenous variables, dimension m), and W (endogenous variables, dimension q). Total possible regressors K = m + q + 1.
  - Dynamic panel model for individual i, time t:
    - y_it = β y_{i,t-1} + x'_it θ_x + w'_it θ_w + u_it
    - u_it = α_i + v_it
  - Observed across N individuals and T periods, with T small relative to N.
  - Assumptions: E(v_it) = 0; v_it not serially correlated; exogenous x_it orthogonal to v_is for all i,t,s; endogenous w_it may be correlated with contemporaneous and past v_is but instruments exist from lags.
- Goal: perform Bayesian model selection/averaging when the parametric likelihood is not fully specified, using only moment conditions.

### IV. Estimation and moment conditions (system GMM context)
- Use system GMM (level and first-difference equations) to build instrument sets and moment conditions.
- First-difference equation instruments and moment conditions:
  - Lagged dependent variable  Δy_{i,t-1}: instruments y_{i,t-2}, ..., y_{i,0} giving moment conditions E(y_{i,t-s} Δv_{it}) = 0 for t = 2,3,...,T; s = 2,3,...,t (requires T ≥ 2).
  - Exogenous x_{l,it}: instrument itself, giving E(x_{l,it} Δv_{it}) = 0 for t = 2,3,...,T; l = 1,...,m.
  - Endogenous w_{l,it}: instruments w_{l,i,t-2}, ..., w_{l,i1}, giving E(w_{l,i,t-s} Δv_{it}) = 0 for t = 3,4,...,T; s = 2,...,t-1 (requires T ≥ 3).
- Counts of moment conditions implied (first-difference equation):
  - Lagged dependent variable: T(T-1)/2 moment conditions.
  - Exogenous variables: m (T-1) moment conditions.
  - Endogenous variables: q (T-2)(T-1)/2 moment conditions.
- Level equation instruments and moment conditions:
  - Lagged dependent variable: Δy_{i,t-1} yields E(Δy_{i,t-1} u_it) = 0 for t = 2,3,...,T.
  - Exogenous variables: Δx_{l,it} yields E(Δx_{l,it} u_it) = 0 for t = 2,3,...,T; l = 1,...,m.
  - Endogenous variables: Δw_{l,i,t-1} yields E(Δw_{l,i,t-1} u_it) = 0 for t = 3,4,...,T; l = 1,...,q (requires T ≥ 3).
- Additional linear moment conditions (Ahn and Schmidt (1995)) available under homoskedastic v_it assumptions:
  - E(y_{i,t} u_{i,t} - y_{i,t-1} u_{i,t-1}) = 0 for t = 2,3,...,T.
- Aggregation for exogenous variables:
  - Summing across periods produces one moment condition per exogenous variable:
    - Σ_{t=2}^T E(x_{l,it} Δv_{it}) + Σ_{t=2}^T E(Δx_{l,it} u_it) = 0, for l = 1,...,m.
- Matrix representation:
  - All moment conditions can be written as E[ (G'_i U_i)' ] = 0. Definitions of G_i and U_i are provided in Appendix I.

### V. Limited information criterion: constructing likelihoods from moments
- Problem: GMM estimation provides moment-based information but not a fully specified parametric likelihood p(D | θ, Mj).
- Approach:
  - Build a limited-information likelihood using only the moment conditions and the linear structure of the model.
  - Use asymptotic properties of GMM estimators (central limit theorem) to obtain a "true" limited information likelihood suitable for Bayesian inference.
  - Combine this likelihood with a unit information prior (Kass and Wasserman (1995)) to produce a posterior with a BIC-like approximation, reducing computational burden for model posteriors.
- Comparison with related methods:
  - Differs from Kim (2002), Tsangarides (2004), Hong and Preston (2008) who approximate marginal likelihoods by quasi-likelihoods justified only asymptotically.
  - Similar in spirit to Schennach (2005) and Ragusa (2008) insofar as the likelihood is constructed from moment information, but the paper's approach is simpler in construction.

### VI. Implementation notes and scope
- The limited information likelihood can, in principle, be used with any prior; this paper focuses on the unit information prior combined with a uniform model prior over model space.
- The LIBMA framework allows:
  - Model selection via computed marginal likelihoods from moment conditions.
  - Model averaging across dynamic panel specifications including endogenous regressors and lagged dependent variables.
- The paper evaluates LIBMA via Monte Carlo experiments and applies it to a dynamic gravity model for bilateral trade (details are presented in later sections).

*Source: _wp11230 - 1. Model Selection . . . . . . . . . . . . . . . . . . . . . . . . . .   15*

### 1. It is continuous on;

### _wp11230 - 1. It is continuous on;

### Theoretical assumptions and asymptotic result
- Assumptions:
  - 1. It is continuous on;
  - 2. E[g(i; )] exists and is Önite for every 2;
  - 3. E[g(i; )] is continuous on .
  - Moment conditions E[g(i; )] = 0 hold for a unique unknown 0 2.
  - E[g(i; 0)g0(i; 0)] and S(0) ≡ lim N!1 Var[N1=2 bgN(0)] exist and are Önite positive deÖnite matrices.
- Lemma 1 (as stated):
  - Under the above assumptions, N1=2 bgN(0) d ¡! N(0; S(0)).
  - Interpretation: N1=2 bgN(0) converges in distribution to a multivariate Normal distribution.

### Moment conditions and model setup
- For model (5), individual moment conditions:
  - g(i; ) = G0i (eyi − ezi )
  - Definitions:
    - i = {eyi; ezi}
    - ~zi = −eyi; −1 exi ew i (layout preserved as in source)
    -  = ( x w )0
    - G i is the matrix deÖned in (14).
- Dependent variable vectors:
  - eyi =
    − y i1
    y i2
     y iT
     y i2
     y i3
     y iT
    0
  - eyi;−1 =
    − y i0
    y i1
     y i;T−1
     y i1
    y i2
     y i;T−1
    0
- Exogenous variable matrix exi and endogenous matrix ewi are defined in block form (structure preserved from source).

### Sample moment, stationarity, and applicability of Lemma 1
- bgN(0) = N−1 ∑N i=1 G0i eyi − N−1 ∑N i=1 G0i ezi 0.
- Assumptions on data:
  - {yi0, xi·, zi·, ui·} is strictly stationary and ergodic with Önite second moment.
  - E[@g(i; )/@] = −E[~z0i G i] is Önite and has full rank by choice of moment conditions.
  - !i ≡ g(i; 0) is stationary and independent with E[!i] = 0, E[!i !0i] exists and is Önite positive deÖnite.
- Consequence:
  - lim N!1 Var[N1=2 ^gN(0)] exists, is Önite, and positive deÖnite (Hansen 1982).
  - Therefore Lemma 1 applies to the dynamic panel data model.

### Likelihood approximation and unit information prior
- Likelihood expression (up to proportionality) using Lemma 1 and Laplace approximation:
  - Model likelihood ∝ ∫ exp(−1/2 N bg0N() S−1() bgN()) p() d.
- Using prior p() second order di§erentiable around ^0 and Laplace approximation leads to:
  - Approximate model likelihood ∝ exp(−1/2 N bg0N(b0) S−1(b0) bgN(b0) − K/2 log N),
    - where b(0) ≡ arg min  N bg0N() S()−1 bgN() is the GMM estimate of 0 with weighting matrix S()−1.
  - K is the dimension of vector .
- Unit information prior specification:
  - Prior p() = N(b(0); @2 −1/2 bg0N S−1 bgN @2|_{=b(0)}).
  - Precision (inverse variance) of b(0) from GMM is @2 −1/2 N bg0N S−1 bgN @2|_{=b(0)}.
  - Unit information prior is interpreted as weakly informative based on observed data (Kass and Wasserman (1995) suggestion preserved).

### Model comparison and posterior odds (BIC-like form)
- For model Mj with kj nonzero elements and estimate b(0;j):
  - Approximate model likelihood becomes proportional to exp(−1/2 N bg0N(b0;j) S−1(b0;j) bgN(b0;j) − kj/2 log N).
- Moment conditions for model Mj:
  - E[G0i (eyi − (ez i)Mj (0)Mj )] = 0 (subscript Mj selects entries corresponding to model Mj).
- Posterior odds ratio of two models M1 and M2:
  - p(M1 | N−1 ∑N i=1 G0i eyi) / p(M2 | N−1 ∑N i=1 G0i eyi)
    = p(M1)/p(M2) × exp(−1/2 N bg0N(b0;1) S−1(b0;1) bgN(b0;1)
      + 1/2 N bg0N(b0;2) S−1(b0;2) bgN(b0;2)
      − (k1 − k2)/2 log N).
  - This has the same form as BIC for fully specified models.
- Implementation notes:
  - Iterative GMM estimation with moment conditions E[G0i (eyi − (ez i)Mj 0;j )] = 0 used to approximate Bayesian factors.
  - A consistent estimate of the weighting matrix replaces S−1(b(0)) in practice.
  - Prior over model space taken as Uniform: p(M1) = p(M2) = ... = p(MK) = 1/K.

### Monte Carlo Simulation and Results — Overview
- Purpose: assess performance of LIBMA (posterior model probabilities, inclusion probabilities, parameter statistics).
- Simulation design:
  - Universe: 9 potential explanatory variables — 6 exogenous variables, 2 endogenous variables, and the lagged dependent variable.
  - Number of periods T = 4 (kept constant).
  - Number of individuals N varied: N = 200; 500; and 2000.
  - For exogenous variables x1it, x2it, x3it, x4it:
    - (x1it x2it x3it x4it)' = (0:3  0:4  0:8  0:5)' + rt with rt ∼ N(0; I4) for t = 0,1, ..., T; i = 1, ..., N.
  - Correlation structure for x5it, x6it with x1it, x2it:
    - (x5it x6it)' = ((x1it x2it)' − (0:3  0:4)') (−0:1) (1  2; 0 −1  1  1) + (1:5  1:8)' + rt with rt ∼ N(0; I2).
  - Endogenous variables (w1it w2it) process:
    - (w1it w2it)' = 0:71 (w1i;t−1 w2i;t−1)' + 6:7 vit (−1  1)' + rt for t = 1,2, ..., T.
    - (w1i0 w2i0)' = 6:7 vi0 (−1  1)' + r0.
    - vit ∼ N(0; 2v) and rt ∼ N(0; I2) for t = 0,1, ..., T.
  - Dependent variable:
    - For t = 0:
      - yi0 = 1/(1−) (xi0 x + wi0 w + i + vi0) with vi0 ∼ N(0; 2v) and i ∼ N(0; 2).
      - x = (−0:05  0  0 −0:05  0  0:05  0)' and w = (0  0:13)'.
    - For t = 1,2, ..., T:
      - yit =  y i;t−1 + x x it + w w it + i + vit with vit ∼ N(0; 2v) and i ∼ N(0; 2).
  - Robustness check: v it distribution relaxed from Normal to discrete constructed via Nv points uniformly sampled on [−1;1], exponential weights, adjusted to mean zero and variance 2v (equivalent to uniform sampling from a simplex in Nv-dimensional space).
- Simulation replications:
  - 1000 instances of the DGP generated.
  - Reported statistics: medians, means, variances, and quartiles.
- Parameters varied:
  - N = 200; 500; 2000.
  - Coefficient of lagged dependent variable  = 0:95 and 0:50.
  - Error variances considered: 2v = 0:05; 0:10; and 0:20 while 2 = 0:10.
  - Also examine non-Normal vit in robustness set.

### Simulation results — Model selection (key findings)
- Posterior probability of the true model (Table 1 summary as described):
  - For N = 200: mean posterior probability of the true model ranges from 0:031 to 0:218 depending on parameter values.
  - For N = 2000: mean posterior probability of the true model ranges from 0:633 to 0:655.
  - Median posterior model probabilities for N = 2000 range from 0:690 to 0:705.
  - Distribution of posterior probabilities becomes skewed toward 1 as N increases.
- Relative measure (Table 2): ratio of posterior probability of true model to highest posterior of other models
  - For N ≥ 500: ratio is above unity for all cases considered (true model on average favored).
  - For N = 200: ratio decreases from 1:591 and 1:039 to 0:422 and 0:249 as 2v increases from 0:05 to 0:20.
  - Ratios increase with N, reaching values above 6:5 for N = 2000.
- Recovery rate of true model (Table 3):
  - For N = 200: recovery rate varies from 7 percent to 59 percent, decreasing as 2v increases from 0:05 to 0:20.
  - For N = 500: success rate ranges from 51 percent to 83 percent.
  - For N = 2000: recovery rate ranges from 91 to 94 percent.
- Note: The theoretical R2 of generated model varies between 0.50 and 0.60.

### Simulation results — Model averaging (key findings)
- Inclusion probabilities (Table 4 summary):
  - Prior probability of inclusion for each variable = 0:50 under uniform model prior.
  - For N ≥ 500: median inclusion probability for all relevant explanatory variables > 0:95 in all cases considered.
  - As N increases, posterior inclusion probabilities approach 1 for relevant variables.
  - For irrelevant variables, median posterior probability of inclusion decreases with sample size; upper bound < 0:07 for all cases when N = 2000.
  - Even when true-model recovery is poor (example: N = 200;  = 0:95; 2v = 0:20 with recovery 12 percent), inclusion probabilities still differentiate relevant vs non-relevant variables.
- Parameter estimation (Table 5 summary and Figure 3 description):
  - Median estimated parameters (averaged over 1000 replications) are very close to true parameter values.
  - As N increases, variance of estimates decreases and medians converge to true values.
  - Variance over 1000 replications is often less than 10−4 in many cases.
- Overall assessment:
  - While model selection properties improve with sample size, the strength of the methodology is its performance in Bayesian Model Averaging context.

*Italic: Source: _wp11230 - 1. It is continuous on; (PDF chapter/section content provided).*

### 3.  Robustness Checks Using non-Gaussian Errors

### 3.  Robustness Checks Using non-Gaussian Errors

### Robustness-check setup and overview
- Normality assumption for the error term v_it is relaxed to assess robustness.
- Overall results (Tables 6-10) are reported as very similar to those in Tables 1-5.
- Tables 6 and 7: posterior model probabilities for the true model and the ratio of the posterior model probability of the true model to the highest posterior probability of all other models, respectively.

### Posterior model probabilities (Table 6) — key findings
- Mean posterior probabilities of the true model increase with the sample size.
- Variation across different combinations of parameters decreases as sample size increases.
- For N = 2000:
  - Median posterior model probabilities range from 0:684 to 0:708.
- Distribution of posterior probabilities of the true model becomes skewed toward 1 (quartiles in Table 6 and density plots in Figure 4).

### Ratio of true-model probability to best alternative (Table 7)
- For sample sizes N = 500 and above, the ratio is above unity in all cases considered, indicating the correct model is on average favored over all other models.
- For N = 200, the ratio decreases:
  - from 1.587 and 1.205 to 0.350 and 0.254, respectively, as the variance of the random error term increases from 0.05 to 0.20.
- Average ratios increase with sample size, reaching values above 6.4 for N = 2000.

### Model recovery performance (Table 8)
- Model recovery under non-Gaussian errors is described as still good and very similar to Table 3.
- For N = 200:
  - Recovery rate varies from 7 percent to 59 percent and decreases as the variance of the random error term increases from 0.05 to 0.20.
- For N = 500:
  - Success rate ranges from 51 percent to 85 percent.
- For N = 2000:
  - Recovery rate ranges from 92 to 93 percent (variation much smaller).

### Posterior inclusion probabilities and parameter estimates (Tables 9 and 10)
- Table 9: posterior inclusion probabilities using LIBMA compared to the true model.
  - For samples N ≥ 500, the median value of the inclusion probability for all the relevant explanatory variables is greater than 0.90 in all cases considered.
  - As sample size increases, posterior inclusion probabilities approach 1 for all relevant variables.
  - For variables not contained in the true model, the median posterior probability of inclusion decreases with sample size.
    - Upper bound for these non-relevant variables is less than 0.073 for all cases when N = 2000.
  - Even when true-model recovery is poor (example: N = 200;  = 0:95; ^2_v = 0:20 where recovery = 8 percent), inclusion probabilities still differentiate relevant and non-relevant variables.
- Table 10: estimated parameter medians and variances are very close to those in Table 5.
  - Methodology performs well in estimating parameters; performance improves as sample size increases.
  - Figure 6 (Appendix A) box plots (case  = 0:95 and _v = 0:1) show variance of parameter distributions decreases as sample size increases and medians move toward true values.

---

### V.  LIBMA Application to a Dynamic Trade Gravity Model

### Background and motivation
- Gravity models examine determinants of bilateral trade; augmented versions examine exchange rate regimes, exchange rate volatility, and FTAs.
- Issues noted:
  - Reverse causality between trade and currency unions creates uncertainty about effect sizes.
  - Debate: FTAs may be trade creating or trade diverting; results depend on regressors used.
  - Most gravity models are static and may ignore trade persistence due to sunk costs or habit formation.
- Complexity: several determinants potentially endogenous and considerable model uncertainty about which explanatory variables to include.
- LIBMA (limited information BMA) is proposed as well suited to estimate a dynamic trade gravity model with model uncertainty and endogenous regressors.

### Model specification and dataset (equations (18) and (19))
- Static gravity: X_ij = Y_i Y_j (T_ij / (P_i P_j))^{1-ψ} with ψ > 1 (equation (18)).
- Dynamic augmented specification (equation (19)):
  - log(X_ijt) = φ_0 + β log(X_ijt−1) + Σ_{k=1}^K φ_k Z_{mijt} + γ_1 CU_ijt + γ_2 DirPeg_ijt + γ_3 IndirPeg_ijt + γ_4 VolSR_ijt + γ_5 VolLR_ijt + Σ_{m=1}^M ρ_cr_m FTA_mijt + Σ_{m=1}^M ρ_dv_m dFTA_mit + τ_t + u_ijt
  - Definitions:
    - Z: vector of traditional time varying and invariant trade determinants.
    - CU_ijt = 1 if i and j share the same currency.
    - DirPeg_ijt = 1 if i's exchange rate is pegged to j or vice versa (but i and j not in same currency union).
    - IndirPeg_ijt = 1 if i is indirectly related to j through a peg with an anchor country.
    - VolSR_ijt and VolLR_ijt: real exchange rate volatility over short-run and long-run horizons.
    - FTA_ijt: vector of FTA dummies (1 if i and j members of same FTA in a year).
    - FTA_it: vector of FTAs where i is a member in a year.
    - τ_t: year-specific effects.
    - u_ij ~ N(0, σ^2).
- Data:
  - Dataset taken from Qureshi and Tsangarides (2010) extended with individual FTA entries from the WTO Regional Trade Agreements database.
  - Time intervals: eight-year intervals yield up to six panels (1960-1967, 1968-1975, 1976-1983, 1984-1991, 1992-1999, 2000-2007); maximum possible number of periods is 5 due to lagged dependent variable.
  - Identified 42 proxies and time effects corresponding to averaging span.
  - Dataset covers 159 countries over 1960-2007, yielding 9,628 individual country pairs (rather than 159*158/2 = 12;561 due to missing observations).
  - Baseline estimation: 9;628 country pairs with 22;875 observations over 1960-2007 and an average of 3 (out of maximum 5) observations per country pair.

### Estimation methods compared
- Estimation of (19) performed with:
  - OLS: does not account for model uncertainty or endogeneity.
  - System-GMM (SGMM): accounts for endogeneity but not model uncertainty.
  - Bayesian Model Averaging for linear regression models (BMA): accounts for model uncertainty but not endogeneity.
  - LIBMA: accounts for both model uncertainty and endogeneity.
- Table 11 reports estimation results for OLS, SGMM, BMA, and LIBMA:
  - OLS and SGMM: estimated means and standard errors; significance at 1, 5, 10 percent.
  - BMA and LIBMA: posterior inclusion probabilities p(Z_i | D); posterior means and standard deviations.
  - Robust variables: posterior inclusion probability above prior threshold (p(Z_i | D) ≥ 0:50).

### Results and comparison across estimators
- LIBMA identifies 25 variables out of 42 as robust, including:
  - lagged trade,
  - several exchange rate regime variables (currency union, indirect peg, long-run exchange rate volatility),
  - some trade-creating and trade-diverting FTAs,
  - sub-Saharan dummy variable (proxy for heterogeneity) — uniquely identified as robust by LIBMA.
- OLS:
  - Finds 33 out of 42 variables statistically significant at least at the 10 percent level (24 of those significant at the 1 percent level).
  - Two potentially endogenous variables (indirect peg and long-run volatility) are incorrectly estimated by OLS:
    - indirect peg not significant under OLS but labeled robust by LIBMA,
    - long-run volatility significant under OLS but not robust under LIBMA.
  - Many variables significant in OLS (including several FTAs with both trade-creating and trade-diverting estimates) disappear when model uncertainty and endogeneity are accounted for with LIBMA.
  - Estimated means are imprecisely estimated for several variables that LIBMA deems robust and OLS finds significant (e.g., several FTAs, currency union, short-run volatility).
- SGMM:
  - With endogeneity accounted for, SGMM identifies fewer statistically significant variables than OLS.
  - However, SGMM still does not properly identify robust determinants compared to LIBMA:
    - Among exchange rate regime and trade variables, only lagged trade is both statistically significant in SGMM and robust in LIBMA.
    - Several FTAs significant in SGMM have very low inclusion probabilities once model uncertainty is considered in LIBMA.
- BMA:
  - Accounting for model uncertainty alone identifies a subset of OLS significant variables as robust.
  - Compared with LIBMA, BMA fails to identify two potentially endogenous variables (indirect peg and short-run volatility) as robust.
  - Several FTAs identified as robust by BMA are not robust when endogeneity is incorporated in LIBMA.
- Summary conclusion:
  - Important differences in identified bilateral trade determinants arise when using LIBMA versus OLS, SGMM, or BMA.
  - Differences are attributed to LIBMA’s joint incorporation of model uncertainty and endogeneity, which the other methods do not jointly address.
  - The application underscores the importance of accounting for both endogeneity and model uncertainty in estimating a dynamic gravity model for trade.

---

### VI.  Conclusion

- Proposed LIBMA: a limited information methodology within Bayesian Model Averaging for panel data models with lagged dependent variable regressors and endogenous regressors.
- LIBMA features:
  - Incorporates a GMM estimator for dynamic panel data models into a Bayesian Model Averaging framework to explicitly account for model uncertainty.
  - Is a limited information technique based on moment restrictions rather than a full stochastic specification.
  - Explicitly controls for endogeneity by replacing likelihood and exact marginal likelihood expressions with a limited information construct modeled on GMM estimation and a limited information criterion as an approximation to actual marginal likelihoods.
  - Is applied in a panel setting, expanding applicability to many use cases.
- Simulation and application findings:
  - Asymptotically, LIBMA performs very well for addressing model uncertainty in dynamic panel data models with endogenous regressors.
  - Application to a dynamic gravity model of trade illustrates applicability where model uncertainty and endogeneity are present in short panels.
- Future research suggestion:
  - Explore using LIBMA for applications constrained by small sample sizes.

*Source: _wp11230 - 3.  Robustness Checks Using non-Gaussian Errors*

### References

### _wp11230 - References

### Appendix I — Representation of Moment Equations for Dynamic Panel Model
- Instrument matrices for first-differenced (FD) and level equations are defined explicitly:
  - Yi is a (T−1)×T(T−1)/2 matrix of lagged dependent variables used as instruments for the FD equation.
  - Wi is a (T−1)×q(T−2)(T−1)/2 matrix of endogenous variables representing instruments for the FD equation.
  - DY i is the T×(T−1) instruments matrix consisting of first differences of the dependent variable for the level equation.
  - DW i is the T×q(T−2) instruments matrix consisting of first differences of the endogenous variables for the level equation.
  - Xi and DX i denote (T−1)×m and T×m matrices, respectively, of exogenous and first-differenced exogenous variables from moment conditions (7) and (11).
  - Y¯ i is the T×(T−1) instrument matrix used for moment conditions derived from the homoskedasticity restriction.
  - ui is the T×1 vector of error terms; Dv i is the (T−1)×1 vector of first-differenced idiosyncratic errors as defined in model (5).
- Summary matrices combining moment conditions:
  - Ui is a (2T−1)×1 matrix defined as Ui = [u′ i , Dv′ i ]′.
  - Gi is a (2T−1)×(T+m−2 + (T+1)((T−2)q+T)/2) matrix defined as Gi = [DX i , DY i , 0, DW i , 0, Y¯ i , X i , 0, Y i , 0, W i , 0].

### Simulation / Monte Carlo Results — Key Statistics (Tables)
- Notes applying to tables:
  - εi ~ N(0,σ2 ε) with σ2 ε = 0.10.
  - vit ~ N(0,σ2 v).

- Table 1 — Posterior probability of the true model (LIBMA summary statistics)
  - Columns reflect ρ = 0.95 and ρ = 0.50; rows for σ2 v = 0.05, 0.10, 0.20 and samples N=200, 500, 2000.
  - N=200 (ρ=0.95): Mean 0.218  0.112  0.052; Variance 0.021  0.011  0.005; Q1 0.091  0.028  0.010; Median 0.213  0.078  0.027; Q3 0.340  0.169  0.066.
  - N=200 (ρ=0.50): Mean 0.140  0.094  0.031; Variance 0.023  0.016  0.004; Q1 0.008  0.003  0.000; Median 0.077  0.033  0.004; Q3 0.257  0.141  0.030.
  - N=500 (ρ=0.95): Mean 0.448  0.419  0.275; Variance 0.025  0.025  0.027; Q1 0.363  0.310  0.132; Median 0.485  0.455  0.266; Q3 0.574  0.547  0.412.
  - N=500 (ρ=0.50): Mean 0.429  0.403  0.257; Variance 0.030  0.033  0.035; Q1 0.319  0.264  0.085; Median 0.466  0.440  0.234; Q3 0.568  0.562  0.420.
  - N=2000 (ρ=0.95): Mean 0.633  0.652  0.655; Variance 0.027  0.023  0.020; Q1 0.585  0.603  0.609; Median 0.690  0.705  0.702; Q3 0.747  0.757  0.753.
  - N=2000 (ρ=0.50): Mean 0.646  0.646  0.644; Variance 0.023  0.027  0.025; Q1 0.584  0.601  0.588; Median 0.699  0.702  0.694; Q3 0.754  0.757  0.753.

- Table 2 — Posterior probability ratio of true model versus best among the rest (LIBMA summary statistics)
  - N=200 (ρ=0.95): Mean 1.591   0.856   0.422; Variance 1.757   0.922   0.438; Q1 0.443   0.189   0.074; Median 1.259   0.503   0.189; Q3 2.480   1.195   0.436.
  - N=200 (ρ=0.50): Mean 1.039   0.761   0.249; Variance 1.661   1.348   0.312; Q1 0.039   0.017   0.001; Median 0.425   0.220   0.024; Q3 1.726   1.012   0.210.
  - N=500 (ρ=0.95): Mean 3.254   3.113   1.975; Variance 4.709   4.618   3.286; Q1 1.407   1.313   0.541; Median 2.977   2.788   1.402; Q3 4.910   4.632   2.953.
  - N=500 (ρ=0.50): Mean 3.034   2.965   1.812; Variance 4.735   5.073   3.647; Q1 1.146   1.024   0.328; Median 2.687   2.526   1.054; Q3 4.670   4.571   2.832.
  - N=2000 (ρ=0.95): Mean 6.534   7.164   7.030; Variance 17.875  19.696  18.042; Q1 3.066   3.348   3.440; Median 6.040   6.854   6.728; Q3 9.424  10.634  10.298.
  - N=2000 (ρ=0.50): Mean 6.930   6.990   6.770; Variance 19.164  19.708  18.381; Q1 3.096   3.316   3.185; Median 6.538   6.414   6.128; Q3 10.380  10.308  10.054.

- Table 3 — Probability of retrieving the true model (LIBMA summary statistics)
  - N=200 (ρ=0.95, σ2 v = 0.05  0.10  0.20): % Correct 59    29    12
  - N=200 (ρ=0.50, σ2 v = 0.05  0.10  0.20): % Correct 35    25    7
  - N=500 (ρ=0.95): % Correct 83    80    59
  - N=500 (ρ=0.50): % Correct 78    76    51
  - N=2000 (ρ=0.95): % Correct 91    93    94
  - N=2000 (ρ=0.50): % Correct 93    93    93

- Table 4 — Model recovery: medians and variances of posterior inclusion probability for each variable (selected entries)
  - Settings: ρ = 0.95 and 0.50; σ2 v = 0.05, 0.10, 0.20; Samples N = 200, 500, 2000.
  - For N=200, ρ=0.95, σ2 v = 0.05 — medians and variances:
    - yt-1: med 1.00000  var 0.00000
    - x1 (true=1): med 0.98188  var 0.03253
    - x4 (true=1): med 0.97537  var 0.03739
    - x6 (true=1): med 0.97151  var 0.04275
    - w2 (true=1): med 0.98348  var 0.07007
    - several true-zero variables (x2, x3, x5, w1) have medians around 0.18–0.19 with small variances.
  - For N=500 and N=2000, inclusion medians for true included variables (yt-1, x1, x4, x6, w2) approach 1.00000 with variances approaching 0.00000 in many settings.
  - For true-zero variables, medians fall below 0.13 for larger N, with small variances (e.g., N=2000 x2 med 0.06983 var 0.01262).

- Table 5 — Model recovery: medians and variances of estimated parameter values (selected entries)
  - True parameter values and estimated medians/variances for various N, ρ, σ2 v.
  - Example (N=200, ρ=0.95, σ2 v = 0.05):
    - yt-1 (True 0.95): med 0.95174  var 0.00003
    - x1 (True 0.05): med 0.04854  var 0.00027
    - x4 (True −0.05): med −0.04656  var 0.00029
    - x6 (True 0.05): med 0.04800  var 0.00034
    - w2 (True 0.13): med 0.14005  var 0.00194
  - As N increases to 500 and 2000, medians of estimated parameters converge closer to true values and variances decline:
    - N=500, yt-1 med 0.95115 var 0.00001; x1 med 0.04975 var 0.00006; w2 med 0.13953 var 0.00012.
    - N=2000, yt-1 med 0.95018 var 0.00000; x1 med 0.05050 var 0.00001; w2 med 0.13280 var 0.00003.

- Table 6 — Posterior probability of the true model (alternate LIBMA summary statistics)
  - Similar structure to Table 1 with slightly different numerical entries:
  - N=200 (ρ=0.95): Mean 0.224  0.109  0.044; Variance 0.021  0.012  0.004; Q1 0.094  0.025  0.008; Median 0.222  0.071  0.023; Q3 0.342  0.161  0.053.
  - N=200 (ρ=0.50): Mean 0.162  0.084  0.031; Variance 0.025  0.015  0.005; Q1 0.019  0.002  0.000; Median 0.115  0.023  0.003; Q3 0.283  0.118  0.023.
  - N=500 (ρ=0.95): Mean 0.459  0.415  0.253; Variance 0.023  0.025  0.028; Q1 0.371  0.315  0.110; Median 0.497  0.447  0.231; Q3 0.580  0.543  0.386.
  - N=500 (ρ=0.50): Mean 0.437  0.376  0.243; Variance 0.032  0.036  0.035; Q1 0.322  0.233  0.069; Median 0.487  0.412  0.211; Q3 0.582  0.534  0.397.
  - N=2000 (ρ=0.95): Mean 0.653  0.632  0.643; Variance 0.021  0.025  0.025; Q1 0.607  0.573  0.585; Median 0.703  0.684  0.703; Q3 0.754  0.747  0.751.
  - N=2000 (ρ=0.50): Mean 0.640  0.643  0.651; Variance 0.026  0.025  0.024; Q1 0.592  0.595  0.598; Median 0.700  0.699  0.708; Q3 0.753  0.754  0.757.

*Source: Appendix I and tables from the content unit _wp11230 - References.*

### 1. The error terms are constructed using discrete distributions.

### _wp11230 - 1. The error terms are constructed using discrete distributions.

### Summary findings from simulation tables (Tables 7–10)
- Table 7: Posterior probability ratio of true model versus best among the rest (LIBMA summary statistics)
  - Sample N=200
    - Mean: 1.587, 0.828, 0.350; 1.205, 0.662, 0.254
    - Variance: 1.767, 1.043, 0.348; 1.813, 1.168, 0.420
    - Q1: 0.463, 0.168, 0.053; 0.099, 0.011, 0.001
    - Median: 1.235, 0.423, 0.173; 0.639, 0.141, 0.016
    - Q3: 2.499, 1.081, 0.357; 1.975, 0.812, 0.162
  - Sample N=500
    - Mean: 3.384, 3.050, 1.783; 3.258, 2.721, 1.715
    - Variance: 4.835, 4.355, 3.257; 5.489, 4.810, 3.590
    - Q1: 1.532, 1.318, 0.437; 1.205, 0.867, 0.258
    - Median: 3.065, 2.705, 1.118; 2.862, 2.216, 1.044
    - Q3: 5.035, 4.584, 2.630; 5.069, 4.278, 2.487
  - Sample N=2000
    - Mean: 6.980, 6.410, 6.992; 6.828, 6.824, 7.081
    - Variance: 18.234, 17.732, 19.169; 19.496, 18.475, 19.192
    - Q1: 3.360, 2.956, 3.035; 3.138, 3.257, 3.221
    - Median: 6.676, 5.738, 6.824; 6.335, 6.464, 6.762
    - Q3: 10.195, 9.520, 10.229; 9.999, 10.021, 10.578
  - Notes: 1. The error terms are constructed using discrete distributions.

- Table 8: Probability of retrieving the true model (% Correct)
  - Sample N=200: % Correct 56, 27, 84, 22, 7  (format in source shows "56    27842    227" — interpreted as reported values in the table)
  - Sample N=500: % Correct 85, 80, 53, 79, 72, 51 (source line: "85    80    5379    72    51")
  - Sample N=2000: % Correct 93, 92, 93, 92, 92, 93 (source: "93    92    9392    92    93")
  - Notes: 1. The error terms are constructed using discrete distributions.

- Table 9: Model recovery — medians and variances of posterior inclusion probability for each variable (selected highlights)
  - Sample N=200 (ψ2_v and ρ settings across columns)
    - yt-1 (True=1): med 1.00000 var 0.00000 (repeated across settings; one cell var 0.00103)
    - x1 (True=1): med 0.98533 var 0.03332; med 0.84439 var 0.07089; med 0.54259 var 0.08866; med 0.94050 var 0.05592; med 0.69674 var 0.08352; med 0.46005 var 0.08605
    - x4 (True=1): med 0.98708 var 0.02926; med 0.83962 var 0.07696; med 0.49877 var 0.08407; med 0.94699 var 0.05413; med 0.63426 var 0.08962; med 0.37623 var 0.07540
    - x6 (True=1): med 0.96697 var 0.04240; med 0.78891 var 0.08121; med 0.47078 var 0.08273; med 0.91206 var 0.05738; med 0.64015 var 0.08609; med 0.33444 var 0.07223
    - w2 (True=1): med 0.98807 var 0.05995; med 0.96496 var 0.07656; med 0.95082 var 0.06951; med 0.64884 var 0.15207; med 0.27934 var 0.15240; med 0.09806 var 0.12459
  - Sample N=500 (selected)
    - yt-1: med 1.00000 var 0.00000 across settings
    - x1: med 1.00000 var 0.00000; med 0.99975 var 0.00718; med 0.96429 var 0.04964; med 1.00000 var 0.00017; med 0.99984 var 0.01012; med 0.98394 var 0.04787
    - x6: med 1.00000 var 0.00025; med 0.99865 var 0.01006; med 0.93790 var 0.06465; med 1.00000 var 0.00151; med 0.99912 var 0.01641; med 0.90998 var 0.08230
    - w2: med 1.00000 var 0.00000; med 1.00000 var 0.00059; med 0.99995 var 0.00698; med 0.99997 var 0.01091; med 0.99912 var 0.04088; med 0.98931 var 0.08415
  - Sample N=2000 (selected)
    - yt-1, x1, x4, x6, w2: med 1.00000 var 0.00000 across almost all settings
    - Irrelevant regressors (x2, x3, x5, w1): med approx 0.064–0.073 with small variances (e.g., x2 med 0.06637 var 0.00923; x3 med 0.06416 var 0.01034; w1 med 0.06465 var 0.01128)
  - Notes: 1. The error terms are constructed using discrete distributions (see Section 4.1.).

- Table 10: Model recovery — medians and variances of estimated parameter values (selected highlights)
  - Sample N=200 (True values and medians/variances)
    - yt-1 True 0.95: med 0.95265 var 0.00003; med 0.95088 var 0.00001; med 0.95004 var 0.00001
    - x1 True 0.05: med 0.04877 var 0.00029; med 0.04260 var 0.00059; med 0.02708 var 0.00092
    - x4 True -0.05: med -0.05008 var 0.00028; med -0.04235 var 0.00060; med -0.02513 var 0.00083
    - x6 True 0.05: med 0.04781 var 0.00035; med 0.04159 var 0.00066; med 0.02405 var 0.00089
    - w2 True 0.13: med 0.14234 var 0.00168; med 0.13564 var 0.00206; med 0.13346 var 0.00206
  - Sample N=500 (selected)
    - yt-1 True 0.95: med 0.95153 var 0.00001; med 0.95077 var 0.00000; med 0.95017 var 0.00000
    - x1 True 0.05: med 0.05043 var 0.00006; med 0.05011 var 0.00014; med 0.04933 var 0.00042
    - x4 True -0.05: med -0.05098 var 0.00006; med -0.04964 var 0.00015; med -0.04733 var 0.00043
    - x6 True 0.05: med 0.05027 var 0.00007; med 0.04868 var 0.00017; med 0.04719 var 0.00048
    - w2 True 0.13: med 0.13947 var 0.00010; med 0.13760 var 0.00016; med 0.13619 var 0.00033
  - Sample N=2000 (selected)
    - yt-1 True 0.95: med 0.95070 var 0.00000; med 0.95034 var 0.00000; med 0.95015 var 0.00000
    - x1 True 0.05: med 0.05046 var 0.00001; med 0.05039 var 0.00003; med 0.05011 var 0.00005
    - x4 True -0.05: med -0.05108 var 0.00001; med -0.04958 var 0.00003; med -0.04965 var 0.00005
    - x6 True 0.05: med 0.05037 var 0.00002; med 0.04906 var 0.00003; med 0.04932 var 0.00005
    - w2 True 0.13: med 0.13350 var 0.00003; med 0.13164 var 0.00003; med 0.13103 var 0.00003
  - Notes: 1. The error terms are constructed using discrete distributions.

### Empirical application — Dynamic gravity model estimation (Table 11)
- Comparison methods: OLS, SGMM, BMA, LIBMA; reported: Mean, St.Error, P(incl.) for BMA/LIBMA where applicable.
- Regime and lagged trade
  - 1 Trade_ijt-1: OLS Mean 0.6210 (0.0074) ***, SGMM Mean 0.3844 (0.0734) ***, BMA Mean 0.6252 (0.0048) P(incl.) 1.0000, LIBMA Mean 0.4287 (0.0192) P(incl.) 1.0000
  - 2 CU_ijt: OLS 0.4170 (0.0808) ***, SGMM 0.0595 (0.2200) P(incl.) 0.9960, BMA 0.3650 (0.0859) P(incl.) 0.9941, LIBMA 0.6113 (0.2730)
  - 5 VolLR_ijt: OLS -0.2540 (0.0790) ***, SGMM -0.3380 (0.3310) P(incl.) 0.9900, BMA -0.3844 (0.0917) P(incl.) 0.0261, LIBMA -0.0033 (0.0042)
  - 6 VolSR_ijt: OLS -0.1600 (0.0461) ***, SGMM -0.0848 (0.0621) P(incl.) 0.2630, BMA -0.1664 (0.0606) P(incl.) 0.8362, LIBMA -0.2828 (0.0674)
- Core gravity and heterogeneity (selected)
  - 7 log(GDP_it GDP_jt): OLS 0.3920 (0.0101) ***, SGMM 0.3785 (0.0892) ***, BMA 0.3974 (0.0086) P(incl.) 1.0000, LIBMA 0.7184 (0.0355) P(incl.) 0.9963
  - 8 log(GDPPC_it GDPPC_jt): OLS 0.0231 (0.0094) **, SGMM 0.3995 (0.1080) ***, BMA 0.0251 (0.0088) P(incl.) 0.1070, LIBMA 0.0861 (0.0193)
  - 9 Border_ij: OLS 0.2560 (0.0652) ***, SGMM 0.0206 (0.2980) P(incl.) 0.9950, BMA 0.2776 (0.0654) P(incl.) 0.8059, LIBMA 0.4237 (0.0898)
  - 17 log(Distance_ij): OLS -0.5150 (0.0168) ***, SGMM -0.6323 (0.1690) ***, BMA -0.5128 (0.0151) P(incl.) 1.0000, LIBMA -0.8137 (0.0361) P(incl.) 0.9975
  - 15 Landlocked_ij: OLS -0.1830 (0.0200) ***, SGMM 0.0510 (0.1180) P(incl.) 1.0000, BMA -0.1997 (0.0191) P(incl.) 0.0553, LIBMA -0.0063 (0.0017)
- FTA trade creation and diversion (selected)
  - 19 AFTA_ijt: OLS 0.3990 (0.1310) ***, SGMM 1.2670 (2.1740) P(incl.) 0.0010, BMA 0.4155 (0.1912) P(incl.) 0.4043, LIBMA 0.2074 (0.1227)
  - 21 AP_ijt: OLS 0.4480 (0.1840) **, SGMM 2.0460 (0.7580) *** P(incl.) 0.0000, BMA 0.6799 (0.3212) P(incl.) 0.4527, LIBMA 0.5452 (0.1207)
  - 22 BILATERAL_ijt: OLS -0.0138 (0.0404), SGMM -1.1593 (0.3480) *** P(incl.) 0.0000, BMA 0.0000 (0.0000) P(incl.) 0.8461, LIBMA -0.3339 (0.0932)
- Diagnostic statistics reported
  - AR(1) p-value 0.00
  - AR(2) p-value 0.14
  - Hansen test p-value 0.25
- Notes for Table 11:
  - 1. For OLS and SGMM asymptotic standard errors corrected for heteroskedasticity are reported in parentheses. Asterisks indicate statistical significance at 10% (ì*î), 5% (ì**î), and 1% (ì***î).
  - 2. For BMA and LIBMA ì+î indicates inclusion probability above 0.50.

### Figures and graphical diagnostics (Figures 1–6)
- Figures 1–6 present posterior densities and box plots for probabilities and parameter estimates corresponding to the simulation tables (Tables 1, 2, 5, 6, 7, 10).
- Repeated notes across figures:
  - For the idiosyncratic error term, t  N(0; 2) where 2 = 0:10.
  - The error term is normally distributed v_it  N(0; 2v).
  - Several figures reiterate: "The error terms are constructed using discrete distributions (see section IV.A.)."

### Methodological note emphasized throughout
- The source repeatedly notes: "The error terms are constructed using discrete distributions." This specification applies across simulation exercises and is explicitly referenced in the notes to Tables 7–10 and Figures 4–6.

*Italic: Source: _wp11230 - 1. The error terms are constructed using discrete distributions.*

---


_Source: https://www.imf.org/-/media/websites/imf/imported-full-text-pdf/external/pubs/ft/wp/2011/_wp11230.pdf_
