## _wp09202

## Source details

**Canonical URL:** [_wp09202](https://www.imf.org/-/media/websites/imf/imported-full-text-pdf/external/pubs/ft/wp/2009/_wp09202.pdf)

## Other formats

- [Markdown version](/-/media/websites/imf/imported-full-text-pdf/external/pubs/ft/wp/2009/_wp09202.pdf.md)
- [Structured JSON version](/-/media/websites/imf/imported-full-text-pdf/external/pubs/ft/wp/2009/_wp09202.pdf.json)

---

### Introduction and role of Zellner’s g in Bayesian Model Averaging (BMA)
- BMA addresses model uncertainty by weighting all potential covariate combinations with posterior model probabilities (PMP):
  - p(M_s | y, X) ∝ p(y | M_s, X) p(M_s)
- Under the Normal-Gamma conjugate framework with Zellner’s g prior:
  - β | σ^2 ∼ N(0, g σ^2 (X′X)^{-1})
  - Marginal likelihood (model M_s) under a fixed g is proportional to:
    - (1 + g)^{-k_s/2} (1 − g/(1+g) R^2_s)^{-(N−1)/2}
- Conceptual points:
  - Larger g concentrates posterior mass on fewer models; smaller g spreads PMPs more evenly.
  - The “supermodel effect”: fixed large g can concentrate posterior mass on very few (or a single) model(s), possibly independent of data-generating truth.
  - A nondegenerate prior on g (hyperprior) lets the data influence g, reducing prior sensitivity and providing data-dependent shrinkage.

### BMA under Zellner’s g prior: framework and consequences
- Model setup:
  - y = 1 α_s + X_s β_s + ε with ε ∼ N(0, σ^2 I)
  - Model space: 2^K possible covariate combinations
- Common (improper) priors:
  - p(α) ∝ 1
  - p(σ) ∝ σ^{-1}
- Zellner’s g prior:
  - (β_s | σ^2, M_s, g) ∼ N(0, σ^2 g (X′_s X_s)^{-1})
  - Posterior mean: E(β_s | y, X, M_s, g) = (g/(1+g)) ˆβ_s
  - Marginal likelihood (fixed g) proportional to:
    - (̆y′ ̆y)^{-(N−1)/2} (1 + g)^{-k_s/2} (1 − g/(1+g) R^2_s)^{-(N−1)/2}
- Bayes factor between models M_s and M_j:
  - B(M_s : M_j) = (1 + g)^{(k_j − k_s)/2} ( (1 − g/(1+g) R^2_s) / (1 − g/(1+g) R^2_j) )^{-(N−1)/2}
  - Larger g drives the second factor away from unity and amplifies PMP concentration.
- Model averaging:
  - p(θ | y, X) = Σ_{j=1}^{2^K} p(θ | y, X, M_j) p(M_j | y, X)

### Popular fixed-g specifications and motivations
- Desiderata:
  - Consistency: choose g = w(N) with lim_{N→∞} w(N) = ∞ and lim_{N→∞} w′(N)/w(N) = 0 to ensure p(M_T | Y) → 1
  - Model-size penalty: calibrate g to mimic information criteria (BIC, RIC)
- Default g choices (following Liang et al. (2008)):
  - g-RIC: g = K^2
  - g-UIP: g = N
  - g-BRIC: g = max(N, K^2)
  - Empirical Bayes – Local (EBL): g_s = arg max_g p(y | M_s, X, g), often expressed as g_s = max(0, F_s − 1) with
    - F_s = R^2_s (N−1−k_s) / ((1−R^2_s) k_s)
- Remarks:
  - Model priors often uniform p(M_s) = 2^{-K} or beta-binomial formulations to reflect prior expected model size.
  - The penalty factor (1 + g)^{-k_s/2} can be offset via model priors if desired.

### The hyper-g prior: Beta prior on the shrinkage factor (Liang et al. (2008))
- Specification:
  - g/(1 + g) ∼ Beta(1, a/2 − 1)  with a ∈ (2, ∞)
  - Equivalent: p(g) = (a−2)/2 (1 + g)^{-a/2}
  - E(g/(1+g)) = 2/a
  - a = 4 → g/(1+g) uniform; a → 2 concentrates prior mass near 1; a > 4 concentrates near 0
- Closed-form posterior / marginal expressions (involving Gaussian hypergeometric functions 2F1):
  - Posterior of g given M_s:
    - p(g | y, X_s, M_s) ∝ (1 + g)^{-(k_s + a)/2} (1 − g/(1+g) R^2_s)^{-(N−1)/2} × constant involving 2F1((N−1)/2,1,(k_s + a)/2,R^2_s)
  - Marginal likelihood under hyper-g:
    - p(y | X_s, M_s) ∝ (̆y′ ̆y)^{-(N−1)/2} (a−2) (k_s + a − 2)/2  2F1( (N−1)/2, 1, (k_s + a)/2, R^2_s )
  - Posterior expected shrinkage factor:
    - E( g/(1+g) | y, X_s, M_s ) = [ 2 / (k_s + a) ]  2F1( (N−1)/2, 2, (k_s + a)/2 + 1, R^2_s )  /  2F1( (N−1)/2, 1, (k_s + a)/2, R^2_s )
  - Posterior expected response:
    - E(y | X_s, M_s) = 1 E(α_s | X_s, M_s) + E( g/(1+g) | y, X_s, M_s ) X_s ˆβ_s
- Posterior covariance and higher moments:
  - Cov(β_s | y, X_s, M_s) given by expression involving 2F1 ratios.
  - Second moment of shrinkage factor:
    - E( (g/(1+g))^2 | y, X_s, M_s ) = [ 8 (k_s + a) (k_s + a + 2) / ( (k_s + a)^2 ) ]  2F1( (N−1)/2, 3, (k_s + a)/2 + 2, R^2_s ) / 2F1( (N−1)/2, 1, (k_s + a)/2, R^2_s )

### Practical implementation and numerical simplifications
- Posterior moments involve ratios of Gaussian hypergeometric functions 2F1; direct computation can be numerically challenging.
- Using Gauss’ relations, many moments are expressed in terms of F^*_s ≡ 2F1( (N−1)/2, 1, (k_s + a)/2, R^2_s ), reducing hypergeometric evaluations.
- Definitions:
  - ̄N ≡ N − 3
  - ̄θ_s ≡ k_s + a − 2
- Algebraic reformulations for R^2_s > 0:
  - E( g/(1+g) | y, X_s, M_s ) =
    - (1 / R^2_s) (̄N − ̄θ_s) ( ̄θ_s F^*_s − ̄θ_s + ̄N R^2_s )^{-1}
  - Cov(β | y, X_s, M_s) =
    - (̆y′ ̆y)/(N−2) (X′ X)^{-1} ̄N ( ̄N − ̄θ_s − 1 ) 2^{-1} (1 − R^2_s)^{-1} R^{-2}_s × ( (1 + 2 ̄N R^2_s/(1 − R^2_s)) ̄θ_s F^*_s + ((̄N − 2) R^2_s − ̄θ_s) )
  - E( (g/(1+g))^2 | y, X_s, M_s ) =
    - 1/(R^2_s)^2 (̄N − ̄θ_s)(̄N − (̄θ_s + 2))^{-1} × ( (( (̄N − 2) R^2_s − (̄θ_s + 2) ) ̄θ_s F^*_s + (̄N R^2_s − ̄θ_s)^2 − 2 (̄N (R^2_s)^2 − ̄θ_s) ) )
- Numerical benefits:
  - Each model requires computing F^*_s only once; other posterior quantities are algebraic functions of F^*_s, improving speed and accuracy over repeated Laplace approximations.
- Relationship to Empirical Bayes – Local (EBL):
  - Equations resemble EBL; the ̄θ_s / F^*_s term (equivalently a−2 times BF(M_s : M_0)) ensures non-negativity and moderates influence of models with very low ̄θ_s / F^*_s.

### Hyper-g interpretation, inequalities, and prior calibration
- Interpretation:
  - Posterior of shrinkage g/(1+g) interpretable in terms of goodness-of-fit.
  - Equation (13) shows model-specific expected shrinkage close to 1 − 1/ȞF_s, with ȞF_s an adjusted OLS F-statistic:
    - ȞF_s = R^2_s (N̄ − ̄θ_s) / ((1 − R^2_s) ̄θ_s)
  - Larger shrinkage corresponds to more variance explained by M_s.
- Inequality linking shrinkage, R^2, and model size (holds if some Bayes factors ≫ 1):
  - 1 / (1 − E(g/(1+g) | y, X)) ≤ R^2_F (1 − R^2_F) (N − E(k | y, X) − a − 1) / (E(k | y, X) + a − 2)  (equation (16))
  - R^2_F is the OLS R-squared of the full model with K regressors; E(k | y, X) is expected posterior model size.
  - Right-hand side is a pseudo F-statistic; forms an upper bound for BMA goodness-of-fit.
- Prior calibration via a:
  - Emulate popular g by a = 2 + 2/w(N) with w(N) > 0 and lim_{n→∞} w(N) = ∞ so E(g/(1+g)) = w(N)/(1 + w(N)).
  - Ensures consistency (cf. Fernández et al. (2001a)).
  - Proposed specifications:
    - HG-UIP: a = 2 + 2/N ⇒ E(g/(1+g)) = N/(1 + N); 95% prior mass on shrinkage in [1 − 0.95/N, 1].
    - HG-RIC: a = 2 + 2/K^2 ⇒ E(g/(1+g)) = K^2/(1 + K^2); 95% prior mass in [1 − 0.95/K^2, 1].
  - Posterior expressions are quite insensitive to a; most formulations yield a close-to-2 value for a and similar posterior statistics.

### Fixed versus flexible prior settings (overview used in simulations)
- Fixed prior settings (degenerate g values) examined:
  - g-RIC (g = K^2)
  - g-UIP (g = N)
  - g-E(g/(1+g) | Y) where g/(1+g) fixed to posterior mean under HG-4
- Flexible prior settings (model-specific, data-dependent g):
  - EBL (local empirical Bayes)
  - HG-3 (a = 3)
  - HG-4 (a = 4)
  - HG-RIC (a = 2 + 2/K^2)
  - HG-UIP (a = 2 + 2/N)

### Simulation study design (Section IV)
- Two Monte Carlo data-generating setups:
  - Setup "A":
    - y = 4 + 2x1 − x5 + 1.5x7 + x11 + 0.5x13 + σε
  - Setup "B":
    - y = 0.2y1 + 0.1y2 + 0.1y3 + 0.3y4 + 0.3y5 (y1,...,y5 are five partially nested models, each with intercept 4 and various x coefficients and σε)
- Covariates and sample:
  - K = 15, N = 100 observations
  - 10 potential explanatory variables x1,...,x10 drawn standard normal
  - Additional 5 variables generated by multiplying first five regressors by [0.3, 0.5, 0.7, 0.9, 1.1] to induce correlation
- Signal-to-noise ratios via σ levels:
  - σ = 1/2, σ = 1, σ = 2.5, σ = 5
- Full model space enumeration: 2^K models used for posterior inference

### Key simulation findings: inclusion, model probabilities, and prediction
- Posterior inclusion probabilities (PIPs) — Setting "A", averaged over 50 Monte Carlo draws:
  - For σ = 1/2 and σ = 1, fixed and flexible priors yield similar results.
  - For σ = 2.5, PIPs of data-generating coefficients differ in magnitude but preserve qualitative interpretation.
  - For σ = 5:
    - Flexible priors spread posterior mass more evenly across variables and models.
    - g-RIC shows strong support for the first variable with PIP for β1 ≈ 80%, other variables negligible.
    - Flexible priors still identify true variables; many covariates have PIP close to 0.5 under flexible priors reflecting high noise.
- Posterior model probability (PMP) behavior (Tables 6 and 7):
  - More information in data allows hyperprior to uncover the data-generating process with higher precision.
  - Increasing noise deteriorates selection ability of BMA for all priors.
  - In higher noise, all specifications favor a model different from the true DGP; flexible priors show dilution of posterior mass (surge in the PMP ratio of true to best model).
  - Flexible priors assign smaller PMP to the best model and reflect increased uncertainty by spreading mass more evenly.
- Visual diagnostics (Figures and QQ-plots):
  - Flexible priors concentrate mass and uncover DGP with high precision when information is high; as noise rises, differences between data-dependent priors and g-RIC increase.
- Prediction (out-of-sample rmse on 30 observations, averaged over 50 MC steps; rmse normalized to g-RIC):
  - Setting "A":
    - g-RIC excels in nearly all signal-to-noise settings by concentrating on a single (here correct) data-generating model.
    - For σ = 1/2, flexible priors can concentrate mass even more tightly than g-RIC and yield better predictions (rmse).
  - Setting "B":
    - Flexible priors nearly dominate across signal-to-noise settings.
    - HG-3 and empirical Bayes show superior predictive abilities; empirical Bayes outperforms g-RIC for all signal-to-noise setups.
    - Contrast: single-model DGPs favor large fixed g priors; more complex DGPs favor flexible priors.

### Application: Growth determinants revisited (Section V)
- Data and setup:
  - 41 potential growth determinants for 72 countries (data set from Fernández et al. (2001b) as described in Sala-i-Martin (1997))
  - Uniform model priors used
  - MCMC: 3,000,000 posterior draws after burn-in of 2,000,000 draws
- Interpretation of PIPs:
  - Variables with PIP > 0.5 often labeled ’robustly related’ to the dependent variable, but threshold depends on data information and model prior penalty.
  - A PIP slightly above 0.5 coupled with a small posterior mean of g/(1+g) differs qualitatively from the same PIP with large E(g/(1+g) | y, X).
- Empirical results:
  - Flexible priors identify additional growth determinants compared to g-RIC; differences manifest in posterior mean model sizes.
  - Flexible priors distribute posterior mass more evenly than fixed priors due to data noise.
  - Small differences within class of flexible priors; g-E(g/(1+g) | Y) and g-UIP among fixed priors are closest to flexible priors.
  - Specific coefficient differences relative to g-RIC observed (examples listed in source: regional dummy Hindu, HighEnroll, PublicEducpt).
- Posterior shrinkage factor in growth exercise:
  - Varies from 0.999 (g-RIC) to 0.951 (HG-4)
  - Fixing g risks ignoring information in data and exerts non-negligible influence on posterior results

### Conclusions and recommendations
- Two main arguments:
  - Decouple model size considerations from g’s scaling feature; incorporate model size into model prior formulation so elicitation of g does not interfere with prior model-size desiderata.
  - Fixing g to arbitrary values can produce unintended consequences on posterior model probabilities: larger g can cause posterior mass to concentrate on a few best-performing ’super models’ (supermodel effect), effectively performing single-model selection.
- The g-RIC prior is particularly prone to supermodel behavior.
- Recommendation:
  - Place a prior distribution on g (hyperprior) to allow data-dependent shrinkage and to adjust prior weight according to data quality.
  - The hyper-g prior is advocated because:
    - It admits closed-form solutions for many quantities of interest.
    - It allows BMA consistency.
    - Its hyperparameter lets one specify prior beliefs on coefficient variance without unintended effects on posterior model mass.
- The paper provides posterior expressions enabling fully Bayesian inference and practical numerical implementation for the hyper-g prior.

### Technical appendix: key analytical findings
- A.1 Consistency of hyper-g:
  - Consistency definition: if only Model M_s is true, require plim_{n→∞} p(M_s | y, X_s) = 1 and plim_{n→∞} p(M_j | y, X_s) = 0 ∀ M_j ≠ M_s
  - Liang et al. (2008) prove this for hyper-g except when true model is null M_0; for null-case an integral evaluates to (a−2)/(k_j + a−2); choosing a = 2 + w(N) with w(N) > 0 and lim_{N→∞} w(N) = 0 yields consistency.
- A.2 Relationship hyper-g and EBL:
  - Laplace approximation yields ˆg = max( R^2 (N−1−k−a) / ((1−R^2)(k+a)) − 1, 0 ); ˆg = 0 iff k + a ≥ R^2 (N−1)
  - As a → 2, BF_h ≈ BF_EBL × k-based model prior; the k-based prior term is bounded and its impact is "virtually negligible" relative to BF_EBL when (N−1) R^2 > k + a.
- A.3 Shrinkage factor and goodness-of-fit:
  - Posterior expected shrinkage has representation involving model-averaged terms and a non-negative remainder ϖ; practical closeness to upper bound depends on posterior variance of model size and parsimoniousness of model priors.
- A.4 Posterior predictive distribution under hyper-g:
  - Conditional predictive: ˆy | ˆX, X, y, g ∼ t_l( ̄y + s ˆX ˆβ, Σ, N−1 ) with s = g/(1+g), Σ = ( I_l + s ˆX (X′X)^{-1} ˆX′ ) ( y′y / (N−1) ) (1 − s R^2)
  - Integrating over s yields an integrand involving 2F1; no closed-form solution—numerical integration recommended.
- A.5 Beta-binomial prior over model space:
  - (a) Flat prior p(M_j) = 2^{−K}
  - (b) Covariate inclusion probability θ → p(M_j) = θ^{k_j} (1 − θ)^{K − k_j}
  - Treating θ as random with beta-binomial(a, b) (a = 1) mitigates strong prior influence; anchoring prior expected model size m sets b implicitly via b = (K − m)/m.
- A.6 Charts and tables:
  - Monte Carlo results across σ = 1/2, 1, 2.5, 5 for setups "A" and "B".
  - Reported numeric summaries include PIPs, E(k | Y) (examples: 5.027, 5.629, 7.711), E( g/(1+g) | Y) (examples: 0.996, 0.990, 0.998), PMP summaries (Min, Mean, Max, St.Dev.), PMP ratios, and relative RMSE entries (example minima and means reported: Min 0.1349, Mean 0.4618 in Table 8 examples).
  - Growth determinants results include variable-level PIPs and fully standardized posterior means (example entries: GDP60 0.9989, Confucian 0.9880, LifeExp 0.9314, EquipInv 0.9271, SubSahara 0.7288, Muslim 0.6524, RuleofLaw 0.4855).

*Source: _wp09202 - References19*

### References19

### References19

### Introduction: Bayesian Model Averaging (BMA) and the role of Zellner’s g
- BMA addresses model uncertainty by weighting all potential covariate combinations with posterior model probabilities (PMP):
  - p(M_s | y, X) ∝ p(y | M_s, X) p(M_s)
- Under the Normal-Gamma conjugate framework with Zellner’s g prior:
  - β | σ^2 ∼ N(0, g σ^2 (X′X)^{-1})
  - Marginal likelihood (model M_s) under a fixed g is proportional to:
    - (1 + g)^{-k_s/2} (1 − g/(1+g) R^2_s)^{-(N−1)/2}  (equation (1))
- Key conceptual points:
  - Larger g concentrates posterior mass on fewer models; smaller g spreads PMPs more evenly.
  - The “supermodel effect”: fixed large g can concentrate posterior mass on very few (or a single) model(s), independent of whether they are data-generating.
  - Using a nondegenerate prior on g (a hyperprior) can “let the data choose” g, reducing prior sensitivity and providing data-dependent shrinkage.

### Bayesian Model Averaging under Zellner’s g prior (framework and consequences)
- Model setup:
  - y = 1 α_s + X_s β_s + ε with ε ∼ N(0, σ^2 I)
  - Model space: 2^K possible covariate combinations
- Common (improper) priors used:
  - p(α) ∝ 1
  - p(σ) ∝ σ^{-1}
- Zellner’s g prior for coefficients:
  - (β_s | σ^2, M_s, g) ∼ N(0, σ^2 g (X′_s X_s)^{-1})
  - Posterior mean: E(β_s | y, X, M_s, g) = (g/(1+g)) ˆβ_s
  - Marginal likelihood (equation (3)):
    - p(y | M_s, g) ∝ (̆y′ ̆y)^{-(N−1)/2} (1 + g)^{-k_s/2} (1 − g/(1+g) R^2_s)^{-(N−1)/2}
- Bayes factor between models M_s and M_j (equation (4)):
  - B(M_s : M_j) = (1 + g)^{(k_j − k_s)/2} ( (1 − g/(1+g) R^2_s) / (1 − g/(1+g) R^2_j) )^{-(N−1)/2}
  - Denote the second factor by D_sj; larger g drives D_sj away from unity and amplifies concentration of PMPs.
- Model averaging for any statistic θ:
  - p(θ | y, X) = Σ_{j=1}^{2^K} p(θ | y, X, M_j) p(M_j | y, X)

### Popular fixed-g specifications and their motivations
- Considerations driving g choices:
  - Consistency: choose g = w(N) with lim_{N→∞} w(N) = ∞ and lim_{N→∞} w′(N)/w(N) = 0 to ensure p(M_T | Y) → 1
  - Model-size penalty: calibrate g to mimic information criteria (e.g., BIC, RIC)
- Frequently used default g choices (as summarized, following Liang et al. (2008)):
  - Risk Inflation Criterion Prior (g-RIC): g = K^2
  - Unit Information Prior (g-UIP): g = N
  - Benchmark Prior (g-BRIC): g = max(N, K^2)
  - Empirical Bayes – Local (EBL): g_s = arg max_g p(y | M_s, X, g), often expressed as g_s = max(0, F_s − 1) with
    - F_s = R^2_s (N−1−k_s) / ((1−R^2_s) k_s)
- Remarks:
  - Model priors p(M_s) are often uniform p(M_s) = 2^{-K}, or beta-binomial formulations can be used to reflect desired prior expected model size.
  - The penalty factor (1 + g)^{-k_s/2} in marginal likelihood (equation (5)) can be neutralized or adjusted via model priors if desired.

### The hyper-g prior: Beta prior on the shrinkage factor
- Hyper-g prior specification (Liang et al. (2008)):
  - g/(1 + g) ∼ Beta(1, a/2 − 1)  with a ∈ (2, ∞)
  - Equivalent prior density on g: p(g) = (a−2)/2 (1 + g)^{-a/2}
  - E(g/(1+g)) = 2/a
  - a = 4 → prior of g/(1+g) is uniform; a → 2 concentrates prior mass near shrinkage factor 1; a > 4 concentrates near 0
- Posterior and marginal expressions (closed-form involving Gaussian hypergeometric functions):
  - Posterior of g given model M_s (equation (6)):
    - p(g | y, X_s, M_s) = [ (k_s + a − 2)/2 ] 2F1( (N−1)/2, 1, (k_s + a)/2, R^2_s ) (1 + g)^{-(k_s + a)/2} (1 − g/(1+g) R^2_s)^{-(N−1)/2}
  - Marginal likelihood under hyper-g (equation (7)):
    - p(y | X_s, M_s) ∝ (̆y′ ̆y)^{-(N−1)/2} a−2  (k_s + a − 2)/2  2F1( (N−1)/2, 1, (k_s + a)/2, R^2_s )
  - Posterior expected shrinkage factor (equation (8)):
    - E( g/(1+g) | y, X_s, M_s ) = [ 2 / (k_s + a) ]  2F1( (N−1)/2, 2, (k_s + a)/2 + 1, R^2_s )  /  2F1( (N−1)/2, 1, (k_s + a)/2, R^2_s )
  - Posterior expected model response (equation (9)):
    - E(y | X_s, M_s) = 1 E(α_s | X_s, M_s) + E( g/(1+g) | y, X_s, M_s ) X_s ˆβ_s
- Posterior covariance of β_s and full posterior of β_s:
  - Cov(β_s | y, X_s, M_s) given in equation (10):
    - Cov (β_s | y, X_s, M_s) = [ 2 / (k_s + a) ] (̆y′ ̆y)/(N−2)  [ 2F1( (N−3)/2, 2, (k_s + a)/2 + 1, R^2_s ) / 2F1( (N−1)/2, 1, (k_s + a)/2, R^2_s ) ] (X′_s X_s)^{-1}
  - Posterior density p(β_s | y, X_s, M_s) expressed via a hypergeometric-ratio form (equation (11)).
- Second moment of the shrinkage factor (equation (12)):
  - E( (g/(1+g))^2 | y, X_s, M_s ) = [ 8 (k_s + a) (k_s + a + 2) / ( (k_s + a)^2 ) ]  2F1( (N−1)/2, 3, (k_s + a)/2 + 2, R^2_s ) / 2F1( (N−1)/2, 1, (k_s + a)/2, R^2_s )

### Practical implementation, algebraic simplifications, and numerical considerations
- Posterior moments involve ratios of Gaussian hypergeometric functions 2F1; direct computation can be numerically challenging.
- Using Gauss’ relations for contiguous hypergeometric functions, many posterior moments can be expressed in terms of a single quantity F^*_s ≡ 2F1( (N−1)/2, 1, (k_s + a)/2, R^2_s ), reducing the number of hypergeometric evaluations required.
- Define:
  - ̄N ≡ N − 3
  - ̄θ_s ≡ k_s + a − 2
- Algebraic reformulations (valid for R^2_s > 0) produce compact expressions:
  - E( g/(1+g) | y, X_s, M_s ) (equation (13)):
    - = (1 / R^2_s) (̄N − ̄θ_s) ( ̄θ_s F^*_s − ̄θ_s + ̄N R^2_s )^{-1}
  - Cov(β | y, X_s, M_s) (equation (14)):
    - = (̆y′ ̆y)/(N−2) (X′ X)^{-1} ̄N ( ̄N − ̄θ_s − 1 ) 2^{-1} (1 − R^2_s)^{-1} R^{-2}_s × ( (1 + 2 ̄N R^2_s/(1 − R^2_s)) ̄θ_s F^*_s + ((̄N − 2) R^2_s − ̄θ_s) )
  - E( (g/(1+g))^2 | y, X_s, M_s ) (equation (15)):
    - = 1/(R^2_s)^2 (̄N − ̄θ_s)(̄N − (̄θ_s + 2))^{-1} × ( (( (̄N − 2) R^2_s − (̄θ_s + 2) ) ̄θ_s F^*_s + (̄N R^2_s − ̄θ_s)^2 − 2 (̄N (R^2_s)^2 − ̄θ_s) ) )
- Numerical benefits:
  - Each model requires computing F^*_s only once; subsequent posterior quantities are algebraic functions of F^*_s, improving computational speed and accuracy compared to repeated Laplace approximations of hypergeometric function ratios.
- Relationship to Empirical Bayes – Local (EBL):
  - Equations (13)–(15) resemble EBL expressions; the key difference is the term ̄θ_s / F^*_s (equivalently a−2 times the Bayes Factor BF(M_s : M_0)), which ensures non-negativity and moderates the disproportionate influence of models with very low ̄θ_s / F^*_s that would otherwise receive extreme PMPs.

*Source: _wp09202 - References19*

### Section IV illustrates this effect in showing that

### _wp09202 - Section IV illustrates this effect in showing that

### Hyper-g prior and interpretation of the shrinkage factor
- The posterior distribution of the shrinkage factor g/(1+g) can be interpreted in terms of goodness-of-fit.
- Equation (13) presents its model-specific expected value as close to 1 − 1/ȞF_s, where ȞF_s represents an adjusted OLS F-statistic.
- For the model M_s:
  - ȞF_s = R^2_s (N̄ − ̄θ_s) / ((1 − R^2_s) ̄θ_s).
  - Larger values of the shrinkage factor correspond to more variance explained by the model M_s.
- The model-averaged expected value E(g/(1+g) | y, X) may be interpreted likewise.

### Inequalities linking the shrinkage factor, R^2, and model size
- As long as there are some Bayes factors considerably larger than one, the following inequality holds:
  - 1 / (1 − E(g/(1+g) | y, X)) ≤ R^2_F (1 − R^2_F) (N − E(k | y, X) − a − 1) / (E(k | y, X) + a − 2)  (equation (16))
  - R^2_F is the OLS R-squared of the ’full model’ with K regressors.
  - E(k | y, X) is the expected posterior model size.
  - The right-hand side constitutes a pseudo F-statistic relating R^2_F with the ’number of parameters’ E(k | y, X) + a − 2 and forms an upper bound for the ’goodness-of-fit’ achievable by BMA.
- Rearranged inequality relating to classical R-squared:
  - R^2_F ≥ (E(k | y, X) + a − 2) / (N − E(g/(1+g) | y, X) (N − E(k | y, X) − a − 1))

### Prior calibration via the hyperparameter a
- The hyperparameter a can be trimmed to represent prior beliefs on the shrinkage factor.
- Popular settings for g can be emulated by a = 2 + 2/w(N), with w(N) > 0 and lim_{n→∞} w(N) = ∞, positioning the prior expected value at E(g/(1+g)) = w(N)/(1 + w(N)).
- Consistency in the sense of Fernández et al. (2001a, p.6) is ensured for the ’hyper-g’ prior in this construction (cf. section A.1).
- Proposed specifications for prior beliefs on the shrinkage factor:
  - HG-UIP: a = 2 + 2/N corresponds to the ’g-UIP’-shrinkage factor with E(g/(1+g)) = N/(1 + N).
    - Then 95% of the prior mass on the shrinkage factor is contained in the interval [1 − 0.95/N, 1].
  - HG-RIC: a = 2 + 2/K^2 corresponds to ’g-RIC’-shrinkage with E(g/(1+g)) = K^2/(1 + K^2).
    - Then 95% of the prior mass is contained in the interval [1 − 0.95/K^2, 1].
- Posterior expressions are quite insensitive to the value of a; most formulations lead to a close-to-2 value for a, yielding virtually identical posterior statistics.

### Fixed versus flexible prior settings (overview from Section IV)
- Two classes of prior settings:
  - Fixed prior settings: degenerate ones fixing g values.
    - Fixed priors examined: g-RIC (g = K^2), g-UIP (g = N), g-E(g/(1+g) | Y) where g/(1+g) is set to the posterior mean under HG-4.
  - Flexible prior settings: model-specific and data-dependent g prior structures.
    - Flexible priors examined: EB L (local empirical Bayes estimate of g), HG-3 (hyper-g with a = 3), HG-4 (hyper-g with a = 4), HG-RIC (a = 2 + 2/K^2), HG-UIP (a = 2 + 2/N).
- Table 1 defines these 8 prior structures (fixed and flexible).

### Simulation study design (Section IV)
- Two Monte Carlo data-generating setups:
  - Setup "A":
    - y = 4 + 2x1 − x5 + 1.5x7 + x11 + 0.5x13 + σε
  - Setup "B":
    - y = 0.2y1 + 0.1y2 + 0.1y3 + 0.3y4 + 0.3y5, with y1,...,y5 defined as five partially nested models (each with intercept 4 and various x coefficients and σε).
- Covariates:
  - K = 15, N = 100 observations.
  - 10 potential explanatory variables x1,...,x10 drawn standard normal.
  - Additional 5 variables generated by multiplying first five regressors by [0.3, 0.5, 0.7, 0.9, 1.1] to induce correlation.
- Signal-to-noise ratios tested via four σ levels:
  - σ = 1/2, σ = 1, σ = 2.5, σ = 5.
- Enumerate full model space of 2^K models to base posterior inference on the full set.

### Key simulation findings: variable inclusion, model probabilities, and prediction
- Posterior inclusion probabilities (PIPs) for setting "A" averaged over 50 Monte Carlo draws:
  - For low noise (σ = 1/2 and σ = 1), results do not differ considerably between fixed and flexible priors.
  - For σ = 2.5, PIPs of data-generating coefficients exhibit magnitude differences but same qualitative interpretation.
  - For σ = 5:
    - Flexible priors spread posterior mass more evenly across variables and models.
    - g-RIC shows strong support for the first variable with a large PIP for β1 of approximately 80%, with remaining variables receiving negligible posterior support.
    - Flexible priors still identify all true variables; many covariates have PIP close to 0.5 under flexible priors reflecting high noise.
- Posterior model probability (PMP) behavior (Tables 6 and 7 summary):
  - More information in the data leads the hyperprior to uncover the data-generating process with higher precision.
  - Increasing noise deteriorates selection ability of BMA for all priors.
  - In higher noise, all specifications favor a model different from the true data-generating model; flexible priors show dilution of posterior mass (surge in the PMP ratio of true to best model).
  - Flexible priors assign smaller PMP to the best model and reflect surge of uncertainty by spreading mass more evenly.
- Figures and QQ-plots:
  - Flexible priors concentrate mass and uncover the data-generating model with high precision when information content is high; as noise increases, differences between data-dependent priors and g-RIC increase.
- Prediction (out-of-sample rmse):
  - rmse computed on 30 out-of-sample observations, averaged over 50 Monte Carlo steps; rmse statistics normalized with respect to g-RIC (values below 1 indicate better predictive performance than g-RIC).
  - Setting "A":
    - g-RIC excels in nearly all signal-to-noise settings by concentrating on a single (and in this case correct) data-generating model.
    - For σ = 1/2, flexible priors concentrate mass even more tightly than g-RIC and yield better predictions (rmse).
  - Setting "B":
    - Flexible priors nearly dominate throughout all signal-to-noise settings.
    - HG-3 and empirical Bayes demonstrate superior predictive abilities; empirical Bayes outperforms g-RIC for all signal-to-noise setups.
    - Contrast with literature simulation exercises where single-model DGPs favor large fixed g priors.

### Application: Growth determinants revisited (Section V)
- Data: 41 potential growth determinants for 72 countries (data set from Fernández et al. (2001b) described in Sala-i-Martin (1997)).
- Model priors: uniform model priors employed (instead of beta-binomial).
- MCMC: 3,000,000 posterior draws after burn-in of 2,000,000 draws.
- Interpretation of PIPs:
  - Variables with PIP > 0.5 often identified as ’robustly related’ to the dependent variable, but the threshold should vary with data information and model prior penalty.
  - A PIP slightly above 0.5 coupled with a rather small posterior mean of g/(1+g) cannot be interpreted the same as when E(g/(1+g) | y, X) is large.
- Empirical results:
  - Flexible priors identify a range of additional growth determinants compared to g-RIC, manifested in differences of posterior mean model sizes.
  - Flexible priors distribute posterior mass more evenly than fixed priors due to data noise.
  - Small differences within the class of flexible priors; g-E(g/(1+g) | Y) and g-UIP fixed priors are closest to flexible priors among fixed priors.
  - Some variables differ in posterior means and standardized coefficients relative to g-RIC (examples: regional dummy Hindu, HighEnroll, PublicEducpt).
- Posterior shrinkage factor observed in growth exercise:
  - Varies from 0.999 (g-RIC) to 0.951 (HG-4).
  - Fixing g risks ignoring information in the data and exerts non-negligible influence on posterior results.

### Conclusions and recommendations
- Two main arguments advanced:
  - Model size considerations should be decoupled from the scaling feature of g and instead be incorporated into model prior formulation; elicitation of g should not interfere with prior desiderata on model size.
  - Fixing g to arbitrary values may have unintended consequences on posterior model probabilities: larger g causes posterior mass to concentrate on a few best-performing ’super models’ (the supermodel effect), effectively acting as model selection favoring a single model.
- The g-RIC prior is particularly prone to supermodel behavior.
- Recommendation:
  - Put a prior distribution on g (hyperprior) to allow data-dependent shrinkage and adjust the weight of prior beliefs according to data quality.
  - The hyper-g prior (Liang et al. (2008)) is advocated because:
    - It admits closed-form solutions for almost any quantity of interest.
    - It allows for BMA consistency.
    - Its hyperparameter allows formulating prior beliefs on coefficient variance without risking unintended consequences on posterior model mass.
- The paper provides additional posterior expressions enabling fully Bayesian inference and sound numerical implementation for the hyper-g prior.

*Source: Excerpt from the provided PDF content unit.*

### Section  IV  contrasts  various  formulations  of  fixed  and  hyper-g  priors  in  simulations,  con-

### Section IV contrasts various formulations of fixed and hyper-g priors in simulations

### Simulation design and focus
- Contrasts various formulations of fixed and hyper-g priors in simulations, concentrating on predictive performance under varying signal-to-noise ratios.
- Emphasis on how prior choice interacts with the data generating process being a single model in the candidate space versus more complex settings.

### Key findings on predictive performance
- Fixed priors (especially the g-RIC) perform considerably well when the data generating process rests on a single model that is part of the candidate model space.
- In more complex settings, flexible prior structures (hyper-g priors) show pronounced virtues:
  - Flexible priors outperform fixed g settings (in particular g-RIC) in terms of forecasting accuracy.
  - Flexible priors exhibit a more stable structure of posterior model and inclusion probabilities as noise varies.

### Application to a prominent growth data set
- Applying the same priors to a prominent growth data set illustrates the simulation considerations.
- Fixing g runs the risk of grossly over- or understating the importance of some variables:
  - Example: the degree of openness is not as important to growth as one may think under the g-RIC prior.
- Fixing g to values larger than implied by flexible priors leads to stronger discrimination among posterior inclusion probabilities, which may incite overconfidence in BMA results.
- The magnitudes of several coefficients differ markedly between fixed and hyper-g priors, but are negligible among the hyper-g prior structures.

### Conclusion
- The hyper-g prior offers a sound, fully Bayesian approach that features the virtues of prior input and predictive gains without incurring the risk of misspecification.

*Source: https://www.imf.org/-/media/websites/imf/imported-full-text-pdf/external/pubs/ft/wp/2009/_wp09202.pdf*

### References

### _wp09202 - References

### References (selected)
- Abramowitz, M. and Stegun, I. (1972). Handbook of Mathematical Functions with Formulas, Graphs, and Mathematical Tables. National Bureau of Standards Applied Mathematics Series 55. Tenth Printing.
- Barbieri, M. M. and Berger, J. O. (2003). Optimal Predictive Model Selection. Ann. Statist., 32:870–897.
- Bernardo, J. and Smith, A. (1994). Bayesian Theory. John Wiley and Sons, New York.
- Brown, P., Vannucci, M., and Fearn, T. (1998). Multivariate Bayesian Variable Selection and Prediction. Journal of the Royal Statistical Society B, 60:627–641.
- Chipman, H., George, E., and McCulloch, R. (2001). The Practical Implementation of Bayesian Model Selection. Institute of Mathematical Statistics Lecture Notes-Monograph Series Vol. 38. Beachwood, Ohio.
- Crespo Cuaresma, J. and Doppelhofer, G. (2007). Nonlinearities in Cross-Country Growth Regressions: A Bayesian Averaging of Thresholds (BAT) Approach. Journal of Macroeconomics, 29:541–554.
- Cui, W. and George, E. (2008). Empirical Bayes vs. fully Bayes variable selection. Journal of Statistical Planning and Inference, 138:4:888–900.
- Eicher, T., Papageorgiou, C., and Raftery, A. (2009). Determining growth determinants: default priors and predictive performance in Bayesian model averaging. Journal of Applied Econometrics, forthcoming.
- Eklund, J. and Karlsson, S. (2007). Forecast Combination and Model Averaging using Predictive Measures. Econometric Reviews, 26:329–362.
- Fernández, C., Ley, E., and Steel, M. F. (2001a). Benchmark Priors for Bayesian Model Averaging. Journal of Econometrics, 100:381–427.
- Fern ández, C., Ley, E., and Steel, M. F. (2001b). Model Uncertainty in Cross-Country Growth Regressions. Journal of Applied Econometrics, 16:563–576.
- Foster, D. P. and George, E. I. (1994). The Risk Inflation Criterion for Multiple Regression. The Annals of Statistics, 22:1947–1975.
- Gelman, A., Carlin, J. B., Stern, S. H., and Rubin, B. D. (1995). Bayesian Data Analysis. Chapman & Hall.
- George, E. and Foster, D. (2000). Calibration and empirical Bayes variable selection. Biometrika, 87(4):731–747.
- Guptar, A. K. and Nagar, D. K. (2000). Matrix Variate Distributions. Chapman & Hall /CRC, Monographs and Surveys in Pure and Applied Mathematics, 104.
- Hansen, M. and Yu, B. (2001). Model selection and the principle of minimum description length. Journal of the American Statistical Association, 96(454):746–774.
- Hoeting, J. A., Madigan, D., Raftery, A. E., and Volinsky, C. T. (1999). Bayesian Model Averaging: A Tutorial. Statistical Science, 14, No. 4:382–417.
- Kass, R. and Raftery, A. (1995). Bayes Factors. Journal of the American Statistical Association, 90:773–795.
- Kass, R. and Wasserman, L. (1995). A reference Bayesian test for nested hypotheses and its relationship to the Schwarz criterion. Journal of the American Statistical Association, pages 928–934.
- Koop, G. and Potter, S. (2003). Forecasting in Large Macroeconomic Panels Using Bayesian Model Averaging. FRB NY Staff Report, 163.
- Laud, P. and Ibrahim, J. (1995). Predictive model selection. Journal of the Royal Statistical Society, Series B 57:247–262.
- Ley, E. and Steel, M. F. (2009). On the Effect of Prior Assumptions in Bayesian Model Averaging with Applications to Growth Regressions. Journal of Applied Econometrics, 24:4:651–674.
- Liang, F., Paulo, R., Molina, G., Clyde, M. A., and Berger, J. O. (2008). Mixtures of g Priors for Bayesian Variable Selection. Journal of the American Statistical Association, 103:410–423.
- Masanjala, W. and Papageorgiou, C. (2008). Rough and Lonely Road to Prosperity: A Re-examination of the Sources of Growth in Africa Using Bayesian Model Averaging. Journal of Applied Econometrics, 23:671–682.
- R Development Core Team (2008). R: A Language and Environment for Statistical Computing. R Foundation for Statistical Computing, Vienna, Austria, ISBN 3-900051-07-0 edition. http://www.R-project.org.
- Raftery, A. E. (1995). Bayesian Model Selection in Social Research. Sociological Methodology, 25:111–163.
- Sala-i-Martin, X. (1997). I Just Ran 2 Million Regressions. American Economic Review, 87:178–183.
- Sala-i-Martin, X., Doppelhofer, G., and Miller, R. I. (2004). Determinants of Long-Term Growth: A Bayesian Averaging of Classical Estimates (BACE) Approach. American Economic Review, 94:813–835.
- Strachan, R. and van Dijk, H. (2004). Exceptions to Bartletts Paradox. Keele Economic Research Papers.
- Zellner, A. (1986). Bayesian Inference and Decision Techniques: Essays in Honor of Bruno de Finetti, chapter On Assessing Prior Distributions and Bayesian Regression Analysis with g-Prior Distributions. North-Holland: Amsterdam.
- Zellner, A. (2008). Comments on Mixtures of g-priors for Bayesian Variable Selection by F. Liang, R. Paulo, G. Molina, MA Clyde and JO Berger.

### Technical Appendix: Major analytical findings and results
- A.1 Consistency of the Hyper-g Prior
  - Definition adopted: if only Model M_s is true, consistency requires:
    - plim_{n→∞} p(M_s | y, X_s) = 1
    - plim_{n→∞} p(M_j | y, X_s) = 0 ∀ M_j ≠ M_s
  - Liang et al. (2008, Appendix B) prove this for the hyper-g prior except when the true model M_s is the null model M_0.
  - For the null-case they derive inequality (A.1):
    - p(M_j | y, X_s) / p(M_0 | y) ≥ ∫_0^∞ (1 + g)^{-k_j/2} p(g) dg
  - Under the hyper-g setting (with a > 2) the integral becomes:
    - ∫_0^∞ (1 + g)^{-k_j/2} p(g) dg = (a−2)/(k_j + a−2)
  - If a = 2 + w(N) with w(N) > 0 and lim_{N→∞} w(N) = 0, then the integral vanishes and consistency is concluded.

- A.2 Relationship between Hyper-g Prior and Empirical Bayes Local (EBL)
  - Liang et al. (2008) propose an Laplace approximation for the posterior model likelihood under hyper-g (equation (17)).
  - Define BF_h (null-based Bayes factor under hyper-g):
    - BF_h = (a−2)/2 ∫_0^∞ (1 + g)^{(N−1−k−a)/2} (1 + g(1 − R^2))^{−(N−1)/2} dg
  - Using Laplace with h(g) = 1/2[(N−1−k−a) log(1+g) − (N−1) log(1 + (1−R^2) g)] yields maximizer:
    - ˆg = max( R^2 (N−1−k−a) / ((1−R^2)(k+a)) − 1, 0 ) — equivalently ˆg = max(R^2 (N−1−k−a)/( (1−R^2)(k+a) ) − 1, 0); ˆg = 0 iff k + a ≥ R^2 (N−1)
  - Approximate BF_h expressions (cases with ˆg > 0 and algebraic manipulations) show that as a → 2, BF_h ≈ BF_EBL × a^{k}-based model prior (model prior not depending on z).
  - The k-based model prior term is bounded:
    - (1 + a)^{−a/2} ≤ ( (k + a)/k )^{k/2} ( (N−1−k−a)/(N−1−k) )^{(N−1−k)/2} < 1
  - The k-based model prior tends to upweight very small or very large models and downweight intermediate sizes; however its impact is "virtually negligible" relative to BF_EBL when (N−1) R^2 > k + a.
  - If k = 0 then BF_EBL = BF_h = 1.

- A.3 The Shrinkage Factor and Goodness-of-Fit
  - Posterior expected shrinkage E( g/(1+g) | y, X ) has representation (A.2):
    - E( g/(1+g) | y, X ) = ϖ + (2/K) ∑_{j=1}^K p(M_s | y, X) [ ̄N R^2_s − ̄θ_s R^2_s ( ̄N − ̄θ_s )^{-1} ] with ϖ defined in the text
  - The term ϖ is non-negative for 'bad' models and rapidly vanishes for models with higher signal-to-noise ratios; for fixed K it vanishes as N → ∞.
  - For K + a < N and ̄θ_s ≤ ̄N the following inequality is shown (derivation summarized):
    - E_M( g/(1+g) | y, X ) − ϖ = 1 − E_M( (1 − R^2_s)/R^2_s θ_s ( ̄N − ̄θ_s )^{-1} ) ≤ R^2_F ̄N − E_M( ̄θ_s ) / ( R^2_F ( ̄N − E_M( ̄θ_s ) ) )
  - Practical implication: closeness to upper bound determined by posterior variance of model size and parsimoniousness of model priors; ϖ is typically negligible if R^2_F > (K + a)/N (> (K + a − 2)/(N − 3)).

- A.4 Posterior Predictive Distribution under Hyper-g
  - Posterior predictive distribution of ˆy conditional on ˆX, X, y, g:
    - ˆy | ˆX, X, y, g ∼ t_l( ̄y + s ˆX ˆβ, Σ, N−1 )
    - Σ = ( I_l + s ˆX (X′X)^{-1} ˆX′ ) ( y′y / (N−1) ) (1 − s R^2)
    - s denotes shrinkage s = g/(1+g); R^2 is centered R-squared of y on X
  - Integrating over s yields an integrand (displayed in the text) involving a hypergeometric 2F1 term and an integral over s from 0 to 1.
  - There is no closed-form solution to the integral nor to its Laplace approximation; numerical integration is recommended.

- A.5 Beta-binomial Prior over Model Space
  - Two typical prior specifications:
    - a) Flat prior over all models → p(M_j) = 2^{−K}
    - b) Prior that each covariate enters with probability θ → p(M_j) = θ^{k_j} (1 − θ)^{K − k_j}
  - Ley and Steel (2009) recommend treating θ as random and placing a hyperprior so model size follows a beta-binomial(a, b) with a = 1:
    - P(k = k_j) = Γ(1 + b) / ( Γ(1) + Γ(b) + Γ(1 + b + K) ) × ( K choose k_j ) Γ(1 + k_j) Γ(b + K − k_j), k_j = 0, ..., K  (equation (A.3))
    - (Note: b implicitly defined through b = (K − m)/m when anchoring prior expected model size m)
  - Ley and Steel (2009) findings:
    - Fixing θ = 1/2 concentrates prior mass on models with K/2 regressors.
    - Treating θ as random (beta-binomial prior) mitigates strong prior influence; choice of prior expected model size m has no influential impact on posterior inference and the model prior is effectively non-informative.

- A.6 Charts and Tables (Overview of supplied figures and tables)
  - Figures and tables provide Monte Carlo simulation results comparing different g-prior choices and priors over model space across settings labeled "A" and "B" with various signal-to-noise ratios σ = 1/2, 1, 2.5, 5.
  - Key reported numeric summaries include:
    - Posterior Inclusion Probabilities (PIP) tables for coefficients β_1 ... β_15 across g-RIC, g-UIP, g-E(g/(1+g)|Y), EBL, HG-3, HG-4, HG-RIC, HG-UIP settings, averaged over 50 Monte Carlo Steps (tables 2–5).
    - Expected model size E(k | Y) and E( g/(1+g) | Y) reported in multiple tables (examples: E(k|Y) reported values such as 5.027, 5.629, 7.711, etc.; E( g/(1+g) | Y) reported values such as 0.996, 0.990, 0.998, etc. in respective panels).
    - Summary statistics of posterior model probabilities for the true model (Table 6) with Min, Mean, Max, St.Dev. panels for σ = 1/2, 1, 2.5, 5.
    - Summary statistics of ratio of posterior model probabilities of true model and best model (Table 7).
    - Relative Root Mean Squared Error (out-of-sample forecasts over 30 forecasts averaged over 50 MC steps) relative to g-RIC (Table 8). Values below 1 indicate superior predictive performance to g-RIC; example entries include 0.1349 minimum, mean 0.4618, etc., across priors.
    - Variable-level posterior inclusion comparison table (Table 9) and fully standardized posterior means (Table 10) for growth determinants exercise; reported entries include many numeric values for individual covariates (examples: GDP60 0.9989, Confucian 0.9880, LifeExp 0.9314, EquipInv 0.9271, SubSahara 0.7288, Muslim 0.6524, RuleofLaw 0.4855, etc.).
  - Visualization suite (Figures 1–5) includes:
    - Cumulated posterior model probabilities across model index for various σ settings (Figures 1 and 3).
    - QQ-plots of cumulated posterior mass comparing different g choices against g-RIC (Figures 2 and 4).
    - Panels showing cumulative posterior mass, posterior inclusion probabilities (PIP), standardized coefficients, and posterior means for the growth determinants exercise (Figure 5).

*Source: _wp09202 - References (PDF content provided)*

---


_Source: https://www.imf.org/-/media/websites/imf/imported-full-text-pdf/external/pubs/ft/wp/2009/_wp09202.pdf_
