## Identifying Reform Priorities: The Role of Non-linearities (WP/20/278)

## Source details

**Canonical URL:** [Identifying Reform Priorities: The Role of Non-linearities (WP/20/278)](https://www.imf.org/-/media/files/publications/wp/2020/english/wpiea2020278-print-pdf.pdf)

## Other formats

- [Markdown version](/-/media/files/publications/wp/2020/english/wpiea2020278-print-pdf.pdf.md)
- [Structured JSON version](/-/media/files/publications/wp/2020/english/wpiea2020278-print-pdf.pdf.json)

---

### Abstract and central question
- Can countries improve their business climate through reforms in specific policy areas?
- Kraay and Tawara (2013) find little agreement across models about which policy indicators matter most when using linear BMA.
- This paper revisits the puzzle using the same data but replaces linear Bayesian Model Averaging (BMA) with a Random Forest (RF) algorithm and evaluates variable importance with Shapley values.

### Data and empirical setup
- Outcome variables:
  - seven business climate indices constructed from seven sources (DRI=Global Insight Global Risk Service; EIU=Economist Intelligence Unit; GAD=Cerberus Corporate Intelligence Gray Area Dynamics; GCS=Global Competitiveness Report; PRS=Political Risk Services; WMO=Global Insight Business Risk Conditions).
  - Results reported for six outcome variables due to confidentiality of CPIA data.
- Explanatory variables:
  - 38 World Bank Doing Business (DB) indicators, rescaled to [0,1]; higher DB scores indicate better performance.
- Cross-section: estimations use data for 2009; country coverage varies across outcome variables.
- BMA prior: KT impose a prior that the number of variables in the true model is 10.
- Random Forest hyperparameters:
  - number of trees set to 5000;
  - parameter m_try obtained through cross-validation;
  - each tree is grown exhaustively until no more splits are possible.
- Variable importance measure for RF: absolute Shapley values averaged over all observations.

### Methodological contrast: linear BMA versus Random Forest
- BMA (as in KT):
  - Assumes linear decision rule y_c,i = x_c′ β_i + ε_c,i.
  - Uses posterior inclusion probabilities (PIP) to rank variable importance.
  - KT result: no DB indicator is among the top 10 in all outcome-variable regressions (high instability).
- Regression trees and Random Forest:
  - Trees split the sample by variable thresholds, capturing interactions and threshold effects.
  - RF builds many trees on bootstrapped samples and random subsets of variables to improve robustness.
  - Shapley values allocate contribution of each predictor to individual predictions; averaging absolute Shapley values yields global importance measures.

### Main empirical findings: stability and consensus under Random Forest
- RF produces much stronger consensus across outcome variables than BMA.
- Four DB indicators are among the top 10 contributors across all six RF models:
  - Recovery rates for creditors in insolvency cases (cents on the dollar for creditors).
  - Time exporters take to complete border formalities.
  - Time importers take to complete border formalities.
  - Cost of starting a business.
- Comparison of stability:
  - Under BMA, not a single DB indicator is in the top 10 for all outcome variables.
  - Under RF, four DB indicators are in the top 10 for all outcome variables.
  - RF places 21 DB variables never in the top 10 (versus 12 under BMA).
- RF’s improved stability manifests at both the top and lower ends of the ranking distribution of variable importance.

### Construction of composite index and Shapley explanation
- Composite index ̃y constructed as:
  - ̃y_c = sum_i ( y_{c,i} − mean(y_i) ) / std.dev.(y_i)
  - missing y_{c,i} replaced with mean(y_i) if at least one outcome value is available for that country.
  - sample for ̃y is the union of the six samples used previously.
- RF re-estimated using ̃y as dependent variable.
- Shapley values c,j measure contribution of indicator j to deviation of ̃y_c from sample average.
- Charts produced by the SHAP package by Lundberg and Lee (2017).

### Key patterns, non-linearities, and threshold effects
- Top indicators (consistent across specifications) include:
  - recovery rate for creditors in bankruptcy procedures;
  - time to export;
  - time to import.
- Shapley-value patterns:
  - Countries with high indicator scores (red) tend to have better ̃y than countries with low scores (blue), but relationships are often non-linear and points cluster.
  - Indicators ranked relatively low overall can be important for specific countries (example: cost of firing employees, ranked 19th, but substantially reducing ̃y for a handful of cases).
  - Using BMA, influential outliers can raise a variable’s PIP.
- Examples of non-linear threshold patterns:
  - Recovery rates in insolvency resolution:
    - for values below 0.6 the slope is only slightly positive;
    - there is a threshold somewhere between 0.6 and 0.8 where perceived business climate improves drastically;
    - for values above 0.8 further improvements have no impact (framework deemed “good enough”).
  - Time to clear border formalities (export/import): ̃y is close to linear in these indicators.
  - Cost of obtaining a construction permit: matters much more once the indicator is above 0.5.
  - Similar thresholds appear for:
    - cost of starting a business;
    - documents required to export;
    - legal rights of borrowers and lenders;
    - time to start a business.
- Policy-relevant implication of non-linearities:
  - For countries with very low DB indicator values, the indicator explains poor perception of the business environment, but crossing thresholds can require large reform efforts (not necessarily a low-hanging fruit).
  - Imposing a linear functional form would make estimated slopes dependent on the share of observations below threshold values (e.g., below 0.6), so small changes to country coverage can lead to big changes in estimated variable importance.

### Interaction analysis and reform sequencing implications
- Interaction strength measured via H-statistics (Friedman and Popescu, 2008):
  - H-statistic for indicator j measures variation in the prediction function explained by interactions between j and all other indicators (excluding j), as a share of total variation explained by j.
  - Figure 5 reports H-statistics and finds interactions are very weak: for any DB indicator, less than 10 percent of explanatory power is obtained from interactions.
- Policy implication:
  - Low degree of complementarity means sequencing of reforms may not be a first-order issue, since payoff from reforms in one area does not depend heavily on the status of reforms in other areas.
- Note on RF interactions:
  - RF considers an interaction only if at least one of the two variables matters in itself (without the other); in most cases this is a mild and plausible restriction.

### Variables of limited importance (stability at the bottom)
- 21 variables are never among the top 10. These include:
  - labor and profit tax rates;
  - employment regulations;
  - protections for minority shareholders;
  - mechanisms to register property and enforce contracts.
- Caveat: due to non-linearities, a variable’s importance can vary substantially across observations.
- A variable’s high importance for explaining country c’s business climate score does not imply it should be a reform priority for country c because reform priorities should reflect marginal effects from changing a variable, not average effects.

### Causality, bias, and limitations
- Causal interpretation caveats:
  - RF inherits causality concerns from BMA; switching methods does not eliminate these concerns.
  - Two remaining concerns:
    - omitted variable bias: the 38 DB variables may not be comprehensive enough (example: 2009 DB indicators do not cover ease of getting electricity).
    - parameter bias induced by methods to avoid overfitting: for RF this bias makes measured importance of top variables more likely to be diluted than exaggerated.
  - With these caveats, results do allow for a causal interpretation, but magnitude caution is warranted.
- Small-sample caveat:
  - Absence of evidence of an effect does not imply evidence of absence of an effect; limited variation in some variables could explain their apparent lack of importance.

### RF hyper-parameter selection and sample statistics
- Repeated cross-validation procedure to find optimal m_try:
  1. sample divided into 10 random subsamples;
  2. guess a value for m_try;
  3. for each subsample estimate RF on remaining observations and measure out-of-sample performance as average RMSE over the 10 subsamples;
  4. repeat for m_try between 2 and 20;
  5. repeat steps 1–4 ten times;
  6. use the m_try that delivers best average RMSE over the 100 held-out samples to estimate RF on full sample.
- Table 3: Hyper-parameters (as reported)
  - Outcome variable columns: DRI, EIU, GAD, GCS, PRS, WMO, ̃y
  - #observations: 137, 144, 158, 130, 133, 178, 178
  - m_try: 510546127

### Policy-relevant conclusions and recommendations
- Perceptions of a country’s business environment can be tied to specific DB indicators under policymakers’ control.
- Generally, first-order importance indicators:
  - efficient insolvency procedures (recovery rates);
  - speedy border formalities (time to export and import);
  - low startup costs (cost/time/documents to start a business).
- Exact order of priorities depends on country-specific circumstances; identifying country-specific priorities is enabled by non-linear modeling techniques like Random Forest.
- Using Random Forest instead of linear techniques reduces sensitivity of explanations to how the overall business environment is measured.

*Source: WP/20/278 — Identifying Reform Priorities: The Role of Non-linearities (Klaus-Peter Hellwig, December 2020).*

### Section 1

### Identifying Reform Priorities: The Role of Non-linearities

### Abstract and central question
- Can countries improve their business climate through reforms in specific policy areas?
- Kraay and Tawara (2013) regress seven different business climate indices on 38 policy indicators and find little agreement across models about which policy indicators matter most.
- This paper revisits the puzzle using the same data but replaces linear Bayesian Model Averaging (BMA) models with a Random Forest (RF) algorithm and evaluates variable importance with Shapley values.

### Data and empirical setup
- Outcome variables: seven business climate indices constructed from seven sources (DRI=Global Insight Global Risk Service; EIU=Economist Intelligence Unit; GAD=Cerberus Corporate Intelligence Gray Area Dynamics; GCS=Global Competitiveness Report; PRS=Political Risk Services; WMO=Global Insight Business Risk Conditions). Results reported for six outcome variables due to confidentiality of CPIA data.
- Explanatory variables: 38 World Bank Doing Business (DB) indicators, rescaled to [0,1]; higher DB scores indicate better performance.
- Cross-section: estimations use data for 2009; country coverage varies across outcome variables.
- BMA prior: KT impose a prior that the number of variables in the true model is 10.
- Random Forest hyperparameters: number of trees set to 5000; parameter m_try obtained through cross-validation; each tree is grown exhaustively until no more splits are possible.
- Variable importance measure for RF: absolute Shapley values averaged over all observations.

### Methodology: linear vs tree-based approaches
- BMA (as in KT):
  - Assumes linear decision rule y_c,i = x_c′ β_i + ε_c,i.
  - Uses posterior inclusion probabilities (PIP) to rank variable importance across candidate models.
  - KT result: no DB indicator is among the top 10 in all outcome-variable regressions (high instability across the seven indices).
- Regression trees and Random Forest:
  - Trees split the sample by variable thresholds, naturally capturing interactions and threshold effects.
  - RF builds many trees on bootstrapped samples and random subsets of variables to improve robustness.
  - Shapley values allocate contribution of each predictor to individual predictions; averaging absolute Shapley values across observations yields global importance measures.

### Main findings: stability and consensus under Random Forest
- RF produces much stronger consensus across outcome variables than BMA.
- Four DB indicators are among the top 10 contributors across all six RF models:
  - Recovery rates for creditors in insolvency cases (cents on the dollar for creditors).
  - Time exporters take to complete border formalities.
  - Time importers take to complete border formalities.
  - Cost of starting a business.
- Comparison of stability:
  - Under BMA, not a single DB indicator is in the top 10 for all outcome variables.
  - Under RF, four DB indicators are in the top 10 for all outcome variables.
  - RF places 21 DB variables never in the top 10 (versus 12 under BMA).
- RF’s improved stability manifests at both the top and lower ends of the ranking distribution of variable importance.

### Interpretation: role of non-linearities and interactions
- Regression trees capture:
  - Threshold effects (e.g., an indicator matters only above or below a critical value).
  - Interactions (e.g., the importance of creditors’ recovery rates is conditional on startup costs).
- Consequence: reform priorities are country-specific because marginal effects of reforms depend on where a country sits relative to thresholds and on interactions among policy areas.
- Example (illustrative tree):
  - First split: cost of starting a business at threshold 0.43.
  - Left branch (low startup cost): next split by time to clear export formalities.
  - Right branch (high startup cost): next split by recovery rates in insolvency procedures.

### Policy implications and reform sequencing
- Policy priorities inferred from linear methods may be misleading if the true mapping from DB indicators to perceived business climate is non-linear and interactive.
- Reform sequencing matters: because interactions create conditional importance, reforms that are effective in one country may be less effective or ineffective elsewhere unless prerequisite reforms or threshold crossings occur.
- The RF results consistently identify improvements in insolvency recovery rates, shorter border formalities for importers and exporters, and lower costs to start a business as robust reform priorities across different measures of perceived business climate.

### Conclusion
- Replacing linear BMA with Random Forest reduces sensitivity of variable-importance rankings to the choice of business climate index and yields a stronger consensus on priority reform areas.
- Marginal effects of reforms are heterogeneous across countries; understanding non-linearities and interactions is essential for identifying country-specific reform priorities.

*Source: WP/20/278 — Identifying Reform Priorities: The Role of Non-linearities (Klaus-Peter Hellwig, December 2020).*

### Section 2

### Section 2 — The role of non-linearities

### Stability of Random Forest (RF) importance and variables of limited importance
- 21 variables are never among the top 10. These include:
  - labor and profit tax rates
  - employment regulations
  - protections for minority shareholders
  - mechanisms to register property and enforce contracts
- Table 2 reports each variable’s average importance across all observations, but:
  - due to non-linearities, a variable’s importance can vary substantially across observations
  - a variable’s high importance for explaining country c’s business climate score does not imply it should be a reform priority for country c because reform priorities should reflect marginal effects from changing a variable, not average effects
  - identifying reform priorities therefore requires looking at each country’s unique circumstances

### Construction of composite business climate index and explanation method
- Composite index ̃y is constructed as:
  - ̃y_c = sum_i ( y_{c,i} − mean(y_i) ) / std.dev.(y_i)
  - whenever y_{c,i} is missing, it is replaced with mean(y_i), as long as at least one outcome value is available for that country
  - the sample for ̃y is the union of the six samples used previously
- RF is re-estimated using ̃y as the left-hand side variable
- Shapley values c,j measure the contribution of indicator j in explaining the deviation of ̃y_c from the sample average
- Charts in this section are produced by the SHAP package by Lundberg and Lee (2017)

### Key empirical findings on variable importance and shapes of relationships
- Top indicators (consistent across specifications) include:
  - recovery rate for creditors in bankruptcy procedures
  - time to export
  - time to import
- Shapley-value patterns:
  - countries with high indicator scores (red) tend to have better ̃y than countries with low scores (blue), but relationships are often non-linear and points cluster
  - indicators ranked relatively low overall can be important for specific countries (example: cost of firing employees, ranked 19th, but substantially reducing ̃y for a handful of cases)
  - when using BMA, influential outliers can raise a variable’s PIP (posterior inclusion probability)
- Examples of non-linear threshold patterns:
  - recovery rates in insolvency resolution:
    - for values below 0.6 the slope is only slightly positive
    - there is a threshold somewhere between 0.6 and 0.8 where perceived business climate improves drastically
    - for values above 0.8 further improvements have no impact (framework deemed “good enough”)
  - time to clear border formalities (export/import): ̃y is close to linear in these indicators
  - cost of obtaining a construction permit: matters much more once the indicator is above 0.5
  - similar thresholds appear for:
    - cost of starting a business
    - documents required to export
    - legal rights of borrowers and lenders
    - time to start a business
- Policy-relevant implication of non-linearities:
  - for countries with very low DB indicator values, the indicator explains poor perception of the business environment, but crossing thresholds can require large reform efforts (not necessarily a low-hanging fruit)
  - imposing a linear functional form would make estimated slopes dependent on the share of observations below threshold values (e.g., below 0.6), so small changes to country coverage can lead to big changes in estimated variable importance

### Causality, bias, and limitations
- Causal interpretation caveats:
  - RF inherits causality concerns from BMA; switching methods does not eliminate these concerns
  - two remaining concerns:
    - omitted variable bias: the 38 DB variables may not be comprehensive enough to cover all important aspects (example: 2009 DB indicators do not cover ease of getting electricity)
    - parameter bias induced by methods to avoid overfitting: for RF this bias makes measured importance of top variables more likely to be diluted than exaggerated
  - with these caveats, results do allow for a causal interpretation, but magnitude caution is warranted
- Small-sample caveat:
  - absence of evidence of an effect does not imply evidence of absence of an effect; limited variation in some variables could explain their apparent lack of importance

### Interaction analysis and implications for reform sequencing
- Interaction strength measured via H-statistics (Friedman and Popescu, 2008):
  - H-statistic for indicator j measures variation in the prediction function explained by interactions between j and all other indicators (excluding j), as a share of total variation explained by j
  - Figure 5 reports H-statistics and finds interactions are very weak:
    - for any DB indicator, less than 10 percent of explanatory power is obtained from interactions
  - Policy implication:
    - low degree of complementarity means sequencing of reforms may not be a first-order issue, since payoff from reforms in one area does not depend heavily on the status of reforms in other areas
  - Note on RF interactions:
    - RF considers an interaction only if at least one of the two variables matters in itself (without the other); in most cases this is a mild and plausible restriction

### RF hyper-parameter selection and sample statistics
- Repeated cross-validation procedure used to find optimal m_try:
  1. sample divided into 10 random subsamples
  2. guess a value for m_try
  3. for each subsample estimate RF on remaining observations and measure out-of-sample performance as average RMSE over the 10 subsamples
  4. repeat for m_try between 2 and 20
  5. repeat steps 1–4 ten times
  6. use the m_try that delivers best average RMSE over the 100 held-out samples to estimate RF on full sample
- Table 3: Hyper-parameters (as reported)
  - Outcome variable columns: DRI, EIU, GAD, GCS, PRS, WMO, ̃y
  - #observations: 137, 144, 158, 130, 133, 178, 178
  - m_try: 510546127

### Conclusion summary (policy-relevant points)
- Perceptions of a country’s business environment can be tied to specific DB indicators under policymakers’ control
- Generally, first-order importance indicators:
  - efficient insolvency procedures (recovery rates)
  - speedy border formalities (time to export and import)
  - low startup costs (cost/time/documents to start a business)
- Exact order of priorities depends on country-specific circumstances; identifying country-specific priorities is enabled by non-linear modeling techniques like Random Forest
- Using Random Forest instead of linear techniques reduces sensitivity of explanations to how the overall business environment is measured

*Source: wpiea2020278-print-pdf - Section 2*

---


_Source: https://www.imf.org/-/media/files/publications/wp/2020/english/wpiea2020278-print-pdf.pdf_
