## wp18182

## Source details

**Canonical URL:** [wp18182](https://www.imf.org/-/media/files/publications/wp/2018/wp18182.pdf)

## Other formats

- [Markdown version](/-/media/files/publications/wp/2018/wp18182.pdf.md)
- [Structured JSON version](/-/media/files/publications/wp/2018/wp18182.pdf.json)

---

### II. DATA AND DESCRIPTIVE STATISTICS
- Sample: panel of EU members and candidates from 1970 to 2017.
- Treatment variable: adoption of the 3 percent fiscal deficit ceiling (first introduced with the Maastricht Treaty in 1992).
- Outcome variable: nominal budget balance as a percent of GDP.
- First-stage covariates used to model probability of adoption:
  - lags of the government balance and government debt
  - inflation
  - GDP per capita
  - GDP per capita growth
  - indicator variable for federal states
  - government fragmentation
  - trade with EU-11 countries
- Measurement note: structural budget balances are difficult and subject to significant measurement error; structural balance is prone to ex-post revisions resulting from the measurement bias of potential GDP.
- Descriptive comparisons:
  - Mean nominal budget balance (percent of GDP):
    - Observations with the 3 percent deficit rule (rulers): -2.7 percent of GDP.
    - Observations without the rule (non-rulers): -3 percent of GDP.
  - Density differences: distributions differ beyond means; rulers’ distribution more concentrated between -3 and 3 percent of GDP with thinner tails.
- Identification caveat: adoption of the 3 percent deficit ceiling coincides with EU accession for a number of countries; without further assumptions these two effects cannot be separately identified in such cases.

### Adoption pattern and sample composition
- Adoption waves:
  - 1992: twelve countries adopted the rule (Belgium, Denmark, France, Germany, Greece, Ireland, Italy, Luxembourg, Netherlands, Portugal, Spain and the UK).
  - 1995: Austria, Finland and Sweden adopted with the fifth EU enlargement.
  - 2004–2007: Eastern European countries (plus Malta and Cyprus) introduced the rule.
  - 2013: Croatia entered the European Union and is the last country in the sample to do so.
- Euro adoption note: Bulgaria, Croatia, the Czech Republic, Denmark, Hungary, Poland, Romania, Sweden and the United Kingdom did not adopt the Euro.
- Never-adopted in sample: Albania, Iceland, Macedonia, Montenegro, Serbia, Switzerland and Turkey.

### III. ESTIMATION UNDER SELECTION ON OBSERVABLES: INVERSE PROBABILITY WEIGHTING
- Framework: treatment effects with potential outcomes Y_ct(FR_ct) and ATET as parameter of interest:
  - ATET = E[ Y_ct(1) − Y_ct(0) | FR_ct = 1 ] = E[ Y_ct(1) | FR_ct = 1 ] − E[ Y_ct(0) | FR_ct = 1 ].
- Identification assumption required (selection on observables):
  - Y_ct(FR_ct) ⟂ FR_ct | X_ct (potential outcomes independent of treatment conditional on observed covariates X_ct).
- Practical implication: rulers and non-rulers differ on many dimensions; methods that adjust for selection are necessary.

### Overlap, IPW strategy and first-stage logit
- Overlap/uncounfoundness assumption: ℙ(퐹푅=1 | 푋푐푡 )<1 for all X_ct (positive probability of adoption for any covariate value).
- Two-step IPW procedure (Hirano, Imbens, Ridder 2003):
  - Step 1: model probability of adoption with a logit; adoption coded as 1 only in the year of adoption and set to missing in subsequent years when the rule is present.
  - Step 2: estimate causal effect conditioning on estimated propensity scores.
- Weights (Equation (3)):
  - 푤̂푐푡 = 퐹푅푐푡 −(1−퐹푅푐푡) 𝑃̂푐푡 1−𝑃̂푐푡, where 𝑃̂푐푡 are propensity scores and 퐹푅푐푡 is the FR dummy.
- First-stage predictors included:
  - 3 lags of the deficit and 1 lag of the overall debt-to-GDP ratio
  - age dependency ratio
  - log of GDP per capita
  - past inflation
  - trade intensity with EU-11
  - euro membership dummy
  - degree of political stability
  - federal country dummy
  - Non-linear terms (squared/cubic) and interactions tested; results qualitatively unaffected.
- Baseline first-stage model: IPW4 achieves highest predictive power (pseudo R-squared and AUROC) and is used as baseline.

### Balance diagnostics and overlap checks
- Covariate balancing:
  - Reweighted series show standardized differences closer to zero.
  - Formal test (Imai and Ratkovic 2014): for all four models the null that covariates in the reweighted sample are mean-balanced cannot be rejected.
  - Variance ratios: reweighting does not always improve variance ratios; may reflect small sample size.
- Overlap diagnostics:
  - Propensity score densities show no excessive mass near 0 or 1 and significant overlap across four first-stage models (Figure 2).
  - Estimated inverse probability weights: some control observations with large weights (Iceland in 2016, Slovakia in 2002 and Luxembourg in 1989) but no single extreme outlier that could singlehandedly drive estimates (Figure 3).
- Kernel bandwidth: baseline bandwidth ℎ = 0.75 chosen using Stata’s optimal bandwidth selection for the treatment group; results robust to a wide range of bandwidths.
- Alternative weighting: entropy weighting (Hainmueller 2012) implemented; Section V shows robustness to alternative estimators.

### Effects on the distribution of deficits and bunching estimation
- Kernel estimator for potential outcome densities (Equation (5)):
  - 푓̂(푦 | 퐹푅) = 1 푁푇 1 ℎ ∑∑푤̂푐푡 퐾(푌푐푡(퐹푅푐푡) − 푦 ℎ ), Epanechnikov kernel, ℎ = 0.75, weights 푤̂푐푡.
- Distributional treatment effects (baseline IPW4):
  - Average (mean) effect on government balances: 1.5 percent of GDP (average improvement).
  - 95th percentile effect: improved balances by 3.1 percent of GDP.
  - 5th percentile effect: reduced balances by 1.2 percent of GDP.
  - Interpretation: evidence of a “magnet” effect—drawing high-deficit observations toward the ceiling while reducing balances for those with low deficits (strong performers).
  - Counterfactual distribution shows excess density around 5 percent of GDP at the top end relative to treatment and no-ruler distributions.
- Robustness: counterfactual kernel density estimates using alternative first-stage specifications display very similar shapes.

### Bunching estimation approach
- Definition of bunching B:
  - B = ∫[f(y|FR=1) − f(y|FR=0)] dy over y− to y+, where y− and y+ are bounds where treated and counterfactual densities cross.
- Empirical estimator (Equation (6)):
  - α̂k, B̂k = arg min αk, Bk ∑ ŵct · c,t [1(Yct ∈ Rk) − αk − Bk FRct]^2, for k ∈ {l, b, u}, with Rk denoting three bunching ranges.
- Estimation details:
  - Total bunching B and excess mass in top and bottom ranges estimated using GMM to account for first-stage estimation of weights and obtain correct standard errors.
  - Bounds taken from estimated intersections of treated and counterfactual densities.

### Main bunching results (preferred specification IPW4, Table 4)
- Bunching area (B̂):
  - IPW1: 0.193**
  - IPW2: 0.192**
  - IPW3: 0.195**
  - IPW4: 0.189**
  - Standard errors (IPW4): (0.080)
- Top excess mass (IPW4): -0.025 (0.024)
- Bottom excess mass (IPW4): -0.163* (0.084)
- Bunching range (IPW4):
  - Lower bound: -4.32
  - Upper bound: 3.30
- Interpretation:
  - Almost 20 percent of the sample is attracted towards the bunching area between -4.3 and 3.3 percent of GDP (B̂ ≈ 0.189).
  - From this pool of bunchers, 86 percent come from the bottom of the distribution below -4.3, while 14 percent come from the top above 3.3 percent of GDP.
  - Bunching estimate statistically significant at the 5 percent level.

### Average treatment effects (Table A3)
- ATET (in percent of GDP) and standard errors:
  - DiD1: 1.028 (0.613)
  - DiD2: 1.774*** (0.623)
  - IPW1: 1.451** (0.687)
  - IPW2: 1.332** (0.639)
  - IPW3: 1.549** (0.696)
  - IPW4: 1.449** (0.637)
- Observations used:
  - DiD columns: 1,291 observations
  - IPW1 through IPW4: 1,172; 1,164; 1,160; 1,152 observations respectively
- Note: DiD1 includes year fixed effects; DiD2 adds country fixed effects. IPW columns report ATET using inverse probability weights.

### Robustness checks (bunching and weighting alternatives)
- Alternative first-stage weighting schemes produce similar results:
  - Logit: Bunching area 0.189** ({0.018})
  - Entropy 1st moment: 0.173* ({0.051})
  - Entropy 1-2nd moments: 0.167 (bootstrap failed to converge >60% of iterations; no p-value reported)
  - Local logit: 0.169** ({0.043})
  - Bottom and top excess mass estimates broadly consistent with baseline.
- Inference approach for robustness checks: pairs cluster bootstrap-t with 1,000 replications; p-values reported in curly brackets.
- Bootstrap note: algorithm fails to converge almost 60 percent of the time for entropy-based estimates matching first and second moments; p-values omitted for that column.
- Clustering by adoption wave (9 waves) with pairs cluster bootstrap-t:
  - IPW4 bunching area: 0.189** with p-value {0.028}
  - Bottom excess mass remains negative and generally significant.

### Subperiod bunching results (Table A4)
- Panel A: Excluding Global Financial Crisis (2008 to 2012)
  - Bunching area:
    - IPW 1: 0.261***
    - IPW 2: 0.268***
    - IPW 3: 0.270***
    - IPW 4: 0.264***
    - Standard errors: (0.076), (0.077), (0.082), (0.092)
  - Top (IPW4): -0.052* (0.030)
  - Bottom (IPW4): -0.212** (0.094)
  - Bunching range (IPW4):
    - Lower bound: -3.28
    - Upper bound: 4.35
  - Observations (IPW4): 708
- Panel B: Excluding 1970s and 1980s (1970 to 1989)
  - Bunching area:
    - IPW 1: 0.148 (0.101)
    - IPW 2: 0.169* (0.098)
    - IPW 3: 0.124 (0.084)
    - IPW 4: 0.151* (0.080)
  - Top (IPW4): -0.003 (0.003)
  - Bottom (IPW4): -0.148* (0.079)
  - Bunching range (IPW4):
    - Lower bound: -2.55
    - Upper bound: 7.48
  - Observations (IPW4): 874

### Country-specific impacts and heterogeneity
- Rank invariance used to recover individualized effects under assumption that country orderings are preserved in counterfactual.
  - Example: Italy in 1996 at 10th percentile with deficit of 6.75 percent of GDP; counterfactual at same percentile would be 7.8 percent of GDP.
- Average impact by country (summary from Figure 7):
  - All 28 adopting countries saw improvement of their average annual deficit across the sample period relative to the counterfactual.
  - Across the sample, government balances improved by 1.5 percent of GDP on average.
  - Range: improvement slightly over 3 percent of GDP for Greece; less than 0.25 percent of GDP for Luxembourg.
- Heterogeneity by debt level (Figure 9):
  - Impact generally larger for countries with higher debt, reaching around 3 percent of GDP in annual deficit adjustment for high-debt countries.

### Interpretation, policy implications and open questions
- Distributional effects:
  - The 3 percent deficit rule reshapes the entire distribution of fiscal balances: constrains large deficits and reduces very large surpluses, producing concentration around the threshold (magnet/bunching effect).
- Policy implication:
  - Policymakers should carefully consider costs and benefits of introducing numerical targets because impacts can be complex and depend on countries’ current fiscal positions.
- Methodological contribution:
  - Combines treatment effects methods with bunching estimation to detect bunching behavior in macroeconomic data — a first application of these combined methods in this context.
- Open questions for future research:
  - Channels by which fiscal rules affect policy are not investigated here; possible channels include reference-point effects on voters (reelection probabilities) and differential financial market responses around the threshold.

*Source: IMF Working Paper wp18182 (excerpt).*

### Section V presents robustness exercises. Section VI shows country specific results under

### wp18182 - Section V presents robustness exercises. Section VI shows country specific results under

### II. DATA AND DESCRIPTIVE STATISTICS
- Sample: panel of EU members and candidates from 1970 to 2017.
- Treatment variable: adoption of the 3 percent fiscal deficit ceiling (first introduced with the Maastricht Treaty in 1992).
- Outcome variable: nominal budget balance as a percent of GDP.
- First-stage covariates used to model probability of adoption:
  - lags of the government balance and government debt
  - inflation
  - GDP per capita
  - GDP per capita growth
  - indicator variable for federal states
  - government fragmentation
  - trade with EU-11 countries
- Appendix Table A2 (descriptive statistics, data sources, variable construction) referenced for details.
- Measurement note: structural budget balances are difficult and subject to significant measurement error; structural balance is prone to ex-post revisions resulting from the measurement bias of potential GDP.

### Adoption pattern and sample composition
- Adoption waves:
  - 1992: twelve countries adopted the rule following the Maastricht Treaty (Belgium, Denmark, France, Germany, Greece, Ireland, Italy, Luxembourg, Netherlands, Portugal, Spain and the UK).
  - 1995: Austria, Finland and Sweden adopted the rule with the fifth EU enlargement.
  - 2004–2007: Eastern European countries (plus Malta and Cyprus) introduced the rule.
  - 2013: Croatia entered the European Union and is the last country in the sample to do so.
- Euro adoption note: Bulgaria, Croatia, the Czech Republic, Denmark, Hungary, Poland, Romania, Sweden and the United Kingdom did not adopt the Euro.
- Never-adopted list in the sample: Albania, Iceland, Macedonia, Montenegro, Serbia, Switzerland and Turkey.
- Identification caveat: adoption of the 3 percent deficit ceiling coincides with EU accession for a number of countries; without further assumptions these two effects cannot be separately identified in such cases.

### Descriptive statistics and naive comparisons
- Mean nominal budget balance (percent of GDP):
  - Observations with the 3 percent deficit rule (rulers): -2.7 percent of GDP.
  - Observations without the rule (non-rulers): -3 percent of GDP.
- Distributional observations:
  - Kernel density estimates show distributions differ beyond their means.
  - For observations with the fiscal rule, distribution is more concentrated between -3 and 3 percent of GDP with thinner tails compared to observations without the rule.
- Interpretation caution: simple comparisons of means or distributions between rulers and non-rulers are biased because rulers and non-rulers differ in macroeconomic and institutional characteristics.

### III. ESTIMATION UNDER SELECTION ON OBSERVABLES: INVERSE PROBABILITY WEIGHTING
- Framework: treatment effects framework using potential outcome notation.
- Potential outcome definition for unit (country c, time t) with fiscal rule indicator FR ∈ {0,1}:
  - Y_ct(FR_ct) = { Y_ct(0) if FR_ct = 0; Y_ct(1) if FR_ct = 1 }  (equation (1) in source)
- Fundamental problem of causal inference: for any unit, the potential outcome associated with the alternative treatment cannot be observed (Rubin 1974); need to build a counterfactual for Y_ct(0) when FR_ct = 1.
- Quantity of interest: average treatment effect on the treated (ATET):
  - ATET = E[ Y_ct(1) − Y_ct(0) | FR_ct = 1 ] = E[ Y_ct(1) | FR_ct = 1 ] − E[ Y_ct(0) | FR_ct = 1 ]  (equation (2) in source)
  - The second term is the average budget balance for rule adopters had they not adopted the fiscal rule (the counterfactual).
- Randomized assignment note: if adoption were randomized, ATET could be estimated by comparing means; adoption is not exogenous, so naive comparisons are biased due to omitted variables and self-selection into rule adoption.
- Practical implication: rulers and non-rulers differ on many dimensions (see Table 1), necessitating methods that adjust for selection.

### Identification assumptions (selection on observables)
- Required assumption to estimate causal effects (and ATET) under selection on observables:
  1. Selection on observables: Y_ct(FR_ct) ⟂ FR_ct | X_ct.
     - Interpretation: potential outcomes Y_ct(0), Y_ct(1) are jointly independent of treatment status FR_ct conditional on observed covariates X_ct.
- Methodological references: Wooldridge 2010; Imbens and Rubin 2015.
- Estimation focus: ATET is preferred over ATE because the required assumptions are weaker for ATET (see Wooldridge 2010); to estimate ATE under selection on observables, stronger conditional independence of unobservables is required.

*Source: IMF Working Paper wp18182 (excerpt).*

### 2.      Overlap: 0<푃

### 2.      Overlap: 0<푃

### Inverse probability weighting (IPW) strategy and identification
- Identification relies on the overlap/uncounfoundness assumption: ℙ(퐹푅=1 | 푋푐푡 )<1, i.e., for any value of covariates 푋푐푡 countries have a positive probability of adopting a fiscal rule.
- To address selection into rule adoption the paper uses the inverse probability weighting scheme of Hirano, Imbens, and Ridder 2003 in a two-step procedure:
  - Step 1: model the probability of FR adoption with a logit model that accounts for predictors of adoption.
  - Step 2: estimate the causal effect of FR adoption on the distribution of government balances conditioning on the estimated propensity to adopt the FR.
- Weights used to correct for missing potential outcomes (Equation (3)):
  - 푤̂푐푡 = 퐹푅푐푡 −(1−퐹푅푐푡) 𝑃̂푐푡 1−𝑃̂푐푡
  - where 𝑃̂푐푡 are propensity scores from the logit, and 퐹푅푐푡 is a dummy equal to 1 when the rule is in place.
- Interpretation of weights: non-ruler observations estimated to be more likely to adopt the FR receive larger weights, producing a control group more comparable to rulers.

### First-stage logit model and covariates
- The probability modeled is adoption (not presence): ℙ(퐹푅푐푡 =1 | 퐹푅푐푡−1 =0) = Λ(α + βX푐푡) (Equation (4)). Adoption is coded as one only in the year of FR adoption and set to missing in subsequent years when the rule is present.
- Predictors included (guided by prior literature):
  - Past fiscal behavior: 3 lags of the deficit and 1 lag of the overall debt-to-GDP ratio.
  - Structural demographic feature: age dependency ratio.
  - Macro variables: log of GDP per capita, past inflation, trade intensity with EU-11, euro membership dummy.
  - Institutional/political variables: degree of political stability, federal country dummy.
- Non-linear checks: squared and cubic terms in deficit, debt and growth; interaction effects for positive balances, balances above -3 percent and debt levels above 60 percent. Results are qualitatively unaffected by inclusion of these non-linear terms.
- Model selection: model IPW4 achieves the highest predictive power (pseudo R-squared and AUROC), and is used as baseline.

### Balance diagnostics and overlap checks
- Covariate balancing:
  - Reweighted series (using weights from Equation (3)) show improved covariate balancing: standardized differences between treatment and control means are closer to zero.
  - Formal covariate balancing test implemented using Imai and Ratkovic 2014 (Stata’s tebalance overid). For all four models, the null that covariates in the reweighted sample are mean-balanced cannot be rejected.
  - Variance ratios: reweighting does not always improve variance ratios (in some cases ratios move further from one); this may reflect relatively small sample size.
  - Robustness: alternative entropy weighting (Hainmuller 2012) implemented; Section V shows results robust to alternative estimators.
- Overlap diagnostics:
  - Visual inspection of estimated propensity score densities for rulers and non-rulers (Figure 2) shows no excessive mass near 0 or 1 for any of the four first-stage models, and significant overlap between groups.
  - Estimated inverse probability weights plotted (Figure 3) show some control observations with particularly large weights (Iceland in 2016, Slovakia in 2002 and Luxembourg in 1989), but no control observation is an extreme outlier that could singlehandedly drive distribution estimates.
- Practical modeling choice: bandwidth of 0.75 used in baseline kernel density estimates (chosen using Stata’s optimal bandwidth selection for the treatment group); baseline results robust to a wide range of bandwidths.

### Effects on the distribution of deficits and bunching estimation
- Estimation approach:
  - Kernel estimator for potential outcome densities (Equation (5)):
    - 푓̂(푦 | 퐹푅) = 1 푁푇 1 ℎ ∑∑푤̂푐푡 퐾(푌푐푡(퐹푅푐푡) − 푦 ℎ )
    - Epanechnikov kernel function 퐾, bandwidth ℎ = 0.75, weights 푤̂푐푡 from Equation (3).
  - Focus on PDFs to analyze bunching behavior; weights from IPW4 used for baseline; robustness checks use other first-stage models.
- Key distributional findings:
  - Average (mean) effect on government balances is positive and estimated at 1.5 percent of GDP (middle blue solid line to the right of middle dashed red line in Figure 4).
  - Effects at extreme quantiles are of opposite signs:
    - At the 95th percentile (countries with very high deficits), FR presence improved balances by 3.1 percent of GDP.
    - At the 5th percentile (countries with very low deficits / high balances), FR presence reduced balances by 1.2 percent of GDP.
  - Interpretation: evidence of a “magnet” effect of the 3 percent deficit ceiling on government balances — drawing high-deficit observations toward the ceiling while reducing balances for those already with low deficits.
  - The counterfactual distribution shape differs from the raw data distribution, notably with excess density around 5 percent of GDP at the top end compared to treatment and no-ruler distributions.
- Robustness:
  - Counterfactual kernel density estimates using alternative first-stage specifications display very similar shapes, indicating results are not driven by specific modelling assumptions in the first stage.

*Source: wp18182 - 2.      Overlap: 0<푃 (IMF Working Paper)*

### Appendix Table A3 provides the 퐴푇퐸푇 estimates for the four different IPW models as well as simple difference-in-

### wp18182 - Appendix Table A3 provides the 퐴푇퐸푇 estimates for the four different IPW models as well as simple difference-in-

### Bunching evidence and estimation approach
- Definition of bunching B:
  - B = ∫[f(y|FR=1) − f(y|FR=0)] dy over y− to y+, where f(y|FR) are treated and counterfactual densities and y+ and y− are the bounds where the two densities cross.
- Empirical estimator (Equation (6)):
  - α̂k, B̂k = arg min αk, Bk ∑ ŵct · c,t [1(Yct ∈ Rk) − αk − Bk FRct]^2, for k ∈ {l, b, u}
  - Rk denotes three bunching ranges partitioned by intersections (ŷ−, ŷ+); FRct dummy as defined; weights ŵct from selection model in Equation (4).
- Estimation details:
  - We estimate total bunching B and excess mass in top and bottom ranges using GMM to account for first-stage estimation of weights ŵct and obtain correct standard errors.
  - Bounds of bunching ranges are taken from estimated intersections of treated and counterfactual densities (Figures referenced in text).

### Main bunching results (Table 4, preferred specification IPW4)
- Bunching area (B̂):
  - IPW1: 0.193**
  - IPW2: 0.192**
  - IPW3: 0.195**
  - IPW4: 0.189**
  - Standard errors (IPW4): (0.080)
- Top excess mass:
  - IPW4: -0.025 (0.024)
- Bottom excess mass:
  - IPW4: -0.163* (0.084)
- Bunching range (IPW4):
  - Lower bound: -4.32
  - Upper bound: 3.30
- Interpretation from preferred specification (column (4), IPW4):
  - Almost 20 percent of the sample is attracted towards the bunching area located between -4.3 and 3.3 percent of GDP (B̂ ≈ 0.189).
  - From this pool of bunchers, 86 percent come from the bottom of the distribution below -4.3, while 14 percent come from the top above 3.3 percent of GDP.
  - Statistical significance: the bunching estimate is statistically significant at the 5 percent level.

### Distributional treatment effects and density patterns
- Density crossings and magnet effect:
  - Treated and counterfactual densities cross in two points (IPW4): at -4.3 and 3.3 (percent of GDP), producing a concentration of observations between these intersections.
  - Result: the 3 percent deficit rule acts as a “magnet”, increasing the number of observations around the threshold and reducing both large deficits and large surpluses.
- Effects across distribution:
  - Large positive effects at the bottom of the distribution (largest deficits).
  - Negative effects at the top when government balances exceed 2 percent of GDP (strong performers reduce fiscal effort).
  - Observations just above the -3 percent limit (in compliance) still tighten government balances by around 1 percent of GDP.
  - For observations with deficit exceeding 10 percent of GDP, the FR improves fiscal outcomes by more than 4 percent of GDP.

### Average treatment effects (Table A3)
- ATET (in percent of GDP):
  - DiD1: 1.028 (0.613)
  - DiD2: 1.774*** (0.623)
  - IPW1: 1.451** (0.687)
  - IPW2: 1.332** (0.639)
  - IPW3: 1.549** (0.696)
  - IPW4: 1.449** (0.637)
- Observations used:
  - DiD columns: 1,291 observations
  - IPW1 through IPW4: 1,172; 1,164; 1,160; 1,152 observations respectively
- Note: DiD1 includes year fixed effects; DiD2 adds country fixed effects. IPW columns report average treatment effect on the treated using inverse probability weights derived from the first-stage models.

### Robustness checks
- Alternative first-stage weighting schemes:
  - Entropy weighting (Hainmueller 2012) matching first moments and first and second moments; entropy-based adjusted series show similar shapes to IPW with logit first stage.
  - Local logit (non-parametric propensity score) also yields very similar baseline results.
- Bunching estimates under alternative weighting (Table 5):
  - Logit: Bunching area 0.189** ({0.018})
  - Entropy 1st moment: 0.173* ({0.051})
  - Entropy 1-2nd moments: 0.167 (bootstrap failed to converge >60% of iterations; no p-value reported)
  - Local logit: 0.169** ({0.043})
  - Bottom and top excess mass estimates also generally consistent with baseline pattern.
  - Inference for robustness checks: pairs cluster bootstrap-t with 1,000 replications; p-values reported in curly brackets.
  - Note: bootstrap algorithm fails to converge almost 60 percent of the time for entropy-based estimates matching first and second moments; hence p-values omitted for that column.
- Clustering scheme robustness (Table 6):
  - Clustering by adoption wave (9 waves) and inference via pairs cluster bootstrap-t:
    - Bunching area estimates remain similar in magnitude and statistical significance:
      - IPW4 bunching area: 0.189** with p-value {0.028}
    - Bottom excess mass remains negative and generally significant at conventional levels.
  - Rationale: addresses potential unobserved correlation across countries that adopted or intended to adopt the rule at the same time.

### Country-specific and heterogeneous impacts
- Rank invariance assumption to recover individualized effects:
  - Under rank invariance, ordering of countries in the treatment group is assumed to remain the same in the counterfactual group.
  - Example: Italy in 1996 at 10th percentile with deficit of 6.75 percent of GDP; counterfactual deficit at same percentile would be 7.8 percent of GDP.
- Average impact by country (Figure 7, summary):
  - All 28 countries that adopted the 3 percent deficit rule in the sample saw an improvement of their average annual deficit across the sample period relative to the counterfactual.
  - Across the sample, government balances improved by 1.5 percent of GDP on average.
  - Range: improvement slightly over 3 percent of GDP for Greece; less than 0.25 percent of GDP for Luxembourg.
- Heterogeneity by debt level (Figure 9):
  - Impact generally larger for countries with higher levels of debt, reaching around 3 percent of GDP in annual deficit adjustment for high-debt countries.
  - Suggests fiscal rule facilitated more ambitious fiscal adjustments where debt was higher.

### Interpretation and policy implications
- Effects of fiscal rules extend beyond mean outcomes and reshape the entire distribution of fiscal balances in complex ways.
- The 3 percent deficit rule:
  - Was effective in constraining large deficits, including for countries that do not formally comply with the threshold.
  - Reduced very large surpluses among strong performers, moving them closer to the 3 percent ceiling.
  - Increased concentration of observations around the threshold (magnet/bunching effect).
- Policy recommendation (implication drawn in text):
  - Policymakers should carefully consider the costs and benefits of introducing numerical targets because impacts can be complex and interact with countries’ current fiscal positions in unexpected ways.
- Methodological contribution:
  - Demonstrates combining treatment effects methods with bunching estimation to detect bunching behavior in macroeconomic data—a first application using these combined methods in this context.
- Open questions (for future research noted in text):
  - Channels by which fiscal rules affect policy are not investigated here: possible channels include reference-point effects on voters (reelection probabilities) and differential financial market responses to outcomes around the threshold.

*Source: wp18182 (excerpt provided).*

### 2. Standard errors clustered at the country level in parentheses. * p<10%, **p<5%, *** p<1%.

### Table A4. Bunching Results for Different Subperiods

### Panel A: Excluding Global Financial Crisis (2008 to 2012)
- Bunching area:
  - IPW 1: 0.261***
  - IPW 2: 0.268***
  - IPW 3: 0.270***
  - IPW 4: 0.264***
  - Standard errors (clustered at the country level): (0.076), (0.077), (0.082), (0.092)
- Top:
  - IPW 1: -0.061 (0.042)
  - IPW 2: -0.079 (0.050)
  - IPW 3: -0.060* (0.037)
  - IPW 4: -0.052* (0.030)
- Bottom:
  - IPW 1: -0.200** (0.080)
  - IPW 2: -0.190** (0.078)
  - IPW 3: -0.210** (0.083)
  - IPW 4: -0.212** (0.094)
- Bunching range:
  - Lower bound:
    - IPW 1: -6.20
    - IPW 2: -6.51
    - IPW 3: -6.31
    - IPW 4: -3.28
  - Upper bound:
    - IPW 1: 3.41
    - IPW 2: 3.51
    - IPW 3: 3.61
    - IPW 4: 4.35
- Observations:
  - IPW 1: 724
  - IPW 2: 719
  - IPW 3: 713
  - IPW 4: 708

### Panel B: Excluding 1970s and 1980s (1970 to 1989)
- Bunching area:
  - IPW 1: 0.148 (0.101)
  - IPW 2: 0.169* (0.098)
  - IPW 3: 0.124 (0.084)
  - IPW 4: 0.151* (0.080)
- Top:
  - IPW 1: -0.004 (0.004)
  - IPW 2: -0.004 (0.004)
  - IPW 3: -0.004 (0.004)
  - IPW 4: -0.003 (0.003)
- Bottom:
  - IPW 1: -0.144 (0.101)
  - IPW 2: -0.165* (0.098)
  - IPW 3: -0.120 (0.084)
  - IPW 4: -0.148* (0.079)
- Bunching range:
  - Lower bound:
    - IPW 1: -4.53
    - IPW 2: -4.43
    - IPW 3: -2.55
    - IPW 4: -2.55
  - Upper bound:
    - IPW 1: 7.27
    - IPW 2: 7.27
    - IPW 3: 8.00
    - IPW 4: 7.48
- Observations:
  - IPW 1: 878
  - IPW 2: 875
  - IPW 3: 877
  - IPW 4: 874

### Notes on Estimation and Significance
- The table reports the coefficients from Equation (6) and the bounds of the estimated bunching area.
- In Panel A, the estimation sample excludes the Global Financial Crisis from 2008 to 2012.
- In Panel B, the sample excludes years from 1970 to 1989.
- Standard errors clustered at the country level in parentheses.
- Significance markers: * p<10%, ** p<5%, *** p<1%.

### Figure A1. Vertical Difference between Treatment and Counterfactual Densities
- Plots vertical difference between the kernel densities for the treated group with FR and the counterfactual group using estimates from Equation (5).
- Probability weights calculated according to first stage models IPW1-IPW4 from Table 2.
- Bandwidth for the Epanechnikov kernel: 0.75.
- Vertical line shows the FR limit at -3 percent of GDP.

### Figure A2. Counterfactual Time Paths for General Government Balance
- Plots actual and counterfactual government balances assuming rank invariance.
- Counterfactual time series obtained using first stage model IPW4 from Table 2.

*Source: wp18182 - 2. Standard errors clustered at the country level in parentheses. * p<10%, **p<5%, *** p<1%.*

---


_Source: https://www.imf.org/-/media/files/publications/wp/2018/wp18182.pdf_
