## wpiea2024221-print-pdf

## Source details

**Canonical URL:** [wpiea2024221-print-pdf](https://www.imf.org/-/media/files/publications/wp/2024/english/wpiea2024221-print-pdf.pdf)

## Other formats

- [Markdown version](/-/media/files/publications/wp/2024/english/wpiea2024221-print-pdf.pdf.md)
- [Structured JSON version](/-/media/files/publications/wp/2024/english/wpiea2024221-print-pdf.pdf.json)

---

### I. Introduction — scope and focus
- Study covers IMF-supported programs with low-income countries (LICs) over the period 2009-2022.
- Compares programs in fragile and conflict-affected states (FCS) versus non-FCS LICs.
- Focus: "quantitative tailoring" — whether macroeconomic targets (growth, inflation, fiscal consolidation, external adjustment) are tailored to fragility/context by being more realistic and less optimistic.
- Key definitions and usage:
  - LICs: countries eligible for financing under the IMF’s Poverty Reduction and Growth Trust (PRGT).
  - FCS designation follows IMF practice; about one fifth of IMF members are FCS and almost half of LICs are FCS.
  - "Program objectives", "projections", and "targets" refer to three-year quantitative forecasts taken from a program’s macro framework at program approval.
  - Optimism = difference between projections (targets) and outcomes.
  - Correlation = degree of association between targets and outcomes.
- The paper avoids causal claims due to pervasive endogeneity and lack of credible instruments.

### Main findings (overview)
- Limited Quantitative Tailoring:
  - "Program targets do not appear to differ significantly between FCS and non-FCS LICs, be it in absolute terms or relative to initial conditions."
- Considerable Optimism:
  - "Targets tend to be missed in all dimensions other than inflation."
  - Footnote: "Among on-track programs, targets are missed in all dimensions other than inflation and reserves."
- Weak Correlations:
  - "For variables other than growth and inflation, we cannot reject the null hypothesis that targets and outcomes are statistically independent."
  - "Country and program-independent targets equal to the mean or median outcomes of all other programs would have outperformed program projections as predictors of actual outcomes in dimensions other than inflation."
- No causal attribution is made for these findings.

### II. Methodology, sample, and variables
- Sample inclusion criteria:
  - LIC programs approved during fiscal years 2009 to 2022 with planned duration ≥ 1.5 years.
  - Ambitiousness sample: 84 programs across 43 countries (75 ECFs and 9 Standby Credit Facilities).
  - Optimism/correlation sample (three-year horizon, outcomes available): maximally 62 programs across 37 countries.
- Time-on-track definition:
  - Program considered “on track” if it had a successful review 2.5 years or more after start.
  - Of 84 programs, outcomes available for 62. Of these 54 had planned durations ≥ 2.5 years: 29 FCS (12 off track = 41 percent off track), 25 non-FCS (4 off track = 16 percent off track). In this sample, FCS were 2.6 times more likely to go off track than non-FCS.
- Variables and measurement (sign conventions preserved):
  - Real GDP Growth: Change — %Δ per annum — initial = geometric avg T-2 to T; target/outcome = geometric avg T+1 to T+3.
  - CPI Inflation: Change — %Δ per annum — same averaging convention as growth.
  - Revenue XOXG: Flow — % of non-oil GDP — initial = flow during T; target/outcome = Δ between T and T+3 (positive denotes increase).
  - Primary Current Expenditure (PCE): Flow — % of non-oil GDP — initial = flow during T; target/outcome = Δ between T and T+3 (positive corresponds to a decrease).
  - Wage Bill: Flow — % of non-oil GDP — same conventions as PCE.
  - Primary Balance (PB) and PB XOXG: Flow — % of GDP — initial = flow during T; target/outcome = Δ between T and T+3 (positive = improvement).
  - Debt: Stock — % of GDP — stock at end of T; target/outcome = Δ between T and T+3 (positive = decrease).
  - Current Account Balance (CAB): Flow — % of GDP — initial = flow during T; target/outcome = Δ between T and T+3.
  - Reserves: Stock — months of imports — stock at end of T; target/outcome = Δ between T and T+3.
- Data sources and handling:
  - Main sources: IMF Financial Data Query Tool, World Economic Outlook (WEO), Country Staff Reports.
  - Initial conditions and objectives: first WEO vintage after program approval; outcomes: WEO October 2023 vintage (supplemented from Staff Reports where missing).
  - GDP re-benchmarking handled via un-re-benchmarking procedure (Appendix G) using geometric means of pre-sample year ratios.
  - Outcomes of programs affected by civil war (CAF2012, YEM2014, SDN2021) removed as outliers.
  - Adjustments made for HIPC debt relief; other debt relief not adjusted due to data limitations.
  - Some candidate variables (social spending, domestic vs foreign-currency debt) dropped due to insufficient comparable data.

### III. Quantitative tailoring — absolute ambitiousness (key statistics)
- Sample sizes: #Obs reported as (F, NF) for FCS (F) and non-FCS (NF).
- Growth (45, 39):
  - Initial Condition: F = 3.4 (3.3); NF = 4.1 (2.3).
  - Target: F = 5.4 (2.5); NF = 5.7 (1.9).
  - Interpretation: average growth targets for FCS (5.4 percent per annum, s.d. 2.5 p.p.) are statistically indistinguishable at the 10 percent level from non-FCS targets (5.7 percent, s.d. 1.9).
- Inflation (44, 39):
  - Initial Condition: F = 7.0 (5.8); NF = 5.6 (3.9).
  - Target: F = 5.5 (3.4); NF = 5.0 (2.6).
- Revenue XOXG (45, 39):
  - Initial Condition: F = 14 (5.8) < NF = 21 (7.8) (significant as reported).
  - Target: F = 1.5 (1.6); NF = 1.3 (1.5).
- PCE (44, 38) and Target (42, 38):
  - Initial: F = 17 (9.2); NF = 19 (8.4).
  - Target: F = .7 (1.8); NF = 1.0 (1.7).
- Wage Bill (44, 39) and Target (43, 39):
  - Initial: F = 7.4 (2.9); NF = 7.7 (3.3).
  - Target: F = .3 (.7); NF = .4 (.9).
- Primary Balance (PB) (45, 39):
  - Initial: F = -1.0 (3.6) > NF = -2.7 (4.6) (directional significance noted).
  - Target: F = .7 (2.7) < NF = 1.9 (2.5) (significant as reported).
- PB XOXG (45, 39):
  - Initial: F = -12 (7.8) < NF = -7.4 (2.8).
  - Target: F = 2.0 (4.2); NF = 2.5 (3.0).
- Debt (45, 39):
  - Initial stock: F = 59 (45); NF = 52 (21).
  - Target Δ: F = 7.1 (13) > NF = 2.2 (7.5) (significant as reported) — programs in FCS have more ambitious debt-reduction targets in absolute terms.
- CAB (38, 33):
  - Initial: F = -7.2 (11); NF = -11 (6.4).
  - Target: F = -.3 (5.6) < NF = 2.5 (3.5) (significant as reported).
- Reserves (44, 39):
  - Initial: F = 4.3 (3.3); NF = 3.7 (1.7).
  - Target: F = .4 (1.2); NF = .3 (1.1).
- Summary on absolute ambitiousness:
  - Little difference between FCS and non-FCS in average targeted growth and inflation and in programmed adjustments for revenue, PCE, wage bill, PB XOXG, and reserve coverage.
  - Debt targets more ambitious for FCS; PB and CAB targets less ambitious for FCS.
  - Standard deviations for FCS tend to be greater than or equal to non-FCS.
  - Using medians rather than means does not materially change the assessment.

### IV. Quantitative tailoring — relative ambitiousness (regression framework and results)
- Relative ambitiousness measured by regressing targets on initial conditions separately for FCS and non-FCS and comparing slopes and intercepts:
  - tt_i = β0 + β1 s_i + ε_i, tt_i = target, s_i = initial condition at T.
  - If slopes statistically indistinguishable, impose identical slope and compare vertical distance between parallel lines (negative distance → FCS relatively less ambitious).
- Key findings:
  - Initial conditions: Growth, Inflation, PCE, Wage Bill, Debt, CAB, and Reserves similar for FCS and non-FCS on average; PB better in FCS; PB XOXG and Revenues XOXG worse in FCS (e.g., Revenues XOXG: 14 versus 21 percent of non-oil GDP).
  - Dispersion: initial conditions for FCS more dispersed.
  - PB XOXG: slopes FCS = -.41 (s.e. .05), non-FCS = -.27 (s.e. .10); p-value for slope difference = .22 (not significant). Imposing identical slopes, FCS regression line lies 2.2 p.p. of GDP below non-FCS (statistically significant at 1 percent) → FCS targets relatively less ambitious for PB XOXG by 2.2 p.p. of GDP.
  - Accounting for oil and grant revenues (programmed average annual oil revenues: 4.7 versus 3.0 percent of GDP for FCS and non-FCS; grant revenues: 6.4 percent versus 3.1 percent of GDP) and focusing on the regular primary balance makes targeted relative adjustments for FCS and non-FCS statistically indistinguishable.
  - Regression slopes significant and sensible (worse initial conditions → larger targeted adjustments) for most variables; exceptions: Reserves for non-FCS and Revenues XOXG for both groups (slopes ≈ 0).
  - Table 2 “Δ Relative Ambition” highlights:
    - PB XOXG: -2.2*** (.64) (FCS less relatively ambitious)
    - Debt: 3.5** (1.7) (FCS more relatively ambitious)
    - CAB: -1.5* (.79) (FCS less relatively ambitious)
  - Overall conclusion: apart from PB XOXG (and possibly CAB at 10 percent), little evidence of systematic quantitative tailoring after accounting for initial conditions; FCS more ambitious debt targets may reflect greater expected non-HIPC debt relief.

### V. Optimism — magnitudes and patterns (targets − outcomes)
- Definition: Optimism = target − outcome (positive → outcome falls short of target). For Inflation, outcome exceeding target is a “shortfall.”
- Sample construction: paired observations only; programs started less than four years prior excluded for optimism/correlation analysis; for non-growth/inflation variables include only programs that targeted non-negative adjustments unless stated.
- General result: IMF-supported programs are optimistic in nearly all variables except Inflation; this holds for full sample and on-track programs.
- Selected arithmetic-average optimism magnitudes (standard deviations in parentheses; percent-of-target in square brackets where reported):
  - Growth:
    - F (30 paired): Target 6.0 (2.7); Outcome 3.7 (3.5); Optimism 2.4*** [39%].
    - NF (32 paired): Target 5.6 (1.9); Outcome 4.7 (1.7); Optimism .9*** [16%].
    - On-track (39 paired): Target 5.7 (2.0); Outcome 4.5 (2.1); Optimism 1.2*** [21%].
  - Inflation:
    - F (30): Target 5.5 (3.3); Outcome 6.0 (4.1); Optimism .5 (8%) — not significant.
    - NF (32): Target 4.9 (2.3); Outcome 5.3 (5.4); Optimism .5 (10%) — not significant.
  - Revenue XOXG (Target ≥ 0):
    - F (25): Target 1.9 (1.2); Outcome .7 (2.2); Optimism 1.1** (62%).
    - NF (26): Target 1.6 (1.2); Outcome .9 (1.6); Optimism .7** (44%).
    - On-track (32): Target 1.7 (1.2); Outcome 1.0 (1.9); Optimism .6** [38%].
  - PCE (Target ≥ 0):
    - All (41): Target 1.7 (1.4); Outcome -.6 (2.2); Optimism 2.3*** (135%).
    - On-track (26): Target 1.9 (1.4); Outcome -.7 (2.5); Optimism 2.6*** [139%].
  - Wage Bill (Target ≥ 0):
    - All (44): Target .7 (.7); Outcome -.0 (.9); Optimism .7*** [102%].
    - On-track (29): Target .7 (.7); Outcome .1 (.9); Optimism .6*** [86%].
  - PB (Target ≥ 0):
    - All (42): Target 2.6 (2.4); Outcome -.2 (3.2); Optimism 2.8*** (107%).
    - On-track (26): Target 3.1 (2.7); Outcome -.8 (3.2); Optimism 3.9*** [126%].
  - PB XOXG (Target ≥ 0):
    - F (22): Target 3.8 (4.6); Outcome 2.2 (4.5); Optimism 1.6* (43%).
    - NF (27): Target 3.5 (2.7); Outcome .9 (2.9); Optimism 2.6*** (75%).
    - On-track (30): Target 3.3 (2.9); Outcome .7 (4.1); Optimism 2.6*** [79%].
  - Debt (Target ≥ 0):
    - All (36): Target 9.6 (8.4); Outcome -3.2 (14); Optimism 13*** (133%).
    - On-track (21): Target 7.7 (5.3); Outcome - .5 (12); Optimism 8.2*** [106%].
  - CAB (Target ≥ 0):
    - All (31): Target 3.6 (3.7); Outcome .9 (4.8); Optimism 2.7*** (75%).
    - On-track (20): Target 4.3 (4.3); Outcome .7 (5.0); Optimism 3.6*** [83%].
  - Reserves (Target ≥ 0):
    - All (43): Target .8 (.9); Outcome .3 (1.4); Optimism .5** (67%).
    - On-track (28): Target .8 (.8); Outcome .5 (1.2); Optimism .32 [41%].
- Robustness checks:
  - Dropping programs approved since 2017 reduces observations by 11 but leaves results essentially unchanged except for revenue and primary balances.
  - Medians vs means do not materially change conclusions; exceptions: PB XOXG for FCS in full sample and Reserves in a subsample where medians cannot reject calibration.
- Differential optimism FCS vs non-FCS:
  - Growth: FCS optimism 2.4 p.p.; non-FCS optimism .9 p.p.; difference significant at 10 percent.
  - For Inflation, Revenue XOXG, PB XOXG: no differential optimism detected.
  - Many variables have small subsamples for FCS vs non-FCS comparisons, limiting power.

### VI. Correlation between targets and outcomes — main evidence
- Marginal accuracy:
  - Regression slopes between target and outcome strictly less than 1 for all variables other than Inflation, indicating optimism on the margin (unit slope tests p-values usually < .01 and always < .05).
  - For CPI Inflation in non-FCS the slope is strictly greater than 1 (still read as optimism in sense of marginal bias).
- Independence for most variables:
  - Except Growth and Inflation, targets and outcomes appear statistically independent, even for on-track programs.
  - Non-parametric rank correlations (Spearman’s 휌 and Kendall’s τ) generally fail to reject independence for variables other than Growth and Inflation, "even at 10 percent significance."
  - Pearson correlations and regression coefficients corroborate rank-correlation findings with two exceptions in the full sample.
  - Possible non-linearity for Growth in FCS suggested by simultaneous rejection patterns.
- Forecast performance (Theil’s U2 and R2):
  - Theil’s U2 using median is significantly below 1 only for Inflation.
  - In full sample U2 close to 1 for PB XOXG in FCS; otherwise U2 > 1; most U2 values between 1.2 and 1.4.
  - Implication: median-based naïve forecasts often outperform program targets in dimensions other than Inflation.
  - R2: except for Inflation, targets generate negative R2s (targets deviate more from outcomes than the uniform average), consistent with U2 > 1.
- Subsample and robustness:
  - On-track subsample: naïve median tends to outperform targets except for Inflation; for on-track full sample Theil’s U2 median significantly below 1 only for Inflation.
  - Dropping COVID-affected programs does not change main conclusions: null of unit slope rejected; independence cannot be rejected for most variables except Growth and Inflation in some cases.
  - Handicap (time-lagged naïve forecasts) and including negative adjustment targets leave main conclusions unchanged: means/medians outperform targets for all dimensions other than Inflation (with select exceptions noted in appendices).

### VII. Role of growth optimism in explaining optimism elsewhere (Appendix F)
- Conceptual identity: a growth shortfall affects ratios X/Y depending on elasticity η_{X,Y}; growth shortfall helps explain optimism about ratio r iff η_{X,Y} > 1.
- Empirical correlations (full sample, Appendix F Table 1):
  - Pearson correlations between forecast errors (Growth vs other variables): Revenue XOXG .38**; PB XOXG .62***; Debt -.08; Inflation .15.
  - R2 between Growth forecast errors and other-variable forecast errors: Revenue XOXG .14; PB XOXG .38; Debt .01; Inflation .02.
- Interpretation:
  - Growth optimism helps explain optimism about Debt reduction (consistent with snowball effects).
  - To a lesser extent, growth optimism helps explain optimism about Inflation and PCE.
  - Growth optimism does not explain optimism for many other variables (CAB, Reserves largely not significant).
  - Empirical point estimates suggest revenue elasticity close to 1 in the sample (small and not always significant correlations).

### VIII. Repeat vs initial programs (Appendix E)
- Definition: Repeat programs (R) have a preceding program in the sample; non-repeat (NR) otherwise.
- Sample sizes:
  - Ambitiousness: 84 programs, 43 countries (43 NR, 41 R).
  - Optimism/correlation: maximally 62 programs, 37 countries (37 NR, 25 R).
- Key comparisons:
  - Little material difference between NR and R programs in ambitiousness, optimism, or correlation based on Tables 1–5.
  - Selected stats (examples):
    - Growth target: NR 6.1 (2.5); R 5.0 (1.8).
    - Growth optimism: NR paired 36: optimism 1.7***; R paired 26: optimism 1.5***.
    - Correlation measures vary but do not indicate systematic learning across repeat programs; Appendix E concludes little material support for learning across repeat programs.

### IX. Robustness, caveats, and methodological considerations
- Independence and inference:
  - t-tests invoke CLT; contemporaneous programs may be correlated (shared world forecasts). Generalized CLT for near-independence invoked; non-parametric Wilcoxon and Mann-Whitney tests yield similar conclusions.
- Non-linearity risks:
  - Linear regression may produce Type I/II errors if true relationships are non-linear (convexity risks flagged, e.g., Revenues XOXG).
  - Tentative conclusion of tailoring for Wage Bill may be a Type I error.
- GDP re-benchmarking:
  - Un-re-benchmarking introduces noise; Appendix G documents procedure and empirical distribution: 26 programs ratio = 1; 40 programs ratio within [0.95,1.05]; 18 programs outside [0.95,1.05].
- Repeat customers and dependence:
  - Many countries appear multiple times (e.g., Sierra Leone appears four times in sample of 84); spacing and staff turnover reduce but do not eliminate correlated errors.
- Burden of proof:
  - Distinction emphasized between rejecting a null and accepting an alternative. Findings on average and marginal optimism are strong; tailoring and independence findings are weaker (inability to reject nulls).

### X. Policy implications and concluding observations
- Findings support IMF’s Strategy for Fragile and Conflict-Affected States (2022) emphasis on realism:
  - Before 2022 FCS Strategy, limited differentiation in program ambition between FCS and non-FCS indicates uniformity in program design.
  - Macroeconomic projections show optimism in several variables; Inflation projections are generally accurate.
  - Independence of outcomes and targets for most variables suggests potential gains from country-independent targets (median-based) to improve forecast accuracy—except for Inflation.
- Suggested direction:
  - Systematically extend the Strategy’s focus on realism and calibration across all LIC programs, including consideration of country-independent benchmarking for variables other than Inflation.

*Source: IMF Working Paper — IMF-Supported Programs in Low-Income Countries: Fragile versus Non-Fragile States (sections: Introduction; Main Findings; Literature Review; Research Questions and Data overview; 3. Correlation; Appendix B–H excerpts).*

### References .............................................................................................................

### IMF-Supported Programs in Low-Income Countries: Fragile versus Non-Fragile States

### I. Introduction — scope and focus
- Study covers IMF-supported programs with low-income countries (LICs) over the period 2009-2022.
- Compares programs in fragile and conflict-affected states (FCS) versus non-FCS LICs.
- Focus: "quantitative tailoring" — whether macroeconomic targets (growth, inflation, fiscal consolidation, external adjustment) are tailored to fragility/context by being more realistic and less optimistic.
- Definitions and usage:
  - LICs are countries that qualify for financing under the IMF’s Poverty Reduction and Growth Trust (PRGT).
  - FCS designation follows IMF practice; about one fifth of IMF members are FCS and almost half of LICs are FCS.
  - Terms "program objectives", "projections", and "targets" are used interchangeably to refer to three-year quantitative forecasts taken from a program’s macro framework at program approval.
  - Optimism = difference between projections (targets) and outcomes.
  - Correlation = degree of association between targets and outcomes.
- The paper avoids causal claims due to pervasive endogeneity and lack of credible instruments.

### Main findings (as reported)
- Limited Quantitative Tailoring:
  - "Program targets do not appear to differ significantly between FCS and non-FCS LICs, be it in absolute terms or relative to initial conditions."
- Considerable Optimism:
  - "Targets tend to be missed in all dimensions other than inflation."
  - Footnote: "Among on-track programs, targets are missed in all dimensions other than inflation and reserves."
- Weak Correlations:
  - "For variables other than growth and inflation, we cannot reject the null hypothesis that targets and outcomes are statistically independent."
  - "Country and program-independent targets equal to the mean or median outcomes of all other programs would have outperformed program projections as predictors of actual outcomes in dimensions other than inflation."
- The study does not assign causality for these findings.

### Literature review — positioning of this paper
- Builds on earlier IMF and academic work on fragile states and program design:
  - Prior findings noted: (i) borrowing frequency of FCS LICs is similar to non-FCS LICs; (ii) completion rates for FCS lending programs are markedly lower—30 versus 75 percent; (iii) level of conditionality shows little variation between FCS and non-FCS.
- Relation to prior studies:
  - Kunduz (2018): finds IMF-supported programs increase official development assistance in FCS relative to non-FCS, more volatility and marginally lower growth in FCS, and growth in FCS tends to increase by approximately one percentage point after program approval.
  - Kim et al. (2021): documented optimism in program design; current work is broadly consistent and extends the temporal scope and focuses on PRGT programs split by FCS status.
  - Celasun et al. (2021): longer-term WEO growth forecasts (two to five years ahead) tend to be over-optimistic and often less accurate than simple forecasts; this paper replicates that finding for LIC programs and extends it beyond growth.
- Notes the broader literature on fragility and economic performance, which splits into:
  - Studies of fragility’s detrimental effects on economic performance (e.g., Collier, Rodrik, Cerra and Saxena).
  - Studies of how economic performance affects fragility and conflict, with mixed findings and methodological challenges (e.g., use of commodity prices as plausibly exogenous instruments).

### II. Research questions, methodology, and data (overview)
- Research questions listed:
  1. Quantitative Tailoring: To what extent do macroeconomic objectives of programs with FCS differ from those with non-FCS?
  2. Optimism: How optimistic are these objectives? Is there a difference between FCS and non-FCS, and between programs that remain on track and those that do not?
- Macro variables of interest include growth, inflation, fiscal consolidation, and external adjustment variables.
- Dataset covers program macro frameworks at time of program approval and compares original three-year targets to outcomes; targets may be adjusted during programs but the analysis focuses on original values.
- Additional methodological details, appendices and figures and tables are part of the chapter (listed in the content inventory), including appendices on initial conditions versus targets, medians vs means, excluding COVID-affected programs, including negative adjustments, initial vs repeat programs, growth optimism, GDP re-benchmarking, and country classifications; and figures/tables on program approval dates, durations, comparisons of initial conditions vs targeted increases, targets vs outcomes, repeat customers timing, and various tabulated statistics (variables, ambitiousness measures, optimism, correlations).

*Source: IMF Working Paper — IMF-Supported Programs in Low-Income Countries: Fragile versus Non-Fragile States (sections: Introduction; Main Findings; Literature Review; Research Questions and Data overview).*

### 3. Correlation: How correlated are program targets and outcomes?

### 3. Correlation: How correlated are program targets and outcomes?

### Methodology and Program Inclusion Criteria
- Sample inclusion criteria:
  - Programs had to be with a LIC (eligible for PRGT financing), approved during fiscal years 2009 to 2022, and have a planned duration of 1.5 years or more.
  - These criteria yielded 84 programs across 43 countries for the ambitiousness assessment: 75 ECFs and 9 Standby Credit Facilities.
  - Using a three-year horizon (standard program duration for LICs), outcome variables initiated after mid-2019 were not available at data collection, reducing observations for optimism and correlation to maximally 62 programs across 37 countries.
- Time-on-track definition:
  - A program was considered “on track” if it had a successful review 2.5 years or more after its start (programs are typically on a six-month review cycle).
  - Of 84 programs, outcomes are available for 62. Of these, 8 had planned durations <2.5 years. Of the remaining 54, 29 programs were with FCS, 12 of which went off track (41 percent); 25 programs were with non-FCS, 4 of which went off track (16 percent). Hence, in this sample, FCS were 2.6 times more likely to go off track than non-FCS.

### Variables and Measurement
- Focus areas: growth, inflation, government finances, and balance of payments.
- Variables included:
  - Real GDP Growth (“Real GDP Growth”) — Change — %Δ per annum — initial, target, outcome measured as geometric averages (avg growth T-2 to T; target and outcome avg growth T+1 to T+3).
  - CPI Inflation (“CPI Inflation”) — Change — %Δ per annum — measured as geometric averages over the same periods as growth.
  - Revenues XOXG (“Revenue XOXG”) — Flow — % of non-oil GDP — initial = flow during T; target/outcome = Δ between T and T+3. Positive target/outcome denotes adjustment in the ‘right’ direction (increase).
  - Primary Current Expenditure (“PCE”) — Flow — % of non-oil GDP — initial = flow during T; target/outcome = Δ between T and T+3. For PCE, a positive number corresponds to a decrease.
  - Wage Bill — Flow — % of non-oil GDP — same conventions as PCE.
  - Primary Balance (“PB”) and PB XOXG — Flow — % of GDP — initial = flow during T; target/outcome = Δ between T and T+3 (positive = improvement).
  - Debt — Stock — % of GDP — stock at end of T; target/outcome = Δ between T and T+3 (positive = decrease).
  - Current Account Balance (“CAB”) — Flow — % of GDP — initial = flow during T; target/outcome = Δ between T and T+3.
  - Reserves — Stock — months of imports — stock at end of T; target/outcome = Δ between T and T+3.
- For change variables (growth, inflation): initial condition equals average during T-2 to T (as estimated at program adoption); target and outcome correspond to forecasted and realized averages for T+1 to T+3.
- For flows/stocks: targets and outcomes are projected and realized differences between T+3 and T; signing convention described above.

### Data Sources and Data Issues
- Main sources: IMF Financial Data Query Tool, World Economic Outlook (WEO) database, and Country Staff Reports.
- For initial conditions and objectives: first WEO vintage after program approval (reflecting program macro framework).
- For outcomes: WEO October 2023 vintage; missing data supplemented from Staff Reports.
- Notable data handling:
  - GDP re-benchmarking/rebasing: assumed re-benchmarking changed GDP levels but left ratios between old and new series constant across years; assumption verified where possible (Appendix G referenced).
  - Adjusted debt stock and revenue data for HIPC debt relief; other debt relief not adjusted due to lack of comparable cross-country data and smaller magnitude.
  - Outcomes of programs affected by civil war (CAF2012, YEM2014, SDN2021) were removed as non-representative outliers.
  - Some initially considered variables (social spending, domestic vs foreign currency debt) were dropped due to insufficient comparable data.

### Analytical Objectives and Metrics
- Using the defined dataset, the analysis assesses:
  - Quantitative Tailoring: Compare ambitiousness of targets for FCS versus non-FCS and study relationship between targets and initial conditions.
  - Optimism: Estimate average differences between targets and outcomes; examine whether optimism about growth explains optimism in other dimensions.
  - Correlation: Study correlation between targets and outcomes; test for independence between them. Calculate R2 of Targets as predictors of Outcomes and Theil’s U2 (using mean/median outcomes for the “naïve” model).

### Quantitative Tailoring — Absolute Ambitiousness (Key Findings from Table 2)
- Sample sizes reported in Table 2 are shown as #Obs (F, NF) for FCS (F) and non-FCS (NF).
- Growth:
  - Initial Condition (45, 39): F = 3.4 (3.3) ≈ NF = 4.1 (2.3).
  - Target (45, 39): F = 5.4 (2.5) ≈ NF = 5.7 (1.9).
  - Interpretation: Average growth targets for FCS (5.4 percent per annum, s.d. 2.5 p.p.) are statistically indistinguishable at the 10 percent level from non-FCS targets (5.7 percent, s.d. 1.9); numbers based on 45 FCS programs and 39 non-FCS programs.
- Inflation:
  - Initial Condition (44, 39): F = 7.0 (5.8) ≈ NF = 5.6 (3.9).
  - Target (44, 39): F = 5.5 (3.4) ≈ NF = 5.0 (2.6).
- Revenue XOXG:
  - Initial Condition (45, 39): F = 14 (5.8) < NF = 21 (7.8) (significant at *** and * as indicated).
  - Target (45, 39): F = 1.5 (1.6) ≈ NF = 1.3 (1.5).
- Primary Current Expenditure (PCE):
  - Initial Condition (44, 38): F = 17 (9.2) ≈ NF = 19 (8.4).
  - Target (42, 38): F = .7 (1.8) ≈ NF = 1.0 (1.7).
- Wage Bill:
  - Initial Condition (44, 39): F = 7.4 (2.9) ≈ NF = 7.7 (3.3).
  - Target (43, 39): F = .3 (.7) ≈ NF = .4 (.9).
- Primary Balance (PB):
  - Initial Condition (45, 39): F = -1.0 (3.6) > NF = -2.7 (4.6) (significant at ** for one direction, ≈ for the other as shown in table).
  - Target (45, 39): F = .7 (2.7) < NF = 1.9 (2.5) (significant at **).
- PB XOXG:
  - Initial Condition (45, 39): F = -12 (7.8) < NF = -7.4 (2.8) (significant at *** and *** in different comparisons).
  - Target (45, 39): F = 2.0 (4.2) ≈ NF = 2.5 (3.0), with significance markers in the table.
- Debt (stock, % of GDP):
  - Initial Condition (45, 39): F = 59 (45) ≈ NF = 52 (21).
  - Target (45, 39): F = 7.1 (13) > NF = 2.2 (7.5) (significant at ** and *** in the comparisons reported).
  - Interpretation: In absolute terms, programs in FCS have more ambitious debt-reduction targets than those in non-FCS.
- Current Account Balance (CAB):
  - Initial Condition (38, 33): F = -7.2 (11) ≈ NF = -11 (6.4).
  - Target (38, 33): F = -.3 (5.6) < NF = 2.5 (3.5) (significant at **).
- Reserves:
  - Initial Condition (44, 39): F = 4.3 (3.3) ≈ NF = 3.7 (1.7).
  - Target (44, 39): F = .4 (1.2) ≈ NF = .3 (1.1).

### Summary Assessment on Absolute Ambitiousness
- Overall, programs with FCS and non-FCS differ little in terms of absolute ambitiousness:
  - No significant difference in average targeted growth and inflation.
  - No significant difference in programmed adjustments for revenue, primary current expenditure, wage bill, PB XOXG, and reserve coverage.
  - Debt targets are more ambitious for FCS.
  - Primary balance (PB) and current account (CAB) targets are less ambitious for FCS.
- Standard deviations in programs with FCS tend to be greater than or equal to those with non-FCS.
- Using medians rather than means does not materially change this assessment (median tests reported in the source).

*Source: wpiea2024221-print-pdf - 3. Correlation: How correlated are program targets and outcomes? (IMF staff analysis, dataset and tables as provided in the source PDF).*

### Appendix B. The conclusion about absolute ambitiousness remains unchanged: there is little difference

### Appendix B. The conclusion about absolute ambitiousness remains unchanged: there is little difference between FCS and non-FCS.

### Relative Ambition — motivation and approach
- Absolute differences in targets ignore economies’ initial conditions; equal absolute ambitiousness can imply lower relative ambitiousness for economies starting from worse initial conditions.
- Quantitative measure of relative ambitiousness:
  - For each variable, regress targets on initial conditions separately for FCS and non-FCS:
    - tt_i = β0 + β1 s_i + ε_i, where tt_i is the target, s_i the initial condition at time T.
  - Test whether slopes (β1) differ across FCS and non-FCS.
  - If slopes are statistically indistinguishable, rerun joint regression with identical slope but different intercepts → two parallel lines.
  - The vertical distance between parallel lines = measure of relative ambitiousness (negative means FCS targets are relatively less ambitious).
  - If slopes differ, interpreting intercept differences is problematic; slope differences themselves are potential evidence of tailoring.

### Methodology details illustrated
- Figure 3 example:
  - Left panel: separate regression lines for FCS (red dots) and non-FCS (green dots).
  - Right panel: identical slopes imposed, different intercepts allowed; distance between lines = relative ambitiousness.
- Interpretation note: comparing intercepts when slopes differ may not be meaningful.

### Findings — relative ambitiousness (summary of regression evidence)
- Initial conditions: on average Growth, Inflation, PCE, Wage Bill, Debt, CAB, and Reserves are similar for FCS and non-FCS; PB is better in FCS; PB XOXG and Revenues XOXG are worse in FCS.
  - Revenues XOXG: 14 versus 21 percent of non-oil GDP for FCS and non-FCS, respectively.
- Dispersion: initial conditions for FCS appear more dispersed than for non-FCS.
- PB XOXG:
  - Slopes: FCS slope = -.41 (standard error .05), non-FCS slope = -.27 (standard error .10); both statistically and economically significant; with p-value = .22 slopes are not significantly different.
  - With identical slopes imposed, the FCS regression line lies 2.2 p.p. of GDP below non-FCS, significant at the 1 percent confidence level → targeted PB XOXG increases are relatively less ambitious for FCS by 2.2 p.p. of GDP.
  - Accounting for oil and grant revenues (programmed average annual oil revenues: 4.7 versus 3.0 percent of GDP for FCS and non-FCS; grant revenues: 6.4 percent versus 3.1 percent of GDP) and focusing on the regular primary balance makes targeted relative adjustments for FCS and non-FCS statistically indistinguishable (see Table 3).
- Regression slopes between initial conditions and targets:
  - Statistically and economically significant for most variables; signs sensible (worse initial conditions → larger targeted adjustments).
  - Exceptions: Reserves for non-FCS and Revenues XOXG for both FCS and non-FCS (slopes indistinguishable from zero).
  - Slopes for FCS and non-FCS are mostly indistinguishable; exception: Wage Bill (slopes differ at 10 percent level), complicating interpretation for Wage Bill.
- Table 2 “Δ Relative Ambition” highlights:
  - PB XOXG: -2.2*** (.64) (FCS less relatively ambitious)
  - Debt: 3.5** (1.7) (FCS more relatively ambitious)
  - CAB: -1.5* (.79) (FCS less relatively ambitious, significance at 10 percent)
  - Other variables: differences small and statistically insignificant.
- Overall conclusion on quantitative tailoring:
  - Apart from PB XOXG (and possibly CAB at 10 percent significance), little evidence of systematic quantitative tailoring in program design after accounting for initial conditions.
  - FCS’s more ambitious debt reduction targets might be attributable to greater expected non-HIPC debt relief.
  - Note: debt and revenue data adjusted to remove the impact of HIPC debt relief; other smaller debt relief not adjusted due to data limitations.

### Optimism — definition, scope, and main results
- Definition: Optimism = target − outcome. Signing convention: positive number → outcome falls short of target. For Inflation, outcome exceeding target is a “shortfall.”
- Sample construction for optimism analysis:
  - Paired observations only (exclude programs started less than four years prior to data collection).
  - For variables other than growth and inflation, include only programs that targeted a non-negative adjustment.
  - Tables 4 and 5 report optimism for full sample and on-track programs, respectively. When # paired observations < 20 for F or NF, programs are pooled (“All”).
  - On-track programs: successful review 2.5 years or more after program approval.
- General result:
  - IM F-supported programs are optimistic in nearly all variables except inflation; this holds for full sample and for on-track programs.
  - Dropping programs possibly affected by COVID does not materially change conclusions for most variables.
- Specific magnitudes (from Table 3 — Arithmetic averages; standard deviations in parentheses; optimism expressed also as percent of target in square brackets):
  - Growth:
    - F (30 paired): Target 6.0 (2.7); Outcome 3.7 (3.5); Optimism 2.4*** [39%].
    - NF (32 paired): Target 5.6 (1.9); Outcome 4.7 (1.7); Optimism .9*** [16%].
    - Combined full-sample average growth optimism reported as 1.6 p.p. (not shown in table).
    - On-track programs (Table 5): #Obs 39; Target 5.7 (2.0); Outcome 4.5 (2.1); Optimism 1.2*** [21%].
  - Inflation:
    - F (30): Target 5.5 (3.3); Outcome 6.0 (4.1); Optimism .5 (8%) — not statistically significant.
    - NF (32): Target 4.9 (2.3); Outcome 5.3 (5.4); Optimism .5 (10%) — not statistically significant.
  - Revenue XOXG (Target ≥ 0):
    - F (25): Target 1.9 (1.2); Outcome .7 (2.2); Optimism 1.1** (62%) .
    - NF (26): Target 1.6 (1.2); Outcome .9 (1.6); Optimism .7** (44%).
    - On-track (Table 5): #Obs 32; Target 1.7 (1.2); Outcome 1.0 (1.9); Optimism .6** [38%].
  - PCE (Target ≥ 0):
    - All (41): Target 1.7 (1.4); Outcome -.6 (2.2); Optimism 2.3*** (135%).
    - On-track (26): Target 1.9 (1.4); Outcome -.7 (2.5); Optimism 2.6*** [139%].
  - Wage Bill (Target ≥ 0):
    - All (44): Target .7 (.7); Outcome -.0 (.9); Optimism .7*** [102%].
    - On-track (29): Target .7 (.7); Outcome .1 (.9); Optimism .6*** [86%].
  - PB (Target ≥ 0):
    - All (42): Target 2.6 (2.4); Outcome -.2 (3.2); Optimism 2.8*** (107%).
    - On-track (26): Target 3.1 (2.7); Outcome -.8 (3.2); Optimism 3.9*** [126%].
  - PB XOXG (Target ≥ 0):
    - F (22): Target 3.8 (4.6); Outcome 2.2 (4.5); Optimism 1.6* (43%).
    - NF (27): Target 3.5 (2.7); Outcome .9 (2.9); Optimism 2.6*** (75%).
    - On-track (30): Target 3.3 (2.9); Outcome .7 (4.1); Optimism 2.6*** [79%].
  - Debt (Target ≥ 0):
    - All (36): Target 9.6 (8.4); Outcome -3.2 (14); Optimism 13*** (133%).
    - On-track (21): Target 7.7 (5.3); Outcome - .5 (12); Optimism 8.2*** [106%].
  - CAB (Target ≥ 0):
    - All (31): Target 3.6 (3.7); Outcome .9 (4.8); Optimism 2.7*** (75%).
    - On-track (20): Target 4.3 (4.3); Outcome .7 (5.0); Optimism 3.6*** [83%].
  - Reserves (Target ≥ 0):
    - All (43): Target .8 (.9); Outcome .3 (1.4); Optimism .5** (67%).
    - On-track (28): Target .8 (.8); Outcome .5 (1.2); Optimism .32 [41%].
- Robustness checks:
  - Dropping programs approved since 2017 (to avoid COVID distortions) reduces observations by 11 but leaves results essentially unchanged in all dimensions other than revenue and primary balances.
  - Using medians instead of means does not materially change conclusions; exceptions noted: PB XOXG for FCS in full sample and Reserves in subsample — medians cannot reject calibration for these two.
- Differential optimism across FCS vs non-FCS:
  - Growth: FCS growth optimism 2.4 p.p. below target (GDP 7.4 percent lower than expected over three years); differential versus non-FCS (.9 p.p.) significant at 10 percent.
  - For Inflation, Revenue XOXG, and PB XOXG: no differential optimism between FCS and non-FCS.
  - For many variables and for the on-track subsample, sample sizes fall below prespecified minimum of 20 for FCS or non-FCS, limiting comparisons.

### Role of growth optimism in explaining optimism elsewhere
- Investigated whether optimism about growth explains optimism in other variables (see Appendix F for details).
- Key intuition:
  - A shortfall in growth affects ratios expressed as percent of GDP via denominator and can affect numerators (e.g., revenues).
  - Under unit elasticity, targeted GDP ratios and outcomes are growth independent; elasticity close to unity may hold for revenues but less likely for expenditures in short run.
- Empirical takeaway:
  - Growth optimism helps explain optimism about Debt reduction.
  - To a lesser extent, growth optimism helps explain optimism about Inflation and PCE.
  - Growth optimism does not explain optimism for other variables.

### Correlation between targets and outcomes — framing
- Despite level optimism, targets might be accurate on the margin (i.e., higher targets associated with higher outcomes).
- Focus for correlation analysis: programs with non-negative adjustment targets.
- Further analysis continues beyond Appendix B (section begins V in source).

*Source: IMF Working Paper — Appendix B of “IMF-Supported Programs in Low-Income Countries: Fragile versus Non-Fragile States.”*

### conclusions might be driven by countries achieving—or even systematically undershooting—projected

### IMF-Supported Programs in Low-Income Countries: Fragile versus Non-Fragile States

### Targets versus outcomes: patterns and marginal optimism
- Regression analysis shows the slope between target and outcome is strictly less than 1 for all variables other than Inflation, indicating optimism on the margin.
- The p-values for tests of unit slope "usually lie below .01 and always below .05."
- For CPI Inflation in non-FCS, the slope is strictly greater than 1 (still interpreted as optimism in the sense described).
- Visual convention in Figure 4: 45-degree line (black dashed); regression lines in red for FCS, green for non-FCS; if paired observations < 20 for either type, regressions pooled and shown in blue.

### Correlation evidence: independence for most variables
- Except for Growth and Inflation, targets and outcomes appear statistically independent, even for on-track programs.
- Non-parametric rank correlations (Spearman’s 휌휌 and Kendall’s 휏휏) generally fail to reject independence for variables other than Growth and Inflation, "even at 10 percent significance."
- Pearson correlation and regression coefficients largely corroborate rank-correlation findings, with two exceptions in the full sample noted in the source.
- Note on nonlinearity: for Growth in FCS the simultaneous rejection patterns suggest a possible non-linear relationship.

### Forecast accuracy: Theil's U2 and R2 findings
- Theil's U2 compares target forecasts to naïve forecasts (mean or median of other programs). U2 definition preserved as presented in the source.
- Main empirical findings:
  - Theil's U2 using the median is significantly below 1 only for Inflation.
  - In the full sample only, U2 is close to 1 for PB XOXG in FCS; otherwise U2 strictly exceeds 1.
  - Most U2 values range between 1.2 and 1.4.
  - Implication: forecast errors using program targets are significantly greater than those using the median-based naïve model (i.e., naïve median often outperforms program targets).
  - The naïve model using the mean performs equally well in most cases, with one exception: Growth in non-FCS in the full sample has U2 = 1.09 using the median but U2 = 0.57 using the mean.
- R2 results:
  - Except for Inflation, targets generate negative R2s (targets deviate more from outcomes than the uniform average).
  - Negative R2s are consistent with U2 > 1 results.

### Subsamples, robustness checks, and special cases
- Pooling and on-track subsample:
  - Tables 6 and 7 report paired observation counts and numerous statistics; where # paired observations ≥ 20 the source distinguishes FCS and non-FCS.
  - For on-track programs, naïve median tends to outperform targets except for Inflation; in the full on-track sample, Theil's U2 using the median is significantly below 1 only for Inflation.
- COVID-19 and sample exclusions:
  - Dropping programs affected by COVID does not materially affect main conclusions: the null hypothesis of unit slope is still confidently rejected; independence between targets and outcomes cannot be rejected for most variables except Growth and Inflation in some cases.
  - Number of programs whose outcomes have been influenced by the epidemic: 11.
- Handicap (time-lagged naïve forecasts):
  - Naïve forecasts recalculated using only outcomes known at program approval or only programs approved more than 3 years earlier (U2 Handicap) do not change main conclusions: means and medians continue to outperform targets in dimensions other than Inflation, with one exception where Targets outperform naïve forecasts for Revenue XOXG in the on-track subsample.
- Inclusion of programs with negative adjustment targets:
  - Means and medians continue to outperform targets in forecasting for all dimensions other than Inflation; outperformance is strict except for PB XOXG (per Appendix D, Tables 3 and 4).

### Fragile vs non-fragile states: similarities and differences
- Limited difference between FCS and non-FCS in average ambitiousness and optimism.
- Table 6 observations:
  - Spearman’s 휌휌 and Kendall’s 휏휏 are essentially identical for FCS and non-FCS for variables with sufficient observations: Growth, Inflation, Revenue XOXG, and PB XOXG.
  - Pearson correlations and slopes differ for Growth and PB XOXG:
    - Growth targets and outcomes are more correlated for non-FCS.
    - PB XOXG targets and outcomes are more correlated for FCS.
- Sample composition for tailoring analysis: 84 programs with 43 countries (14 countries appear once, 18 twice, 10 three times, 1 country four times (Sierra Leone)).
- For other assessments: 62 programs with 37 countries (14 once, 19 twice, 4 countries three times).

### Caveats, limitations, and methodological considerations
- Independence of observations:
  - t-tests rely on CLT; independence may be violated for contemporaneous programs (e.g., shared world growth forecasts).
  - Generalized CLT requiring only near-independence for temporally distant programs is invoked; non-parametric Wilcoxon signed-rank and Mann-Whitney U tests in Appendices B, C, and E yield similar conclusions.
- Non-linearity:
  - Linear specification for evaluating tailoring could generate Type I or Type II errors if the true relationship between initial conditions and targets is non-linear (convex).
  - Example: average Revenues XOXG at T were 14 versus 21 percent of GDP for FCS versus non-FCS (see Table 2); this difference could interact with non-linearity.
  - The tentative conclusion of quantitative tailoring for the Wage Bill might be a Type I error; average Wage Bill at T is "around half a p.p. of GDP lower for FCS than for non-FCS."
- Un-re-benchmarking GDP:
  - Adjusting realized GDP for program years to remove effects of re-benchmarking (October 2023 WEO adjustments) introduces noise and potential distortions that complicate interpretation.
- Repeat customers:
  - Many countries appear multiple times; for the tailoring sample (84 programs) there are 43 countries; for other assessments (62 programs) there are 37 countries.
  - Programs with planned duration < 1.5 years were dropped; spacing and staff turnover (typical tenure 2-3 years) reduce correlated error-term concerns.
  - Appendix E finds little material support for the hypothesis of learning across repeat programs.
- Burden of proof:
  - The report emphasizes the distinction between (i) rejecting a null hypothesis and (ii) accepting an alternative. Findings about average and marginal optimism fall into the stronger category (established), whereas findings on tailoring and independence reflect inability to reject nulls (weaker).

### Policy implications and concluding observations
- The findings support the rationale for the IMF’s Strategy for Fragile and Conflict-Affected States (2022): avoid overly optimistic assumptions and apply more realistic macro frameworks across LIC programs.
- Three main insights:
  - Before the 2022 FCS Strategy, there was limited differentiation in program ambition between FCS and non-FCS, suggesting uniformity in IMF program design.
  - Macroeconomic projections in IMF-supported programs show a tendency toward optimism in several variables; Inflation projections are generally accurate.
  - Independence of outcomes and targets for most variables points to potential gains from country-independent targets (e.g., median-based) to improve forecast accuracy—except for Inflation.
- Suggested direction: systematically extend the Strategy’s focus on realism and calibration across all LIC programs, including consideration of country-independent benchmarking for variables other than Inflation.

*Source: IMF Working Paper — IMF-Supported Programs in Low-Income Countries: Fragile versus Non-Fragile States (excerpts provided).*

### References

### References

### Key bibliographic items
- Lists studies on IMF engagement, program design, fragility, conflict, and macroeconomic outcomes, including works by Anselin and O’Loughlin; Bal Gunduz et al. 2013; Baqir, Ramcharan, and Sahay 2005; Bellemare 2015; Brückner and Ciccone 2010; Buhaug and Gleditsch 2008; Cerra and Saxena 2008; Collier 1999; Easterly and Rebelo 1993; Rodrik 1999; and multiple IMF and Independent Evaluation Office publications (2008–2022).
- Background and evaluation documents cited include: IMF Occasional Paper 13/277; IMF Working Paper WP/21/216; IMF Working Paper WP/23/68; Independent Evaluation Office background papers and reviews; IMF policy and staff guidance notes on fragile situations and program design.

### Appendix A — Initial conditions versus Targets
- Figure 1: regression lines of targets versus outcomes presented separately for FCS and non-FCS; right panel shows regression lines with equality of slopes and different intercepts. The distance between the parallel lines is used as the measure of relative ambitiousness.

### Appendix B — Medians Instead of Means (Table highlights)
- Growth (45, 39)
  - Initial Condition median: 3.5 [1.1, 5.0]
  - Target median: 5.1 [4.2, 6.3]
  - Non-FCS Target median: 6.0 [4.5, 6.7]
- Inflation (44, 39)
  - Initial Condition median: 6.3 [2.5, 9.8]
  - Target median: 5.2 [2.5, 7.7]
- Revenue (45, 39)
  - Initial Condition median: 13 [10, 16]
  - Target median (Non-FCS): 20 [14, 25]
  - Statistical sign: "< ∗∗∗" comparing F with NF
- Primary Balance (PB) (45, 39)
  - Initial Condition median: -1.4 [- 2.8, .2]
  - Non-FCS Initial median: -2.4 [- 4.1, -1.6]
  - PB Target median comparison: F .6 [- .9, 1.8] vs NF 1.5 [.5, 2.9] with "< ∗∗"
- Debt (45, 39) initial median: 48 [35, 71]; Non-FCS: 49 [38, 60]

- Statistical methods: To compare F with NF used Mann-Whitney U Test (nonparametric 2-sample test for equality of medians of unmatched data).

### Appendix B — Optimism (Table 2 highlights)
- Growth optimism (paired)
  - F (30): Target median 5.6 [4.5, 7.0]; Outcome median 3.7 [2.6, 4.6]; reported difference 1.9*** (and 1.2*** approximately)
  - NF (32): Target median 5.9 [4.4, 6.7]; Outcome median 4.7 [4.0, 5.7]
- Debt Target ≥ 0 (All, 36 paired)
  - Target median: 8.3 [3.4, 13]
  - Outcome median: -5.5 [- 12, 4.8]
  - Optimism reported: 14*** 
- CAB Target ≥ 0 (All, 31)
  - Target median: 2.4 [1.2, 5.2]
  - Outcome median: .5 [- 1.5, 3.8]
  - Optimism: 1.9**
- Reserves Target ≥ 0 (All, 43)
  - Target median: .5 [.2, 1.0]
  - Outcome median: .4 [- .3, .9]
  - Optimism: .1*

- Statistical methods: Wilcoxon signed-rank test for Targets vs Outcomes (paired); Mann-Whitney U Test for F vs NF.

### Appendix B — On-Track Only (Table 3 highlights)
- Growth (39 paired)
  - Target median: 5.5 [4.7, 6.7]
  - Outcome median: 4.5 [3.7, 5.6]
  - Optimism: 1.0***
- PB Target ≥ 0 (26 paired)
  - Target median: 2.1 [1.4, 4.5]
  - Outcome median: - .3 [-.2.6, 1.3]
  - Optimism: 2.3***
- Debt Target ≥ 0 (21 paired)
  - Target median: 6.3 [3.6, 12]
  - Outcome median: - .1 [-8.9, 7.1]
  - Optimism: 6.5**

### Appendix C — Excluding COVID-Affected Programs (selected highlights)
- Pre-2017 programs (Arithmetic averages)
  - Growth F (23): Target 6.4 (2.8); Outcome 4.0 (3.8); Optimism 2.4** (38%) and ≈ .9** (15%) for comparison
  - Debt Target ≥ 0 All (27): Target 8.6 (6.0); Outcome -2.4 (15); Optimism 11*** (128%)
  - PCE Target ≥ 0 All (32): Target 1.7 (1.3); Outcome - .5 (2.4); Optimism 2.2*** (127%)
- Statistical methods for pre-2017: paired t-test for Targets vs Outcomes; independent t-test to compare F vs NF.

- Pre-2017 programs (Median highlights)
  - Growth F (23): Target 5.8 [4.5, 7.6]; Outcome 3.8 [2.7, 5.0]; Optimism 2.0**
  - Debt Target ≥ 0 All (27): Target 8.3 [3.6, 12]; Outcome -3.3 [-12, 7.1]; Optimism 12***

### Appendix C — Correlations and Theil measures (selected entries)
- Table 5 (Pre-2017 programs)
  - Growth NF (27): Pearson Corr. .44; Spearman 휌 .51***; Kendall 휏 .36*; R2 .27*
  - Inflation F (23): Pearson Corr. .9***; Spearman 휌 .79***; Kendall 휏 .77***; R2 .60***
  - PB XOXG All: Pearson Corr. .51***; Spearman 휌 .46***; Kendall 휏 .36**; R2 .25**

- Table 6 (Pre-2017 On-Track Only)
  - Inflation All (32): Slope 1.47***; Pearson Corr. .74***; Spearman 휌 .79***; R2 .62***
  - PB XOXG All (24): Pearson Corr. .50*; Spearman 휌 .40*; R2 .10

### Appendix D — Including Negative Adjustments (selected highlights)
- Table 1 (Optimism with Negative Adjustment Targets, arithmetic averages)
  - Revenue XOXG F (30): Target 1.3 (1.8); Outcome .5 (2.2); Optimism 1.1** (61%) and ≈ .2 (15%)
  - Debt F (28): Target 4.8 (8.9); Outcome -3.3 (12); Optimism 8.1*** (168%) and ≈ 9.7*** (410%) vs NF
  - PB F (30): Target .6 (2.8); Outcome - .4 (3.4); Optimism 1.0* (175%); NF PB target 2.0 (2.7) with outcome - .6 (2.8) and PB NF optimism 2.6*** (130%)
- Table 2 (On-Track Only with Negative Adjustments)
  - Growth (39 paired): Target 5.7 (2.0); Outcome 4.5 (2.1); Optimism 1.2*** [21%]
  - PB Target ≥ 0 (39 paired): Target 1.6 (3.1); Outcome -1.1 (3.1); Optimism 2.7*** [165%]
  - Debt Target ≥ 0 (39 paired): Target 2.7 (7.1); Outcome -3.2 (11); Optimism 5.8*** [218%]

- Correlation tables with negative adjustments (selected)
  - PB XOXG All (62): Slope .55***; Pearson Corr. .49***; Spearman 휌 .39**; R2 .27**
  - Wage Bill All (59): Slope .47***; Pearson Corr. .36***; Spearman 휌 .38**; R2 .27**
  - On-Track Only PB XOXG All (39): Slope .62; Pearson Corr. .51***; Spearman 휌 .39**; R2 .26**

- Statistical notes: paired t-tests, Wilcoxon signed-rank tests, Mann-Whitney U Tests, and correlation measures (Pearson, Spearman, Kendall) used throughout to assess targets versus outcomes, optimism, and correlations.

*Source: wpiea2024221-print-pdf - References*

### Appendix E. Initial vs Repeat Programs

### Appendix E. Initial vs Repeat Programs

### Overview
- Repeat programs (R) are defined as programs with countries that have a preceding program in the sample; all other programs are initial or non-repeat (NR).
- Sample sizes:
  - Ambitiousness assessment: 84 programs across 43 countries (43 NR, 41 R).
  - Optimism and correlation assessments: maximally 62 programs across 37 countries (37 NR, 25 R).
- Conclusion stated in the text: Tables 1 to 5 reveal very little material differences between NR versus R programs in terms of ambitiousness, optimism, or correlation.

### Absolute Ambitiousness (Tables 1–2)
- Arithmetic averages (Table 1) and medians [inter-quartile range] (Table 2) are reported for initial condition and target levels across variables. Selected exact values:
  - Growth initial condition (43, 41): NR 4.0 (2.8); R 3.4 (2.9). Growth target (43, 41): NR 6.1 (2.5) > ∗∗ R 5.0 (1.8).
  - Inflation initial condition (42, 41): NR 6.8 (5.1) ≈ R 5.8 (4.9). Inflation target (42, 41): NR 5.3 (3.1) ≈ R 5.3 (3.1).
  - Revenue XOXG initial condition (43, 41): NR 18 (8.3) ≈ R 16 (7.0). Revenue XOXG target (43, 41): NR 1.5 (1.9) ≈ R 1.3 1.2.
  - PCE initial condition (41, 41): NR 19 (8.9) ≈ R 18 (8.8). PCE target (40, 40): NR .7 (2.1) ≈ R .9 (1.3).
  - Wage Bill initial condition (42, 41): NR 7.6 (3.6) ≈ R 7.5 (2.5). Wage Bill target (42, 40): NR .3 (.9) ≈ R .4 (.6).
  - PB initial condition (43, 41): NR -2.2 (3.4) ≈ R -1.4 (3.2). PB target (43, 41): NR 1.3 (2.8) ≈ R 1.2 (2.6).
  - PB XOXG initial condition (43, 41): NR -11 (7.6) ≈ R -9.0 (6.0). PB XOXG target (43, 41): NR 2.6 (4.6) ≈ R 2.0 (2.4).
  - Debt initial condition (43, 41): NR 61 (45) ≈ R 50 (22). Debt target (41, 41): NR 4.3 (8.8) ≈ R 3.0 (7.9).
  - CAB initial condition (31, 40): NR -10 (9.9) ≈ R -7.6 (9.1). CAB target (31, 40): NR 1.7 (5.3) ≈ R .5 (4.6).
  - Reserves initial condition (42, 41): NR 4.0 (1.9) ≈ R 4.1 (3.3). Reserves target (42, 41): NR .2 (1.0) ≈ R .4 (1.2).
- Statistical testing notes:
  - Mann-Whitney U Test used to compare F with NF for medians of unmatched data (noted under both Tables 1 and 2).

### Optimism (Tables 3–4)
- Optimism defined as Target minus Outcome; comparisons use paired observations (#Obs (paired)). If #Obs (paired) < 20 for R or NR, programs are pooled (“All”).
- Selected arithmetic-average optimism results (Table 3):
  - Growth:
    - NR 36 paired: Target 6.2 (2.6) > ∗∗∗; Outcome 4.5 (3.5); Optimism 1.7*** ≈ 1.5*** (comparison NR vs R).
    - R 26 paired: Target 5.3 (1.7) > ∗∗∗; Outcome 3.8 (2.1).
  - Inflation:
    - NR 36 paired: Target 5.4 (3.0) ≈ Outcome 5.9 (4.0); Optimism .5 ≈ .5.
    - R 26 paired: Target 4.8 (2.5) ≈ Outcome 5.3 (3.0).
  - Revenue XOXG (Target ≥ 0):
    - NR 29 paired: Target 2.0 (1.3) > ∗∗; Outcome .9 (2.2); Optimism 1.1** ≈ .7** (NR vs R).
    - R 22 paired: Target 1.4 (.8) > ∗∗; Outcome .7 (1.5).
  - PCE (Target ≥ 0, All 41): Target 1.7 (1.4) > ∗∗∗; Outcome -.6 (2.2); Optimism 2.3***.
  - Wage Bill (Target ≥ 0, All 44): Target .7 (.7) > ∗∗∗; Outcome -.0 (.9); Optimism .7***.
  - PB (Target ≥ 0, All 42): Target 2.6 (2.4) > ∗∗∗; Outcome -.2 (3.2); Optimism 2.8***.
  - PB XOXG (Target ≥ 0):
    - NR 29 paired: Target 4.2 (4.3) > ∗∗; Outcome 2.0 (5.1); Optimism 2.2** ≈ 2.5*** (NR vs R).
    - R 20 paired: Target 2.9 (2.3) > ∗∗∗; Outcome .8 (3.6).
  - Debt (Target ≥ 0, All 36): Target 9.6 (8.4) > ∗∗∗; Outcome -3.2 (14); Optimism 13***.
  - CAB (Target ≥ 0, All 31): Target 3.6 (3.7) > ∗∗∗; Outcome .9 (4.8); Optimism 2.7***.
  - Reserves (Target ≥ 0):
    - NR 23 paired: Target .7 (.9) ≈ Outcome .4 (1.5); Optimism .3 ≈ .8** (NR vs R).
    - R 20 paired: Target .9 (1.0) > ∗∗; Outcome .1 (1.4).
- Median results (Table 4) largely mirror arithmetic findings (selected medians and IQRs reported in the table).
- Statistical testing notes:
  - Paired t-test used to compare Targets with Outcomes (arithmetic averages).
  - For medians, Wilcoxon signed-rank test used to compare Targets with Outcomes; Mann-Whitney U Test used to compare optimism for R vs NR.

### Correlations (Table 5)
- Table reports measures of monotone relationships between targets and outcomes and forecasting performance.
- Selected exact entries (by variable and program type):
  - Growth:
    - NR 36 paired: Slope .32*; Pearson Corr. .27; Spearman 휌 .30*; Kendall 휏 .22*; R2 -.53; Theil U2 .94; Mean 1.34; Median 1.34.
    - R 26 paired: Slope -.20; Pearson Corr. -.16; Spearman 휌 .40**; Kendall 휏 .34**; R2 -1.47; Theil U2 .96; Mean 1.54; Median 1.54.
  - Inflation:
    - NR 36 paired: Slope .95***; Pearson Corr. .72***; Spearman 휌 .80***; Kendall 휏 .60***; R2 .51; Theil U2 .53; Mean .69.
    - R 26 paired: Slope 1.9***; Pearson Corr. .83***; Spearman 휌 .85***; Kendall 휏 .69***; R2 .53; Theil U2 .57; Mean .68.
  - Revenue XOXG (Target ≥ 0):
    - NR 29 paired: Slope .02; Pearson Corr. .02; Spearman 휌 -.02; Kendall 휏 -.01; R2 -.62; Theil U2 1.21; Mean 1.26.
    - R 22 paired: Slope .71; Pearson Corr. .40*; Spearman 휌 .4; Kendall 휏 .31; R2 -.09; Theil U2 .97; Mean 1.04.
  - PB XOXG (Target ≥ 0):
    - NR 29 paired: Slope .79*; Pearson Corr. .46*; Spearman 휌 .25; Kendall 휏 .17; R2 -.31; Theil U2 1.10; Mean 1.12.
    - R 20 paired: Slope .42**; Pearson Corr. .51**; Spearman 휌 .44*; Kendall 휏 .31*; R2 -.12; Theil U2 1.03; Mean 1.01.
  - Debt (Target ≥ 0, All 35): Slope -.15; Pearson Corr. -.07; Spearman 휌 -.12; Kendall 휏 -.07; R2 -1.21; Theil U2 1.44; Mean 1.47.
  - Reserves (Target ≥ 0):
    - NR 23 paired: Slope .06; Pearson Corr. .03; Spearman 휌 .00; Kendall 휏 -.04; R2 -.32; Theil U2 1.11; Mean 1.14.
    - R 20 paired: Slope -.07; Pearson Corr. -.05; Spearman 휌 .11; Kendall 휏 .05; R2 -.94; Theil U2 1.37; Mean 1.34.
- Interpretation notes in the text:
  - Theil U2 > 1 means naive forecasts (using mean/median outcomes of other programs) do better on average than using program targets.
  - R2 gives proportion of variance in outcomes explained by program targets; negative R2 indicates targets deviate more from outcomes than the uniform average.

---

### Appendix F. How Much Does Growth Optimism Explain?

### Conceptual framework
- Variables:
  - Y = nominal GDP; X = another nominal variable (e.g., revenue or expenditure).
  - Deflating by price level P yields real values x and y.
  - r denotes ratio X/Y.
  - Elasticities 휂_{X,Y} and 휂_{x,y} with respect to Y and y; 휂_{X,Y} = 휂_{x,y}.
- Key identity and implication (equation (1) in text):
  - d dY/Y < = > 0 ⇔ 휂_{X,Y} < = > 1.
  - A negative surprise for growth helps to explain optimism about variable r iff 휂_{X,Y} > 1.

### Empirical correlations (Tables 1–2 in Appendix F)
- Table 1 (full sample) reports R2 and correlation measures between forecast errors for Growth and other variables:
  - R2: Growth .04; Inflation .02; Revenue XOXG .14; PCE .00; Wage Bill .00; PB .00; PB XOXG .38; Debt .01; CAB .00; Reserves .00.
  - Pearson Corr.: Growth -.20; Inflation .15; Revenue XOXG .38**; PCE .05; Wage Bill -.04; PB .01; PB XOXG .62***; Debt -.08; CAB .02; Reserves (not listed in Pearson row).
  - Spearman 휌: Growth -.36***; Inflation .20; Revenue XOXG .36**; PCE .07; Wage Bill -.04; PB .05; PB XOXG .55***; Debt -.01; CAB .03.
  - Kendall 휏: Growth -.26***; Inflation .13; Revenue XOXG .26**; PCE .05; Wage Bill -.01; PB .03; PB XOXG .40***; Debt -.00; CAB .3.
- Table 2 (on-track programs only) reports lower magnitudes generally; selected entries:
  - R2: Growth .01; Inflation .00; Revenue XOXG .06; PCE .00; Wage Bill .05; PB .02; PB XOXG .44; Debt .01; CAB .04; Reserves .00.
  - Pearson Corr.: Growth -.11; Inflation .07; Revenue XOXG .25; PCE .05; Wage Bill .22; PB .14; PB XOXG .67**; Debt -.11; CAB -.20.
  - Spearman 휌 and Kendall 휏 similarly reported with smaller magnitudes and significance levels.

### Interpretation and narrative findings
- Revenue XOXG:
  - Empirical point estimate of the correlation between forecast errors for growth and revenue XOXG is small and not statistically significant in full sample → suggests elasticity around 1.
  - Textual reasoning: average elasticity of government revenue with respect to GDP shocks typically slightly greater than one in some studies but LICs often have lower elasticities; to first order elasticity of 1 implies revenues as a percentage of GDP would not differ materially if growth had been as predicted.
- Government current expenditure (PCE):
  - Elasticity typically somewhat less than one in many contexts → implies a negative relationship between unexpectedly higher growth and expenditure as a percentage of GDP (i.e., realized growth associated with expenditure reduction as % GDP).
  - Empirical result: full sample correlation for PCE is .38 (significant); on-track programs correlation is .25 (not statistically significant).
- Primary Balance (PB) and PB XOXG:
  - Zero correlations for PB (XOXG) in full sample discussed as possibly due to offsetting effects from fixed public investment; on-track programs show positive but not statistically significant correlations.
- Debt, inflation, CAB, and reserves:
  - Strong positive correlations between forecast errors for growth and debt reduction consistent with “snowball effect” of growth shocks on debt-to-GDP ratio.
  - Negative correlations between forecast errors for growth and inflation suggest growth shocks in the sample came predominantly from the supply side.
  - CAB and Reserves correlations not statistically significant.
- Overall summary statement: growth optimism is an important contributor to optimism about Debt reduction, and to a lesser degree for PCE.

---

### Appendix G. GDP Re-benchmarking

### Methodology to compare initial targets and outcomes when GDP series were re-benchmarked
- Data vintages used:
  - Initial conditions and program targets: first WEO vintage published after program adoption.
  - Outcomes: October 2023 WEO.
- Un-re-benchmarking procedure:
  1. Calculation of Ratios: For pre-sample period years 2003 to 2007, compute ratios between nominal GDP in the 2023 WEO and nominal GDP in the first WEO after program approval (five ratios).
  2. Geometric Mean: For each country, take the geometric mean of the five ratios (fifth root of the product).
  3. Un-re-benchmarking: Divide the nominal GDP for T+3 from October 2023 WEO by this geometric mean to estimate nominal GDP under the original benchmark; this permits comparison between outcomes and targets.
- Empirical distribution in the study of 84 programs:
  - 26 programs had a ratio equal to 1 (no re-benchmarking).
  - 40 programs had ratios within the [0.95,1.05] range (limited re-benchmarking).
  - 18 programs had more sizable re-benchmarking (ratios outside [0.95,1.05]).
- Non-oil GDP application:
  - 29 programs had ratios equal to 1.
  - 42 programs had geometric ratios within the [0.95,1.05] range.

---

### Appendix H. Country Classifications

### FCS Status 2008–2023 (Table 1)
- Source: IMF. Between 2008 and 2022, IMF had its own FCS classification; since 2023 it is harmonized with the World Bank’s FCS classification.
- Table lists each country with a string of digits for years 2008–2023 indicating FCS status across years (as provided in the source table).

### PRGT-Eligibility 2008–2023 (Table 2)
- Source: IMF.
- Table lists each country with sequences of digits for years 2009–2022 indicating PRGT-eligibility across years (as provided in the source table).

*Source: wpiea2024221-print-pdf — Appendix E. Initial vs Repeat Programs (IMF Working Paper content provided).*

---


_Source: https://www.imf.org/-/media/files/publications/wp/2024/english/wpiea2024221-print-pdf.pdf_
