## wpiea2025109-print-pdf

## Source details

**Canonical URL:** [wpiea2025109-print-pdf](https://www.imf.org/-/media/files/publications/wp/2025/english/wpiea2025109-print-pdf.pdf)

## Other formats

- [Markdown version](/-/media/files/publications/wp/2025/english/wpiea2025109-print-pdf.pdf.md)
- [Structured JSON version](/-/media/files/publications/wp/2025/english/wpiea2025109-print-pdf.pdf.json)

---

### 1. Introduction — role, framework, and key insights
- Central bank communication shapes market expectations, influences economic decisions, and enhances accountability and transparency.
- Paper develops an automated sentence-level classification tool analyzing communications across four dimensions:
  - Topic (e.g., monetary policy, financial stability, climate change).
  - Communication stance (forward-looking vs backward-looking).
  - Audience (financial markets, businesses, households, governments, international stakeholders).
  - Policy sentiment (hawkish, dovish, neutral, risk-highlighting, confidence-building).
- Dataset scope and coverage:
  - 74,882 documents.
  - Approximately 21 million sentences.
  - 169 central banks.
  - Time span from 1884 to 2025.
- Sentiment taxonomy extended to include risk-highlighting and confidence-building alongside hawkish, dovish, and neutral.
- Key empirical insights (selected):
  - Advanced economies emphasize financial stability more; emerging economies emphasize fiscal policy.
  - Adoption of inflation targeting shifted messages toward forward-looking statements on inflation, interest rates, and economic activity.
  - Financial sector is primary audience; outreach to international stakeholders and businesses increased.
  - Accommodative policies accompanied by extensive dovish messaging; restrictive policies conveyed through concise, hawkish statements.

### 2. Dataset construction, processing, and descriptive statistics
- Document collection and formats:
  - Sources: 169 central bank websites.
  - File formats: PDFs, .docx, .doc, .txt, HTML; OCR applied to image-based documents.
  - Sentence-level processing with language detection and language-specific segmentation.
- Collected document counts and timeframes:
  - Annual Report: 3,879 docs; Begin Year 1884; End Year 2024.
  - Monetary Policy Report: 4,671 docs; Begin Year 1993; End Year 2025.
  - Financial Stability Report: 2,092 docs; Begin Year 1996; End Year 2025.
  - Monetary Policy Decision: 14,238 docs; Begin Year 1936; End Year 2025.
  - Speeches: 36,725 docs; Begin Year 1986; End Year 2025.
  - Other Documents: 11,437 docs; Begin Year 1993; End Year 2025.
- Corpus-level metrics:
  - Total central bank documents: 74,882.
  - Regular documents: 24,880.
  - Corpus size: over 80 GB.
  - Processed PDF pages: more than 1.3 million.
  - Processed sentences: approximately 21 million.
- Language distribution (sentences):
  - English: 87.4 percent.
  - Spanish: 6.6 percent.
  - Portuguese: 1.7 percent.
  - French: 1.4 percent.
- Document length and outlet trends:
  - Annual reports grew significantly in length, especially among low-income and pegged economies.
  - Monetary policy decisions were streamlined from 2000 until COVID-19; narrative detail increased after global inflation surge.
  - Speeches show a U-shaped pattern; longer in advanced and inflation-targeting economies.

### 3. Pro-forma readability and syntactic complexity
- Lexical readability (English only) — Flesch-Kincaid Ease Score:
  - Formula: Score = 206.835 − 1.015 × (Total Words / Total Sentences) − 84.6 × (Total Syllables / Total Words)
  - Interpretation thresholds: above 60 → 8th-grade; 30–50 → college-level; below 30 → highly technical.
  - Findings:
    - Pegged regimes and low-income economies: higher Flesch Reading Ease (simpler language).
    - Inflation-targeting and advanced economies: lower readability (more complex vocabulary).
- Syntactic complexity — dependency depth:
  - Metric: longest path from the root of a sentence’s syntactic tree to any terminal node.
  - Findings:
    - Advanced economies: more complex vocabulary but simpler sentence constructions (lower dependency depth).
    - Low-income and pegged economies: deeper sentence structures despite simpler vocabulary.
    - Since 2020, advanced economies simplified sentence structure; low-income economies increased sentence complexity starting in 2022.
- Lexical vs syntactic relationship:
  - Advanced economies combine low Flesch Ease with low dependency depth.
  - Low-income and many non-inflation-targeting economies: higher readability but greater syntactic depth.

### 4. Semantic classifier design, training, and performance
- Model choice and multilingual design:
  - Encoder-only sentence transformer selected (multilingual BGE, bge-m3) producing 1024-dimensional embeddings.
  - Multilingual capability to analyze over 100 languages and mitigate translation distortions.
- Labeled dataset:
  - 1,200 annotated sentences total:
    - ~240 synthetic instances (20 percent).
    - ~240 expert-crafted examples (20 percent).
    - 720 real-document sentences (60 percent).
  - Training/validation split:
    - Training = 840 examples (240 synthetic + 240 expert-crafted + 360 real).
    - Hold-out validation = 360 real sentences.
- Classification architecture:
  - Block factorization: P(ytopic, ystance, yaudience, ysentiment | x) ≈ P(ytopic, ystance | x) · P(yaudience, ysentiment | x).
  - Classifiers:
    - Classifier A: jointly predicts topic and communication stance.
    - Classifier B: jointly predicts audience and sentiment.
  - Additional “Metadata” class: ≈ 6.3 percent of total for topic/stance classification; 14.1 percent for audience/sentiment classification.
- Fine-tuning pipeline:
  1. Siamese contrastive training with hard triplet loss to produce embeddings.
  2. Dense neural network (softmax output) mapping embeddings to labels.
  - Hyperparameter search ranges:
    - learning rate: [10^{-5}, 10^{-2}].
    - L2-norm regularization: [10^{-6}, 10^{0}].
  - Optimizer: Adam variant with decoupled weight decay; warm-up = 10 percent of an epoch.
- Benchmarking vs ChatGPT 4o (validation set metrics):
  - Topic:
    - Our Classifier: Accuracy 0.689; Precision (Macro) 0.699; Recall (Macro) 0.645; F1 (Macro) 0.650; Cohen’s Kappa 0.666.
    - ChatGPT 4o: Accuracy 0.731; Precision (Macro) 0.593; Recall (Macro) 0.535; F1 (Macro) 0.551; Cohen’s Kappa 0.711.
  - Communication stance:
    - Our Classifier: Accuracy 0.924; Precision (Macro) 0.925; Recall (Macro) 0.910; F1 (Macro) 0.917; Cohen’s Kappa 0.834.
    - ChatGPT 4o: Accuracy 0.828; Precision (Macro) 0.591; Recall (Macro) 0.551; F1 (Macro) 0.570; Cohen’s Kappa 0.658.
  - Audience:
    - Our Classifier: Accuracy 0.706; Precision (Macro) 0.713; Recall (Macro) 0.699; F1 (Macro) 0.700; Cohen’s Kappa 0.622.
    - ChatGPT 4o: Accuracy 0.506; Precision (Macro) 0.529; Recall (Macro) 0.463; F1 (Macro) 0.437; Cohen’s Kappa 0.400.
  - Sentiment:
    - Our Classifier: Accuracy 0.700; Precision (Macro) 0.612; Recall (Macro) 0.595; F1 (Macro) 0.594; Cohen’s Kappa 0.589.
    - ChatGPT 4o: Accuracy 0.704; Precision (Macro) 0.557; Recall (Macro) 0.393; F1 (Macro) 0.418; Cohen’s Kappa 0.592.
- Key benchmarking conclusions:
  - Our classifier achieves higher macro-level metrics (notably macro F1) across most dimensions, improving detection of minority/infrequent classes.
  - ChatGPT 4o attains higher micro-level accuracy on topic driven by frequent classes but underperforms on macro-F1.
  - Cost considerations for generative LLMs (pricing as of April 2025, per text):
    - GPT-4o: $2.50 per million input tokens and $10.00 per million output tokens.
    - GPT-4.5: $75.00 per million input tokens and $150.00 per million output tokens.
    - Estimated cost processing 21 million prompt–response interactions:
      - GPT-4o: approximately $20,160.
      - GPT-4.5: approximately $565,530.

### 5. Semantic mapping and topical dynamics
- Embedding projection and temporal span:
  - Sentences embedded into 1024-dimensional vectors then reduced to two dimensions via t-SNE; data span “1884 to 2025.”
- Observed clustering:
  - Clear topic clusters with overlaps (e.g., forward-looking supervision and financial stability).
  - Emerging topics (climate change, technological innovation) show less dispersion.
- Topic composition by outlet and development level:
  - Monetary policy dominant in decisions and reports.
  - Financial stability reports focus on financial stability, supervision, and regulation.
  - Advanced economies: more attention to financial stability.
  - Emerging and low-income economies: greater importance on fiscal policy.
- Monetary policy subtopics and IT adoption:
  - Advanced economies: more forward-looking communication and emphasis on interest rates.
  - Inflation-targeting adoption associated with shift away from exchange rate discussions toward inflation, interest rates, and economic activity.
  - IT-adopting economies listed include Brazil, Chile, Georgia, Republic of Kazakhstan, Republic of Korea, Mexico, Republic of Moldova, New Zealand, Paraguay, Peru, the Philippines, Russian Federation, Seychelles, Sri Lanka, Uganda, and Ukraine.
- Forward-lookingness Score:
  - Formula: Forward-lookingness Score_{c,d,t} = #Forward-Looking_{c,d,t} / (#Forward-Looking_{c,d,t} + #Backward-Looking_{c,d,t})
  - Trends:
    - Forward-looking communication increased across most outlets; speeches and monetary policy decisions have highest scores.
    - Advanced economies consistently above-median forward-looking; pegged and low-income economies generally below.
    - Crisis periods associated with declines in forward-lookingness in crisis-related communication.

### 6. Communication metrics — definitions and construction
- Notation: H = hawkish count; D = dovish; C = confidence-building; R = risk-highlighting; N = neutral.
- Metrics (as defined):
  - Net Policy Sentiment (NPS): NPS = (H − D) / (H + D)  (range [−1, 1]).
  - Straightforwardness Index (SI): SI = (N + |H − D|) / (N + H + D)  (range [0,1]).
  - Explanation Index (EI): EI = (C + R + N) / (H + D).
  - Net Confidence Index (NCI): NCI = (C − R) / (C + R)  (range [−1, 1]).
- Forward/backward decompositions:
  - NPS_{fwd} = (H_{fwd} − D_{fwd}) / (H_{fwd} + D_{fwd}); NPS_{bwd} analogous.
  - Aggregation weights ω_{fwd} and ω_{bwd} based on counts (see text).

### 7. Empirical associations with policy rates, market rates, and OIS
- Panel estimation (monthly) specification example:
  - Policy Rate_{i,t} = α_i + λ_t + η Policy Rate_{i,t−1} + β NPS_{i,t} + X′_{i,t} γ + ε_{i,t}
  - Controls include CPI inflation, exchange rate (USD/local), Explanation Index, Straightforwardness Index, Net Confidence Index.
- Main panel results (Table 5 highlights):
  - Specification (I): total NPS coefficient = 0.031*** (0.005).
    - With sample SD of policy rate = 6.36 percentage points, one-standard-deviation increase in total NPS associated with ≈ 0.20 p.p. increase in policy rate (0.031 × 6.36 ≈ 0.20 p.p.).
  - Specification (II) forward/backward decomposition:
    - Forward-looking NPS = 0.026*** (0.004).
    - Backward-looking NPS = 0.015*** (0.004).
  - Policy rate persistence: Policy Rate_{i,t−1} ≈ 0.970*** (0.005).
  - Observations reported: 52095 (Spec I), 67391 (Spec II), 33807 (Spec III), 3807 (Spec IV).
  - R^2 examples: 0.970, 0.969, 0.980, 0.980.
- Predictive analysis (one-period-ahead, Table 7 selected coefficients; standardized variables):
  - ΔPolicy Rate_{i,t+1} coefficients:
    - Total NPS (Col I): 0.101*** (0.022).
    - Forward-looking (Col II): 0.083*** (0.017).
    - Backward-looking (Col II): 0.062*** (0.014).
    - Gap (Fwd−Bwd) (Col III): −0.096*** (0.022).
  - ΔT-Bill Rate (forward-looking): 0.162*** (0.033) (Col V).
  - ΔT-Bond Rate (forward-looking): 0.105* (0.058) (Col VIII).
  - Interpretation: forward-looking sentiment predicts one-period-ahead changes in policy and short-term market rates; effect sizes translate into basis-point movements (examples in text).
- OIS regressions (Table 9, standardized coefficients; sample: 15 economies, January 2015 to February 2025):
  - Net Policy Sentiment (Forward) coefficients by tenor:
    - 1 month: 0.017** (0.007).
    - 3 months: 0.024*** (0.007).
    - 6 months: 0.029*** (0.008).
    - 12 months: 0.036*** (0.008).
  - Backward-looking sentiment not statistically significant across tenors.
  - Policy Rate strongly anchors OIS across tenors (e.g., 1 month: 1.949*** (0.017)).
  - Observations by tenor: 491, 504, 556, 605.
  - R2 by tenor: 0.991, 0.983, 0.976, 0.968.
- Horizon and persistence (multi-horizon regressions):
  - Forward-looking sentiment significant for future policy rate changes up to several leads; significance for T-bill rates concentrated in first five leads.
  - Backward-looking sentiment lacks consistent predictive power for future market rates.

### 8. Heterogeneity, audience targeting, and net confidence dynamics
- Heterogeneity by monetary framework (Table 6 summary):
  - Inflation-targeting regimes:
    - Total NPS = 0.029*** (0.006).
    - Forward-looking NPS = 0.024*** (0.005).
    - Backward NPS = 0.014*** (0.004).
    - Strongest association between NPS and policy rates.
  - Exchange rate anchor regimes: no significant relationship (e.g., Total NPS = −0.068 (0.065)).
- Heterogeneity by level of development (Table 8 highlights; standardized coefficients):
  - ΔPolicy Rate — Other Levels: 0.079*** (0.018).
  - ΔPolicy Rate — Advanced Economies: 0.106** (0.039).
  - ΔT-Bond Interest Rate — Other Levels: 0.109* (0.053).
  - In advanced economies, forward-looking effects on market rates are weaker or absent for some specifications.
- Audience targeting and net confidence (Figures 28–31 summary):
  - Net Confidence Index (NCI) persistently negative across groups, indicating more risk-highlighting than confidence-building overall.
  - Audience-specific stylized facts:
    - General public communications are framed most confidently.
    - Government-directed communications emphasize risks more heavily.
    - Business and financial sector communications are more neutral and balanced.
  - Variance decomposition: audience dummies explain the largest share of variation in audience-specific net confidence indices; time and country effects explain less.
- Granger-causality between NCI components and VIX (Table 10 summary; monthly data 1990–2025, 420 observations):
  - ΔForward-looking NCI → VIX:
    - Significant F-statistics at lags 2 and 4 (e.g., Lag 2: F-stat 3.563, p-value 0.029; Lag 4: F-stat 3.151, p-value 0.014).
    - Conclusion: forward-looking NCI Granger-causes VIX for up to six months in selected lags.
  - VIX → ΔForward-looking NCI: no statistically significant lags.
  - ΔBackward-looking NCI → VIX: significant at lag 1 (F-stat 6.230, p-value 0.013) only.
  - VIX → ΔBackward-looking NCI: significant at lags 2–6 (e.g., Lag 2: F-stat 3.366, p-value 0.035), indicating lagged VIX predicts adjustments in backward-looking confidence.
  - Interpretation: forward-looking communication contains predictive signals about future market volatility, while backward-looking narratives tend to react to realized volatility with delays.

### 9. Classifier robustness, multilabel incidence, and confidence
- Cross-lingual consistency (Appendix B, translation-consistency sample):
  - Jensen–Shannon Divergence (JSD) values by language and dimension (selected):
    - Topic: none exceeding 0.16; Spanish (ES) = 0.07; Russian (RU) = 0.14; French (FR) and Arabic (AR) near 0.16.
    - Communication stance: Arabic (AR) and French (FR) just below 0.04; Spanish (ES) near 0.00.
    - Audience: Russian (RU) highest JSD = 0.12; Spanish (ES) JSD = 0.01.
    - Sentiment: Russian (RU) highest JSD = 0.15; Spanish (ES) near 0.01.
  - Summary: low divergence values indicate strong cross-lingual stability; some divergence for Russian and Arabic.
- Prediction confidence (Appendix C):
  - More than half of sentences classified with marginal probabilities above 60 percent for topic.
  - More than half above 70 percent for audience and sentiment.
  - Nearly 80 percent above 70 percent for communication stance.
  - Metadata class and substantive classes like “MP - inflation,” “forward-looking,” “financial sector,” and “risk-highlighting” have high marginal probabilities.
- Multilabel sentences (Appendix D):
  - Measurement: two distinct classes within same dimension if both have predicted marginal probabilities > 25 percent.
  - Topic co-occurrence: most pairs appear together in fewer than 0.04 percent of sentences.
  - Audience co-occurrence notable: Business Sector — Financial Sector = 1.89 percent.
  - Sentiment co-occurrence examples: Neutral/Balanced with Confidence-building = 3.62 percent; Neutral/Balanced with Risk-highlighting = 3.14 percent.
  - Conclusion: co-occurrence modest; single-label probabilistic approach adopted for practicality.

### 10. Practical implications and recommended research directions
- Practical uses for central banks and policymakers:
  - Enhance accountability, transparency, and expectation management through quantitative assessment of messaging.
  - Benchmark communications against historical trends or peer institutions.
  - Support improvements in monetary policy transmission by measuring communication effects.
- Recommended future research:
  - Empirically examine causal impact of communication—particularly forward-guidance proxy—on macroeconomic variables and monetary policy transmission.
  - Relate Net Confidence Index to financial stress indicators (e.g., country-level financial stress indices).
  - Study FX-related communication and exchange rate dynamics.
  - Extend methodology to social media and other non-traditional channels for timelier, targeted messaging analysis.

*Source: wpiea2025109-print-pdf — content derived from the working paper material provided.*

### 1.  Introduction

### 1. Introduction

### Role and challenges of central bank communication
- Central bank communication shapes market expectations, influences economic decisions, and enhances central bank accountability and transparency (Blinder et al., 2008; Woodford, 2005).
- Communication has expanded beyond monetary policy to cover many emerging topics, reflecting growing responsibilities and heightened public scrutiny.
- Transforming complex policy language into actionable insights is challenging due to policy discourse intricacies and the public’s unfamiliarity with central banking.
- Systematic and quantitative approaches are increasingly essential to assess clarity and consistency, align messages with policy objectives, benchmark performance, and respond to public concerns.

### Framework developed in this paper
- The paper develops an automated sentence-level classification tool that analyzes central bank communications along four key dimensions:
  - Topic (e.g., monetary policy, financial stability, climate change).
  - Communication stance (forward-looking vs backward-looking).
  - Audience (financial markets, businesses, households, governments, international stakeholders).
  - Policy sentiment (hawkish, dovish, neutral, risk-highlighting, confidence-building).
- The classifier operates at the sentence level to capture intra-document content shifts and provide fine-grained analysis.

### Methodological innovation and relation to prior work
- Integrates four dimensions into a unified, semantically rich framework; addresses limitations of prior dictionary and document-level approaches by leveraging large language models (LLMs).
- Prior studies applied embeddings or finetuned LLMs for narrow tasks or single-country datasets; this framework jointly classifies multiple dimensions within a single architecture trained on a large, multilingual corpus from most central banks worldwide.
- The approach enables direct textual measurement of forward-looking communication (forward guidance and other prospective signals), distinguishing these from backward-looking assessments.
- The semantic measure captures unconventional tools (asset purchase programs, liquidity operations, balance sheet strategies) beyond conventional interest rate guidance.

### Sentiment taxonomy enhancement
- Extends conventional hawkish-dovish-neutral sentiment by adding:
  - Risk-highlighting.
  - Confidence-building.
- Rationale: prevents misclassification of risk-related or reassurance language that can be strategically different from simple positive/negative sentiment.

### Dataset and multilingual scope
- Constructed, compiled, and processed a comprehensive dataset with:
  - 74,882 documents.
  - Approximately 21 million sentences.
  - 169 central banks.
  - Time span from 1884 to 2025.
- Appended Campiglio et al. (2025)’s consolidated dataset with speeches from 131 central banks covering 1986 to 2023.
- Most documents are in English, but the corpus includes multiple languages and a multilingual sentence transformer enables analysis across more than 100 languages.

### Model choice and technical design
- Starts with a moderate-size sentence transformer LLM; a general-purpose open-source language model is fine-tuned for central bank communication.
- Selected an encoder-only sentence transformer to extract dense-vector context-aware sentence representations suitable for downstream classification.
- Multilingual capability allows inclusive global analysis, including central banks publishing only in local languages.

### Key empirical insights
- Topic emphasis varies across development stages:
  - Advanced economies emphasize financial stability more.
  - Emerging economies focus more on fiscal policy.
- Structural change following adoption of inflation targeting:
  - Backward-looking discussions on exchange rates give way to more forward-looking statements on inflation, interest rates, and economic activity.
- Audience shifts:
  - Financial sector is the primary recipient, with increased engagement with international stakeholders and businesses.
- Sentiment asymmetries:
  - Accommodative policies are often accompanied by extensive dovish messaging.
  - Restrictive policies are conveyed through more concise, hawkish statements.

### Proposed semantic textual metrics
- Four metrics are proposed and decomposed into forward- and backward-looking components:
  - Net policy sentiment
    - Captures balance between hawkish and dovish signals.
    - Forward-looking component predicts future policy rate adjustments and market-based interest rates, with pronounced effects for longer-term OIS contracts.
    - Backward-looking component correlates with contemporaneous policy rates but is less associated with forward-looking market variables.
  - Straightforwardness index
    - Measures extent of unidirectional stance signals versus conditional/hedged language.
    - Straightforwardness declines sharply during systemic stress (global financial crisis, COVID-19 pandemic).
    - Advanced and inflation-targeting economies exhibit lower straightforwardness; emerging and low-income economies favor more explicit statements.
    - Forward-looking communication is consistently less straightforward than backward-looking statements.
  - Explanation index
    - Quantifies how central banks justify and contextualize policy decisions.
    - Explanation rises sharply during tightening and normalization phases—especially in advanced economies—and declines during easing cycles.
    - Explanation levels are higher on average in low-income and pegged exchange rate economies but converge during global monetary cycles.
  - Net confidence index
    - Captures balance between confidence-building and risk-highlighting language.
    - Backward-looking confidence is correlated with implicit market volatility (VIX) predicting subsequent shifts.
    - Forward-looking confidence predicts future market volatility, indicating risk communication actively shapes expectations.

### Cross-country heterogeneity and audience targeting
- Advanced economies heavily use forward-looking net policy sentiment, especially during crises.
- Emerging and low-income economies rely more on backward-looking narratives shaped by prevailing conditions and structural limitations.
- Audience differentiation is structural and persistent:
  - Communications to the general public emphasize confidence-building to reinforce trust and anchor expectations.
  - Communications to governments use more risk-oriented tone to highlight vulnerabilities relevant for fiscal prudence.
  - Communications to business and financial sectors are more balanced to avoid misinterpretation and destabilizing market reactions.
- During crises:
  - Messaging shifts toward building confidence for financial sector and government audiences.
  - Messaging to the general public and international stakeholders becomes more cautious and risk-focused.
- Differentiated tone across audiences remains stable over decades, indicating deliberate, enduring communication strategy.

### Structure of the paper
- Section 2: sample selection criteria, dataset compilation, preprocessing procedures.
- Section 3: analysis of textual form measures (readability and syntactic complexity).
- Section 4: semantic analysis using a fine-tuned classifier; categorization by topic, communication stance, audience, and sentiment.
- Section 5: definitions and applications of semantic textual metrics—net policy sentiment, straightforwardness index, explanation index, net confidence index.
- Section 6: concluding remarks and directions for future research.

*Source: wpiea2025109-print-pdf - 1.  Introduction*

### 2.  The Central Bank Communication Dataset

### 2. The Central Bank Communication Dataset

### 2.1. Data Collection
- Document classification:
  - Regular outlets: annual reports, monetary policy decisions (statements, press releases, and minutes), monetary policy reports, financial stability reports.
  - Non-regular outlets: speeches, specialized reports, press releases.
- Collected from 169 central bank websites to ensure broad representation.
- Language handling:
  - Most documents are in English; dataset also contains documents in several other languages.
  - Only official English translations explicitly published by the central bank were used; otherwise the original local language is processed.
  - Majority of speeches originate from Campiglio et al. (2025) compilation; dataset augments speeches from January 2024 onward using the consolidated dataset maintained by the BIS.
- File formats ingested and preprocessing pipeline:
  - PDFs, Word documents (.docx, .doc), plain text files (.txt), HTML files.
  - Image-based documents (scanned PDFs, embedded images) processed with OCR.
  - HTML documents post-processed to extract core textual content and remove navigation menus, banners, and non-relevant elements.
- Sentence-level processing:
  - Classification performed at the sentence level rather than document level.
  - Language detection applied, followed by language-specific segmentation models to account for language-specific sentence boundary idiosyncrasies (ambiguous abbreviations, varying sentence-ending markers).
  - Minimal textual modifications beyond standardization to preserve semantic content for transformer-based models.

- Collected document counts and timeframe (from source table):
  - Annual Report: 3,879 docs; Begin Year 1884; End Year 2024.
  - Monetary Policy Report: 4,671 docs; Begin Year 1993; End Year 2025.
  - Financial Stability Report: 2,092 docs; Begin Year 1996; End Year 2025.
  - Monetary Policy Decision: 14,238 docs; Begin Year 1936; End Year 2025.
  - Speeches: 36,725 docs; Begin Year 1986; End Year 2025.
  - Other Documents: 11,437 docs; Begin Year 1993; End Year 2025.
- Notes on nomenclature and contents:
  - Monetary policy reports are called inflation reports in some economies; some countries publish both.
  - Financial stability reports are called financial stability reviews or financial stability surveillance in some economies.
  - Monetary Policy Decisions include press releases, official statements, and meeting minutes.
  - Other Documents encompass bulletins, balance of payments reports, economic and monetary reports, monthly reviews, economic outlooks, monetary policy data, macroeconomic reviews, market operations, and monetary policy specialized publications.

### 2.2. Data Exploratory Analysis — Corpus Scope and Coverage
- Corpus size and composition:
  - Total central bank documents in dataset: 74,882.
  - Regular documents: 24,880.
  - Corpus size: over 80 GB of textual information.
  - Processed PDF pages: more than 1.3 million (excluding other file formats).
  - Processed sentences: approximately 21 million.
- Language distribution (sentences):
  - English: 87.4 percent.
  - Spanish: 6.6 percent.
  - Portuguese: 1.7 percent.
  - French: 1.4 percent.
  - Other languages include Arabic, Russian, Chinese, and others (dataset spans over 30 languages).
- Temporal and cross-country coverage:
  - Bulgaria: most extensive time series of annual reports, dating back to 1884.
  - United States: longest recorded time series of monetary policy minutes and policy actions, dating back to 1936, with relatively high frequency.
  - United Kingdom: most extensive series of financial stability reports, dating back to 1996.
  - Earliest monetary policy reports in dataset: Sweden and the United Kingdom, both dating back to 1993.
- Observed temporal imbalances:
  - Some central banks maintain long, consistent records allowing for decades-long studies of evolving communication strategies.
- Publication trends and institutional drivers (summary of figures and narrative):
  - Increasing number of central banks publishing monetary policy reports (MPRs) and financial stability reports (FSRs) over time.
  - Monetar y policy reports increased in prevalence with the adoption of inflation-targeting frameworks.
  - Number of central banks publishing financial stability reports jumped after the global financial crisis.
  - Financial stability reports are more prevalent than monetary policy reports in economies with exchange rate anchor regimes (example economies cited: Brunei Darussalam, Nepal, Singapore), where limited discretion in monetary policy reduces the need for dedicated MPRs.

- Document length trends (Figure 2 summary):
  - Document types analyzed: annual reports, financial stability reports, monetary policy reports, monetary policy decisions, speeches.
  - Annual reports grew significantly in length over time, particularly among low-income and pegged economies.
  - Advanced, emerging, and inflation-targeting economies have streamlined annual reports since 2018.
  - Financial stability and monetary policy reports: longer formats in advanced and inflation-targeting economies; shorter but gradually expanding in other economies.
  - Monetary policy decisions were being streamlined in advanced and inflation-targeting economies from 2000 until the COVID-19 pandemic; this trend reversed after the global inflation surge, leading to more detailed narrative explanations.
  - Speeches exhibited a U-shaped pattern: declined initially, then increased as central banks used speeches more for forward guidance and stakeholder engagement.
  - Speech length is consistently higher in advanced and inflation-targeting economies.

### 3. Pro-Forma Analysis of Central Bank Communications
- Analytical focus:
  - Lexical readability metrics (e.g., Flesch-Kincaid Ease scores).
  - Structural complexity indicators (e.g., syntactic dependency depth).
  - Limitations: these pro-forma measures assess surface accessibility and syntactic parsing effort but do not capture semantic content, policy intent, or rhetorical strategy.

- Lexical readability (Figure 3) — English documents only:
  - Metric: Flesch-Kincaid Ease Score; formula provided in source:
    - Score = 206.835 − 1.015 × (Total Words / Total Sentences) − 84.6 × (Total Syllables / Total Words)
    - Score range: 0 to 100 (higher = better readability).
    - Interpretation in English-language texts: scores above 60 correspond to an 8th-grade reading level; values between 30 and 50 suggest college-level complexity; scores below 30 indicate highly technical or specialized content.
  - Findings:
    - Pegged exchange rate regimes and low-income economies display higher Flesch Reading Ease scores (simpler language), driven by shorter sentences and simpler word choices.
    - Inflation-targeting and advanced economies score lower in readability (more complex vocabulary), particularly in monetary policy and financial stability reports.
    - Widening readability gap across development levels suggests increased complexity of mandates is associated with more complex public communication.
  - Caveat: cross-linguistic structural variations affect direct comparability; only English documents were used for this part.

- Syntactic complexity — dependency depth (Figure 4):
  - Metric: syntactic dependency depth measured as the longest path from the root of a sentence’s syntactic tree to any terminal node; formal definition provided in source.
  - Illustrative examples:
    - Simple sentence “Central banks adjust interest rates” has a depth of 2.
    - More complex sentence “Central banks, in response to inflationary pressures, swiftly adjust interest rates to maintain stability” exhibits a depth of 4.
  - Findings:
    - Advanced economies use more complex vocabulary but tend to structure messages with simpler sentence constructions (lower dependency depth) across annual reports, financial stability reports, and monetary policy reports.
    - Low-income and pegged economies often exhibit deeper sentence structures despite addressing less technically complex issues, suggesting more convoluted exposition.
    - The gap in sentence structure complexity widened in monetary policy decisions since 2020:
      - Advanced economies simplified sentence structure amid post-pandemic uncertainty.
      - Low-income economies increased sentence complexity, with a sharp rise beginning in 2022 that coincides with the global surge in inflation.
    - Implication: structural simplicity paired with technically rich content is associated with more effective communication; institutional capacity constraints in low-income economies may undermine transparency during macroeconomic stress.

- Lexical vs syntactic complexity relationship (Figure 5):
  - Observed patterns:
    - Advanced economies: combine more complex vocabulary (lower Flesch Reading Ease) with clearer sentence structure (lower dependency depth).
    - Low-income and many non-inflation-targeting economies: simpler vocabulary but more convoluted sentence constructions, which may impede clarity and communicative effectiveness.
  - Visualization approach: scatter plots of lexical (Flesch-Kincaid Ease) vs syntactic (dependency depth) complexity by development level and monetary policy framework; each dot represents country-specific average values across communication types.

- Limitations of pro-forma indicators:
  - Lexical and syntactic measures do not capture meaning, intent, or rhetorical strategy.
  - The next analytic step introduced in the source is a semantic analysis framework based on a large language model fine-tuned for central bank communication to assess policy narratives, forward guidance, and institutional messaging strategies.  

*Source: 2. The Central Bank Communication Dataset (content unit).*

### 4.  Semantic Analysis of Central Bank Communications

### 4. Semantic Analysis of Central Bank Communications

### Methodology (overview)
- Develop an LLM-based sentence-level central bank communication classification methodology and apply it to a large multilingual corpus of central bank communications.
- Key methodological steps: selecting an LLM; constructing a labeled dataset; designing a classification structure; fine-tuning and validating the model; benchmarking out-of-sample performance; applying the classifier to the full corpus.

### Selection of the LLM
- Rationale for model choice:
  - Prefer an encoder-only sentence transformer over decoder-only autoregressive models (e.g., GPT) for sentence-level semantic tasks (classification, textual entailment, document similarity).
  - Sentence transformers produce fixed-size, semantically meaningful embeddings directly optimized for classification and similarity (Reimers & Gurevych, 2019).
  - Bidirectional encoder architectures enable full contextualization by attending to both left and right contexts (Devlin et al., 2019).
  - Encoder-only models offer greater computational and cost efficiency for large-scale inference (millions of sentences).
  - Easier adaptability over time through incremental or complete fine-tuning compared to instruction tuning or prompt re-engineering for decoder-only models.
- Multilingual considerations:
  - Central bank documents include native-language reports; ignoring multilingualism risks selection bias (example: Mexico and Chile annual reports dating back to 1925 and 1926 with English only from 2000).
  - Selected model: multilingual BGE (bge-m3) sentence transformer trained for cross-lingual sentence embeddings that align semantically similar statements regardless of language (Chen et al., 2024).
  - Cross-lingual consistency mitigates distortions and biases introduced by translation models (Conneau et al., 2020). Appendix B explores cross-lingual consistency.

### Labeled Dataset Construction
- Supervised fine-tuning dataset composition:
  - Total labeled sentences: 1,200 annotated sentences.
  - Initial synthetic generation: approximately 240 instances (20 percent) generated by a generative AI model to satisfy (i) semantic differentiation across labels and (ii) intra-label diversity. (See Prompt A.1.)
  - Expert-crafted examples: approximately 240 examples (20 percent) created by three domain experts after reviewing chatbot output.
  - Real-document sentences: 720 examples (60 percent) extracted from central bank documents and manually annotated by domain experts.
- Training/validation split:
  - Training set includes: all generated sentences (240), expert-constructed sentences (240), and half of real extracted sentences (360) → total training examples = 840.
  - Hold-out validation set: remaining 360 real sentences (validation composed exclusively of real central bank content to align with deployment distribution).
- Motivation for hybrid strategy:
  - Synthetic and expert-crafted sentences provide semantic coverage and balanced taxonomy representation.
  - Real sentences provide empirical grounding and mirror deployment distribution.
  - Cites representation learning and domain generalization literature supporting training on broader/augmented datasets while evaluating on target distribution (Ben-David et al., 2010; Hendrycks et al., 2021; Recht et al., 2019).

### Classification Structure
- Label set Y = (ytopic, ystance, yaudience, ysentiment).
- Full joint estimation P(ytopic, ystance, yaudience, ysentiment | x) infeasible due to combinatorial explosion of sparse four-way combinations.
- Adopted block factorization:
  - P(ytopic, ystance, yaudience, ysentiment | x) ≈ P(ytopic, ystance | x) · P(yaudience, ysentiment | x). (Equation (1))
- Choice of paired classifiers:
  - Classifier A jointly predicts topic and communication stance.
  - Classifier B jointly predicts audience and sentiment.
  - Selection informed by empirical co-occurrence patterns (Z-scores of pairwise empirical coupling); topic + guidance (stance) exhibited the highest absolute Z-score, followed by topic + audience, then audience + sentiment. The disjoint pair with highest combined absolute Z-score chosen: ⟨topic, guidance⟩ and ⟨audience, sentiment⟩.
- Taxonomy and labeling rules:
  - Topic taxonomy covers monetary policy (subdivided: interest rates, inflation, balance sheet size including asset purchase programs), financial stability, supervision and regulation, payments, structural economic issues, etc.
  - Topic labeling based on economic origin rather than effect (example: sentence referencing borrowing costs and pressure on firms classified under Monetary Policy (Interest Rate) rather than Financial Stability).
  - Communication stance categorized as backward-looking or forward-looking; stance determined by substantive content orientation, not solely verb tense (example: “Last year’s findings emphasized the need for continued efforts…” → Financial Inclusion (Forward-Looking)).
  - Audience taxonomy captures recipients such as financial markets/financial sector, business sector, general public, government officials, international partners.
  - Sentiment categories include nuanced tones (e.g., dovish, hawkish, neutral/balanced, confidence-building, risk-highlighting).
- Metadata class:
  - An additional “Metadata” class captures sentences without economic meaning (figure captions, formalities, acknowledgments, boilerplate disclaimers, procedural statements).
  - In final classified results: “Metadata” comprised approximately 6.3 percent of total for topic and communication stance classification, and 14.1 percent for audience and sentiment classification.

### Fine-Tuning Setup
- Two-phase fine-tuning pipeline (Tunstall et al., 2022):
  1. Siamese contrastive training of the sentence transformer to learn fixed-length, dense sentence embeddings (bge-m3 → 1024-dimensional vector space).
     - Contrastive objective (hard triplet loss) optimizes cosine similarity to bring semantically similar pairs closer and push dissimilar pairs apart.
     - Using sentence pairs yields up to N(N−1)/2 training pairs from dataset of size N (contrastive learning enriches optimization signal).
     - Example: training dataset of 840 examples corresponds to 352,380 pairs if no up-sampling were done.
  2. Train a dense neural network (softmax output layer) mapping embeddings to classification labels; output neurons correspond to number of classes.
- Training details and hyperparameter tuning:
  - Distance metric: cosine similarity.
  - Loss function for embedding training: hard triplet (contrastive).
  - No oversampling or undersampling within pair generation since class distribution in labeled dataset is balanced.
  - Hyperparameters tuned:
    - learning rate searched within [10^{-5}, 10^{-2}],
    - L2-norm regularization searched within [10^{-6}, 10^{0}].
  - Optimizer: variant of Adam that decouples weight decay from adaptive gradient updates.
  - Learning rate scheduler: warm-up phase equal to 10 percent of an epoch followed by decay.
  - Model evaluation every 500 steps on validation set; select model minimizing embedding loss in the first phase and maximizing out-of-sample accuracy in the second phase.

### Out-of-Sample Performance (benchmarking vs ChatGPT 4o)
- Weakly supervised evaluation protocol:
  - ChatGPT 4o instructed to assign one label for each of the four sentence-level dimensions (topic, communication stance, audience, sentiment) from fixed allowable classes; invalid responses repeated until valid output obtained. (See Prompt A.2.)
- Cost considerations (processing full dataset ~21 million sentences):
  - Pricing as of April 2025:
    - GPT-4o: $2.50 per million input tokens and $10.00 per million output tokens.
    - GPT-4.5: $75.00 per million input tokens and $150.00 per million output tokens.
  - Assumptions: each prompt contains 250 words (~333 tokens) and each response contains 10 words (~13 tokens).
  - Cost per prompt–response interaction:
    - GPT-4o: $0.00096.
    - GPT-4.5: $0.02693.
  - Processing 21 million interactions cost estimates:
    - GPT-4o: approximately $20,160.
    - GPT-4.5: approximately $565,530.
  - Practical limitations of using generative LLMs at scale: batching may reduce API calls but increases hallucination risk; reprocessing entire dataset required for schema changes; outputs are non-reproducible and non-transferable across users/time.
  - Argument: trained classification model is reproducible, version-controlled, and shareable for transparent consistent results (contrast with commercial LLM limitations; Gambacorta et al., 2024).
- Comparative performance (validation set) — Table 2 metrics (Accuracy, Precision, Recall, F1; Macro and Micro; Cohen’s Kappa):
  - Topic
    - Our Classifier: Accuracy 0.689; Precision (Macro) 0.699; Recall (Macro) 0.645; F1 (Macro) 0.650; Precision (Micro) 0.689; Recall (Micro) 0.689; F1 (Micro) 0.689; Cohen’s Kappa 0.666.
    - ChatGPT 4o: Accuracy 0.731; Precision (Macro) 0.593; Recall (Macro) 0.535; F1 (Macro) 0.551; Precision (Micro) 0.731; Recall (Micro) 0.731; F1 (Micro) 0.731; Cohen’s Kappa 0.711.
  - Communication stance
    - Our Classifier: Accuracy 0.924; Precision (Macro) 0.925; Recall (Macro) 0.910; F1 (Macro) 0.917; Precision (Micro) 0.924; Recall (Micro) 0.924; F1 (Micro) 0.924; Cohen’s Kappa 0.834.
    - ChatGPT 4o: Accuracy 0.828; Precision (Macro) 0.591; Recall (Macro) 0.551; F1 (Macro) 0.570; Precision (Micro) 0.828; Recall (Micro) 0.828; F1 (Micro) 0.828; Cohen’s Kappa 0.658.
  - Audience
    - Our Classifier: Accuracy 0.706; Precision (Macro) 0.713; Recall (Macro) 0.699; F1 (Macro) 0.700; Precision (Micro) 0.706; Recall (Micro) 0.706; F1 (Micro) 0.706; Cohen’s Kappa 0.622.
    - ChatGPT 4o: Accuracy 0.506; Precision (Macro) 0.529; Recall (Macro) 0.463; F1 (Macro) 0.437; Precision (Micro) 0.506; Recall (Micro) 0.506; F1 (Micro) 0.506; Cohen’s Kappa 0.400.
  - Sentiment
    - Our Classifier: Accuracy 0.700; Precision (Macro) 0.612; Recall (Macro) 0.595; F1 (Macro) 0.594; Precision (Micro) 0.700; Recall (Micro) 0.700; F1 (Micro) 0.700; Cohen’s Kappa 0.589.
    - ChatGPT 4o: Accuracy 0.704; Precision (Macro) 0.557; Recall (Macro) 0.393; F1 (Macro) 0.418; Precision (Micro) 0.704; Recall (Micro) 0.704; F1 (Micro) 0.704; Cohen’s Kappa 0.592.
- Key out-of-sample findings:
  - Our classifier achieves higher macro-level metrics (notably macro F1) across most dimensions, indicating better performance on minority/infrequent classes.
  - Largest performance gaps in communication stance and audience dimensions where domain-specific fine-tuning is most beneficial.
  - ChatGPT 4o attains slightly higher accuracy and micro-F1 on topic due to strong performance on frequent classes, but has markedly lower macro-F1.
  - Sentiment: both models similar on accuracy and micro-F1, but our classifier achieves considerably higher macro-F1, indicating better detection of less prevalent sentiment categories (e.g., hawkish, dovish).
  - Error analysis: many misclassifications stem from subtle overlaps across closely related policy categories (e.g., open market operations vs inflation vs balance sheet size; fiscal vs monetary policy references; financial stability vs supervision and regulation).

### Empirical Application on the Central Bank Communications Dataset
- Unit of analysis: sentence within a document.
- Aggregation: semiannual level for visual inspections; within-year interpolation used when necessary for less frequent documents (e.g., annual reports). (See footnote explanation.)
- Example application: Table 3 — classification output for the ECB’s monetary policy decision published on December 12, 2024. Selected sentence-level outputs (each row: Sentence → Topic → Communication Stance → Audience → Sentiment):
  - "PRESS RELEASE Monetary policy decisions 12 December 2024 The Governing Council today decided to lower the three key ECB interest rates by 25 basis points." → MP - interest rate | Backward-looking | Financial Sector | Dovish
  - "In particular, the decision to lower the deposit facility rate – the rate through which the Governing Council steers the monetary policy stance – is based on its updated assessment of the inflation outlook, the dynamics of underlying inflation and the strength of monetary policy transmission." → MP - interest rate | Forward-looking | Financial Sector | Dovish
  - "The disinflation process is well on track." → MP - inflation | Forward-looking | Financial Sector | Confidence-building
  - "Most measures of underlying inflation suggest that inflation will settle at around the Governing Council’s 2% medium-term target on a sustained basis." → MP - inflation | Forward-looking | Financial Sector | Neutral/ Balanced
  - "Domestic inflation has edged down but remains high, mostly because wages and prices in certain sectors are still adjusting to the past inflation surge with a substantial delay." → MP - inflation | Backward-looking | General Public | Risk-highlighting
  - "Financing conditions are easing, as the Governing Council’s recent interest rate cuts gradually make new borrowing less expensive for firms and households." → MP - interest rate | Forward-looking | Business Sector | Dovish
  - "But they continue to be tight because monetary policy remains restrictive and past interest rate hikes are still transmitting to the outstanding stock of credit." → MP - interest rate | Backward-looking | Business Sector | Risk-highlighting
  - "Staff now expect a slower economic recovery than in the September projections." → MP - economic activity | Forward-looking | Business Sector | Neutral/ Balanced
  - "Although growth picked up in the third quarter of this year, survey indicators suggest it has slowed in the current quarter." → MP - economic activity | Backward-looking | Business Sector | Risk-highlighting
  - "The projected recovery rests mainly on rising real incomes—which should allow households to consume more—and firms increasing investment." → MP - economic activity | Forward-looking | General Public | Confidence-building
  - "Over time, the gradually fading effects of restrictive monetary policy should support a pick-up in domestic demand." → MP - interest rate | Forward-looking | General Public | Dovish
  - "The Governing Council is determined to ensure that inflation stabilises sustainably at its 2% medium-term target." → MP - inflation | Forward-looking | Financial Sector | Neutral/ Balanced
  - "It will follow a data-dependent and meeting-by-meeting approach to determining the appropriate monetary policy stance." → MP - interest rate | Forward-looking | Financial Sector | Confidence-building
  - "The Governing Council is not pre-committing to a particular rate path." → MP - interest rate | Forward-looking | Financial Sector | Neutral/ Balanced

*Italicized source attribution: Content derived from "4. Semantic Analysis of Central Bank Communications" (source PDF: wpiea2025109-print-pdf).*

### 7.5 billion

### wpiea2025109-print-pdf - 7.5 billion

### Classifier performance and sentence-level reliability
- Appendix B: classification framework performs consistently across languages, with only minor differences between original and translated documents.
- Appendix C: predictions are typically made with high confidence; the classifier assigns unambiguous labels in the large majority of cases.
- Appendix D: sentences involving multiple overlapping classifications are relatively rare and do not materially affect interpretation.
- Aggregate conclusion: sentence-level classifications are linguistically robust, highly confident, and predominantly unambiguous, supporting aggregation across countries and over time.

### ECB communication excerpts and monetary policy operational stance
- “7.5 billion per month on average.”
- Governing Council decision: “The Governing Council will discontinue reinvestments under the PEPP at the end of 2024.”
- Governing Council readiness: “The Governing Council stands ready to adjust all of its instruments within its mandate to ensure that inflation stabilises sustainably at its 2% target over the medium term and to preserve the smooth functioning of monetary policy transmission.”
- Transmission Protection Instrument (TPI): “available to counter unwarranted, disorderly market dynamics that pose a serious threat to the transmission of monetary policy across all euro area countries, thus allowing the Governing Council to more effectively deliver on its price stability mandate.”
- Press conference timing: “The President of the ECB will comment on the considerations underlying these decisions at a press conference starting at 14:45 CET today.”
- Contact: “European Central Bank Directorate General Communications Sonnemannstrasse 20 60314 Frankfurt am Main, Germany +49 69 1344 7455 media@ecb.europa.eu Reproduction is permitted provided that the source is acknowledged.”

### Semantic mapping of central bank communications (Figure 9) and topic structure
- Data span: “1884 to 2025.”
- Embedding and projection: original 1024-dimensional sentence embeddings reduced to two dimensions using the t-SNE nonlinear dimensionality reduction technique (van der Maaten & Hinton, 2008); reduction performed in an unsupervised way.
- Each dot represents a sentence; proximity indicates semantic similarity.
- Observed patterns:
  - Clear clustering patterns with well-defined topic boundaries.
  - Overlap between forward-looking supervision and regulation and financial stability.
  - Forward- and backward-looking statements within the same topic typically appear adjacent.
  - “Traditional” topics (inflation, fiscal policy, exchange rates, interest rates) exhibit broad semantic variation.
  - “Emerging” topics (climate change, technological innovation) show less dispersion and repeat similar sentence types.
  - Classes positioned in the middle (emerging topics) are semantically more similar to other classes.

### Topic composition by communication outlet (Figure 10) and historical/topical dynamics
- Monetary policy: dominant topic in decisions and reports, aligning with price stability mandate.
- Financial stability reports: allocate most attention to financial stability, supervision, and regulation.
- Crisis management and fiscal policy: episodic spikes during economic/financial crises.
- Emerging topics (climate change, technological innovation): gained space in most documents in recent years.
- Speeches: display the most thematic diversity and dedicate more attention to emerging issues than other outlets.

### Topic composition by level of development (Figure 11) and country heterogeneity
- Level of development classification: taken from the IMF AREAER dataset (International Monetary Fund, 2025).
- Cross-economy similarities: suite of topics is remarkably similar across economies.
- Notable differences:
  - Advanced economies: more attention to financial stability.
  - Emerging and low-income economies: greater importance on fiscal policy; low-income countries show pronounced fiscal emphasis due to limited market depth and reliance on central bank financing.
- Economy-level heterogeneity: example analyses for United Kingdom and United States show stable long-run monetary policy share but historical surges in specific topics (governance, crisis management, financial stability).

### Monetary policy subtopics, inflation targeting (IT) adoption, and communication stance (Figures 13–15)
- Advanced economies: focus on signaling monetary policy stance; prominent share dedicated to interest rates; more forward-looking communication.
- Emerging and low-income economies: prioritize inflation over interest rates.
- Declining emphasis on exchange rates across all economies; exchange rate discussions remain more relevant in low-income economies.
- IT adoption effects:
  - Shift away from exchange rate discussions toward inflation, interest rates, and economic activity.
  - Increased forward-looking references to economic activity.
- Figure 15: IT-adopting economies listed include Brazil, Chile, Georgia, Republic of Kazakhstan, Republic of Korea, Mexico, Republic of Moldova, New Zealand, Paraguay, Peru, the Philippines, Russian Federation, Seychelles, Sri Lanka, Uganda, and Ukraine.

### Forward-lookingness: definition, trends, and outlets (Figures 16–18)
- Forward-lookingness Score (Equation (2)):
  - Forward-lookingness Score_{c,d,t} = #Forward-Looking_{c,d,t} / (#Forward-Looking_{c,d,t} + #Backward-Looking_{c,d,t})
  - #Forward-Looking and #Backward-Looking are the counts of forward-looking and backward-looking sentences, respectively.
- Figure 16a (by calendar year): forward-looking communication has increased across most outlets; speeches and monetary policy decisions have the highest forward-looking scores.
- Figure 16b (by elapsed months since first publication): aligning by months since first document confirms trend is systematic rather than sample composition artifact.
- Figure 17: advanced economies consistently exhibit above-median forward-looking communication across report types; pegged economies and low-income countries generally remain below the global median, particularly in financial stability reports, monetary policy decisions, and monetary policy reports.
- Figure 18:
  - Core topics: forward-looking communication increased across nearly all traditional/core topics, notably within monetary policy.
  - Crisis periods (dot-com bust, Global Financial Crisis, COVID-19 pandemic): associated with declines in forward-lookingness in crisis-related communication.
  - Emerging topics (climate change, technological innovation, structural economic reform): marked increases in forward-lookingness.

### Audience targeting in communications (Figure 19)
- Primary audiences tracked: Business Sector, Financial Sector, General Public, Government, International Stakeholders.
- Financial sector remains primary audience across groups but its dominance is declining, especially in advanced economies.
- Government-targeted messaging: inverse relationship with level of development; low-income and emerging economies direct a larger share toward government.
- Household targeting: share of communication directed at households has increased across all groups.
- Advanced economies: similar shares directed at households and businesses.
- Emerging and low-income economies: businesses receive more attention than the general public.

### Sentiment composition and risk communication (Figure 20)
- Sentiment categories tracked: Confidence-building, Dovish, Hawkish, Neutral/Balanced, Risk-highlighting.
- Neutral/balanced content: prevalence inversely related to level of development; low-income countries rely heavily on neutral language.
- Risk-highlighting: share broadly proportional to level of development; more mature economies emphasize potential downside risks.
- Confidence-building statements: relatively stable across development groups.
- Dovish vs hawkish asymmetry: dovish content appears more frequently than hawkish across all development groups.

*Source: wpiea2025109-print-pdf - 7.5 billion*

### 5.  Communication Metrics and Their Connection with Financial Variables

### 5.  Communication Metrics and Their Connection with Financial Variables

### 5.1. Methodology
- Textual metrics are constructed from sentence-level classifier outputs to form document-level indicators.
- Metrics defined (using only monetary policy decisions unless noted):
  - Net Policy Sentiment (NPS): 푁푃푆 = (퐻 − 퐷) / (퐻 + 퐷)
  - Straightforwardness Index (SI): 푆퐼 = (푁 + |퐻 − 퐷|) / (푁 + 퐻 + 퐷)
  - Explanation Index (EI): 퐸퐼 = (퐶 + 푅 + 푁) / (퐻 + 퐷)
  - Net Confidence Index (NCI) — evaluated on all regular central bank documents: 푁퐶퐼 = (퐶 − 푅) / (퐶 + 푅)
- Notation: 퐻, 퐷, 퐶, 푅, and 푁 represent counts of hawkish, dovish, confidence-building, risk-highlighting, and neutral sentences, respectively.
- Temporal decomposition:
  - Forward- and backward-looking components for NPS and NCI:
    - 푁푃푆푓푤푑 = (퐻푓푤푑 − 퐷푓푤푑) / (퐻푓푤푑 + 퐷푓푤푑)
    - 푁푃푆푏푤푑 = (퐻푏푤푑 − 퐷푏푤푑) / (퐻푏푤푑 + 퐷푏푤푑)
    - 푁퐶퐼(푠) = (퐶(푠) − 푅(푠)) / (퐶(푠) + 푅(푠)), 푠 ∈ {Forward, Backward}
  - Aggregation weights:
    - 푁푃푆 = 휔푓푤푑 · 푁푃푆푓푤푑 + 휔푏푤푑 · 푁푃푆푏푤푑, where 휔푓푤푑 = (퐻푓푤푑 + 퐷푓푤푑) / (퐻 + 퐷)
    - 푁퐶퐼 = 휔(Forward) · 푁퐶퐼(Forward) + 휔(Backward) · 푁퐶퐼(Backward), where 휔(푠) = (퐶(푠) + 푅(푠)) / (퐶 + 푅)

### 5.1.1. Net Policy Sentiment — interpretation and decomposition
- Purpose:
  - Quantifies directional stance by balancing hawkish and dovish sentences; ranges in [−1, 1].
  - Distinguishes communication stance from implemented monetary policy instruments.
- Forward- versus backward-looking components:
  - Forward-looking NPS (푁푃푆푓푤푑) serves as a text-derived proxy for forward guidance.
  - Backward-looking NPS (푁푃푆푏푤푑) captures retrospective narrative and justifications.
  - Overall NPS is a weighted linear combination of these components using 휔푓푤푑 and 휔푏푤푑.
- Theoretical and empirical relevance:
  - High 휔푓푤푑 indicates emphasis on guiding expectations; high 휔푏푤푑 indicates emphasis on past/current assessment.
  - Disaggregation enables study of ex-post justification versus ex-ante guidance effects on markets.

### 5.1.2. Straightforwardness Index (SI)
- Definition: 푆퐼 = (푁 + |퐻 − 퐷|) / (푁 + 퐻 + 퐷), range [0,1].
- Interpretation:
  - Higher SI → clearer, more unidirectional messaging; SI ≈ 1 implies dominance of one sentiment.
  - Lower SI → coexistence of conflicting signals or presentation of multiple scenarios.
- Decomposition by temporal orientation:
  - 푆퐼(푠) = (푁(푠) + |퐻(푠) − 퐷(푠)|) / (푁(푠) + 퐻(푠) + 퐷(푠)), 푠 ∈ {Forward, Backward}
  - Backward-looking communications are typically more straightforward; forward-looking communications can legitimately have lower SI due to conditionality and scenario-based guidance.
- Cross-country interpretation:
  - Lower forward-looking SI in advanced economies may reflect sophisticated conditional guidance.
  - Very low forward-looking SI in emerging and low-income countries may reflect limited capacity to articulate future guidance or inherent policy uncertainty.

### 5.1.3. Explanation Index (EI)
- Definition: 퐸퐼 = (퐶 + 푅 + 푁) / (퐻 + 퐷)
- Purpose:
  - Measures explanatory content (confidence-building, risk-highlighting, neutral statements) relative to directional statements.
  - Higher EI → more justification and contextualization of policy decisions.
- Comparisons with other proxies:
  - Provides a functional alternative to proxies like document length or readability scores by focusing on sentence roles rather than verbosity or linguistic complexity.
- Practical interpretation:
  - Low EI can reflect deliberate concision (e.g., during high uncertainty like COVID-19) rather than poor communication.

### 5.1.4. Net Confidence Index (NCI)
- Definition: 푁퐶퐼 = (퐶 − 푅) / (퐶 + 푅), range [−1, 1].
- Interpretation:
  - Higher NCI → optimism/confidence; lower NCI → caution/risk emphasis.
- Temporal decomposition:
  - 푁퐶퐼(Forward) and 푁퐶퐼(Backward) measure forward- and backward-looking tone about risks and resilience.
  - Overall NCI expressed as weighted average per Eq. (12).
- Uses across document types:
  - In monetary policy decisions, captures macro-financial concerns (e.g., inflation risks).
  - In financial stability reports, reflects systemic vulnerability assessments.
  - In broader documents, conveys composite institutional confidence across policy areas.

### 5.2. Empirical Application on the Central Bank Communications Data

#### 5.2.1. Net Policy Sentiment — descriptive findings
- Construction details:
  - Series weighted by country-level nominal GDP (U.S. dollars).
  - Curves standardized within each group (advanced, emerging, low-income) to express deviations in standard deviation units.
- Three key descriptive findings:
  - Forward-looking NPS frequently anticipates changes in policy rates, especially in advanced economies.
  - Forward-looking NPS captures policy communication beyond policy rate changes (e.g., during effective lower bound episodes, balance sheet policies, asset purchases).
  - Alignment between NPS and realized policy rates varies by development level:
    - Strong alignment in advanced economies.
    - Weaker relationship in emerging and low-income economies due to limited forward guidance use, weaker transmission, and credibility vulnerabilities.

#### 5.2.2. Country examples and money market comparisons
- United States:
  - Forward- and backward-looking NPS roughly track the federal funds rate once available; notable alignment during 2001, 2008 recessions, and COVID-19 tightening cycle.
- Inflation-targeting emerging economies (selected sample: Brazil, Ghana, Georgia, Iceland, Mexico, Chile, Russia, Uruguay):
  - Forward-looking NPS generally closely tracks short-term interbank money market rates.
  - Observed delayed or muted responses in some cases due to liquidity management, signaling without incurring costs, less liquid markets, and credibility differences.

#### 5.2.3. Panel-data analysis — internal consistency between communication and policy (specification and results)
- Consistency specification (monthly panel):
  - Equation (13): Policy Rate_{i,t} = 훼_i + 휆_t + 휂 Policy Rate_{i,t−1} + 훽 NPS_{i,t} + X′_{i,t} γ + ε_{i,t}
  - Controls: lagged policy rate, CPI inflation, exchange rate (USD/local), Explanation Index, Straightforwardness Index, Net Confidence Index.
  - Variables standardized by country; standard errors clustered at country level; country and time fixed effects included.
  - Timing: NPS recorded as of monetary policy decision date (on or before last day of month); policy rate constructed using end-of-month values (avoids lookahead bias).
- Main results (Table 5 highlights):
  - Specification (I): coefficient on total NPS = 0.031*** (standard error 0.005).
    - Interpretation: with sample standard deviation of policy rate = 6.36 percentage points (Table 1, Appendix E), a one-standard-deviation increase in total NPS associated with ≈ 0.20 p.p. (20 basis points) increase in policy rate (0.031 × 6.36 ≈ 0.20 p.p.).
    - Context: absolute policy rate change is 25 basis points or lower for 75 percent of observed changes in sample.
  - Specification (II) decomposition:
    - Forward-looking NPS coefficient = 0.026*** (0.004).
    - Backward-looking NPS coefficient = 0.015*** (0.004).
    - Forward-looking coefficient relatively larger than backward-looking.
  - Other controls (selected coefficients from Table 5):
    - Policy Rate_{i,t−1} ≈ 0.970*** (0.005) across specs.
    - Net Confidence Index: not statistically significant in main specs (e.g., −0.002 (0.003)).
    - Explanation Index and Decisiveness Index: generally small and not significant in main specs.
  - Sample and robustness:
    - Specifications (III)–(IV) restrict to countries with ≥ 10 years of data; results remain significant.
    - Observations: 52095 (Spec I), 67391 (Spec II), 33807 (Spec III), 3807 (Spec IV) as reported.
    - R^2 reported: 0.970, 0.969, 0.980, 0.980 for respective specs.

#### 5.2.4. Heterogeneity across monetary frameworks (Table 6)
- Regression stratified by framework: Inflation-targeting, Monetary aggregate, Other framework, Exchange rate anchor.
- Key findings:
  - Inflation-targeting regimes:
    - Total NPS coefficient = 0.029*** (0.006) (Spec I).
    - Forward-looking NPS = 0.024*** (0.005) (Spec II).
    - Backward NPS = 0.014*** (0.004) (Spec III).
    - Strongest association between NPS and policy rates found in inflation-targeting regimes.
  - Monetary aggregate and Other frameworks:
    - Mixed results; some statistically significant coefficients but heterogeneity across specs.
    - Example: Monetary aggregate Forward-looking coefficient = 0.080** (0.029) in one spec but small sample heterogeneity noted.
  - Exchange rate anchor regimes:
    - No significant relationship; e.g., Total NPS coefficient = −0.068 (0.065) (Spec IV).
    - Interpretation: policy rate subordinated to exchange rate objective, reducing communicative influence on domestic rate setting.
- Additional selected coefficients (from Table 6):
  - Net Confidence Index: negative and weakly significant in some inflation-targeting specs (e.g., −0.006* (0.003)).
  - Exchange Rate (USD/local): negative and significant in inflation-targeting specs (e.g., −0.012** (0.005)).

#### 5.2.5. Predictive analysis for future rate changes (Eq. 14 and Table 7)
- Predictive specification:
  - ΔRate_{i,t+1} = 훼_i + 휆_t + 훽 Net Policy Sentiment_{i,t} + X′_{i,t} γ + ε_{i,t}
  - ΔRate denotes one-period-ahead change in policy rate, short-term (T-bill) rate, or long-term (T-bond) rate.
  - Variants include total NPS, forward- and backward-looking components, and gap between forward- and backward-looking sentiment.
- Main predictive results (Table 7 summary):
  - Forward-looking sentiment is consistently positive and statistically significant for future changes in:
    - Policy rates (Specs I–III).
    - Short-term T-bill rates (Specs IV–VI), particularly over the first five leads.
  - The gap between forward- and backward-looking sentiment is statistically significant for policy rates, indicating markets respond to relative emphasis on forward-looking communication.
  - Backward-looking sentiment shows no consistent predictive power for future short-term or long-term market rates.
  - Long-term bond rates:
    - Forward-looking sentiment loses statistical significance beyond the first lead; muted effects consistent with term premia and broader macro-fiscal influences diluting short-term signals.
- Interpretation of magnitudes (note from text):
  - Coefficients are in standardized units; example conversions:
    - Coefficient of 0.083 for policy rate changes implies a one-standard-deviation increase in forward-looking sentiment → ≈ 7 basis points change in policy rate, given sample SD of policy rate changes = 0.87 percentage points (Table 1, Appendix E).
    - For T-bill rates, coefficient of 0.062 corresponds to ≈ 8 basis points movement based on sample SD (text notes; numerical conversion referenced to Appendix E).

#### 5.2.6. Horizon and persistence (Figure 24 summary)
- Multi-horizon regressions (1 to 10 policy decisions ahead) show:
  - Forward-looking sentiment:
    - Statistically significant association with future policy rate changes up to several leads, peaking in early horizons.
    - Significant association with short-term market rate (T-bill) changes particularly over the first five leads.
    - Loses significance for long-term bond rate changes beyond the first lead.
  - Backward-looking sentiment:
    - Fails to explain future short-term or long-term market rate movements; weak associations only at very short horizons.

### Key empirical takeaways
- Net Policy Sentiment (NPS) is:
  - Statistically significantly associated with contemporaneous policy rates (Table 5: total NPS coefficient 0.031*** (0.005)).
  - The forward-looking component carries larger coefficients than the backward-looking component (e.g., forward 0.026*** (0.004) vs backward 0.015*** (0.004) in Spec II).
- Predictive content:
  - Forward-looking NPS predicts one-period-ahead changes in policy rates and short-term market rates; effect sizes translate to economically meaningful basis-point movements (examples: ≈7 basis points for policy rate change per one SD increase in sentiment).
- Heterogeneity:
  - Effects are strongest and most consistent in inflation-targeting regimes.
  - Exchange rate anchor regimes show no significant communication–policy link.
- Communication indices (Explanation Index, Straightforwardness Index, Net Confidence Index) generally have smaller or mixed associations with policy rates once controls and fixed effects are included.

*Source: IMF Working Paper — chapter 5 "Communication Metrics and Their Connection with Financial Variables" (wpiea2025109-print-pdf).*

### 1.25 percentage points.

### 1.25 percentage points.

### Forward-looking net policy sentiment and future policy & market rates
- Forward-looking net policy sentiment explains future changes in the policy rate in both advanced and other economies (Specs. I–II).
- From Table 7 (selected coefficients, standardized variables):
  - Total Net Policy Sentiment (ΔPolicy Rate 푖,푡+1, Col I): 0.101*** (0.022)
  - Forward-looking (ΔPolicy Rate 푖,푡+1, Col II): 0.083*** (0.017)
  - Backward-looking (ΔPolicy Rate 푖,푡+1, Col II): 0.062*** (0.014)
  - Gap (Fwd−Bwd) (ΔPolicy Rate 푖,푡+1, Col III): -0.096*** (0.022)
  - Forward-looking (ΔT-Bill Rate 푖,푡+1, Col V): 0.162*** (0.033)
  - Forward-looking (ΔT-Bond Rate 푖,푡+1, Col VIII): 0.105* (0.058)
- Specification and fit information (Table 7):
  - Observations across columns vary (examples listed): 5318, 5318, 5318, 4445, 6445, 6445, 6304, 0304, 0304 (as presented)
  - R2 examples: 0.168, 0.173, 0.173, 0.105, 0.108, 0.108, 0.145, 0.148, 0.148
- Notes from text:
  - Forward-looking statements serve as a reliable signal of central banks’ policy intentions across institutional and macroeconomic contexts.
  - For market-driven interest rates, forward-looking sentiment significantly explains future movements in both short-term (T-bill) and long-term (T-bond) rates only in emerging and low-income economies.
  - In advanced economies, forward-looking sentiment coefficients for market rates are weak or absent (Specs. IV and VI), consistent with markets more efficiently internalizing central bank behavior.
  - Results are associations and should not be interpreted causally.

### Heterogeneity by level of development (Table 8)
- Forward-looking net policy sentiment on future changes by group (all coefficients are standardized):
  - ΔPolicy Rate, Other Levels of Development (Col I): 0.079*** (0.018)
  - ΔPolicy Rate, Advanced Economies (Col II): 0.106** (0.039)
  - ΔT-Bill Interest Rate, Other Levels of Development (Col III): 0.058*** (0.021)
  - ΔT-Bill Interest Rate, Advanced Economies (Col IV): 0.055 (0.031)
  - ΔT-Bond Interest Rate, Other Levels of Development (Col V): 0.109* (0.053)
  - ΔT-Bond Interest Rate, Advanced Economies (Col VI): -0.047 (0.046)
- Backward-looking net policy sentiment (selected):
  - ΔPolicy Rate, Other Levels of Development: 0.064*** (0.016)
  - ΔPolicy Rate, Advanced Economies: 0.071** (0.030)
  - ΔT-Bill Interest Rate, Other Levels of Development: -0.022 (0.020)
  - ΔT-Bond Interest Rate, Other Levels of Development: 0.013 (0.018)
- Control indices (Straightforwardness, Explanation, Net Confidence) and macro controls (Exchange Rate, Inflation) reported with coefficients and standard errors in Table 8.
- Sample and fit (Table 8):
  - Observations examples: 3797, 1518, 3201, 1248, 1690, 1349 (as presented)
  - R2 examples: 0.188, 0.422, 0.158, 0.334, 0.210, 0.516
- Interpretations from text:
  - Forward-looking communication plays a more prominent role in shaping interest rate expectations in economies where financial markets are less developed, less liquid, or more information-constrained.
  - In such contexts, forward-looking sentiment can shift both short- and long-term yields across the entire yield curve.
  - In advanced economies, broader expectation mechanisms reduce the marginal information content of central bank statements.

### Overnight Indexed Swaps (OIS) and net policy sentiment (Table 9)
- OIS regression specification: OIS Rate 푖,푡+1 = 훼푖 + 휆푡 + 훽 Net Policy Sentiment 푖,푡 + 훾 Policy Rate 푖,푡 + X′푖,푡 + 휀푖,푡 (Equation (15)); tenors: 1, 3, 6, 12 months.
- Sample and coverage:
  - Dataset comprises 15 economies with OIS series from Bloomberg: Australia, Canada, People’s Republic of China, the Czech Republic, India, Indonesia, Japan, Malaysia, New Zealand, Republic of Poland, Russian Federation, Sweden, Switzerland, Thailand, and the United States.
  - After merging, retained 491 to 605 observations depending on tenor.
  - Data span from January 2015 to February 2025.
- Key results (Table 9, coefficients standardized; standard errors in parentheses):
  - Net Policy Sentiment (Forward):
    - 1 month (Col I): 0.017** (0.007)
    - 3 months (Col II): 0.024*** (0.007)
    - 6 months (Col III): 0.029*** (0.008)
    - 12 months (Col IV): 0.036*** (0.008)
  - Net Policy Sentiment (Backward): not statistically significant across tenors (examples: 0.017 (0.013), 0.032 (0.020), 0.034 (0.025), 0.036 (0.028))
  - Policy Rate (strong anchor across the curve):
    - 1 month: 1.949*** (0.017)
    - 3 months: 1.951*** (0.057)
    - 6 months: 1.950*** (0.073)
    - 12 months: 1.838*** (0.097)
  - Inflation (CPI):
    - 1 month: -3.207*** (0.715)
    - 3 months: -1.922*** (0.588)
    - 6 months: -1.785** (0.736)
    - 12 months: -0.532 (0.960)
- Fit statistics (Table 9):
  - Observations by tenor: 491, 504, 556, 605
  - R2 by tenor: 0.991, 0.983, 0.976, 0.968
- Interpretations from text:
  - Forward-looking net policy sentiment is positively associated with OIS rates at all horizons, increasing monotonically with tenor and peaking at 12 months.
  - Backward-looking sentiment is not statistically significant across tenors, indicating markets react to signals about future conditions rather than retrospective assessments.
  - Inflation is negatively associated with OIS rates at shorter maturities, suggesting markets may expect future easing following high inflation episodes.
  - Forward-looking sentiment remains significant after conditioning on current policy rates and macro fundamentals, implying communication shifts market expectations beyond observable conditions.

### Straightforwardness Index (Figures 25–26)
- Time coverage: 2000 to 2025; interdecile range shown (10th–90th percentile); global median indicated by dashed brown line.
- Two key findings:
  - Straightforwardness drops during episodes of systemic stress (global financial crisis, COVID-19 shock) across all country groups, reflecting more nuanced and conditional communication in crises.
  - Outside crises, straightforwardness varies systematically:
    - Advanced and inflation-targeting economies exhibit lower median straightforwardness.
    - Low-income and pegged exchange rate economies maintain higher straightforwardness in normal times.
  - From 2021 to 2024, straightforwardness scores increased steadily across most groups; slight decline after 2024.
- Forward vs backward decomposition (Figure 26):
  - Across advanced, emerging, and low-income economies, backward-looking statements are consistently more straightforward than forward-looking statements.
  - Forward-looking communication is less straightforward, conveying alternative paths and conditional scenarios.

### Explanation Index (Figure 27)
- Time coverage: 2000 to 2025; index EI defined as sum of confidence-building, risk-highlighting, and neutral statements normalized by sum of hawkish and dovish statements (Equation (9)).
- Patterns:
  - Explanation intensity rises systematically during policy tightening episodes (e.g., post-COVID-19 synchronized hikes).
  - Explanation scores fall during easing cycles (e.g., onset of COVID-19 in 2020).
  - Low-income and pegged exchange rate economies tend to have higher explanation scores, though cross-group differences have narrowed since 2015.
  - Since 2015, the index remains within a relatively tight range with transitory spikes at global policy turning points.

### Net Confidence Index (Figures 28–29)
- Time coverage: 2000 to 2025 for the index; global GDP-weighted forward/back components from 1990 to 2025 (Figure 29 begins in 1990 to match VIX availability).
- Aggregate patterns:
  - The net confidence index is persistently negative across most country groups and periods, indicating central banks prioritize highlighting risks over building confidence.
  - This tendency is more pronounced in advanced and inflation-targeting economies.
  - Pegged economies exhibit the highest volatility in the net confidence index.
  - The index declines sharply during systemic stress (global financial crisis, COVID-19) and rises in recoveries.
- Forward vs backward global aggregation (Figure 29):
  - Global GDP-weighted aggregation separates net confidence into forward-looking (blue) and backward-looking (red) components.
  - The gap (forward − backward) is used to measure directional risk communication; positive gap values indicate more expressed confidence in future conditions than in current/past ones.
  - The series is compared with the negative of the VIX index (dashed green line) for reference to implied market volatility.

*Source: wpiea2025109-print-pdf - 1.25 percentage points.*

### 2025.  The backward-looking index reflects assessments of current and recent macro-financial conditions, while the

### wpiea2025109-print-pdf - 2025.  The backward-looking index reflects assessments of current and recent macro-financial conditions, while the

### Forward vs. backward-looking sentiment and relation to VIX
- Backward-looking index reflects assessments of current and recent macro-financial conditions; forward-looking index captures expectations about future stability.
- Yellow bars in figure representation denote the gap between forward- and backward-looking sentiment; positive values indicate stronger forward confidence.
- The VIX is standardized and inverted to align visually with the direction of the net confidence index (i.e., higher risk perception corresponds to lower confidence).
- Central banks tend to express more confidence in their forward-looking communication than in backward-looking assessments over the past 30 years.
- Interpretation: optimistic forward tilt may reflect strategic communication objectives to stabilize expectations, reinforce perceptions of policy efficacy, and minimize amplification of uncertainty, consistent with theoretical models emphasizing communication under incomplete information (Woodford, 2005).

### Connection between Net Confidence Index (NCI) components and VIX: research question and methodology
- Research question: directionality — do central banks adapt communication in response to market volatility, or do forward-looking statements shape future risk perceptions?
- Approach: Granger causality analysis on three monthly global time series aggregated from 1990 to 2025:
  - forward-looking component of the Net Confidence Index (NCI)
  - backward-looking component of the Net Confidence Index (NCI)
  - standardized VIX
- Each NCI component is computed as a GDP-weighted average across countries.
- Dataset contains 420 monthly observations for each series.
- Maximum number of lags considered: twelve months.
- Stationarity treatment:
  - For forward-looking NCI: ADF test p-value = 0.265 (fails to reject nonstationarity); KPSS p-value = 0.010 (rejects stationarity) → use differenced series.
  - For backward-looking NCI: ADF p-value = 0.013 (rejects unit root); KPSS p-value = 0.010 (rejects stationarity) → mixed evidence → use differenced series.
  - For VIX: ADF p-value = 0.007; KPSS p-value = 0.290 → treated in levels.
- Note: Granger causality captures predictive content and temporal ordering but does not establish structural causality or control for omitted variables.

### Granger causality findings (Table 10)
- Panel A: ΔForward-looking NCI → VIX
  - Lag 1: F-stat 0.417, p-value 0.519 — Signif.Direction: ΔForward→VIX
  - Lag 2: F-stat 3.563, p-value 0.029 — ** — Signif.Direction: ΔForward→VIX
  - Lag 3: F-stat 1.840, p-value 0.139 — Signif.Direction: ΔForward→VIX
  - Lag 4: F-stat 3.151, p-value 0.014 — ** — Signif.Direction: ΔForward→VIX
  - Lag 5: F-stat 1.980, p-value 0.081 — * — Signif.Direction: ΔForward→VIX
  - Lag 6: F-stat 2.013, p-value 0.063 — * — Signif.Direction: ΔForward→VIX
  - Lags 7–12: F-stats 1.473, 1.313, 1.205, 1.205, 1.217, 1.161 with p-values 0.175, 0.235, 0.290, 0.286, 0.273, 0.309 — Signif.Direction: ΔForward→VIX
- Panel B: VIX → ΔForward-looking NCI
  - Lags 1–12 F-stats and p-values: 0.719 (0.397), 2.074 (0.127), 1.973 (0.117), 1.387 (0.237), 0.964 (0.440), 0.833 (0.545), 0.803 (0.585), 0.682 (0.708), 0.853 (0.567), 1.000 (0.443), 0.976 (0.467), 0.921 (0.525) — Signif.Direction: VIX→ΔForward (no statistically significant lags).
- Panel C: ΔBackward-looking NCI → VIX
  - Lag 1: F-stat 6.230, p-value 0.013 — ** — Signif.Direction: ΔBackward→VIX
  - Lags 2–12: F-stats 0.449, 0.598, 0.498, 0.373, 0.224, 0.869, 0.793, 0.750, 0.780, 0.843, 0.687 with p-values 0.638, 0.616, 0.737, 0.867, 0.969, 0.531, 0.609, 0.663, 0.648, 0.597, 0.764 — Signif.Direction: ΔBackward→VIX (only lag 1 significant).
- Panel D: VIX → ΔBackward-looking NCI
  - Lag 1: F-stat 0.762, p-value 0.383 — Signif.Direction: VIX→ΔBackward
  - Lag 2: F-stat 3.366, p-value 0.035 — ** — Signif.Direction: VIX→ΔBackward
  - Lag 3: F-stat 3.293, p-value 0.021 — ** — Signif.Direction: VIX→ΔBackward
  - Lag 4: F-stat 2.717, p-value 0.029 — ** — Signif.Direction: VIX→ΔBackward
  - Lag 5: F-stat 2.279, p-value 0.046 — ** — Signif.Direction: VIX→ΔBackward
  - Lag 6: F-stat 2.255, p-value 0.037 — ** — Signif.Direction: VIX→ΔBackward
  - Lag 7: F-stat 1.867, p-value 0.073 — * — Signif.Direction: VIX→ΔBackward
  - Lags 8–12: F-stats 1.394, 1.242, 1.347, 1.354, 1.180 with p-values 0.197, 0.268, 0.203, 0.193, 0.295 — Signif.Direction: VIX→ΔBackward (significance concentrated at lags 2–7).
- Interpretation of results:
  - Forward-looking NCI Granger-causes VIX for up to six months (statistically significant F-statistics up to six lags), implying forward-looking communication contains predictive signals about future market volatility.
  - There is no evidence that lagged VIX values predict changes in forward-looking confidence (Panel B).
  - Backward-looking NCI does not predict future VIX movements except at lag 1 (Panel C), while lagged VIX Granger-causes subsequent adjustments in backward-looking confidence, particularly two to seven months after a volatility shock (Panel D). This indicates backward-looking narratives react with delay to market stress.

### Implications for transmission of central bank communication
- Distinguishing proactive (forward-looking) and reactive (backward-looking) communication at sentence level enhances narrative measurement precision and reveals distinct transmission patterns to financial volatility.
- Forward-looking communication appears to carry signals that can shape market risk perceptions proactively.
- Backward-looking communication serves as an ex-post assessment that integrates realized market information with a lag.

### Audience-specific communication: construction and stylized facts
- For every document, five audience-specific net confidence indices are computed: general public, government, business sector, financial sector, and international stakeholders (Figure 30 shows evolution from 2000 to 2025).
- Rationale for starting analysis in 2000: central banks broadened communication practices beyond annual reports in the 2000s.

Key empirical findings on audience tailoring:
- No uniform communication style across audiences; central banks consistently tailor messages to audience informational needs, expertise, and policy relevance.
- Communication to the general public: systematically framed in the most confident and reassuring terms.
- Communication to the government: consistently emphasizes risks more heavily; markedly lower confidence scores.
- Communication to business and financial sectors: more neutral and balanced tone, tending toward risk-aware but favoring clarity, precision, and fact-based language.
- During systemic stress (global financial crisis, COVID-19 shock, market turbulence):
  - Central banks build relatively more confidence in messages to the financial sector and the government.
  - Confidence directed at the general public, business sector, and international stakeholders tends to decline.
- These audience-specific patterns are structural and persistent over three decades despite macroeconomic shocks and monetary regime shifts.

### Variance decomposition and tests of persistence in audience tone
- Empirical strategy: panel regressions regressing audience-specific net confidence index on dummies for audience, country, communication outlet, and time dimensions; quantify share of explained variance attributable to each factor.
- Findings from Figure 31 and regression analysis:
  - Audience dummies explain the largest share of variation when considering time and audience fixed effects.
  - Adding country and communication outlet dummies reduces but does not eliminate the audience contribution—the audience effect remains substantial across specifications.
  - Time fixed effects consistently explain only a modest fraction of variation.
  - Country fixed effects explain a relatively small portion of variation.
- Interpretation: tone differentiation across audiences is primarily driven by structural and intentional communication design choices rather than by transitory macroeconomic or country-specific factors.
- Note: communication addressed to international stakeholders exhibits somewhat more variation over time than other audiences, with sentiment becoming more negative in distressed periods.

*Source: wpiea2025109-print-pdf (2025), analysis of Net Confidence Index components and their Granger-causal relationships with the standardized VIX; audience-specific sentiment analysis and variance decomposition.*

### 6.  Conclusions

### 6. Conclusions

### Framework and methodological contributions
- Develops a novel automated classification framework operating at the sentence level that classifies central bank communication along four dimensions: topic, communication stance, sentiment, and audience.
- Uses a fine-tuned multilingual large language model trained explicitly for central bank communication to capture semantic and contextual nuances that traditional dictionary-based methods miss.
- Enables extraction of policy signals with high precision, translating qualitative text into quantitative data for rigorous evaluation of communication effectiveness.
- All prompts reported in the paper were run in the OpenAI API interface using the ChatGPT-4o chatbot model (as of April 2025).

### Dataset and empirical scope
- Applied to an extensive dataset of 74,882 documents from 169 central banks worldwide.
- Demonstrates ability to analyze evolution of central bank messaging across time, countries, and monetary policy regimes.
- Variance decomposition (Figure 31) shows that the audience dimension consistently accounts for the largest share of explained variance in the net confidence index; time effects explain little and country effects contribute modestly, with a residual capturing unexplained heterogeneity.

### Key empirical findings
- Adoption of inflation-targeting frameworks coincides with significant shifts in messaging: backward-looking exchange rate statements give way to forward-looking discussions on inflation, interest rates, and economic activity.
- Central banks strategically tailor messaging across audiences (financial markets, businesses, households, international stakeholders).
- The audience dimension drives most of the explained variation in central bank confidence communication, underscoring targeted tone across stakeholder groups; global shocks and cyclical conditions (time effects) are secondary.

### Novel textual metrics introduced
- Net policy sentiment metric: quantifies overall stance of communication and distinguishes forward- and backward-looking components; forward-looking component serves as a proxy for forward guidance.
  - Empirical result: forward-looking sentiment robustly predicts policy rate changes and influences market interest rates.
- Straightforwardness index: evaluates clarity of messaging.
- Explanation index: evaluates depth of policy justifications.
- Net confidence index: captures balance between confidence-building and risk-highlighting statements, offering insights into uncertainty communication.

### Classifier robustness and multilingual performance (Appendix B)
- Translation-consistency evaluation used a representative sample of non-English documents translated to English with ChatGPT 4o and classified in both original and translated versions; consistency measured with Jensen–Shannon Divergence (JSD).
- Sample counts by language: 74 documents in Portuguese (PT), 37 in French (FR), 69 in Russian (RU), 19 in Romanian (RO), 16 in Spanish (ES), and 7 in Arabic (AR).
- Observed JSD results:
  - Topic dimension: all languages display JSD values below 0.16; French (FR) and Arabic (AR) near 0.16; Russian (RU) at 0.14; Portuguese (PT) at approximately 0.13; Spanish (ES) at 0.07.
  - Communication stance: highest JSD observed for Arabic (AR) and French (FR) at just below 0.04; all other languages below 0.03; Spanish (ES) approaching 0.00.
  - Audience: Russian (RU) highest JSD at 0.12; all other languages fall below 0.06; Spanish (ES) JSD = 0.01.
  - Sentiment: Russian (RU) highest JSD at 0.15; all other languages remain below 0.05; Spanish (ES) near 0.01.
- Summary: low divergence values across languages and dimensions—none exceeding 0.16 and most falling well below 0.10—indicate strong cross-lingual stability, while slight divergences for Russian and Arabic highlight language-specific limits.

### Prediction confidence and classifier reliability (Appendix C)
- Classifier estimates joint probabilities over paired dimensions ⟨topic, communication stance⟩ and ⟨audience, sentiment⟩; confidence analysis focuses on marginal probabilities derived from joint outputs.
- Marginal probability computation examples preserved in the source (equations (16) and (17)).
- Overall marginal-probability results:
  - More than half of all sentences are classified with marginal probabilities above 60 percent for topic.
  - More than half of all sentences are classified with marginal probabilities above 70 percent for audience and sentiment.
  - Nearly 80 percent of sentences are classified with marginal probabilities above 70 percent for communication stance.
  - (Figure note): For instance, approximately 80 percent of sentences classified by topic receive marginal probabilities greater than 80 percent.
- Class-level confidence heterogeneity:
  - “Metadata” classifications consistently exhibit the highest marginal probabilities.
  - High-confidence substantive classes include “MP - inflation” (topic), “forward-looking” (communication stance), “financial sector” (audience), and “risk-highlighting” (sentiment).
  - Lower-confidence or wider-distribution classes include topic labels such as “fiscal policy,” “financial stability,” and “supervision and regulation.”

### Practical implications for central banks and policymakers
- Provides practical tools to enhance accountability for past actions, build transparency of current actions, and shape expectations through systematic assessment and refinement of messaging.
- Enables benchmarking communication against historical trends or peer institutions to align strategies with best practices.
- Strengthens market confidence and can improve monetary policy transmission by enabling quantitative assessment of communication’s role.

### Recommended directions for future research
- Empirically examine the causal impact of communication—particularly the forward-guidance proxy—on macroeconomic variables and monetary policy transmission, and explore whether communication can influence expectations in ways that contribute to price stability.
- Relate the net confidence index to financial stress indicators (e.g., country-level financial stress indices) to study how risk communication shapes market sentiment and stability perceptions.
- Study FX-related communication topics alongside exchange rate dynamics to investigate whether central bank messaging affects currency movements and volatility.
- Extend the methodology to social media and other non-traditional communication channels to capture timely and targeted messaging and broaden applicability beyond monetary policy to financial stability and institutional accountability.

*Source: 6. Conclusions (wpiea2025109-print-pdf)*

### Appendix D.    Multilabel Sentences

### Appendix D.    Multilabel Sentences

### Problem framing and measurement approach
- Central bank sentences can address multiple topics, audiences, and convey multiple sentiments within a single sentence, creating multilabel instances that diffuse classifier-assigned probabilities across categories and lower maximum predicted probabilities.
- Measurement: frequency with which two distinct classes within the same dimension are simultaneously assigned predicted marginal probabilities above 25 percent for the same sentence.
- Dimensions analyzed: topic, audience, and sentiment. Communication stance is excluded because of its binary structure.

### Empirical findings on co-occurrence (summary)
- Overall conclusion: co-occurrence rates are modest and not pervasive enough to undermine interpretability of predicted probabilities; the single-label probabilistic approach is deemed pragmatic and analytically robust given empirical prevalence and practical costs of alternatives.

- Topic co-occurrence (heatmap summary):
  - Most topic pairs appear together in fewer than 0.04 percent of sentences.
  - Heatmap values include multiple explicit 0.00% entries and scattered nonzero entries such as 0.01%, 0.04% in select off-diagonal cells.
  - Labeling note: main diagonal excluded to focus on cross-topic associations.

- Audience co-occurrence (heatmap summary):
  - Co-occurrence rates are slightly higher than topic co-occurrence.
  - Notable pair: Business Sector — Financial Sector: 1.89%.
  - Other audience co-occurrence values presented: 0.69%, 0.22%, 0.73%, 0.01%, 0.41%, 0.33%, 0.63%, 0.18%, 0.11%, 0.16%, 0.03%, 0.19%, 0.02% (as shown in the audience heatmap).

- Sentiment co-occurrence (heatmap summary):
  - Sentiment co-occurrence is more pronounced than topic co-occurrence.
  - Notable pairs:
    - Neutral/Balanced with Confidence-building: 3.62%.
    - Neutral/Balanced with Risk-highlighting: 3.14%.
    - Confidence-building with Not applicable: 1.51%.
  - Additional sentiment heatmap values include: 0.09%, 0.03%, 0.08%, 0.01%, 0.13%, 0.16%, 0.07% (as shown in the sentiment heatmap).

### Interpretation and methodological trade-offs
- Practical costs of adopting multilabel classification:
  - Requires substantially more complex and labor-intensive annotation: annotators must identify all applicable categories rather than selecting a single most-relevant label, raising the threshold for assembling a sufficiently large and representative labeled dataset.
  - Multilabel models often necessitate additional calibration and thresholding choices, complicating interpretation of output probabilities.
- Given the relatively limited empirical prevalence of multi-topic sentences and the annotation/calibration costs, the authors adopt a single-label probabilistic approach that assigns marginal probabilities across categories while maintaining interpretability.

*From "Appendix D. Multilabel Sentences" of Working Paper No. WP/2025/109*

---


_Source: https://www.imf.org/-/media/files/publications/wp/2025/english/wpiea2025109-print-pdf.pdf_
