## wpiea2023241-print-pdf

## Source details

**Canonical URL:** [wpiea2023241-print-pdf](https://www.imf.org/-/media/files/publications/wp/2023/english/wpiea2023241-print-pdf.pdf)

## Other formats

- [Markdown version](/-/media/files/publications/wp/2023/english/wpiea2023241-print-pdf.pdf.md)
- [Structured JSON version](/-/media/files/publications/wp/2023/english/wpiea2023241-print-pdf.pdf.json)

---

### Introduction and AI/ML context
- AI and ML offer improvements in products, services, risk management, compliance, and the development of legislation and regulation.
- Drivers for AI/ML adoption include increased computing power, cheaper storage, parsing and analysis of data, and the “rapid growth of datasets for learning and prediction owing to increased digitization and the adoption of web-based services.”
- Machine learning is described as a branch or subset of AI that uses algorithms to automatically learn from data and form predictive models; machine learning is a pathway to artificial intelligence and is a current application of AI used in day-to-day life.

### The IMF Central Bank Legislation Database (CBLD): scope and structure
- Coverage and access:
  - Contains laws of 175 central banks and monetary unions.
  - Includes datasets from four specific update moments: 2010, 2015, 2020/2021, and 2023.
  - The database was opened to the public in February 2019; access requires a one-time (free) registration via the CBLD’s website (https://data.imf.org/cbld).
- Search and categorization:
  - CBLD has 273 specific search categories that allow granular queries.
  - The 2020/2021 update expanded search categories from 112 to 273.
  - Allows searching by country or by pre-set groups of countries (region, income level, exchange rate arrangements, and membership of a monetary union).
- Document types and language:
  - Includes central bank laws/charters, relevant excerpts of constitutions, payments system laws, banking/supervision laws, AML/CFT laws, resolution laws, financial stability laws, and related statutes.
  - All collected data is in English only—either provided by authorities or translated by the IMF (translated texts carry a disclaimer).

### Data collection and update process
- 2010 and 2015: questionnaire sent to all central bank governors and monetary union presidents requesting relevant central bank legislation.
- 2020/2021 onward: data largely collected by a questionnaire (in 2020) and supplemented (in 2021) by the IMF CBLD team without consulting central banks.
- Exceptions and notes:
  - Countries not included either did not respond, are not IMF members, or IMF could not include relevant legislation.
  - The only noted exception where the central bank law is not published online is Eritrea.
  - For multiple published versions of legislation, IMF CBLD team contacted central bank legal departments for clarification.
- Relevance and added value:
  - Strong added value for low-income countries (LICs) and emerging markets (EMs) with legal capacity constraints.
  - 2020/2021 expansion reflects broadening of central bank mandates since the Global Financial Crisis.

### Alignment with IMF Central Bank Transparency Code (CBT)
- CBT (published in 2020) is a voluntary international standard with 47 principles divided into five groups: (i) transparency over central bank governance; (ii) transparency over central bank policy; (iii) transparency over central bank operations; (iv) transparency over central bank outcomes; and (v) transparency over central bank official relations.
- For most principles, three sets of practices are listed: core, expanded, and comprehensive.
- CBLD search categories align with CBT principles; example categories include 2.20, 5.11, 6.13, 7.04, 8.09, 10.08, and 11.11.

### CBLD dataset characteristics and key datapoints
- 273 CBLD categories.
- Average length of a CBLD text (such as an article): 97 words.
- Median length of a CBLD text: 51 words.
- CBLD dataset is large and largely qualitative (textual), suitable for AI/ML approaches to identify patterns and support policy discussions.

### A. CBLD Coding App — methodology and user workflow
- Manual coding:
  - Each submitted piece was manually coded by a team (two central bank lawyers and an IMF research officer) overseen by an IMF senior expert.
  - Coding annotated every article and often specific sentences against the 273 categories; texts could be assigned multiple categories (up to six in examples).
- Electronic coding app:
  - Prototype built from an open-source project for PDF annotation and published under an MIT license.
  - Hosted on Microsoft Azure, accessible via major web browsers; requires conversion of documents to PDF.
  - Interface allowed highlighting text and assigning one or more CBLD categories; previous annotations were visible and reusable.
- AI/ML recommender integration:
  - App recorded coding actions and used reinforced learning to suggest best categories for new highlighted text.
  - Workflow: highlighted text sent to an Azure server-less API using a pre-computed recommender model to return suggested categories; users accept/reject suggestions to reinforce the model.
  - UI adjusted from reordering by affinity scores to numeric order with top affinities highlighted in orange (intensity proportional to score).
  - Recommender model based on a Naïve Bayes model trained on the previous (2015) CBLD dataset; annotations were periodically re-computed and reloaded into the API.

### B. ML Algorithm — models tested, performance, and interpretability
- Algorithm choice and rationale:
  - Naïve Bayes classifier implemented via scikit-learn in Python selected for short text lengths and multi-label potential.
  - Advantages: scalable, linear estimation time, explainability via inspection of word-class probabilities P(x_i|y), and ability to rank candidate categories.
- Models and preprocessing tested:
  - Stemmers: Porter or Snowball stemmer (via NLTK).
  - Classifiers: MultinomialNB and BernoulliNB from scikit-learn.
  - Weighting: raw word counts and TF-IDF counts.
  - Final set included combinations leading to eight distinct Naïve Bayes models.
- Model evaluation:
  - Data split into training and test datasets; performance reported out-of-sample on the test set.
  - Coefficients provide transparency on which words influence each category.
- Key performance findings (exact figures preserved):
  - Example category 1.01: probability of the correct category being the top prediction ≈ 25 percent.
  - For category 1.01, probability of the correct category appearing within the top five candidates with the best model is more than 90 percent.
  - Across all categories, with the best overall model the chance of suggesting the correct category within the top 10 candidates is about 50 percent.
  - Best overall model: simple word counts with Multinomial Naïve Bayes, followed by Snowball stemmer + counts + MultinomialNB.
  - Model efficiency increases with more training data; most gains for current data appear already realized with existing categorized data.
- Practical implication:
  - Presenting users with a short ranked list (e.g., top 5 candidates) yields high chances of including the correct category and materially speeds up coding.
  - Naïve Bayes provides interpretable word-level coefficients useful for legal analysis and transparency.

### C. Word frequency, N-gram and correlation analyses
- Tokenization and corpus statistics:
  - Total words tokenized across all CBLD 2020/2021 documents: 6,746,772.
  - After filtering numbers, special characters, and stop words, remaining words: 2,763,901.
- Highest-frequency words include: "board", "bank", "central", "legal", "act", "financial".
- N-gram themes (tentative groupings):
  - Governance: governor, governors, deputy, council, staff.
  - Policy areas: banking, securities, foreign + exchange, policy, reserves, monetary, credit, financial, supervisory, debt, coins and notes.
  - Stakeholders and transparency: ministers, government, person, public, banks, international system, market.
- Country-specific bigram examples:
  - Albania: strong connections between "Albania" and "bank"; links to "council" and "supervisory".
  - Italy (Banca d’Italia Act): Eurozone references—"European System of Central Banks (ESCB)", "euro", "treaties"; links indicating shareholders and governing structures.
- Word correlation analysis:
  - Example: United Kingdom vs India laws—words near the y=x line used equally; "treasury" appears more frequently in the BoE Act.
  - Axes formatted in percentage log scale for visualization.
- Specific topical n-gram use-case — “independ” and “autonom” roots:
  - Co-occurrence with words such as "act", "external", "functions", "financial", "audit", "entities", "assessment", "legal".
  - Country coverage of independence/autonomy bigrams: counts range from 1–95 bigrams per country; regions with high counts include Europe, the US, China, and many African countries.

### D. Sentiment Analysis of “Independence”
- Methodology:
  - CBLD categories 2.13 through 2.17 (relating to central bank independence and its forms) were pulled.
  - Manually coded references labeled positive (strengthening independence) or negative (limiting independence).
  - Training dataset composed of several hundred pre-selected sentences based on CBLD categories, plus 100 additional sentences not based on categories and 100 "fuzzy" sentences.
  - Algorithm trained to identify words associated with positive or negative influences on independence and then applied across the entire CBLD dataset.
- Findings and limitations:
  - A Naïve Bayes approach struggles to capture nuanced sentiment because it cannot fully account for word connections and contextual meanings.
  - More complex models (e.g., advanced LLMs such as ChatGPT) could likely enable more accurate automatic sentiment assignment and meaningful analysis, but raise ethical and reliability considerations.
  - Ethical risks to consider with LLM use: social stereotyping and unfair discrimination, false or misleading information, outdated information, and proprietary model constraints.

### User statistics and demand signals
- Registration-based tracking enables analysis of queries, categories, countries searched, and user domains.
- Query volume growth (exact figures preserved):
  - Average individual queries per day in 2021 (between September 7 and December 7): 323.
  - Average individual queries per day (later period): 5,441.
  - Excluding days in 2022 with queries above 10,000 per day, average daily queries in 2022 would be 932.
  - Note: each data point represents single day access; one user might perform multiple actions.
- Regional and user-domain patterns:
  - Many queries from Western hemisphere, Middle East, Europe, and Asia-Pacific.
  - Many users use private email accounts (e.g., Gmail, Yahoo).
- Most-searched categories: governance (decision-making process, institutional arrangement), policy (objectives and functions), accountability (accountability framework and reporting obligations), and legal categories (legal status and name of the central bank law).
- Country-level search counts (2021–2023):
  - Hungary: 40,635 searches.
  - Other frequently searched jurisdictions include Czech Republic, Lithuania, Poland, Belgium, Slovenia, the Central African Monetary Union, the European Monetary Union, Argentina, the United States, and India.
  - Single-search countries during 2021–2023 include Seychelles, Panama, Suriname, Maldives, Morocco, UAE, Lao, Barbados, Venezuela, Antigua and Barbuda, Vietnam, Mozambique, Comoros, Eritrea, and Eswatini.
- Use-case implications:
  - User search patterns can inform refinement of CBLD search categories and identification of missing categories.
  - High-demand topics could be prioritized for improved coding accuracy or targeted NLP analyses.
  - Potential for selected crowdsourcing of coding to engaged users or institutions.

### Conclusions and policy-relevant implications
- AI/ML (especially interpretable models like Naïve Bayes) can substantially accelerate manual coding by surfacing ranked candidate categories; presenting top N candidates (e.g., top 5) yields high inclusion probability for the correct category.
- Word-frequency, n-gram, and correlation analyses reveal thematic emphases—governance, policy functions, and stakeholder transparency—which align with user search behavior.
- Advanced NLP and LLMs offer promise for nuanced tasks (e.g., sentiment analysis, semantic understanding) but require careful consideration of ethical risks and data quality.
- Combining user analytics with ML outputs can guide prioritization of coding efforts, category refinement, and targeted legal-policy research (e.g., studies on central bank independence across regions and regimes).

### Operational lessons, limitations, and recommended next steps
- Early coding was manual and time-consuming; the coding app reduced friction and enabled simultaneous work by multiple team members.
- UI evolution: initial score-based reordering confused users; later color-coding of normalized scores (white to orange) improved usability.
- Integration hurdles: automated conversion scripts sometimes failed, introducing invisible spaces and line breaks that corrupted underlying code; affected text was manually reformatted and reloaded.
- The CBLD AI/ML approach is a first step; annual updates provide opportunities to enhance the app, apply other AI/ML approaches, and identify additional patterns.
- Recommended collaborative actions:
  - IMF should work with selected academic partners and country authorities to better understand AI/ML outcomes and finetune methods.
  - Improvements would benefit external users through enhanced search options and greater access to CBLD data.

*Source: IMF Working Paper "Predicting the Law: Artificial Intelligence Findings from the IMF’s Central Bank Legislation Database" (CBLD sections and AI/ML analysis).*

### Introduction ...........................................................................................................

### Introduction

### AI/ML opportunities and context
- AI and ML offer numerous opportunities for financial sector participants, enabling improvements in products, services, risk management, compliance, and the development of legislation and regulation.
- The FSB (2017) highlights drivers for AI/ML adoption including increased computing power, cheaper storage, parsing and analysis of data, and the “rapid growth of datasets for learning and prediction owing to increased digitization and the adoption of web-based services.”
- Machine learning is described as a branch or subset of AI that uses algorithms to automatically learn from data and form predictive models; machine learning is a pathway to artificial intelligence and is a current application of AI used in day-to-day life.

### The IMF Central Bank Legislation Database (CBLD): scope and structure
- The CBLD is described as the most comprehensive central bank legislation database in the world.
- Coverage:
  - Contains laws of 175 central banks and monetary unions.
  - Includes datasets from four specific update moments: 2010, 2015, 2020/2021, and 2023.
  - The database was opened to the public in February 2019; access requires a one-time (free) registration via the CBLD’s website (https://data.imf.org/cbld).
- Search and categorization:
  - The CBLD has 273 specific search categories that allow granular queries on nearly any topic involving a central bank.
  - The 2020/2021 update expanded search categories from 112 to 273.
  - The CBLD allows searching by country or by pre-set groups of countries (including region, income level, exchange rate arrangements, and membership of a monetary union).
- Document types and language:
  - Documents included are often central bank laws/charters, relevant excerpts of constitutions, and numerous other laws that relate to the central bank (e.g., payments system laws, banking/supervision laws, AML/CFT laws, resolution laws, financial stability laws with focus on microprudential supervision, macroprudential oversight, ELA/LOLR, and resolution).
  - All collected data is in English only—either provided by the authorities or translated by the IMF (in which case a disclaimer would be added noting the text is not an official translation).
  - For most central banks, the number of included documents is 1–3; for some (e.g., European Union members, West African Monetary Union) more than 4 separate laws per authority are included.

### Data collection and update process
- For 2010 and 2015 updates: a questionnaire was sent to all central bank governors and monetary union presidents requesting relevant central bank legislation.
- From the 2020/2021 update onwards: data was largely collected by a similar questionnaire (in 2020) and then supplemented with additional legislation (in 2021) by the IMF CBLD team without consulting central banks, reflecting that for almost all countries the relevant legislation is available online.
- Exceptions and notes:
  - Countries not included either did not respond to the IMF’s questionnaire, are not IMF members, or the IMF CBLD team could not include relevant legislation (e.g., ongoing amendments or time constraints).
  - The only noted exception where the central bank law is not published online is Eritrea.
  - In instances of multiple published versions of legislation, the IMF CBLD team would contact central bank legal departments for clarification.

### Relevance and added value
- The CBLD has strong added value for low-income countries (LICs) and emerging markets (EMs), where legal capacity constraints and limited resources create demand for efficient access to peer-country central bank legal frameworks—especially during legislative amendment processes.
- The 2020/2021 expansion reflects the significant broadening of central bank mandates since the Global Financial Crisis and the consolidation of multiple functions within single central bank organizations.

### Alignment with IMF Central Bank Transparency Code (CBT)
- The CBLD has been designed with the IMF Central Bank Transparency Code (CBT) in mind.
- The CBT (published in 2020) is a voluntary international standard containing 47 principles (most principles include sub-principles) divided into five groups:
  - (i) transparency over central bank governance;
  - (ii) transparency over central bank policy;
  - (iii) transparency over central bank operations;
  - (iv) transparency over central bank outcomes; and
  - (v) transparency over central bank official relations.
- For most principles, three sets of practices are listed: core, expanded, and comprehensive.
- The CBLD and the CBT are aligned in terms of specific search categories for central bank transparency by function; examples of CBLD search categories include 2.20 (general policy on institutional transparency), 5.11 (currency), 6.13 (monetary policy and operations), 7.04 (international reserves), 8.09 (foreign exchange policy), 10.08 (macroprudential policy), and 11.11 (financial supervision).

*IMF Working Paper — Introduction section*

### 12.12 on lender of last resort, 13.11 on financial integrity, 14.07 on consumer protection, 15.21 on resolution,

### wpiea2023241-print-pdf - 12.12 on lender of last resort, 13.11 on financial integrity, 14.07 on consumer protection, 15.21 on resolution,

### Overview
- The Central Bank Legislation Database (CBLD) contains coded central bank legislation annotated against a predefined list of 273 categories.
- CBLD 2020/2021 data collection: central banks uploaded documents; IMF CBLD team supplemented with additional laws (including banking laws and AML/CFT laws).
- The CBLD dataset is large and largely qualitative (textual), enabling AI/ML approaches to identify patterns and support policy discussions on topics such as central bank independence and objectives.
- Key datapoints:
  - 273 CBLD categories.
  - Average length of a CBLD text (such as an article): 97 words.
  - Median length of a CBLD text: 51 words.

### A. CBLD Coding App (methodology and user workflow)
- Manual coding process:
  - Each submitted piece of legislation was manually coded by a team (two central bank lawyers and an IMF research officer) overseen by an IMF senior expert.
  - Coding annotated every article and often specific sentences against the 273 categories; texts could be assigned multiple categories (up to six in examples).
- Electronic coding app:
  - Prototype built from an open-source project for PDF annotation and published under an MIT license.
  - Hosted on Microsoft Azure, accessible via major web browsers, requiring only conversion of documents to PDF.
  - Interface allowed highlighting text and assigning one or more CBLD categories; previous annotations were visible and reusable.
- AI/ML recommender integration:
  - The app recorded coding actions and used reinforced learning to suggest best categories for new highlighted text.
  - Workflow: highlighted text sent to an Azure server-less API, which used a pre-computed recommender model to return suggested categories; users accept/reject suggestions to reinforce the model.
  - UI adjustment: categories initially reordered by affinity scores; after user feedback the interface reverted to numeric order but highlighted top affinities in orange with intensity proportional to score.
  - The recommender model was based on a Naïve Bayes model trained on the previous (2015) CBLD dataset; annotations were periodically re-computed and reloaded into the API.

### B. ML Algorithm (models tested, performance, and interpretability)
- Algorithm choice and rationale:
  - Naïve Bayes classifier implemented via scikit-learn in Python was selected for document classification given short text lengths and multi-label potential.
  - Naïve Bayes advantages: scalable, linear estimation time, explainability via inspection of word-class probabilities P(x_i|y), and ability to rank candidate categories.
- Models and preprocessing tested:
  - Stemmers: Porter or Snowball stemmer (via NLTK).
  - Classifiers: MultinomialNB and BernoulliNB from scikit-learn.
  - Weighting: raw word counts and TF-IDF counts.
  - Final set of models included combinations leading to eight distinct Naïve Bayes models:
    1) Porter/Snowball stemmer combined with raw counts/TF-IDF counts modeled as MultinomialNB;
    2) counts/TF-IDF counts modeled by MultinomialNB/BernoulliNB.
- Model evaluation approach:
  - Data split into training and test datasets; performance estimates reported out-of-sample on the test set.
  - Coefficients (weights) provide transparency on which words influence each category.
- Key performance findings:
  - Example category 1.01: probability of the correct category being the top prediction ≈ 25 percent; probability of the correct category appearing within the top five candidates with the best model is more than 90 percent.
  - Across all categories, with the best overall model the chance of suggesting the correct category within the top 10 candidates is about 50 percent.
  - Best overall model: simple word counts with Multinomial Naïve Bayes, followed by Snowball stemmer + counts + MultinomialNB.
  - Model efficiency increases with more training data; most gains for current data appear already realized with existing categorized data.
- Practical implication:
  - Presenting users with a short ranked list (e.g., top 5 candidates) yields high chances of including the correct category and can materially speed up coding.
  - The Naïve Bayes approach provides interpretable word-level coefficients useful for legal analysis and transparency.

### C. Word Frequency Comparison and N-gram analysis
- Tokenization and corpus statistics:
  - Total words tokenized across all CBLD 2020/2021 documents: 6,746,772.
  - After filtering numbers, special characters, and stop words, remaining words: 2,763,901.
- Highest-frequency words observed include: "board", "bank", "central", "legal", "act", "financial".
- N-gram and word-cloud themes identified (tentative groupings):
  - Governance: governor, governors, deputy, council, staff.
  - Policy areas: banking, securities, foreign + exchange, policy, reserves, monetary, credit, financial, supervisory, debt, coins and notes.
  - Stakeholders and transparency: ministers, government, person, public, banks, international system, market.
- Country-specific bigram examples:
  - Albania: strong connections between "Albania" and "bank"; notable links to "council" and "supervisory".
  - Italy (Banca d’Italia Act): prominent Eurozone references—"European System of Central Banks (ESCB)", "euro", "treaties"; strong connections indicating presence of shareholders and governing structures.
- Word correlation analysis:
  - Example: United Kingdom vs India laws—words near the y=x line used equally in both acts; e.g., "treasury" appears more frequently in the BoE Act.
  - Axes formatted in percentage log scale for visualization.
- Specific topical n-gram use-case — “independ” and “autonom” roots:
  - Bigram correlations show co-occurrence of independence/autonomy with words such as "act", "external", "functions", "financial", "audit", "entities", "assessment", "legal".
  - Country coverage of independence/autonomy bigrams: counts range from 1–95 bigrams per country; regions with high counts include Europe, the US, China, and many African countries.

### D. Sentiment Analysis of “Independence”
- Methodology:
  - CBLD categories 2.13 through 2.17 (relating to central bank independence and its forms) were pulled.
  - Manually coded references were labeled as positive (strengthening independence) or negative (limiting independence).
  - Training dataset composed of several hundred pre-selected sentences based on CBLD categories, plus 100 additional sentences not based on categories and 100 "fuzzy" sentences.
  - Algorithm trained to identify words associated with positive or negative influences on central bank independence and then applied across the entire CBLD dataset.
- Findings and limitations:
  - A Naïve Bayes approach struggles to capture nuanced sentiment because it cannot fully account for word connections and contextual meanings.
  - More complex models (e.g., advanced LLMs such as ChatGPT) could likely enable more accurate automatic sentiment assignment and meaningful analysis, but raise ethical and reliability considerations.
  - Ethical risks to consider with LLM use: social stereotyping and unfair discrimination, false or misleading information, outdated information, and proprietary model constraints.

### User Statistics and demand signals
- Registration-based tracking enables analysis of queries, categories, countries searched, and user domains.
- Query volume growth:
  - Average individual queries per day in 2021 (between September 7 and December 7): 323.
  - Average individual queries per day (later period): 5,441.
  - Excluding days in 2022 with queries above 10,000 per day, average daily queries in 2022 would be 932.
  - Note: each data point represents single day access (searches, browsing, and queries); one user might perform multiple actions.
- Regional distribution of users: many queries from Western hemisphere, Middle East, Europe, and Asia-Pacific; many users use private email accounts (e.g., Gmail, Yahoo).
- Most-searched categories: governance (decision-making process, institutional arrangement), policy (objectives and functions), accountability (accountability framework and reporting obligations), and legal categories (legal status and name of the central bank law).
- Country-level search counts (2021–2023):
  - Hungary: 40,635 searches.
  - Other frequently searched jurisdictions include Czech Republic, Lithuania, Poland, Belgium, Slovenia, the Central African Monetary Union, the European Monetary Union, Argentina, the United States, and India.
  - Single-search countries during 2021–2023 include: Seychelles, Panama, Suriname, Maldives, Morocco, UAE, Lao, Barbados, Venezuela, Antigua and Barbuda, Vietnam, Mozambique, Comoros, Eritrea, and Eswatini.
- Use-case implications:
  - User search patterns can inform refinement of CBLD search categories and identification of missing categories.
  - High-demand topics could be prioritized for improved coding accuracy or targeted NLP analyses.
  - Potential for selected crowdsourcing of coding to engaged users or institutions.

### Conclusions and policy-relevant implications
- AI/ML (especially interpretable models like Naïve Bayes) can substantially accelerate the manual coding of legal texts in the CBLD by surfacing ranked candidate categories; presenting top N candidates (e.g., top 5) yields high inclusion probability for the correct category.
- Word-frequency, n-gram, and correlation analyses reveal thematic emphases in central bank legislation—governance, policy functions, and stakeholder transparency—which align with user search behavior.
- Advanced NLP and LLMs offer promise for more nuanced tasks (e.g., sentiment analysis, semantic understanding), but require careful consideration of ethical risks and data quality issues.
- Combining user analytics with ML outputs can guide prioritization of coding efforts, category refinement, and targeted legal-policy research (e.g., studies on central bank independence across regions and regimes).

*Source: IMF Working Paper "Predicting the Law: Artificial Intelligence Findings from the IMF’s Central Bank Legislation Database" (CBLD sections and AI/ML analysis).*

### Conclusion

### Conclusion

### Importance of the CBLD
- The IMF’s Central Bank Legislation Database (CBLD) is a critical tool for central banks.
- It offers insights into central banking developments, primarily in terms of legislation and regulation.
- These insights can help national legislators and central banks identify possible patterns, topics, and modalities that feed into policy discussions or legislative amendments.

### Challenges in using legal texts
- Central bank and related laws are often not easily accessible to central banks and policy makers due to:
  - Complex terminology and syntax across drafting languages.
  - A structure designed to guide users to specific (im)possibilities, which makes it difficult to identify general patterns or linkages between concepts and stipulations.

### AI/ML tools and the CBLD coding app
- The IMF explored various Artificial Intelligence and Machine Learning approaches for central bank legislation to:
  - Understand, analyze, and identify patterns.
  - Make the incorporation of legislation into the CBLD more reliable and cost-efficient.
- The IMF’s in-house developed coding app was initially intended for IMF central bank lawyers to analyze and code legislation for the CBLD, but its increasing predictive value also allows identification of patterns and interlinkages throughout central bank legislation.
- The newly developed AI coding app enabled:
  - Electronic access to all uploaded legislation.
  - Coding of any text within documents against one or more of the 273 CBLD search categories.
  - Multiple team members to access and work on documents simultaneously, with coding saved automatically and AI-provided suggestions for possible categories.

### Key findings from AI/ML analysis
- Tokenizing individual words across the entire CBLD dataset identified three main themes:
  - (i) central bank governance;
  - (ii) central bank policy; and
  - (iii) central bank stakeholders and transparency.
- New topics within those themes (for instance, fintech or climate change as part of central bank policy) are not clearly emerging in the legislation data.
  - This is explained by the fact that amendments to legislation take time, and legislation is, by definition, behind the curve.

### Operational lessons and limitations
- Early coding process (pre-2020/2021 cycle) was manual and time-consuming, involving paper-and-pen coding and XML generation one law at a time.
- The coding app’s UI and ML features evolved based on user experience:
  - Initial versions reordered suggested codes by score, which confused users; later versions color-coded normalized scores from zero to one (items with a zero score in white; the item with the highest score in orange).
- Integration hurdles:
  - Automated conversion scripts sometimes failed to handle formatting in a small set of documents, introducing invisible spaces and line breaks that corrupted underlying code and affected CBLD search results.
  - The team manually reformatted affected text sections and reloaded them into the application.
- The CBLD AI/ML approach is a first step; annual updates provide opportunities to:
  - Enhance the coding app, apply other AI/ML approaches, and identify additional and more detailed patterns.
  - Identify legal similarities tied to shared legal histories (for instance, jurisdictions with a colonial past or strong trading ties) and deviations over time.

### Policy relevance and next steps
- The AI/ML outcomes could help countries that wish to:
  - Amend their central bank legislation while avoiding legal pitfalls or further enhancing their legal framework.
- Recommended collaborative actions:
  - The IMF should work together with selected academic partners and country authorities to better understand AI/ML outcomes and finetune methods.
  - Improvements would ultimately benefit external users through more enhanced search options and greater access to CBLD data.

### Operational inventory (from Annexes)
- The coding app development and integration details are documented in Annex I.
- The CBLD includes 273 search categories across main topics listed in Annex II.

*Source: Conclusion — Predicting the Law: Artificial Intelligence Findings from the IMF’s Central Bank Legislation Database (Working Paper No. WP/2023/241).*

---


_Source: https://www.imf.org/-/media/files/publications/wp/2023/english/wpiea2023241-print-pdf.pdf_
