## wpiea2025199-source-pdf

## Source details

**Canonical URL:** [wpiea2025199-source-pdf](https://www.imf.org/-/media/files/publications/wp/2025/english/wpiea2025199-source-pdf.pdf)

## Other formats

- [Markdown version](/-/media/files/publications/wp/2025/english/wpiea2025199-source-pdf.pdf.md)
- [Structured JSON version](/-/media/files/publications/wp/2025/english/wpiea2025199-source-pdf.pdf.json)

---

### I. Introduction and scope
- Financial supervisory authorities must upgrade toolkits to keep pace with financial sector innovations and intensified use of digital technology combining vast data, cloud processing, and AI systems.
- Supervisory authorities face challenges in monitoring a rapidly evolving technology landscape and are compelled to harness data-driven tools and implement more efficient supervisory approaches to fulfill mandates.
- Purpose of the paper:
  - Outline essential prerequisites for financial supervisory authorities implementing AI initiatives.
  - Provide a detailed toolkit (governance, risk management, project management, team composition).
  - Build on existing Fund positions; does not offer new policy positions.
  - Section 2 focuses on internal governance, risk management, project management, and project team composition.

### Suptech and AI adoption statistics and trends
- Adoption and survey statistics:
  - "164 financial authorities from 105 countries" have implemented suptech tools.
  - In 2024, "75 percent" of advanced economies and "58 percent" of emerging markets and developing economies had adopted suptech tools; in 2023 these figures were "79 percent" and "54 percent", respectively.
  - Suptech survey sample sizes: "64 authorities responded to the 2023 survey vis-à-vis 164 in 2024."
  - "60 percent" of respondents are exploring how to incorporate AI into supervisory processes.
  - Use of GenAI more than doubled between 2023 and 2024 (from "eight" to "19 percent").
  - A stocktake found "32 out of 42 respondents" experimenting with, using, or developing GenAI tools in financial supervision.

### Opportunities and supervisory use cases of AI
- Data management contributions:
  - AI can assist validation, consolidation, and visualization of supervisory data.
- Examples of supervisory applications:
  - Substitute manual quality controls (completeness, correctness, consistency of formatting and calculation) with AI while adhering to reporting rules.
  - Combine web scraping and AI to detect suspicious granular-data patterns for AML/CFT supervision.
  - Assist market surveillance by identifying content in regulatory filings warranting enforcement.
  - Identify misleading marketing and perform real-time monitoring of market transactions.
  - Enhance prudential supervision of credit and liquidity risks.
  - Facilitate document review during licensing and governance assessments.
  - Predict liquidity problems for participants in financial market infrastructures and design resilience-test scenarios.
- GenAI-specific potentials:
  - Enhances information retrieval, content creation, code generation, debugging, and legacy code optimization.
  - Accelerates deployment of AI in fraud detection, market monitoring, and data management.
  - LLMs can scan board meeting minutes to identify discussions on executive compensation, related party transactions, and risk management to support corporate governance supervision.

### Examples of AI adoption by authorities
- Bank of England: supports macro-financial and macro-prudential surveillance (GDP growth forecasts, banking distress, financial crises).
- Bank of Thailand: analyzes board meeting minutes for prudential supervision and regulatory compliance.
- European Central Bank: NLP and AI to read fit-and-proper questionnaires and flag issues to streamline authorization.
- Malaysian Securities Commission: monitors corporate governance practices and disclosure quality on the Malaysia Stock Exchange.

### Challenges, constraints, and risks in deploying AI
- Key deployment challenges:
  - Developing solutions that are explainable, robust, and prevent discrimination.
  - Limited skilled resources.
  - Incompatibility of legacy systems and existing data and model governance frameworks with AI model specificities.
  - Public sector characteristics: competing objectives, complex organizational structures, procurement transparency requirements.
- Organizational implications:
  - Embedding AI often necessitates suptech strategy, enhanced data collection, and a long-term IT supervisory ecosystem view.
- Resource and model suitability constraints:
  - Deployment of GenAI models requires significant skilled IT personnel and computing power.
  - LLMs trained on diverse sources may be unsuitable for specific supervisory requirements and require substantial adaptation.

### Core project-management approach and rationale
- Methodology recommendation:
  - Use a structured data-centered methodology adapted to AI and big-data projects; CRISP-DM is presented as a widely used baseline.
  - CRISP-DM attributes: explains results to diverse stakeholders and provides granular resource estimation to aid monitoring and accountability.
  - CRISP-DM may require adjustments when data availability is limited or objectives are unclear; perform a D.A.T.A. pre-assessment to reduce adjustments.
  - Favor agile, iterative execution and integrate risk and IT considerations from the outset.

### D.A.T.A. framework (pre-assessment)
- Purpose: predict success likelihood of AI/big-data initiatives and prioritize projects.
- Four sequential yes/no questions:
  - Do we have access to data that is valuable and rare?
  - Can employees use data to create solutions?
  - Can our technology provide the solution?
  - Is our solution in accordance with laws and ethics?
- Scoring and interpretation:
  - "4": high likelihood of success.
  - "3": requires critical adjustments to achieve success.
  - "2": unlikely to succeed.
  - "1": success very unlikely.
  - "0": project should be terminated immediately.
- Empirical motivation: a representative study found that ""85 percent"" of big data initiatives fail because executives cannot accurately assess project risks in advance.

### CRISP-DM adapted for supervisory AI projects: the six phases
- Phases used/adapted:
  1. Project Foundation (Business Understanding broadly interpreted)
  2. Data Understanding
  3. Data Preparation
  4. Modelling
  5. Evaluation
  6. Deployment
- Execution guidance:
  - Phases are iterative; later findings may require revisiting earlier phases.
  - Use agile methodologies and DevOps/MLOps for iterative cycles and continuous improvement.
  - Initial phase should expand scope to include risk and IT integration for supervisory context.

### Project Foundation: objectives, metrics, and planning
- Key tasks:
  - Justify contribution to supervisory business objectives (e.g., risk-based supervision).
  - Select performance metrics: KPIs (business) and technical indicators (model performance).
  - Assess risks (human resources, IT infrastructure, cybersecurity, funding, organizational culture).
  - Produce a detailed project plan specifying business problem, verification criteria, challenge assessment, technical/functional requirements, and resource estimation.
- Recommendations:
  - Avoid vague objectives; define scope and exclusions.
  - Prefer "low cost first" / modest-scope projects to build trust and learning.
  - Examples of KPIs and technical indicators: increase by a certain percentage in daily loan files reviewed; minimum reduction in average days to process fit-and-proper documents; accuracy; precision; recall; root mean squared error.

### Data governance, privacy, and cybersecurity
- Data governance as bedrock for explainability, robustness, and bias mitigation:
  - Policies, processes, and activities to ensure data is secure, private, accurate, available, and usable across lifecycle.
  - Recommended practices: metadata recording, asset inventories/data catalogs, metadata registries, APIs for exploration.
- Privacy techniques and concerns:
  - Anonymization reduces identifiers but "significant privacy risks remain."
  - EU GDPR pseudonymization definition preserved in source.
  - Federated learning and other privacy technologies highlighted for sharing/training without revealing raw data.
- Cybersecurity:
  - Must be considered at design, development/procurement, deployment, and operations stages; "cannot and should not be an afterthought."
  - References ongoing standards work: ISO/IEC 27090 likely to be published in "2025".
  - Adhere to three-lines of defense: business controls, risk management, independent audit.
- Third-party risk management:
  - Require frameworks covering due diligence, monitoring, independent assurance, business continuity, disaster recovery, and exit plans.
  - Clarify when evaluating third-party data/models: data quality, sources, ownership, traceability, model type, learning method, biases, autonomy level, extent of human oversight.
  - Align third-party AI risks with corporate risk appetite and legal agreements.

### Agile, DevOps, DevSecOps, and MLOps in AI projects
- Agile:
  - Iterative sprints, rapid feedback, flexibility, stakeholder onboarding, alignment with business objectives; preferred over Waterfall for AI projects.
- DevOps and DevSecOps:
  - DevOps shortens and improves software lifecycle; useful for deploying supervisory AI tools.
  - DevSecOps integrates security testing continuously; recommended after DevOps and continuous integration.
- MLOps:
  - Applies DevOps principles to ML models and data: automates ingestion, preparation, modeling, evaluation, deployment, monitoring, and retraining.
  - Helps address model degradation when post-deployment data differs from training data.

### Project team composition and role priorities
- Recommended core roles (multi-disciplinary):
  - Business owner(s)
  - Project manager / connector
  - Data engineers
  - Data scientists
  - AI engineers
  - Software engineers
  - Optional/supplementary: DevOps engineers, cloud engineers, user experience engineers, test engineers, Legal, Ethics
- High-level responsibilities:
  - Business owner: defines use case, success criteria, operational tasks, selects KPIs, ensures alignment after deployment, provides system access.
  - Project manager: bridges business and technical teams, sets milestones, ensures timely delivery.
  - Data engineer: prepares/organizes raw data; ensures dataset quality and usability.
  - Data scientist: selects representations, cleans data, trains models, evaluates alignment with KPIs, visualizes results.
  - AI engineer: builds production AI systems; oversees deployment and monitoring.
  - Software engineer: develops/tests software; maintenance and integration; conducts systems risk/reliability analysis.
  - Ethics and Legal: review risk frameworks, draft data principles, ensure legal compliance and contract review.
- Resource-scarce prioritization:
  - Prioritize business owner, data scientist, and software engineer (data scientists handle data prep; software engineers handle production deployment and monitoring).

### Modeling, validation, and evaluation best practices
- Model selection and training:
  - Choose model family by business problem (clustering, regression, classification, time series, textual analysis).
  - Train using partitioned datasets (training/validation/testing); favor cross-validation for small/complex datasets.
- Sampling and validation methods preserved:
  - Holdout method
  - Cross-validation (k-fold)
  - Bootstrap sampling
  - Stratified sampling
  - Rolling time series cross validation (recommended for time series)
- Validation and challenger models:
  - Encourage challenger models to compare production outputs against alternatives.
- Documentation:
  - Maintain thorough documentation of model design choices, data sources, validation techniques, and decision-making for reproducibility, auditing, and knowledge transfer.

### Bias detection, explainability, and evaluation metrics
- Bias detection metrics (non-exhaustive):
  - Confusion Matrix Analysis (true positives, true negatives, false positives, false negatives)
  - Disparate Impact
  - Statistical Parity Difference
  - Equality of Opportunity
  - Predictive Equality
  - Calibration
- Explainability guidance:
  - Prefer simpler models when possible; higher explainability required for high-materiality applications (e.g., stress testing).
  - Explainable AI methods (e.g., LIME, Shapley) can help but are not a panacea; check clarity, simplicity, broadness, completeness, and soundness.
  - Tailor explanation level to recipient (head of supervision vs. data scientist).
- Evaluation decisions:
  - Combine technical indicators and KPIs selected during Project Foundation.
  - Evaluate technical robustness, regulatory constraints, bias, and explainability prior to deployment.
  - Conclude phase with decision to proceed, adjust, or abort.

### Deployment, monitoring, and lifecycle controls
- Deployment tasks:
  - Integrate model into existing IT infrastructure; define user, purpose, real-time requirements, and re-run frequency.
  - Ensure seamless integration with supervisory tools and workflows.
- Monitoring and maintenance:
  - Provide monitoring/maintenance plans addressing model deterioration, distributional shifts, and unexpected inputs.
  - Use statistical distance measures (Kolmogorov-Smirnov, Wassertein) to detect shifts between training and new data distributions.
  - Maintain an inventory of deployed models proportional to impact; high-impact models should include purpose, limitations, validation findings, and governance metadata.
- Documentation and handoff:
  - Produce a summary report narrating alignment with business objectives, value created, and actionable recommendations.
  - Ensure knowledge transfer and accessible explanations for end users and governance stakeholders.
- Deployment decision factors:
  - Conclude Modeling and Evaluation by stating model purpose, synthesizing results, detailing major limitations and key assumptions.
  - Final review to evaluate lessons learned and enhance future efforts.
  - Build data pipelines in accordance with MLOps practices.

### How the methodology contributes to the supervisory process
- Enables evaluation of compliance with regulatory requirements related to model risk management, AI governance, and data governance.
- Helps assess whether supervised entities deployed solutions in accordance with international standards, local regulations, and best practices.
- Deployment benefits:
  - Ensure supervisory skills/practices remain up to date while providing technological resources to process information efficiently.
  - Streamline information processing and enhance skills in data and model governance.
  - Leverage data governance best practices to assess supervised entities’ controls on data quality and IT support for data aggregation and risk reporting.

### Supervisory capacity building and validation
- AI governance experience improves ability to evaluate boards’ understanding of AI solutions and deployment risks.
- Model validation equips supervisors to assess models under normal and stressed conditions.
- Bias mitigation in deployed solutions helps ensure fair treatment in AI-priced products.
- Assess whether supervised entities maintain comprehensive documentation of AI solutions.

### Project design, governance, and implementation best practices
- AI projects require careful reflection, planning, and resources across technology, processes, and people.
- Assess IT infrastructure and human resources before undertaking AI projects.
- Start with small-scope projects to establish a learning curve prior to larger initiatives.
- Ensure knowledge transfer whether projects are in-house or outsourced.
- Project teams should be diverse and adopt iterative approaches to assimilate lessons quickly.
- Select business cases aligned with RBS to secure organizational buy-in.
- AI should augment supervisory capabilities and enable focus on tasks requiring supervisory judgment.
- Adopt data and AI governance frameworks to ensure legal/regulatory compliance, alignment with organizational objectives, and societal values; include checks and balances throughout model lifecycle.

### Cybersecurity, third-party risk, and foundational elements
- Treat cybersecurity as an integral design-to-deployment consideration due to new AI-related cyber risks.
- Evaluate challenges of third-party development/deployment and apply good third-party risk management practices:
  - Policies, processes, risk management activities, due diligence, ongoing monitoring, independent assurance, exit plans, business continuity, disaster recovery.
- Strengthen foundational elements with AI-tailored risk management practices:
  - Explainability proportional to use-case materiality and legal constraints.
  - Combat discrimination with robust data collection, diverse development teams, and continuous monitoring.
  - Provide training on data privacy/governance and reorganize teams to leverage IT expertise.
  - Evaluate data repositories for unstructured data capacity and ensure legacy systems have sufficient processing power.
  - Apply privacy techniques for legal compliance when sharing data with third parties.

### Data-centric project management and lifecycle controls
- Adopt a data-centric methodology incorporating risk-mitigating techniques.
- Prevent failures by clearly defining use case aligned with business objectives for accurate project risk assessment.
- Subsequent stages: proper data treatment, model optimization, and outcome assessment from technical and business perspectives.
- Deploy only if the model meets objectives set in the initial project stage.
- Deployment must include monitoring and maintenance to mitigate decaying performance.

### Considerations for emerging markets and developing economies
- Emerging markets and developing economies lag in AI adoption despite advances in supervisory AI tools.
- Novel AI forms can improve information processing efficiency and facilitate code generation to mitigate insufficient engineering skills, but deployment requires significant resources and careful risk consideration.
- Appropriate project-management methodology assists alignment with supervisory objectives, but organizational buy-in and adequate human and IT resources are essential.

*Source: IMF Working Paper — AI Projects in Financial Supervisory Authorities (References and introductory material; Section 3 and Conclusions).*

### References .............................................................................................................

### References

### Glossary
- AI: Artificial Intelligence
- AML/CFT: Anti-Money Laundering/Combating the Financing of Terrorism
- BIS: Bank for International Settlements
- CCAF: Cambridge Center for Alternative Finance
- CRISP-DM: Cross Industry Standard for Data Mining
- D.A.T.A.: Data, Autonomy, Technology, and Accountability
- DevOps: Development and Operations
- FSI: Financial Stability Institute
- GDPR: General Data Protection Regulation
- IT: Information Technology
- ML: Machine Learning
- MLOps: Machine Learning Operations
- OECD: Organization for Economic Co-operation and Development
- RBS: Risk-Based Supervision

### Figures and Tables referenced in the unit
- Figures
  - 1. Key Supervisory Activities under RBS
  - 3. Project Team
  - 4. The six phases of an AI project
  - 5. Technological infrastructure for advance analytics
- Tables
  - 1. Roles and Responsibilities in an AI Project Team

### I. Introduction — context and scope
- Financial supervisory authorities must upgrade and adapt toolkits to keep pace with financial sector innovations.
- Financial institutions have intensified the use of digital technology by combining vast amounts of data with cloud processing and AI systems to deploy customer centric solutions, streamline business processes, and redesign risk management.
- Supervisory authorities face challenges in monitoring a rapidly evolving technology landscape and are compelled to harness data-driven tools and implement more efficient supervisory approaches to fulfill mandates.

### Suptech and AI adoption statistics and trends
- Emerging research suggests that 164 financial authorities from 105 countries have implemented suptech tools.
- In 2024, 75 percent of advanced economies and 58 percent of emerging markets and developing economies had adopted suptech tools, compared to 79 percent and 54 percent, respectively, in 2023.
- Suptech survey sample sizes: 64 authorities responded to the 2023 survey vis-à-vis 164 in 2024.
- 60 percent of respondents are exploring how to incorporate AI into supervisory processes.
- Use of GenAI more than doubled between 2023 and 2024 (from eight to 19 percent).
- A recent stocktake indicated that 32 out of 42 respondents are experimenting with, using or developing GenAI tools in financial supervision.
- Common GenAI use cases in supervision are grouped into three categories: (i) basic document processing; (ii) knowledge management; and (iii) document review.

### Opportunities and use cases of AI for financial supervisory authorities
- AI systems can contribute to data management tasks such as validation, consolidation, and visualization.
- Examples of supervisory applications:
  - Substitute manual quality controls (completeness, correctness, consistency of formatting and calculation) with AI while adhering to reporting rules.
  - In combination with web scraping, detect suspicious patterns in granular data for AML/CFT supervision.
  - Assist market surveillance by identifying content in investment advisers’ regulatory filings that may warrant enforcement action.
  - Identify misleading marketing and perform real-time monitoring of market transactions.
  - Enhance prudential supervision of risks, notably credit and liquidity risks.
  - Facilitate review of documents during licensing and governance assessments.
  - Predict liquidity problems affecting participants in financial market infrastructures for financial stability purposes and design scenarios to test resilience.
- GenAI-specific potentials:
  - Enhances information retrieval, content creation, code generation, debugging, and legacy code optimization.
  - Accelerates deployment of AI in traditional use cases such as fraud detection, market monitoring, and data management.
  - Large Language Models (LLMs) can scan board meeting minutes to identify discussions on executive compensation, related party transactions, and risk management to support corporate governance supervision.

### Examples of AI adoption by authorities
- Bank of England: AI tools to support macro-financial and macro-prudential surveillance, including GDP growth forecasts, banking distress, and financial crises.
- Bank of Thailand: AI to support prudential supervision, including analyzing board meeting minutes for regulatory compliance.
- European Central Bank: NLP and AI tools to read fit-and-proper questionnaires and flag issues to streamline authorization.
- Malaysian Securities Commission: AI tools to monitor adoption of corporate governance best practices and quality of disclosures of listed companies on the Malaysia Stock Exchange.

### Challenges, constraints, and risks in deploying AI
- Key deployment challenges:
  - Developing solutions that are explainable, robust, and capable of preventing discrimination.
  - Operating with limited skilled resources.
  - Incompatibility of legacy systems and existing data and model governance frameworks with AI model specificities.
  - Public sector characteristics: competing objectives, complex organizational structures with multiple specializations and different priorities, and procurement processes requiring extensive transparency.
- Embedding AI tools often necessitates broader organizational digital transformation efforts:
  - Adopting a suptech strategy.
  - Enhancing data collection practices.
  - Developing a long-term view of the IT supervisory ecosystem.
- Limited preparedness and access may hinder AI adoption, particularly in emerging markets and developing economies.
- Deployment of GenAI models necessitates significant resources, including skilled IT personnel and sufficient computing power.
- LLMs trained on diverse data sources may not be suitable for specific supervisory requirements and may require substantial adaptations.

### Purpose and approach of the paper
- The paper aims to outline essential prerequisites that financial supervisory authorities should consider when implementing AI initiatives and to provide a detailed toolkit for guidance.
- It builds on risk considerations of AI identified in previous IMF publications and explores these considerations from the perspective of financial supervisory authorities.
- The paper does not offer new policy positions but uses existing Fund positions to develop this guide.
- Section 2 outlines key elements of internal governance, risk management, project management and project team composition that financial supervisory authorities should consider when planning to initiate an AI project.

*Source: IMF Working Paper — AI Projects in Financial Supervisory Authorities (References and introductory material).*

### Section 3 emphasizes a project management methodology adapted to the

### Section 3 emphasizes a project management methodology adapted to the iterative nature of AI initiatives, prioritizing the achievement of business objectives while incorporating techniques to mitigate the unique risks associated with AI.

### Core project-management approach and rationale
- Emphasizes a methodology adjusted to AI and big-data specific challenges; CRISP-DM is presented as a structured, widely used methodology with attributes that:
  - Serve project management goals for data-centered projects used for over 20 years.
  - Help explain results to diverse stakeholders.
  - Provide granular estimation of resources to ease monitoring and accountability.
- Notes that CRISP-DM may require adjustments when data availability is limited or business objectives are unclear; recommends an initial assessment using the D.A.T.A. framework to reduce the need for adjustments.
- Recommends an agile, iterative style of execution to reflect AI projects’ cyclic discovery and improvement needs and to integrate risk and IT considerations.

### D.A.T.A. framework (pre-assessment)
- Purpose: predict the success likelihood of AI/big-data initiatives and prioritize projects.
- Four components/questions (sequential, each answered yes/no):
  - Do we have access to data that is valuable and rare?
  - Can employees use data to create solutions?
  - Can our technology provide the solution?
  - Is our solution in accordance with laws and ethics?
- Scoring and interpretation:
  - 4: high likelihood of success.
  - 3: requires critical adjustments to achieve success.
  - 2: unlikely to succeed.
  - 1: success very unlikely.
  - 0: project should be terminated immediately.
- Empirical motivation: a representative study found that "85 percent" of big data initiatives fail because executives cannot accurately assess project risks in advance.

### CRISP-DM adapted for supervisory AI projects: the six phases
- CRISP-DM original phases (interpreted/used here):
  1. Project Foundation (Business Understanding broadly interpreted)
  2. Data Understanding
  3. Data Preparation
  4. Modelling
  5. Evaluation
  6. Deployment
- Execution guidance:
  - Phases are iterative, not strictly sequential; findings in later phases can require revisiting earlier work.
  - Agile methodologies and DevOps/MLOps practices support iterative cycles and continuous improvement.
  - The initial phase should expand scope to include risk and IT integration to reflect supervisory context.

### Project Foundation: objectives, metrics, and planning
- Key tasks:
  - Justify project contribution to supervisory business objectives (e.g., risk-based supervision).
  - Select performance metrics: KPIs (business) and technical indicators (model performance).
  - Assess risks and challenges (human resources, IT infrastructure, cybersecurity, funding, organizational culture).
  - Produce a detailed project plan specifying business problem, verification criteria, challenge assessment, technical and functional requirements, and resource estimation.
- Recommendations:
  - Avoid vague objectives (“do something with all our data”); define what the project will not address.
  - Prefer “low cost first” / modest-scope projects to build trust and learning curves.
  - Examples of KPIs: increase by a certain percentage in daily loan files reviewed; minimum reduction in average days to process fit-and-proper documents.
  - Technical indicators examples: accuracy, precision, recall, root mean squared error.

### Data governance, privacy, and cybersecurity
- Data governance:
  - Identified as the bedrock for explainability, robustness, and bias mitigation.
  - Encompasses policies, processes, and activities to ensure data is secure, private, accurate, available, and usable throughout its lifecycle.
  - Recommends data management practices: metadata recording, asset inventories/data catalogs, metadata registries, APIs for exploration.
- Privacy techniques and concerns:
  - Anonymization reduces identifiers but "significant privacy risks remain."
  - EU GDPR pseudonymization definition preserved in source text.
  - Privacy technologies (e.g., federated learning) highlighted as promising for sharing/training without revealing raw data.
- Cybersecurity:
  - Should be considered at design, development/procurement, deployment, and operations stages; "cannot and should not be an afterthought."
  - References ongoing standards work: ISO/IEC 27090 likely to be published in "2025".
  - Recommends adherence to the three-lines of defense: business controls, risk management, independent audit.
- Third-party risk:
  - If third parties develop/deploy AI, require a third-party risk management framework covering due diligence, monitoring, independent assurance, business continuity, disaster recovery, and exit plans.
  - When evaluating third-party data/models, clarify data quality, sources, ownership, traceability, model type, learning method, biases, autonomy level, and human oversight extent.
  - Align third-party AI risks with corporate risk appetite and legal agreements.

### Agile, DevOps, DevSecOps, and MLOps in AI projects
- Agile methodologies:
  - Based on the "Manifesto for Agile Software Development"; favor iterative sprints, rapid feedback, flexibility, stakeholder onboarding, and alignment with business objectives.
  - Preferred alternative to Waterfall for AI projects; helps identify problems early and reduce costly delays.
- DevOps and DevSecOps:
  - DevOps integrates Development and Operations to shorten and improve the software lifecycle; useful for deploying supervisory AI tools.
  - DevSecOps integrates security testing at every stage; recommended after implementing DevOps and continuous integration.
- MLOps:
  - Applies DevOps principles specifically to ML models and data.
  - Automates data ingestion, preparation, modeling, evaluation, deployment, continuous monitoring, and retraining.
  - Helps address model degradation when post-deployment data differs from training data.

### Project team composition and role priorities
- Core team roles recommended (multi-disciplinary):
  - Business owner(s)
  - Project manager / connector
  - Data engineers
  - Data scientists
  - AI engineers
  - Software engineers
  - Optional/supplementary roles: DevOps engineers, cloud engineers, user experience engineers, test engineers, Legal, Ethics.
- Role responsibilities (high-level):
  - Business owner: defines use case, success criteria, operational tasks, selects KPIs, ensures alignment after deployment, provides system access.
  - Project manager: bridges business and technical teams, sets milestones, ensures timely delivery.
  - Data engineer: prepares/organizes raw data, responsible for dataset quality and usability.
  - Data scientist: selects representations, cleans data, trains models, evaluates alignment with business KPIs, visualizes results.
  - AI engineer: builds production AI systems, oversees deployment and monitoring.
  - Software engineer: develops/tests software, performs maintenance and integration, conducts systems risk/reliability analysis.
  - Ethics and Legal representatives: review risk frameworks, draft data principles, ensure legal compliance and contract review.
- When resources are scarce: prioritize business owner, data scientist, and software engineer (with data scientists handling data preparation and software engineers handling production deployment and monitoring).

### Modeling, validation, and evaluation best practices
- Modeling:
  - Choose model family based on business problem (clustering, regression, classification, time series, textual analysis).
  - Train using partitioned datasets (training/validation/testing); favor cross-validation for small or complex datasets; be aware of sampling methods and when to use them:
    - Holdout method
    - Cross-validation (k-fold)
    - Bootstrap sampling
    - Stratified sampling
    - Rolling time series cross validation (recommended for time series)
- Validation techniques and sampling guidance preserved as stated in source.
- Encourage challenger models to compare production model outputs against alternatives.
- Documentation:
  - Maintain thorough documentation of model design choices, data sources, validation techniques, and decision-making to support reproducibility, auditing, and knowledge transfer.

### Bias detection, explainability, and evaluation metrics
- Bias detection metrics commonly used (non-exhaustive list preserved):
  - Confusion Matrix Analysis (true positives, true negatives, false positives, false negatives).
  - Disparate Impact.
  - Statistical Parity Difference.
  - Equality of Opportunity.
  - Predictive Equality.
  - Calibration.
- Explainability guidance:
  - Prefer simpler models when possible; higher explainability required for high-materiality applications (e.g., stress testing).
  - Explainable AI methods (e.g., LIME, Shapley) can help but are not a panacea; check for clarity, simplicity, broadness, completeness, and soundness.
  - Tailor the level of explanation to the recipient (e.g., head of supervision vs. data scientist).
- Evaluation decisions:
  - Combine technical indicators and KPIs selected during Project Foundation.
  - Evaluate technical robustness, regulatory constraints, bias, and explainability prior to deployment.
  - Conclude phase with decision to proceed, adjust, or abort.

### Deployment and monitoring
- Deployment tasks:
  - Integrate model into existing IT infrastructure; define user, purpose, real-time requirements, and re-run frequency.
  - Ensure seamless integration with supervisory tools and workflows.
- Monitoring and maintenance:
  - Provide a comprehensive monitoring and maintenance plan addressing model deterioration, distributional shifts, and unexpected inputs.
  - Use statistical distance measures (Kolmogorov-Smirnov, Wassertein) to detect shifts between training and new data distributions.
  - Maintain an inventory of deployed models with detail levels proportional to model impact; high-impact models (e.g., stress testing) should include purpose, limitations, validation findings, and governance metadata (who validated, when).
- Documentation and handoff:
  - Produce a summary report that narrates alignment with business objectives, highlights value created, and proposes actionable recommendations.
  - Ensure knowledge transfer and accessible explanations for end users and governance stakeholders.

*Source: https://www.imf.org/-/media/files/publications/wp/2025/english/wpiea2025199-source-pdf.pdf*

### conclusions of the Modeling and Evaluation phases, stating the model purpose, synthesizing results, and

### IV. Conclusions

### Modeling and Evaluation: purpose, synthesis, and limitations
- Conclude the Modeling and Evaluation phases by stating the model purpose, synthesizing results, and detailing major limitations and key assumptions.
- Recognize that the adopted AI governance structure will influence who participates in the deployment decision.
- Require a final review aimed at evaluating what went well, what could have been better, and how to enhance future efforts.
- Emphasize building up data pipelines in accordance with MLOps practices during this phase.

### How the methodology can contribute to the supervisory process
- The AI project management methodology can be used by supervisory authorities to evaluate compliance with regulatory requirements related to model risk management, AI governance, and data governance.
- A thorough understanding of AI project phases and the corresponding governance framework enables assessment of whether supervised entities deployed solutions in accordance with international standards, local regulations, and best practices.
- Deployment of AI solutions by financial supervisory authorities will:
  - Help ensure that supervisory skills and practices remain up to date while providing technological resources to process information from supervised entities efficiently. 40
  - Streamline information processing and enhance supervisory skills in areas such as data and model governance.
  - Leverage best practices in data governance to assess whether supervised entities have adequate controls on data quality throughout its lifecycle and whether IT infrastructure supports data aggregation and risk reporting. 41

### Supervisory capacity building and validation
- Experience in establishing an AI governance structure assists supervisory authorities in evaluating boards’ understanding of AI solutions, their limitations, and associated deployment risks.
- Implementing model validation techniques equips supervisory authorities to assess models’ outcomes and behavior under normal and stressed conditions in supervised entities.
- Implementing bias mitigation strategies in deployed solutions can help ensure consumers receive fair treatment in AI-priced products.
- Supervisory authorities can assess whether supervised entities maintain comprehensive documentation of AI solutions, including design choices, data sources, and decision-making processes.

### Roles of AI in supervision and operational guidance
- AI solutions can assist financial supervisory authorities in keeping up with the rapidly digitalizing financial sector by efficiently processing vast amounts of data from diverse sources and detecting hidden patterns.
- By automating routine tasks (compiling, validating, summarizing information), authorities can focus on exercising supervisory judgment.
- Ongoing adoption trends suggest financial sector authorities will need to further incorporate AI into their supervisory toolkit; AI can transform capital markets—making them more efficient but also more volatile—and authorities must prepare for this new world (IMF 2024).
- AI can facilitate:
  - Real-time monitoring
  - Process automation
  - Forward-looking modelling
  - Increased efficiency and effectiveness of the supervisory process

### Project design, governance, and implementation best practices
- AI projects require careful reflection, planning, and resources, considering technology, processes, and people.
- Authorities should assess IT infrastructure and human resources before undertaking AI projects.
- Begin with small-scope projects to establish a learning curve prior to larger projects.
- Knowledge transfer is critical whether projects are developed in-house or outsourced.
- Project teams should be diverse, with representatives from across the organization and a range of skills.
- Adopt an iterative approach to assimilate lessons learned during execution quickly.
- Select business cases aligned with the objectives of RBS to ensure organizational buy-in.
- AI solutions should augment supervisory capabilities and enable focus on tasks where supervisory judgment is critical.
- Adopt data and AI governance frameworks to ensure legal and regulatory compliance, alignment with organizational objectives, and societal values, with appropriate checks and balances throughout the model lifecycle.
- Depending on materiality, numerous stakeholders with different skills may be involved in the deployment decision; legal, ethical, technical, cultural, and business perspectives should be factored in for responsible adoption.

### Cybersecurity, third-party risk, and foundational elements
- AI introduces new sources of cyber risk; cybersecurity must be considered from design to deployment.
- Evaluate challenges associated with developing or deploying AI solutions that involve third parties.
- Good practices in third-party risk management should be in place, including:
  - Policies
  - Processes
  - Risk management activities
  - Due diligence
  - Ongoing monitoring
  - Independent assurance
  - Exit plans
  - Business continuity arrangements
  - Disaster recovery arrangements
- Strengthen foundational elements by implementing risk management practices tailored to AI applications:
  - Ensure an appropriate level of explainability based on use case materiality and legal constraints, favor interpretable solutions where feasible.
  - To combat discrimination: robust data collection, diverse AI development teams, and continuous monitoring are essential.
  - Provide training on data privacy and governance; reorganize teams to leverage IT expertise as needed.
  - Evaluate data repositories for capacity to handle unstructured data and ensure legacy systems have sufficient processing power; infrastructure updates generally precede AI model deployment.
  - Apply privacy techniques to facilitate compliance with data protection laws when sharing data with third parties.

### Data-centric project management and lifecycle controls
- Adopt a data-centric project management methodology incorporating risk-mitigating techniques.
- Prevent failures by clearly defining the use case aligned with business objectives to enable accurate project risk assessment.
- Subsequent stages: proper data treatment, model optimization, and assessment of outcomes from technical and business perspectives.
- Deploy only if the model meets objectives set in the initial project stage.
- Deployment stage must include monitoring and maintenance to mitigate the risk of decaying performance.

### Considerations for emerging markets and developing economies
- While AI tools for streamlining supervisory processes have advanced significantly, emerging markets and developing economies lag in adoption.
- Novel AI forms can improve information processing efficiency and facilitate code generation to mitigate insufficient engineering skills, but deploying them requires significant resources and careful risk consideration.
- An appropriate project management methodology can assist in aligning AI deployments with supervisory objectives, but organizational buy-in and adequate human and IT resources remain essential.

*Source: IMF Working Paper “AI Projects in Financial Supervisory Authorities”, conclusions section; https://www.imf.org/-/media/files/publications/wp/2025/english/wpiea2025199-source-pdf.pdf*

---


_Source: https://www.imf.org/-/media/files/publications/wp/2025/english/wpiea2025199-source-pdf.pdf_
