## tarea2024010

## Source details

**Canonical URL:** [tarea2024010](https://www.imf.org/-/media/files/publications/tar/2024/english/tarea2024010.pdf)

## Other formats

- [Markdown version](/-/media/files/publications/tar/2024/english/tarea2024010.pdf.md)
- [Structured JSON version](/-/media/files/publications/tar/2024/english/tarea2024010.pdf.json)

---

### Mission activities
- Conducted four workshops and five one-to-one sessions to provide guidance on:
  - Ideal organizational arrangements to support data analytics
  - How to access and utilize better and more data, including third party data
  - Current data workflows and processes with a view to streamlining and improving them
  - How to improve data integrity
  - Using data science and big data analytics to strengthen risk assessment
  - Leveraging the value of country-by-country reports
  - Developing several data analytics/risk assessment models

### Key findings
- SFA has high-quality data and competent staff able to benefit from data analytics.
- Staff engagement and stamina were strong; participants built four new pilot risk assessment models during the mission that showed promising results.
- Example model result: the CIT audit selection model was estimated to increase the strike rate in comprehensive audits (measured as a correction above €5000) from 50-60 percent to roughly 90 percent (while doing the same number of audits).
- The SFA has an established “analysis unit” and more than 50 data applications in use, with an advanced predictive model under development in the IT department.
- The mission team was able to obtain high-quality data in a matter of days: audits, tax returns, Country-by-Country-Reports, business risk rules and balance sheets, all mergeable using unique taxpayer identifiers.

### Recommended senior management actions (five actions)
- a) Fully endorse, prioritize and incentivize the journey towards relying more on data analytics.
- b) Allow time for data analytics to grow (sometimes fail) and keep a stable and dedicated team in place.
- c) Invest in further recruitment and retention of data analysts and retrain operational staff to join the data analytics team.
- d) Prioritize development of high-value projects (deprioritize low-value projects) and vigilantly implement them.
- e) Have a constant focus on business needs and rigorously evaluate the success of projects.

### Technical assistance advice — Background and principles
- “Big Data” does not automatically translate into “Big Improvements”; advanced data analytics requires solving statistical, organizational and technical problems.
- Advanced Data Analytics: use of statistical techniques to make predictions and draw inferences about cause and effect; seeks to make judgments with more and better reliance on data.
- Machine learning models can incorporate all available taxpayer information (CIT, VAT, PIT, social security registers) and self-correct by learning from past mistakes.
- Advantages of advanced analytics are largest where data abundance and large populations meet scarce resources for manual examination; less advantageous when populations are small or data is difficult to make available.
- Effortless and automatic data integration is essential: propose enlarging datasets via internal integration (tax and customs) and external third-party automatic exchange; collect DAC2 and CRS-equivalent data from resident financial reporters; submit VAT special files (issued and received invoices) automatically on monthly or quarterly basis; require mandatory annual balance sheet submission before CIT returns—may require amendments to data exchange laws.
- Advanced analytics methods suggested include pattern recognition, outlier detection, cluster analysis, experimental design, network analysis, and text mining.

### Organizational structure, recruitment, and integration of solutions
- Risk identified: an organizationally detached and non-prioritized data analytics unit (satellite unit) risks being out of touch with end-user needs and lacking resources/authority.
- Measures to ensure integration and effectiveness:
  - Ensure employee flow between business units and the data science team (e.g., embedding “tours”).
  - Prioritize education of end-users; set aside time for end-user training for every application.
  - Aim for rapid deployment to avoid “white elephants”.
  - Rigorously measure impact of data analytics projects from end-user perspective; many existing applications lack systematic assessment.
  - Ensure model predictions are actionable and transparent.
- Mitigating cultural resistance:
  - Senior management endorsement and authority for the data science team.
  - Involve veteran employees and provide training in analytics basics.
  - Incentivize change and avoid punishing expected failures.
  - Translate efficiency gains into improved services rather than workforce reduction.
  - Prove value by competing new approaches with old approaches (e.g., split-sample audit selection).
- Recommended iterative deployment process:
  1. Understand business issues with end-users.
  2. Inventory and use easily accessible data before requesting new data.
  3. Build and evaluate a statistical model with end-users.
  4. Refine model with more/better data.
  5. Deploy when satisfied with theoretical properties.
  6. Evaluate and refine post-deployment.
- Model transparency techniques:
  - Use surrogate models (e.g., linear or decision tree) to proxy complex models.
  - Compute marginal effects (individual conditional expectation) for variables.
  - Make targeted predictions (e.g., predict “adjustment due to un(der)invoicing” rather than generic “adjustment”).

### Staffing, training, and tools
- Organizational stability is critical; frequent restructurings have harmed staff and momentum (example: Compliance Risk Management Unit (CRMU) was barely in place before being closed).
- Centralized analytics unit is appropriate for pooling scarce data analyst resources; unit should work across CIT, VAT, customs, and other areas.
- Analytics unit should have necessary authority and strong collaboration with Special Financial Office (SFO) and IT department.
- Prioritize recruiting/retaining few highly competent data analysts: “Five good data analysts can function as the core of a data analytic unit.”
- Consider special pay scales for specialized staff; do not let language barriers hinder recruitment; part-time university statistics students can be useful.
- Retrain selected operational staff into data analysts to ensure business understanding and recruitment.
- Train other departments in advanced data skills via structured multiday workshops with curriculum, homework and assignments.
- Consider analytical software packages with low entry barriers for onboarding non-trained data scientists (example: KNIME®); other open-source packages mentioned: WEKA, ANACONDA, RATTLE.

### Low-cost/high-reward initiatives in risk assessment
- Prioritize simple adjustments before larger investments in advanced analytics.
- Emphasize size in risk prioritization: size of risk = likelihood * consequence; interventions should weigh both likelihood and potential consequence.
- Audit case selection should weigh business risk scores by taxpayer size (sales, employees, taxes paid, etc.). Current VAT and TP business risk rules will increase overall level of risk by at most 10 percent when moving from a micro firm to a large firm—this understates variance in risk.
- Strike rate measurement recommendations:
  - Introduce thresholds so marginal corrections do not contaminate strike-rate statistics; a correction of €20 is rarely a meaningful audit success.
  - Strike rates from comprehensive audits fall by 10-20 percentage points when applying a €5000 threshold.
- Priority monitoring:
  - The 30 largest taxpayers should be continuously observed; create risk profiles for the 30 largest taxpayers (can be automated using TP indicators).
  - Establish one-to-one client account managers for these taxpayers; prioritize intervention (education, guidance or audit) if risk profiles suggest non-compliance.

### End-user education and operational efficiency
- End-user education should be prioritized to:
  - Ensure expanded usage of existing applications.
  - Limit data analytics resources spent on trivial work (e.g., reduce time Analysis unit spends rearranging data in reports).
- Example finding: business risk indicators set up for TP case selection were accurate in predicting the yield of CIT comprehensive audits. Utilizing existing work and educating end-users can realize such low-hanging fruits.

### Systematic evaluation of past interventions and business risk rules
- Previous audit strategies/business risk rules were not being systematically evaluated.
- Suggested simple evaluation methods:
  - Use past audit results to evaluate business risk rules, e.g., simple scatterplots.
  - Run “horse-races” by applying different approaches in parallel and comparing outcomes.
- Simple statistical improvements:
  - Using simple benchmarks and standard deviations can root business risk in systematic analysis rather than intuition.

### Multifaceted (butterfly) approach to compliance risk management
- Steps:
  - Group non-compliant taxpayers into three clusters: The unaware, the unable and the unwilling.
  - Deter non-compliance through education, service and exemplification targeted at the different clusters.
  - Detect non-compliance using risk modelling.
  - Deal with detected non-compliance based on the connected risk (likelihood*consequence) and the underlying reason (cluster).
- Recommendation: reestablish the Compliance Risk Management (CRM) Unit to support the CRM Committee and ensure 360o assessment of risk mitigation. The data analytics unit must serve the entire CRM spectrum.

### Proof of concept: pilot empirical models constructed
- Four pilot empirical models were constructed and handed over to the SFA (incl. additional models on customs not covered by the mission):
  - An advanced (random forest) CIT audit selection model (incl. deployment workflow)
  - An advanced (random forest) VAT audit selection model
  - Workflow to evaluate existing business risk rules against historical data (VAT/CIT)
  - Country-by-Country-Reporting anomaly detection tool
- All models targeted delivering real-time risk assessment using all available data and an optimal workflow where interventions and outcomes are fed back to improve the model.

### Pilot model performance and verification
- Pilot models outperformed existing risk rules.
  - Example: CIT audit selection model was estimated to increase the strike rate in comprehensive audits (measured as a correction above €5000) from 50-60 percent to 90 percent (while doing the same number of audits).
- Two demonstrated evaluation methods for business risk rules:
  - Correlate business risk rules with historical audit yields.
  - Use advanced audit selection models to inform relevance and weights of different risk rules.
- Country-by-Country-Reporting anomaly tool: compute benchmark profitability for each MNE group and industry and compare to subsidiary profitability in the Republic of Slovenia.
- Model limitations and recommended verification:
  - Use and verify all available data and test competing models.
  - Ensure SFA can fully comprehend models before implementation.
  - Aim for incremental improvement rather than perfection.
  - Minimize time from pilot build to small-scale implementation.
- Credible evaluation requires a randomized double-blind study with random assignment (use RAND in Excel as described).

### Prioritization principles and immediate priorities
- Considerations when prioritizing projects:
  - Data available (fit for data analytics?)
  - Size of closable tax gap (based on actual analysis or intuition)
  - Bidding round (how large resources will each division devote)
  - Political decision (importance beyond revenue)
  - Next best alternative
- Recommended immediate priorities (coming year):
  - Create risk profiles for the 30 largest taxpayers.
  - Create business risk indicators for CIT.
  - Put larger weight on firm size and redefine strike rates to reflect firm size.
  - Evaluate existing business risk rules using historical data.
  - Only after these simple adjustments proceed with more advanced modelling.

### Summary of suggested operational actions (timing and responsibilities preserved)
- Topic: Prioritization/endorsement by senior management
  - 1. Prioritize/endorse high-value work by creating protected time for data analytics projects. Deadline: Short term (within next 12 months). Responsibility: Director General (DG)/ Heads of department.
  - 2. Move towards “top-down” risk assessments in CIT by defining and using existing business risk indicators. Deadline: Short term (within next 12 months). Responsibility: Heads of department/Working Groups (WG).
  - 3. Increase emphasis on size of taxpayers in risk assessments. Deadline: Short term (within next 12 months). Responsibility: Heads of department/WG.
  - 4. Endorse a multifaceted approach to compliance risk management. Deadline: Short term (within next 12 months). Responsibility: Heads of department/Working Groups (WG).
  - 5. Move towards automatic exchange of third-party information from banks, other regulators and other relevant stakeholders. Deadline: Medium term (within two to three years). Responsibility: Heads of department/Relevant stakeholders.
- Topic: Improvements in data analytics and way-of-work (selected items)
  - 1. Readjust risk rules to reflect expected adjustment/risk. Deadline: Short term (within next 12 months). Responsibility: Analysis unit and relevant business owners. Implementation advice: multiply current risk rules with measures of size (turnover, tax paid, employees).
  - 2. Build automated reports for the largest taxpayers. Deadline: Short term (within next 12 months). Responsibility: Analysis unit and large taxpayer unit. Implementation advice: begin by building upon current TP business risk rules and update according to John Middleton’s 2022 report Appendix III; reports should indicate suggestive actions such as “contact taxpayer” or “audit”.
  - 3. Build CIT business risk indicators. Deadline: Short term (within next 12 months). Responsibility: Analysis unit and relevant business owners.
  - 4. Evaluate existing risk rules based on historical data and advanced models -> update accordingly. Deadline: Short term (within next 12 months). Responsibility: Analysis unit and relevant business owners. Implementation advice: See KNIME workflow for inspiration.
  - 5. Expand use of anomaly detection in TP using CbCR-data + other data sources. Deadline: Short term (within next 12 months). Responsibility: Analysis unit and TP unit. Implementation advice: analysis can inform which of the >2,300 cases of firms outside safe-harbor thin capitalization rules to pursue; prioritization should rely on taxpayer size.
  - 6. Advanced data analytics projects. Deadline: Short term (within next 12 months). Responsibility: Analysis unit, IT & relevant business owners. Implementation advice: after low-cost investments proceed with more advanced predictive modelling (such as random forest models).
  - 7. Evaluate implementation of data analytics projects and devise remedying actions if necessary. Deadline: Short term (within next 12 months). Responsibility: Impacted divisions & Analysis unit.
- Topic: Organizational structure
  - 1. Ensure the Analysis unit supports the entire SFA. Deadline: Short term (within next 12 months). Responsibility: DG in consultation with senior management.
  - 2. Reestablish the Compliance Risk Management Unit. Responsibility: DG in consultation with senior management.
- Topic: Staffing and training
  - 1. Prioritize end-user education in all existing applications. Deadline: Short term (within next 12 months). Responsibility: (not specified). Implementation advice: No application should be built without prioritizing end-user education and feedback.
  - 2. Establish learning programs where data scientists teach non-experts. Deadline: Short term (within next 12 months). Responsibility: Analysis unit in conjunction with heads of divisions. Implementation advice: use KNIME® as a low-barrier tool for trainings.
  - 3. Consider improving terms for data analysts to ensure recruitment/retention. Deadline: Short term (within next 12 months). Responsibility: DG/HR.
  - 4. Retrain 1-3 non-data scientists from SFA to join the Analysis unit. Deadline: Short term (within next 12 months). Responsibility: Data Science Team and HR.
  - 5. Start “tours” program embedding non-data analysts in the Analysis unit and vice versa. Deadline: Medium term (within two to three years). Responsibility: DG and relevant heads of division in consultation with Analysis units.

### Deliverables and project ideas provided
- Documents produced and provided:
  - PowerPoint presentations: High-level introduction; Workshop presentation; Exit-meeting presentation.
  - Models (KNIME workflow + data + .xls): CIT audit selection model (KNIME workflow plus data); CIT audit selection model with deployment (KNIME workflow plus data); VAT audit selection model (KNIME workflow plus data); CbCR anomaly detector model (KNIME workflow plus data); CbCR overview (Excel model); Business risk indicator evaluation model (KNIME workflow plus data); Customs consignment selection model (KNIME workflow plus data).
- Appendix I: Gross list of possible projects across Registration, Filing, Reporting, Payment, Service, Customs, Policy.
- Appendix II: The Multifaceted (Butterfly) Approach to Risk (framework for grouping taxpayers by Ready/Willing/Able and tailoring interventions).

### Appendix III — Suggested Workflow for Data Analytics (high-level)
- Import and merge data sources
  - Using unique identifiers all accessible data is merged
  - This process may involve setting up algorithms to read large numbers of pdf files or integrate databases
- Transform data
  - Data is transformed so that models can leverage the full information
  - This process may involve simple transformations such as calculating totals or more complex algorithms such as building networks
- Predict
  - A predictive model (random forest, neural network) is built that generates a likelihood score and estimate of collectible revenue
  - The model is built such that it uses a test and training data set to maximize predictive value
- Deploy and learn
  - Audits are performed based on the prediction
  - By keeping track of when the model is right and wrong the algorithm can update and learn to improve

### Appendix IV — Practical Tips for Predictive Model Development and Use (selected)
- Data Analysis Process:
  - Use a formal process, such as CRISP_DM to document data mining efforts for better analysis, consistent reuse and shared learning.
- Data — Typical data set guidance:
  - Prefer categorical/ordinal/interval data over free text; reduce number of categories where necessary; convert time fields to numeric representations; manipulate mirror datasets and preserve originals; save SQL scripts; use representative sampling for extremely large data sets.
- Modeling / Mining:
  - For rare target variables (<10% strike rate) consider oversampling or SMOTE; avoid duplicating data before partitioning; use Random Forest initially at default settings, then increase trees from default of 100 to 300 to improve performance; consider ensemble approaches and incorporate consequence (median adjustment by turnover range) into ranking (LxC).
- Deployment:
  - For deployment, retrain model on 100% training partition; use a single decision tree (CART) to explain black-box models; stage gate deployment (e.g., 100 -> 400 -> 500 cases); revisit models every six months; include a random case selection component for monitoring.
- Glossary (selected terms preserved): Audit; CART; Case Selection; Risk Filter/Risk Rule; False positives (FP); False negatives (FN); Strike Rate (precision) = TP/(TP+FP); Miss Rate = FN/(TP+FN); Accuracy = TP+TN/(TP+TN+FP+FN); Confusion Matrix; ROC Curve; Cumulative Curve Inflection Point Approach; Data mining; KNIME; Meta Node; Random Forest.

*Source: tarea2024010 — IMF mission report excerpt.*

### 1.      This mission aimed at supporting the Republic of Slovenia Financial Administration

### 1.      This mission aimed at supporting the Republic of Slovenia Financial Administration

### Mission activities
- Conducted four workshops and five one-to-one sessions to provide guidance on:
  - Ideal organizational arrangements to support data analytics
  - How to access and utilize better and more data, including third party data
  - Current data workflows and processes with a view to streamlining and improving them
  - How to improve data integrity
  - Using data science and big data analytics to strengthen risk assessment
  - Leveraging the value of country-by-country reports
  - Developing several data analytics/risk assessment models

### Key findings
- SFA has high-quality data and competent staff able to benefit from data analytics.
- Staff engagement and stamina were strong; participants built four new pilot risk assessment models during the mission that showed promising results.
- Example model result: the CIT audit selection model was estimated to increase the strike rate in comprehensive audits (measured as a correction above €5000) from 50-60 percent to roughly 90 percent (while doing the same number of audits).
- The SFA has an established “analysis unit” and more than 50 data applications in use, with an advanced predictive model under development in the IT department.
- The mission team was able to obtain high-quality data in a matter of days: audits, tax returns, Country-by-Country-Reports, business risk rules and balance sheets, all mergeable using unique taxpayer identifiers.

### Recommended senior management actions (five actions)
- a) Fully endorse, prioritize and incentivize the journey towards relying more on data analytics.
- b) Allow time for data analytics to grow (sometimes fail) and keep a stable and dedicated team in place.
- c) Invest in further recruitment and retention of data analysts and retrain operational staff to join the data analytics team.
- d) Prioritize development of high-value projects (deprioritize low-value projects) and vigilantly implement them.
- e) Have a constant focus on business needs and rigorously evaluate the success of projects.

### Technical assistance advice — Background and principles
- “Big Data” does not automatically translate into “Big Improvements”; advanced data analytics requires solving statistical, organizational and technical problems.
- Advanced Data Analytics: use of statistical techniques to make predictions and draw inferences about cause and effect; seeks to make judgments with more and better reliance on data.
- Machine learning models can incorporate all available taxpayer information (CIT, VAT, PIT, social security registers) and self-correct by learning from past mistakes.
- Advantages of advanced analytics are largest where data abundance and large populations meet scarce resources for manual examination; less advantageous when populations are small or data is difficult to make available.
- Effortless and automatic data integration is essential: propose enlarging datasets via internal integration (tax and customs) and external third-party automatic exchange; collect DAC2 and CRS-equivalent data from resident financial reporters; submit VAT special files (issued and received invoices) automatically on monthly or quarterly basis; require mandatory annual balance sheet submission before CIT returns—may require amendments to data exchange laws.
- Advanced analytics methods suggested include pattern recognition, outlier detection, cluster analysis, experimental design, network analysis, and text mining.

### Organizational structure, recruitment, and integration of solutions
- Risk identified: an organizationally detached and non-prioritized data analytics unit (satellite unit) risks being out of touch with end-user needs and lacking resources/authority.
- Measures to ensure integration and effectiveness:
  - Ensure employee flow between business units and the data science team (e.g., embedding “tours”).
  - Prioritize education of end-users; set aside time for end-user training for every application.
  - Aim for rapid deployment to avoid “white elephants”.
  - Rigorously measure impact of data analytics projects from end-user perspective; many existing applications lack systematic assessment.
  - Ensure model predictions are actionable and transparent.
- Mitigating cultural resistance:
  - Senior management endorsement and authority for the data science team.
  - Involve veteran employees and provide training in analytics basics.
  - Incentivize change and avoid punishing expected failures.
  - Translate efficiency gains into improved services rather than workforce reduction.
  - Prove value by competing new approaches with old approaches (e.g., split-sample audit selection).
- Recommended iterative deployment process:
  1. Understand business issues with end-users.
  2. Inventory and use easily accessible data before requesting new data.
  3. Build and evaluate a statistical model with end-users.
  4. Refine model with more/better data.
  5. Deploy when satisfied with theoretical properties.
  6. Evaluate and refine post-deployment.
- Model transparency techniques:
  - Use surrogate models (e.g., linear or decision tree) to proxy complex models.
  - Compute marginal effects (individual conditional expectation) for variables.
  - Make targeted predictions (e.g., predict “adjustment due to un(der)invoicing” rather than generic “adjustment”).

### Staffing, training, and tools
- Organizational stability is critical; frequent restructurings have harmed staff and momentum (example: Compliance Risk Management Unit (CRMU) was barely in place before being closed).
- Centralized analytics unit is appropriate for pooling scarce data analyst resources; unit should work across CIT, VAT, customs, and other areas.
- Analytics unit should have necessary authority and strong collaboration with Special Financial Office (SFO) and IT department.
- Prioritize recruiting/retaining few highly competent data analysts: “Five good data analysts can function as the core of a data analytic unit.”
- Consider special pay scales for specialized staff; do not let language barriers hinder recruitment; part-time university statistics students can be useful.
- Retrain selected operational staff into data analysts to ensure business understanding and recruitment.
- Train other departments in advanced data skills via structured multiday workshops with curriculum, homework and assignments.
- Consider analytical software packages with low entry barriers for onboarding non-trained data scientists (example: KNIME®); other open-source packages mentioned: WEKA, ANACONDA, RATTLE.

### Low-cost/high-reward initiatives in risk assessment
- Prioritize simple adjustments before larger investments in advanced analytics.
- Emphasize size in risk prioritization: size of risk = likelihood * consequence; interventions should weigh both likelihood and potential consequence.
- Audit case selection should weigh business risk scores by taxpayer size (sales, employees, taxes paid, etc.). Current VAT and TP business risk rules will increase overall level of risk by at most 10 percent when moving from a micro firm to a large firm—this understates variance in risk.
- Strike rate measurement recommendations:
  - Introduce thresholds so marginal corrections do not contaminate strike-rate statistics; a correction of €20 is rarely a meaningful audit success.
  - Strike rates from comprehensive audits fall by 10-20 percentage points when applying a €5000 threshold.
- Priority monitoring:
  - The 30 largest taxpayers should be continuously observed; create risk profiles for the 30 largest taxpayers (can be automated using TP indicators).
  - Establish one-to-one client account managers for these taxpayers; prioritize intervention (education, guidance or audit) if risk profiles suggest non-compliance.

*Source: tarea2024010 - 1.      This mission aimed at supporting the Republic of Slovenia Financial Administration*

### 29.      End-user education should be prioritized both to ensure expanded usage of existing

### tarea2024010 - 29.      End-user education should be prioritized both to ensure expanded usage of existing

### End-user education and operational efficiency
- End-user education should be prioritized to:
  - Ensure expanded usage of existing applications.
  - Limit data analytics resources spent on trivial work (e.g., reduce time Analysis unit spends rearranging data in reports).
- Example finding: business risk indicators set up for TP case selection were accurate in predicting the yield of CIT comprehensive audits. Utilizing existing work and educating end-users can realize such low-hanging fruits.

### Systematic evaluation of past interventions and business risk rules
- Previous audit strategies/business risk rules were not being systematically evaluated.
- Suggested simple evaluation methods:
  - Use past audit results to evaluate business risk rules, e.g., simple scatterplots.
  - Run “horse-races” by applying different approaches in parallel and comparing outcomes.

### Simple statistical improvements to risk rules
- Using simple benchmarks and standard deviations can root business risk in systematic analysis rather than intuition.

### Multifaceted (butterfly) approach to compliance risk management
- The mission emphasizes a multifaceted approach; audit selection models are only one component.
- The “butterfly” approach consists of steps:
  - Group non-compliant taxpayers into three clusters: The unaware, the unable and the unwilling.
  - Deter non-compliance through education, service and exemplification targeted at the different clusters.
  - Detect non-compliance using risk modelling.
  - Deal with detected non-compliance based on the connected risk (likelihood*consequence) and the underlying reason (cluster).
- Recommendation: reestablish the Compliance Risk Management (CRM) Unit to support the CRM Committee and ensure 360o assessment of risk mitigation. The data analytics unit must serve the entire CRM spectrum.

### Proof of concept: pilot empirical models constructed
- Four pilot empirical models were constructed and handed over to the SFA (incl. additional models on customs not covered by the mission):
  - An advanced (random forest) CIT audit selection model (incl. deployment workflow)
  - An advanced (random forest) VAT audit selection model
  - Workflow to evaluate existing business risk rules against historical data (VAT/CIT)
  - Country-by-Country-Reporting anomaly detection tool
- All models targeted delivering real-time risk assessment using all available data and an optimal workflow (see Appendix III referenced in source) where interventions and outcomes are fed back to improve the model.

### Pilot model performance and insights
- Even as pilot models with shortcuts, audit selection models outperformed existing risk rules.
  - Example: CIT audit selection model was estimated to increase the strike rate in comprehensive audits (measured as a correction above €5000) from 50-60 percent to 90 percent (while doing the same number of audits).
- Two demonstrated ways to evaluate business risk rules:
  - Correlate business risk rules with historical audit yields to test accuracy (business risk rules generally proved accurate).
  - Use advanced audit selection models to inform relevance and weights of different risk rules (e.g., variables related to wrongful registration should have larger risk weights).
- Country-by-Country-Reporting anomaly tool: compute benchmark profitability for each MNE group and industry and compare to subsidiary profitability in the Republic of Slovenia.

### Model limitations and recommended verification
- Pilot models have room for improvement and require verification:
  - Use and verify all available data and test competing models.
  - Ensure SFA can fully comprehend models before implementation.
  - Aim for incremental improvement rather than perfection.
- Field implementation is the real test; minimize time from pilot build to small-scale implementation.
- Credible evaluation requires a credible status quo benchmark via a randomized double-blind study:
  - Procedure: (1) SFA computes business risk rules as usual; (2) SFA picks initially 100 cases based on new model prediction; (3) randomly assign audits from new predictive model and standard SFA model to different inspectors (inspectors blinded); (4) track adjustment outcomes in cases picked by new and old models.
  - Record adjustment outcomes and all taxpayer information related to top picks; best achieved by tracking consignment numbers of top picks.
  - Random assignment can be achieved using RAND in Excel as described in the source.

### Prioritization principles for future projects
- Considerations when prioritizing projects:
  - Data available (fit for data analytics?)
  - Size of closable tax gap (based on actual analysis or intuition)
  - Bidding round (how large resources will each division devote)
  - Political decision (importance beyond revenue)
  - Next best alternative
- Recommended immediate priorities (coming year):
  - Create risk profiles for the 30 largest taxpayers.
  - Create business risk indicators for CIT.
  - Put larger weight on firm size and redefine strike rates to reflect firm size.
  - Evaluate existing business risk rules using historical data.
  - Only after these simple adjustments proceed with more advanced modelling.

### Summary of suggested operational actions (timing and responsibilities preserved)
- Topic: Prioritization/endorsement by senior management
  - 1. Prioritize/endorse high-value work by creating protected time for data analytics projects. Deadline: Short term (within next 12 months). Responsibility: Director General (DG)/ Heads of department.
  - 2. Move towards “top-down” risk assessments in CIT by defining and using existing business risk indicators. Deadline: Short term (within next 12 months). Responsibility: Heads of department/Working Groups (WG).
  - 3. Increase emphasis on size of taxpayers in risk assessments. Deadline: Short term (within next 12 months). Responsibility: Heads of department/WG.
  - 4. Endorse a multifaceted approach to compliance risk management. Deadline: Short term (within next 12 months). Responsibility: Heads of department/Working Groups (WG).
  - 5. Move towards automatic exchange of third-party information from banks, other regulators and other relevant stakeholders. Deadline: Medium term (within two to three years). Responsibility: Heads of department/Relevant stakeholders.
- Topic: Improvements in data analytics and way-of-work (selected items)
  - 1. Readjust risk rules to reflect expected adjustment/risk. Deadline: Short term (within next 12 months). Responsibility: Analysis unit and relevant business owners. Implementation advice: multiply current risk rules with measures of size (turnover, tax paid, employees).
  - 2. Build automated reports for the largest taxpayers. Deadline: Short term (within next 12 months). Responsibility: Analysis unit and large taxpayer unit. Implementation advice: begin by building upon current TP business risk rules and update according to John Middleton’s 2022 report Appendix III; reports should indicate suggestive actions such as “contact taxpayer” or “audit”.
  - 3. Build CIT business risk indicators. Deadline: Short term (within next 12 months). Responsibility: Analysis unit and relevant business owners.
  - 4. Evaluate existing risk rules based on historical data and advanced models -> update accordingly. Deadline: Short term (within next 12 months). Responsibility: Analysis unit and relevant business owners. Implementation advice: See KNIME workflow for inspiration.
  - 5. Expand use of anomaly detection in TP using CbCR-data + other data sources. Deadline: Short term (within next 12 months). Responsibility: Analysis unit and TP unit. Implementation advice: analysis can inform which of the >2,300 cases of firms outside safe-harbor thin capitalization rules to pursue; prioritization should rely on taxpayer size.
  - 6. Advanced data analytics projects. Deadline: Short term (within next 12 months). Responsibility: Analysis unit, IT & relevant business owners. Implementation advice: after low-cost investments proceed with more advanced predictive modelling (such as random forest models).
  - 7. Evaluate implementation of data analytics projects and devise remedying actions if necessary. Deadline: Short term (within next 12 months). Responsibility: Impacted divisions & Analysis unit.
- Topic: Organizational structure
  - 1. Ensure the Analysis unit supports the entire SFA. Deadline: Short term (within next 12 months). Responsibility: DG in consultation with senior management.
  - 2. Reestablish the Compliance Risk Management Unit. Responsibility: DG in consultation with senior management.
- Topic: Staffing and training
  - 1. Prioritize end-user education in all existing applications. Deadline: Short term (within next 12 months). Responsibility: (not specified). Implementation advice: No application should be built without prioritizing end-user education and feedback.
  - 2. Establish learning programs where data scientists teach non-experts. Deadline: Short term (within next 12 months). Responsibility: Analysis unit in conjunction with heads of divisions. Implementation advice: use KNIME® as a low-barrier tool for trainings.
  - 3. Consider improving terms for data analysts to ensure recruitment/retention. Deadline: Short term (within next 12 months). Responsibility: DG/HR.
  - 4. Retrain 1-3 non-data scientists from SFA to join the Analysis unit. Deadline: Short term (within next 12 months). Responsibility: Data Science Team and HR.
  - 5. Start “tours” program embedding non-data analysts in the Analysis unit and vice versa. Deadline: Medium term (within two to three years). Responsibility: DG and relevant heads of division in consultation with Analysis units.

### Deliverables and project ideas provided
- Documents produced and provided:
  - PowerPoint presentations: High-level introduction; Workshop presentation; Exit-meeting presentation.
  - Models (KNIME workflow + data + .xls): CIT audit selection model (KNIME workflow plus data); CIT audit selection model with deployment (KNIME workflow plus data); VAT audit selection model (KNIME workflow plus data); CbCR anomaly detector model (KNIME workflow plus data); CbCR overview (Excel model); Business risk indicator evaluation model (KNIME workflow plus data); Customs consignment selection model (KNIME workflow plus data).
- Appendix I: Gross list of possible projects across Registration, Filing, Reporting, Payment, Service, Customs, Policy (examples include predicting non-filers, predicting size of potential adjustment, predicting likelihood and size of upliftment, tax gap measurement, etc.).
- Appendix II: The Multifaceted (Butterfly) Approach to Risk (framework for grouping taxpayers by Ready/Willing/Able and tailoring interventions).

*Content derived from the provided IMF mission report excerpt.*

### Appendix III.  Suggested Workflow for Data Analytics

### Appendix III.  Suggested Workflow for Data Analytics

### The data-driven approach to examination selection
- Merged data sources illustrated: Government registers; Credit card transactions; Tax returns.
- High-level workflow stages:
  - Import and merge data sources
    - Using unique identifiers all accessible data is merged
    - This process may involve setting up algorithms to read large numbers of pdf files or integrate databases
  - Transform data
    - Data is transformed so that models can leverage the full information
    - This process may involve simple transformations such as calculating totals or more complex algorithms such as building networks
  - Predict
    - A predictive model (random forest, neural network) is built that generates a likelihood score and estimate of collectible revenue
    - The model is built such that it uses a test and training data set to maximize predictive value
  - Deploy and learn
    - Audits are performed based on the prediction
    - By keeping track of when the model is right and wrong the algorithm can update and learn to improve

### Appendix IV. Practical Tips for Predictive Model Development and Use (By Stuart Hamilton, IMF expert)

- 1. Data Analysis Process
  - Use a formal process, such as CRISP_DM to document data mining efforts for better analysis, consistent reuse and shared learning.
  - https://www.ibm.com/support/knowledgecenter/en/SS3RA7_15.0.0/com.ibm.spss.crispdm.help/crisp_overview.htm

- 2. Data — Typical data set for predictive analytics for tax audit case selection:
  - Where possible have/use categorical data rather than free text fields. You may need to limit the number of categories for data mining purposes (e.g. 1000 categories in a 1000 cases means no category correlation is possible. Reduce by grouping up to a smaller meaningful subsets. (e.g. Use higher ANZIC Codes from industry categorization.) The ‘Domain Calculator node’ in KNIME can be used to restrict domain numbers.
  - Where possible use ordinal data (categories that have a meaningful order A>B>C etc) rather than just categorical data.
  - Where possible use interval data (numeric with meaningful spacing between data) rather than ordinal data.
  - Where time data may be relevant (e.g. year established for a company or the age of a person) consider converting it into a simple numeric such as Age / Number of Years since establishment. Similarly for overdue accounts – use days overdue rather than date due. Review time data to see if is being considered as a string variable because of its presentation in the data set (e.g. 2015/2016 will be treated as a string. Convert to 2016.)
  - Manipulate and transform mirror data sets - never the original data on the database. Even with data cleansing operations maintain the original data set for evidence purposes and create a new / updated dataset.
  - Save SQL scripts for consistent data retrieval, reuse and process documentation.
  - If faced with extremely large data sets use representative sampling (random or stratified random) to reduce the size of the data set being analyzed.
  - If running analytic software on a relatively small device, a laptop or desktop, consider making data transformations on the server / mainframe. The trade off is between wait-time for the new data set to arrive v execution time on a small device.
  - Always review the data in detail and fully understand how the data was created and input, its possible errors, as well as the distribution of the data.

- 3. Modeling / Mining
  - If the target variable is relatively rare (e.g. <10% strike rate) consider ‘oversampling’ the target or ‘under sampling’ the negative class or generating additional ‘synthetic’ examples using the SMOTE node in KNIME. This is done because the learner needs to ‘see’ sufficient cases to learn.
  - Similarly if the target variable is almost always present (e.g. >80% strike rate) consider using higher thresholds to reduce the apparent strike rate (e.g. move the threshold towards the median strike and the positive class will reduce towards 40% if originally ~80%).
  - Note: Do not duplicate data (oversample) prior to partitioning data into the training and verification sets as this will result in duplicated data being present in both the learner and verification sets and will thus overstate the models predictive ability.
  - Review the data again. Consider restricting the number of categorical variables using the Domain Calculator node. Consider transforming quantitative string data into numeric data using the String to Number node. Consider transforming data via normalization and ratios etc, to better highlight discriminate features. Some modeling approaches work best with normalized interval data. Explore whether such transformations improve the predictions. (e.g. using the KNIME Normalizer node).
  - To reduce the number of low value cases selected iteratively raise the threshold for a ‘strike’ and evaluate the results until a suitable balance between strike rate and caseload is obtained. (The median strike value is often a good starting point.)
  - Use a Rules Engine node to set a threshold for a strike at an appropriate value:
    - (e.g. $ADJUSTMENT$ >= 50000 => "Y"
      $ADJUSTMENT$ < 50000 => "N")
  - Review the data again. Ensure that the variables going into the learner node are not the product of the adjustment.
  - Exclude clearly non-relevant variables such as the tax file number (using the learner nodes configuration dialogue in Random Forest or using a Column Filter node before the decision tree learner).
  - Use a good ‘out of the box’ algorithm, such as Random Forest’ initially at its default settings. Once a promising model has been found increase the number of trees in the configuration dialogue from the default of 100 to 300 to improve model performance. (Additional performance tweaks to Random Forest inputs can be obtained by using the Tree Ensemble Learner / Tree Ensemble Predictor node in place of the basic Random Forest pairing).
  - To reduce ‘noise’ in the dataset evaluate and eliminate variables that don’t provide predictive ability. (The Meta Node provided for Random Forest ‘variable importance’ indicates the relative use of variables in the Random Forest learner.)
  - As expertise develops explore the use of other modeling approaches to see if they can improve predictions over part of the data. (It is usually hard to beat Random Forest in practice but it can be a useful learning / capability build experience for the team.)
  - Consider using ‘ensemble approaches’ (taking the best predictions from multiple models – those with the highest ‘Area Under the Curve’) when appropriate. There is a model parameter optimization workflow on the KNIME examples server.
  - Economic data is usually very highly skewed, as are the outcomes of compliance activities. As binary classifiers create a view of the likelihood of adjustment and not the consequence, it is crucial that a view of the potential size of the adjustment is brought into the risk equation. Using a view of the median adjustment by turnover range as a useful proxy for the potential consequence and then use this to prioritize cases by predicted risk adjusted value (LxC). Test this set of consequence proxies on the verification data set to establish the best ranking approach that maximizes overall revenue recovery.

- 4. Deployment
  - Once a robust predictive model has been built and rigorously evaluated, for the deployment build push the ‘training partition’ to 100% and re-execute the learner node to maximize the deployed models’ predictive ability.
  - Use a single decision tree (CART) to provide a broad explanation of what the more accurate ‘black box’ model is doing. It won’t be exact nor always ‘correct’ but should provide a useful indication. ‘Prune’ the tree to an appropriate level to explain most cases by increasing the lowest number of cases for a split point.
  - Errors in model development can occur, particularly in the early stages of trialing analytics when building capability. Stage gate model deployment to ensure that the results produced are in line with those expected. E.g. do 100 cases and evaluate the results, if ok then do a further 400 and evaluate, if ok do a further 500 etc...
  - Use ‘blind’ testing when practical so that observer bias (for or against) is less likely.
  - A models’ predictive ability degrades over time as the underlying economy changes. Revisit models every six months with additional data to see if they need to be rebuilt or enhanced.
  - Consider using a small random case selection component (e.g. using stratified random sampling) to maintain intelligence on new risks and as a means of monitoring model performance over time. A random case selection component can also assist in estimating tax gaps and prioritizing compliance campaigns.
  - Build risk rules for segments not previously examined or with low representation in the data set as these will usually be excluded by the predictive model as no or limited prior successful / unsuccessful case data exists for it to learn from.

### Glossary of Technical Terms (selected)
- Audit
  - A process used to establish whether the correct amount of tax has been assessed. It involves formal evidence gathering to establish the facts and then the application of relevant law to those facts.
- CART
  - Classification and Regression Tree. A decision tree data mining software algorithm. Usually not the optimal data mining method, with a tendency to ‘over-fit’ the training data, it has the advantage of not being a ‘black box’. The single decision tree rules are explainable.
- Case Selection
  - The process (e.g. via data-mining or subject matter expert rules) used to initially identify a set of taxpayers (positives) that may have compliance risks. Ideally should produce a listing of taxpayers prioritized (ranked) by predicted revenue risk = % likelihood multiplied by Rs consequence [aka Risk Adjusted Value].
- Risk Filter/Risk Rule
  - A set of rules used to select cases for a particular risk. Can be created by subject matter experts or from predictive data mining.
- False positives (FP)
  - Taxpayers that initially appear to have a tax compliance risk, but on review are found to be compliant. Opposite of true positives (TP)
- False negatives (FN)
  - Taxpayers that appear to be compliant but are not. The opposite of true negatives (TN)
- Strike Rate (precision)
  - TP/(TP+FP)
  - The ratio of true positives over the number selected. A function of case selection rationale, efficacy and size, auditor detection capability, and the underlying compliance rate.
- Miss Rate
  - FN/(TP+FN)
  - The ratio of false negatives over the total number of non compliant. A function of case selection rationale, efficacy and size, auditor detection capability, and the underlying compliance rate.
- Accuracy
  - TP+TN/(TP+TN+FP+FN)
  - The ratio of correctly determined cases to the total number of cases.
- Confusion Matrix
  - A table setting out True Positives/False Negatives/False Positives/True Negatives from the selection model. Used in case selection model evaluation.
- Receiver Operating Characteristic (ROC) Curve
  - A plot of how the ratio of True Positives to False Positives (TP:FP) varies over the sample/caseload. The greater the area under the curve (AUC) the better the selection model. Used in model evaluation.
- Cumulative Curve Inflection Point Approach
  - Technique that looks for the ‘turning points’ on matched cumulative curves of turnover and the adjusted tax. The median adjustment for the range between turning points is then used as a proxy for consequence in the risk calculation likelihood x consequence for ranking/prioritizing candidate cases.
- Data mining
  - The use computer algorithms (machine learning programs) to identify patterns and associations (knowledge discovery) in data. Data mining can be descriptive or predictive.
- KNIME
  - KNIME (KoNstanz Information MinEr, pronounced ‘Nime’) is an open source data analytics, reporting and integration platform that runs on OS, MS and Linux operating systems. KNIME integrates various components for exploratory data analysis and data mining through a modular workflow / linked node concept.
- Meta Node
  - A user created node in KNIME allowing workflows to be enclosed within it to reduce the complexity of the overall workflow display and enable the easy ‘packaging’ of reusable processes.
- Random Forest
  - A data-mining / machine learning algorithm that usually provides good ‘out of the box’ performance that is close to optimal. Created by repeated (e.g. 300 times) random sampling of the data and building a decision tree each time. The multiple decision trees (the Forest) then ‘vote’ on the correct categorization of an instance. Relatively fast, copes with missing values and non-numeric data and is resistant to ‘over-fitting’ the training data set.

*Appendix III. Suggested Workflow for Data Analytics — tarea2024010 (PDF chapter).*

---


_Source: https://www.imf.org/-/media/files/publications/tar/2024/english/tarea2024010.pdf_
