## Introduction (wpiea2025060-print-pdf)

## Source details

**Canonical URL:** [Introduction (wpiea2025060-print-pdf)](https://www.imf.org/-/media/files/publications/wp/2025/english/wpiea2025060-print-pdf.pdf)

## Other formats

- [Markdown version](/-/media/files/publications/wp/2025/english/wpiea2025060-print-pdf.pdf.md)
- [Structured JSON version](/-/media/files/publications/wp/2025/english/wpiea2025060-print-pdf.pdf.json)

---

### Growth of Big Data and its role in financial services
- The digital data universe grew from "around two zettabytes in 2010, to approximately one hundred zettabytes in 2024."
- Big Data underpins the "fintech revolution" where commercial success depends on the use of consumer data.
- Relevant technologies and concepts:
  - Distributed ledger technologies (DLT): methods to transfer, store and process data in distributed systems, with ability to establish and maintain trust.
  - Machine learning (subset of AI): relies on high-quality data to train algorithms.
  - Open Finance: controlled sharing of data (data portability).
  - BigTech business models: predicated on generation, use, and regeneration of Big Data.

### Risks of Big Data and trust implications
- Key risks created or amplified by Big Data:
  - Pervasive profiling of individuals.
  - Targeted misinformation campaigns.
  - Data breaches with more severe consequences.
  - Erosion of market integrity and adverse impacts on financial safety and soundness; potential threats to financial stability.
- Consequences of eroded trust:
  - Reduced willingness of individuals and firms to share data, constraining fintech innovation, efficiencies, and inclusive growth.
  - Increased digital exclusion and welfare reduction when users withdraw from digital economy participation.
  - Excluded individuals may face challenges accessing products and services at reasonable cost; data granularity may inadequately capture indicators of financial health and could demand sharing of granular, potentially harmful data.
  - Linked data: data shared by contacts or similar users can reveal information about excluded individuals.
  - Digital privacy paradox: users may state they value privacy but not act to keep data private; privacy expectations vary by context, culture, and demographics.

### Role of regulation, supervision, and privacy technologies
- Observations:
  - Data protection regulation is uneven globally; many emerging and developing economies lack frameworks, and existing frameworks may have gaps relative to internationally recognized guidelines (e.g., OECD guidelines).
  - Robust data protection regulation and proportionate supervision are prerequisites before addressing data ownership or portability (Open Banking / Open Finance).
- Options to preserve privacy and their trade-offs:
  - Non-participation in digital economy: preserves privacy but undesirable (limits access to services and economic growth).
  - Data minimization: shares only necessary data but limits benefits.
  - Anonymization: useful but has considerable drawbacks.
  - Privacy technologies: promising approach to enable data sharing while improving regulatory protections.
- Caveats about privacy technologies:
  - Do not replace robust data protection regulation.
  - Cannot stop subjects from choosing or being encouraged to overshare data or prevent misuse/repurposing beyond original purpose.
  - Could facilitate entry of new firms that lack effective regulatory oversight via increased data portability.
  - Supervisory use (SupTech) could widen capability gaps across regulators and create cross-border arbitrage.
  - Complex technologies may introduce cybersecurity and validation risks and require specialist expertise.

### Taxonomy of privacy technologies (input vs output)
- Broad classification:
  - Input Privacy: reduce data access.
  - Output Privacy: reduce inference from shared data.
- Input technologies highlighted:
  - Homomorphic Encryption
  - Secure Multiparty Computation (SMPC)
  - Federated Learning
- Output technologies highlighted:
  - Zero-Knowledge Proofs (ZKP)
  - Data Masking and tokenization
  - Differential Privacy
  - Synthetic Data

### Input privacy technologies — characteristics, limitations, and use cases
- Homomorphic Encryption
  - Enables processing of encrypted data directly without decryption; operations on ciphertext yield same results as operations on plaintext when decrypted.
  - Strengths: protects data during outsourced computation (e.g., cloud), reduces risk of data leakage at processing facilities.
  - Limitations: currently several orders of magnitude slower than unencrypted operations; commercially unviable for large-scale computations; variants trade functionality for performance (somewhat or partial homomorphic encryption).
  - Potential applications: protecting training data for ML models; Open Finance data sharing in encrypted format; third-party credit scoring or risk assessments; regulatory reporting with encrypted sensitive information.
- Secure Multiparty Computation (SMPC)
  - Enables parties to jointly compute functions over private data without revealing full data to others via secret sharing protocols.
  - Strengths: eliminates need for trusted third party; certain protocols detect dishonest participants.
  - Limitations: cannot protect if majority colludes; computational and telecommunication overheads grow with dataset size.
  - Potential applications: privacy-preserving analytics (trend analysis, benchmarking, risk assessment, collaborative fraud detection), smart contracts, secure cross-border payments, high-security wallets (split private keys), supporting data sovereignty via cross-border transfers without revealing sensitive information.
  - Interaction with homomorphic encryption: homomorphic schemes can enable central processing on encrypted inputs without keys.
- Federated Learning
  - Enables ML model training across decentralized local nodes; raw data remains on local nodes; model updates or gradients aggregated centrally or via decentralized protocols.
  - Variants: centralized federated learning (central server coordinates); decentralized federated learning (peer-to-peer).
  - Strengths: preserves local data privacy; scales well with number of participants and data.
  - Limitations and risks:
    - Vulnerable to inference attacks: model inversion, membership inference, property inference.
    - Memorization risks in complex models increase inference vulnerabilities.
    - Combining federated learning with secure aggregation, SMPC, homomorphic encryption, or differential privacy can strengthen privacy.
  - Use cases: collaborative model development among firms; suspicious transaction detection; third-party infrastructure providers; adding noise pre-submission to improve privacy.

- Comparative summary (as presented):
  - Homomorphic Encryption: High privacy level; High overheads; Low to moderate scalability; Use cases: secure data processing in the cloud, secure voting systems, private information retrieval, open banking/finance.
  - Secure Multiparty Computation: High privacy level; Moderate to high overheads; Moderate scalability; Use cases: collaborative data analysis, joint financial processing.
  - Federated Learning: High privacy level; Low to moderate overheads; High scalability; Use cases: decentralized ML, collaborative learning, privacy-preserving ML and AI.

### Output privacy technologies — characteristics, limitations, and use cases
- Zero-Knowledge Proofs (ZKP)
  - Enables proving assertions without revealing underlying information (e.g., proving age without revealing other PII).
  - Types: interactive (multiple rounds) and non-interactive.
  - Desirable properties: completeness, soundness, zero-knowledge.
  - Strengths: protects both prover and verifier; reduces incentives for data theft.
  - Limitations: resource intensive; scalability issues with large datasets; complex cryptographic constructs requiring scarce expertise; proofs do not vouch for correctness or legitimacy of underlying attributes.
  - Use cases: KYC/KYB procedures, identity verification, anonymous blockchain transactions, credit scoring, privacy-preserving digital transactions, Zcash and Zero-Knowledge Rollups in blockchain.
- Data Masking (including format-preserving encryption and tokenization)
  - Enables obfuscating or replacing sensitive fields with realistic or arbitrary substitutes to preserve application functionality.
  - Types: irreversible masking (break link to original) and reversible masking (difficult but possible to retrieve original).
  - Strengths: simple concept; high scalability; useful for testing, development, business intelligence.
  - Limitations: complex in large, heterogeneous data systems; reversible methods rely on key or token system security; irreversible masking may be impractical in some contexts.
  - Use cases: testing and development, proof-of-concept systems, business intelligence.
- Differential Privacy
  - Defined as a mathematical promise that output does not reveal whether any single datapoint is included; commonly achieved by adding noise.
  - Mechanisms: Laplace and Gaussian mechanisms; local (noise added at individual level) vs global (noise added during aggregation).
  - Key parameters:
    - Epsilon: defines privacy level (Privacy Budget); lower epsilon = stronger privacy but less accuracy.
    - Delta: probability of extreme privacy breach (approximate differential privacy when delta used).
  - Limitations: noise degrades utility especially in small datasets; repeated queries increase cumulative epsilon and reduce privacy; managing privacy budgets is critical.
  - Use cases: statistical/trend analysis, ML/AI, regulatory authorities sharing broad statistics (complaints data, consumer spending patterns, economic surveys), credit ratings demographic analyses.
- Synthetic Data
  - Algorithmically generated artificial data (partially or fully synthetic) that aims to preserve statistical properties of real data.
  - Strengths: useful when real data unavailable or to augment incomplete datasets; can support ML training, testing, stress testing, fraud detection model training, chatbots, product testing.
  - Limitations and privacy risks:
    - Requires real dataset benchmark for fidelity; higher fidelity increases privacy risk.
    - Re-identification risks: membership inference, attribute inference; AI generative models can retain "memory" of raw data.
    - Tradeoff between fidelity and privacy: increasing fidelity decreases privacy and risks model drift over time.
  - Mitigations: remove outliers, enlarge dataset, feature reduction, incorporate differential privacy in generation process.
  - Resourcing issues for authorities: generating synthetic data requires access to data, expertise, and costs — challenging for resource-constrained authorities or small jurisdictions.

- Comparative summary (as presented):
  - Zero Knowledge Proofs: High privacy; Moderate to high overheads; Moderate scalability; Use cases: authentication systems, blockchain, secure voting, digital transactions.
  - Data Masking: Moderate to high privacy; Low overheads; High scalability; Use cases: testing and development, demos, business intelligence.
  - Differential Privacy: Moderate to high privacy (tunable by parameters); Moderate to high overheads; Moderate scalability; Use cases: statistical/trend analysis, ML/AI.
  - Synthetic Data: Moderate to high privacy (not automatically privacy preserving); Moderate to high overheads; High scalability; Use cases: AI/ML model training, testing, development.

### Digital Sandbox and synthetic data (Box 1)
- Digital Sandbox (UK FCA)
  - Launched in October 2021 by the UK Financial Conduct Authority (FCA) in conjunction with the City of London Corporation and expanded into a permanent initiative to accelerate data-driven innovation.
  - Provides access to synthetic data and application programming interfaces (APIs) through an API marketplace to enable firms to safely test and refine new technologies in a controlled environment.
  - Aims to address challenges related to real-world data access and regulatory constraints for firms and start-ups that often lack proprietary data.
- Synthetic data generation and limitations
  - Real financial data contains sensitive and confidential information and, outside of proprietary data, cannot be used to train new algorithms without risk.
  - Data anonymization and/or pseudonymization removes identifying elements but can potentially be reverse engineered or combined with other data sets to re-identify data subjects.
  - Synthetic data generation is technically challenging and resource intensive and may not be feasible for many regulatory authorities today.
- Synthetic Data Expert Group (SDEG)
  - In March 2023, the FCA established the Synthetic Data Expert Group (SDEG).
  - The SDEG consists of 21 experts from financial services, public sector, data and technology vendors, and consumer groups.
  - The group focuses on identifying key issues, use cases, and challenges surrounding synthetic data in the UK financial markets and helps shape responsible use through cross-sector collaboration.

### Supervisory considerations and recommendations
- Outreach, engagement, and understanding trade-offs
  - Supervisory authorities should understand how privacy technologies are developing in their markets to inform policy pacing and assess the need for new policies.
  - Engagement approaches include:
    - Direct engagement with supervised firms using existing supervisory relationships.
    - Demonstration days for regulated and unregulated firms to showcase innovations to regulatory teams (relatively resource light).
    - Innovation Hubs to provide specialist fintech expertise or lead outreach where budgets/resources permit.
  - Sandboxes:
    - Digital sandboxes can fill technical knowledge gaps and may be beneficial for privacy technologies such as federated learning by providing access to necessary data under regulatory oversight.
    - Digital sandboxes can facilitate dissemination of synthetic data and enable proofs-of-concept for homomorphic encryption, particularly fully homomorphic encryption.
    - Regulatory (product testing) sandboxes have potential utility in certain instances but currently show little evidence of consistently delivering on objectives set by authorities.
- Coordination and cooperation
  - An integrated approach considering privacy, competition, innovation, financial stability, and data protection is important.
  - Ensuring compliance with data protection frameworks is primarily the remit of data protection authorities, but financial supervisors have a stake due to operational risks from data breaches that can affect reputation, safety, soundness, and potentially financial stability.
  - Domestic inter-authority coordination examples include the South African Intergovernmental Fintech Working Group and the British Digital Regulation Cooperation Forum.
  - Fintech cooperation agreements can formalize information sharing and allow faster exchange of information and ideas among specialist units, though without legal underpinning they may not be appropriate for sensitive or confidential cases.
  - Agreements enabling aggregated data sharing between conduct and data protection authorities can help detect worsening cybersecurity stances at supervised institutions.
- Cybersecurity and outcome-based supervision
  - Data privacy concerns the collection, use, storage, and deletion of personal data; data security is broader, covering confidentiality, integrity, and availability of all data types.
  - Cybersecurity provides a foundational layer for privacy technologies; datasets made less sensitive still require protection and secure environments (authentication, authorization, firewalls, anti-malware, key management).
  - Supervisors can adopt a technology-agnostic, outcome-based approach: specify desired privacy-preserving properties of processing and allow supervised institutions to select technologies that meet those properties.
  - Supervisory focus should be on whether processing complies with data protection and cybersecurity requirements irrespective of underlying technologies, while being alert to false senses of security absent strong cybersecurity controls.
- Trade-offs, practical adoption, and policy balance
  - Privacy technologies vary in maturity: some are theoretically appealing but practically limited today; others are production grade and in use.
  - General trade-off: the more privacy preserved, the less additional value extraction beyond original use—this can disincentivize firms economically and requires regulatory balancing.
  - Operational trade-offs include overhead and scalability versus scope of use and targeted privacy preservation (e.g., where to use zero-knowledge proofs or homomorphic encryption, and how to tune differential privacy parameters).

### Key supervisory recommendations (summary)
- Understand trade-offs of privacy technologies through outreach and engagement.
- Ensure domestic collaboration and international cooperation to improve knowledge sharing, regulatory certainty, and clarity of mandates.
- Treat privacy technologies as complementary to a base layer of cybersecurity controls.

*Source: wpiea2025060-print-pdf - Introduction (IMF).*

### Introduction ...........................................................................................................

### Introduction

### Major themes (from the chapter structure)
- Big Data, Trust, and the Digital Economy (page 4)
- Risks of Big Data (page 5)
- The role of privacy technologies (page 6)
- Privacy Technologies (page 8)
  - Input Privacy Technologies (page 8)
    - Homomorphic Encryption (page 8)
    - Secure Multiparty Computation (page 11)
    - Federated Learning (page 12)
  - Output Privacy Technologies (page 13)
    - Zero-Knowledge Proofs (page 14)
    - Data Masking (page 15)
    - Differential Privacy (page 17)
    - Synthetic Data (page 19)
- Considerations for Supervisors (page 22)
  - Outreach, engagement and understanding trade-offs (page 22)
  - Coordination and cooperation (page 23)
  - Understanding and managing cyber implications (page 24)
- Conclusion (page 26)
- References (page 28)

### Enumerated privacy-technology topics covered
- Homomorphic Encryption
- Secure Multiparty Computation
- Federated Learning
- Zero-Knowledge Proofs (ZKP)
- Data Masking
- Differential Privacy
- Synthetic Data

### Supervisory considerations highlighted
- Outreach, engagement and understanding trade-offs
- Coordination and cooperation
- Understanding and managing cyber implications

### Glossary entries included in this content unit
- AE: Advanced Economy
- AI: Artificial Intelligence
- API: Application Programming Interface
- BaaS: Banking-as-a-Service
- BCBS: Basel Committee on Banking Supervision
- BigTech: Large Technology Conglomerates Operating Across Markets
- DLT: Distributed Ledger Technology
- EMDE: Emerging Market and Developing Economy
- FCA: Financial Conduct Authority (UK)
- Fintech: Financial Technology
- ICO: Information Commissioner’s Office
- KYC: Know Your Customer
- ZKP: Zero Knowledge Proof

*Source: wpiea2025060-print-pdf - Introduction (IMF).*

### Introduction

### Introduction

### Growth of Big Data and its role in financial services
- The digital data universe grew from "around two zettabytes in 2010, to approximately one hundred zettabytes in 2024."
- Big Data underpins the "fintech revolution" where commercial success depends on the use of consumer data.
- Technologies referenced:
  - Distributed ledger technologies (DLT): methods to transfer, store and process data in distributed systems, with ability to establish and maintain trust.
  - Machine learning (subset of AI): relies on high-quality data to train algorithms.
  - Open Finance: controlled sharing of data (data portability).
  - BigTech business models: predicated on generation, use, and regeneration of Big Data.

### Risks of Big Data and trust implications
- Key risks created or amplified by Big Data:
  - Pervasive profiling of individuals.
  - Targeted misinformation campaigns.
  - Data breaches with more severe consequences.
  - Erosion of market integrity and adverse impacts on financial safety and soundness; potential threats to financial stability.
- Consequences of eroded trust:
  - Reduced willingness of individuals and firms to share data, constraining fintech innovation, efficiencies, and inclusive growth.
  - Increased digital exclusion and welfare reduction when users withdraw from digital economy participation.
  - Excluded individuals may face challenges accessing products and services at reasonable cost; data granularity may inadequately capture indicators of financial health and could demand sharing of granular, potentially harmful data.
  - Linked data: data shared by contacts or similar users can reveal information about excluded individuals.
  - Digital privacy paradox: users may state they value privacy but not act to keep data private; privacy expectations vary by context, culture, and demographics (e.g., older users more likely to value privacy; users more concerned about transactions/financial data than browsing/app usage data).

### Role of financial regulation, supervision, and privacy technologies
- Data protection regulation is uneven globally; many emerging and developing economies lack frameworks, and existing frameworks may have gaps relative to internationally recognized guidelines (e.g., OECD guidelines).
- Robust data protection regulation and proportionate supervision are prerequisites before addressing data ownership or portability (Open Banking / Open Finance).
- Options to preserve privacy and tradeoffs:
  - Non-participation in digital economy: preserves privacy but undesirable (limits access to services and economic growth).
  - Data minimization: shares only necessary data but limits benefits.
  - Anonymization: useful but has considerable drawbacks.
  - Privacy technologies: promising approach to enable data sharing while improving regulatory protections.
- Caveats about privacy technologies:
  - Do not replace robust data protection regulation.
  - Cannot stop subjects from choosing or being encouraged to overshare data or prevent misuse/repurposing beyond original purpose.
  - Could facilitate entry of new firms that lack effective regulatory oversight via increased data portability.
  - Supervisory use (SupTech) could widen capability gaps across regulators and create cross-border arbitrage.
  - Complex technologies may introduce cybersecurity and validation risks and require specialist expertise.

### Privacy Technologies — classification and overview
- Broad classification:
  - Input Privacy: reduce data access.
  - Output Privacy: reduce inference from shared data.
- Input technologies highlighted:
  - Homomorphic encryption
  - Secure Multiparty Computation (SMPC)
  - Federated learning
- Output technologies highlighted:
  - Zero-knowledge proofs (ZKP)
  - Data masking and tokenization
  - Differential privacy
  - Synthetic data

### Input Privacy Technologies — characteristics, trade-offs, and use cases
- Homomorphic Encryption
  - What it enables: processing of encrypted data directly without decryption; operations on ciphertext yield same results as operations on plaintext when decrypted.
  - Strengths: protects data during outsourced computation (e.g., cloud), reduces risk of data leakage at processing facilities.
  - Limitations: currently several orders of magnitude slower than unencrypted operations; commercially unviable for large-scale computations; variants trade functionality for performance (somewhat or partial homomorphic encryption).
  - Potential applications: protecting training data for ML models; Open Finance data sharing in encrypted format; third-party credit scoring or risk assessments; regulatory reporting with encrypted sensitive information.

- Secure Multiparty Computation (SMPC)
  - What it enables: parties jointly compute functions over private data without revealing full data to others via secret sharing protocols.
  - Strengths: eliminates need for trusted third party; certain protocols detect dishonest participants.
  - Limitations: cannot protect if majority colludes; computational and telecommunication overheads grow with dataset size.
  - Potential applications: privacy-preserving analytics (trend analysis, benchmarking, risk assessment, collaborative fraud detection), smart contracts, secure cross-border payments, high-security wallets (split private keys), supporting data sovereignty via cross-border transfers without revealing sensitive information.
  - Interaction with homomorphic encryption: homomorphic schemes can enable central processing on encrypted inputs without keys.

- Federated Learning
  - What it enables: ML model training across decentralized local nodes; raw data remains on local nodes; model updates or gradients aggregated centrally or via decentralized protocols.
  - Variants: centralized federated learning (central server coordinates); decentralized federated learning (peer-to-peer).
  - Strengths: preserves local data privacy; scales well with number of participants and data.
  - Limitations and risks:
    - Vulnerable to inference attacks: model inversion, membership inference, property inference.
    - Memorization risks in complex models increase inference vulnerabilities.
    - Combining federated learning with secure aggregation, SMPC, homomorphic encryption, or differential privacy can strengthen privacy.
  - Use cases in financial services: collaborative model development among firms; suspicious transaction detection; third-party infrastructure providers; adding noise pre-submission to improve privacy.

- Comparative summary (as presented):
  - Homomorphic Encryption: High privacy level; High overheads; Low to moderate scalability; Use cases: secure data processing in the cloud, secure voting systems, private information retrieval, open banking/finance.
  - Secure Multiparty Computation: High privacy level; Moderate to high overheads; Moderate scalability; Use cases: collaborative data analysis, joint financial processing.
  - Federated Learning: High privacy level; Low to moderate overheads; High scalability; Use cases: decentralized ML, collaborative learning, privacy-preserving ML and AI.

### Output Privacy Technologies — characteristics, trade-offs, and use cases
- Zero-Knowledge Proofs (ZKP)
  - What it enables: proving assertions without revealing underlying information (e.g., proving age without revealing other PII).
  - Types: interactive (multiple rounds) and non-interactive.
  - Desirable properties: completeness, soundness, zero-knowledge.
  - Strengths: protects both prover and verifier; reduces incentives for data theft.
  - Limitations: resource intensive; scalability issues with large datasets; complex cryptographic constructs requiring scarce expertise; proofs do not vouch for correctness or legitimacy of underlying attributes (e.g., membership wrongly granted).
  - Use cases: KYC/KYB procedures, identity verification, anonymous blockchain transactions, credit scoring, privacy-preserving digital transactions, Zcash and Zero-Knowledge Rollups in blockchain.

- Data Masking (including format-preserving encryption and tokenization)
  - What it enables: obfuscating or replacing sensitive fields with realistic or arbitrary substitutes to preserve application functionality.
  - Types: irreversible masking (break link to original) and reversible masking (difficult but possible to retrieve original).
  - Tools and methods: dedicated software for static and dynamic masking; format-preserving encryption keeps input format; tokenization issues hinge on token management system security.
  - Strengths: simple concept; high scalability; useful for testing, development, business intelligence.
  - Limitations: complex in large, heterogeneous data systems; reversible methods rely on key or token system security; irreversible masking may be impractical in some contexts.
  - Use cases: testing and development, proof-of-concept systems, business intelligence.

- Differential Privacy
  - What it is: mathematical promise that output does not reveal whether any single datapoint is included; commonly achieved by adding noise.
  - Mechanisms: Laplace and Gaussian mechanisms; local (noise added at individual level) vs global (noise added during aggregation).
  - Key parameters:
    - Epsilon: defines privacy level (Privacy Budget); lower epsilon = stronger privacy but less accuracy.
    - Delta: probability of extreme privacy breach (approximate differential privacy when delta used).
  - Limitations: noise degrades utility especially in small datasets; repeated queries increase cumulative epsilon and reduce privacy; managing privacy budgets is critical.
  - Use cases: statistical/trend analysis, ML/AI, regulatory authorities sharing broad statistics (complaints data, consumer spending patterns, economic surveys), credit ratings demographic analyses.
  - Existing commercial and public-sector use noted.

- Synthetic Data
  - What it is: algorithmically generated artificial data (partially or fully synthetic) that aims to preserve statistical properties of real data.
  - Strengths: useful when real data unavailable or to augment incomplete datasets; can support ML training, testing, stress testing, fraud detection model training, chatbots, product testing.
  - Limitations and privacy risks:
    - Requires real dataset benchmark for fidelity; higher fidelity increases privacy risk.
    - Re-identification risks: membership inference, attribute inference; AI generative models can retain "memory" of raw data.
    - Tradeoff between fidelity and privacy: increasing fidelity decreases privacy and risks model drift over time.
  - Mitigations: remove outliers, enlarge dataset, feature reduction, incorporate differential privacy in generation process.
  - Resourcing issues for authorities: generating synthetic data requires access to data, expertise, and costs — challenging for resource-constrained authorities or small jurisdictions.

- Comparative summary (as presented):
  - Zero Knowledge Proofs: High privacy; Moderate to high overheads; Moderate scalability; Use cases: authentication systems, blockchain, secure voting, digital transactions.
  - Data Masking: Moderate to high privacy; Low overheads; High scalability; Use cases: testing and development, demos, business intelligence.
  - Differential Privacy: Moderate to high privacy (tunable by parameters); Moderate to high overheads; Moderate scalability; Use cases: statistical/trend analysis, ML/AI.
  - Synthetic Data: Moderate to high privacy (not automatically privacy preserving); Moderate to high overheads; High scalability; Use cases: AI/ML model training, testing, development.

### Implications for supervisors and final considerations (Introduction summary)
- The paper is targeted at financial sector supervisors in fintech, information technology, and cybersecurity functional areas and to those involved in designing, developing, or operating key components of the financial system.
- The Introduction explores:
  - Risks of Big Data.
  - Role of financial regulation and supervision in managing those risks.
  - Concept and taxonomy of privacy technologies (input vs output).
- Supervisory considerations emphasized:
  - Privacy technologies could enable more data sharing while improving protections and enabling innovation, but are not a substitute for data protection regulation.
  - Supervisors should be aware of potential new risks from privacy technologies, including regulatory arbitrage, capability divergence across authorities, cybersecurity vulnerabilities, validation gaps, and potential misuse by firms or regulators.
  - Hybrid approaches (combining privacy technologies, e.g., federated learning + SMPC/homomorphic encryption + differential privacy) can mitigate particular weaknesses (e.g., inference attacks) but introduce complexity and resource requirements.

*International Monetary Fund — Introduction (extracted content)*

### Box 1. The Digital Sandbox and Synthetic Data

### Box 1. The Digital Sandbox and Synthetic Data

### Digital Sandbox: purpose and design
- Launched in October 2021 by the UK Financial Conduct Authority (FCA) in conjunction with the City of London Corporation and expanded into a permanent initiative to accelerate data-driven innovation.
- Provides access to synthetic data and application programming interfaces (APIs) through an API marketplace to enable firms to safely test and refine new technologies in a controlled environment.
- Addresses challenges related to real-world data access and regulatory constraints for firms and start-ups that often lack proprietary data.

### Synthetic data generation and limitations
- Real financial data contains sensitive and confidential information and, outside of proprietary data, cannot be used to train new algorithms without risk.
- Data anonymization and/or pseudonymization removes identifying elements but can potentially be reverse engineered or combined with other data sets to re-identify data subjects.
- Synthetic data provides an alternative approach; the FCA generated synthetic data through a “DataSprint” that brought together public and private sectors to explore different methodologies to generate synthetic data.
- Synthetic data generation is technically challenging and resource intensive and may not be feasible for many regulatory authorities today, though evolving technology could increase accessibility.

### Synthetic Data Expert Group (SDEG)
- In March 2023, the FCA established the Synthetic Data Expert Group (SDEG).
- The SDEG consists of 21 experts from financial services, public sector, data and technology vendors, and consumer groups.
- The group focuses on identifying key issues, use cases, and challenges surrounding synthetic data in the UK financial markets and helps shape responsible use through cross-sector collaboration.

### Considerations for supervisors: privacy technologies and risks
- Privacy technologies can unlock innovation by inducing greater trust, allowing data to flow better within and across borders, and enabling more data generation; they are potentially transformative for markets and consumers.
- Outstanding questions remain about deploying privacy technologies at scale; some technologies could provide a false sense of security and many will require combinations of approaches for true privacy.
- Data privacy is not the same as data security; privacy technologies do not replace data security or cybersecurity controls. Cybersecurity should be regarded as a foundation for proper use of privacy technologies when dealing with personal or sensitive data.
- Defining “true privacy” in the digital age is challenging; frictions may arise between efficiency, regulatory requirements, and increased cyber risks from greater data sharing and flows.

### Outreach, engagement, and supervisory tooling
- Supervisory authorities should understand how privacy technologies are developing in their markets to inform policy pacing and assess the need for new policies.
- Engagement approaches include:
  - Direct engagement with supervised firms using existing supervisory relationships.
  - Demonstration days for regulated and unregulated firms to showcase innovations to regulatory teams (relatively resource light).
  - Innovation Hubs to provide specialist fintech expertise or lead outreach where budgets/resources permit.
- Sandboxes:
  - Digital sandboxes can fill technical knowledge gaps and may be beneficial for privacy technologies such as federated learning by providing access to necessary data under regulatory oversight.
  - Digital sandboxes can facilitate dissemination of synthetic data and enable proofs-of-concept for homomorphic encryption, particularly fully homomorphic encryption.
  - Regulatory (product testing) sandboxes have potential utility in certain instances but currently show little evidence of consistently delivering on objectives set by authorities.

### Coordination, cooperation, and information sharing
- An integrated approach considering privacy, competition, innovation, financial stability, and data protection is important.
- Ensuring compliance with data protection frameworks is primarily the remit of data protection authorities, but financial supervisors have a stake due to operational risks from data breaches that can affect reputation, safety, soundness, and potentially financial stability.
- Domestic inter-authority coordination examples include the South African Intergovernmental Fintech Working Group and the British Digital Regulation Cooperation Forum.
- Fintech cooperation agreements can formalize information sharing and allow faster exchange of information and ideas among specialist units, though without legal underpinning they may not be appropriate for sensitive or confidential cases.
- Agreements enabling aggregated data sharing between conduct and data protection authorities can help detect worsening cybersecurity stances at supervised institutions.

### Cybersecurity and outcome-based supervision
- Data privacy concerns the collection, use, storage, and deletion of personal data; data security is broader, covering confidentiality, integrity, and availability of all data types.
- Cybersecurity provides a foundational layer for privacy technologies; datasets made less sensitive still require protection and secure environments (authentication, authorization, firewalls, anti-malware, key management).
- Supervisors can adopt a technology-agnostic, outcome-based approach: specify desired privacy-preserving properties of processing and allow supervised institutions to select technologies that meet those properties.
- Supervisory focus should be on whether processing complies with data protection and cybersecurity requirements irrespective of underlying technologies, while being alert to false senses of security absent strong cybersecurity controls.

### Trade-offs, practical adoption, and policy balance
- Privacy technologies vary in maturity: some are theoretically appealing but practically limited today; others are production grade and in use.
- General trade-off: the more privacy preserved, the less additional value extraction beyond original use—this can disincentivize firms economically and requires regulatory balancing.
- Operational trade-offs include overhead and scalability versus scope of use and targeted privacy preservation (e.g., where to use zero-knowledge proofs or homomorphic encryption, and how to tune differential privacy parameters).

### Key supervisory recommendations
- Understand trade-offs of privacy technologies through outreach and engagement.
- Ensure domestic collaboration and international cooperation to improve knowledge sharing, regulatory certainty, and clarity of mandates.
- Treat privacy technologies as complementary to a base layer of cybersecurity controls.

*Source: Digital Sandbox | FCA; Report: Using Synthetic Data in Financial Services | FCA*

---


_Source: https://www.imf.org/-/media/files/publications/wp/2025/english/wpiea2025060-print-pdf.pdf_
