## 4. Role of Cash in Contingency Planning

## Source details

**Canonical URL:** [4. Role of Cash in Contingency Planning](https://www.imf.org/-/media/files/publications/wp/2021/english/wpiea2021288-print-pdf.pdf)

## Other formats

- [Markdown version](/-/media/files/publications/wp/2021/english/wpiea2021288-print-pdf.pdf.md)
- [Structured JSON version](/-/media/files/publications/wp/2021/english/wpiea2021288-print-pdf.pdf.json)

---

### I. Introduction — context and motivation
- Digitalization has increased operational risks through:
  - New technologies: "application programming interfaces, big data analytics, biometric technology, cloud computing, contactless technologies, digital identification, distributed ledger technology, and the Internet of things."
  - New products and service offerings: "instant payments, central bank digital currencies (CBDC), and stablecoins."
  - New access points: "electronic wallets, open banking, and super apps."
- Rising interdependencies and outsourcing raise resiliency challenges related to third-party service providers, critical infrastructures, and cross-border links.
- Evolving risks (cyber-attacks, natural disasters, pandemics) increase oversight expectations for continuous improvement of operational resilience.
- Paper organization (high level): Section II defines operational resilience; Section III shares international observations; Section IV explores development and operational issues; Section V concludes and suggests future research.

### II. Operational resilience — definition, scenarios, and principles
- Definition and PFMI expectations:
  - "Operational resilience is defined as the ability of a payment system and its service providers to deliver critical operations in disaster situations and extreme circumstances."
  - PFMI guidance summarized: FMI business continuity plans should "incorporate the use of a secondary site" and ensure "critical information technology systems can resume operations within two hours following disruptive events" and "enable the FMI to complete settlement by the end of the day of the disruption."
- Key expected features and sound practices:
  - Redundancy for primary and secondary sites; duplication of software and hardware; data replication consistent with recovery-point objectives.
  - Consider distinct risk profiles between primary and secondary sites, single point of failure mitigation, and alternative arrangements (e.g., manual paper-based procedures).
  - Four sound practices from international experience: (i) identify clearing and settlement activities supporting critical markets; (ii) determine appropriate recovery and resumption objectives; (iii) maintain geographically dispersed resources; (iv) routinely use and test recovery and resumption arrangements.
- Operational resilience elements (BoE, 2021a; Tiernan et al., 2019):
  - Develop an operational resilience framework subject to governance and audit.
  - Identify important services and set impact tolerance for each service.
  - Map interdependencies; test, monitor, and report.
- Regulatory harmonization example: European Commission’s proposed Digital Operational Resilience Act (DORA) would apply to financial entities and ICT third-party service providers.
- Technical complications: technical and organizational complexity; single software-version dependence; external infrastructure dependence; information security vulnerabilities; geographical/jurisdictional/time-zone issues; insufficient incentives for resilience.

#### B. Wide-Scale or Major Disruption Scenarios — ten scenario types
- PFMI definition: event causing "a severe disruption or destruction of transportation, telecommunications, power, or other critical infrastructure components across a metropolitan or other geographic area" or resulting in "a wide-scale evacuation or inaccessibility of the population within normal commuting range of the disruption's origin."
- Ten illustrated risk scenarios:
  - Internal: (1) inadequate arrangements; (2) internal sabotage; (3) workplace violence (active shooter).
  - External: (4) civil disturbances; (5) failure of critical service providers or utilities; (6) pandemics; (7) terrorism and warfare; (8) cyber-attacks and cyber-warfare; (9) natural disasters (bushfires, earthquakes, floods, haze, hurricanes, tornadoes, winter storms); (10) mega disasters.
- Six plausible wide-scale or major disruption scenarios (discussed in depth):
  - Critical service provider or utility failures (telecommunications, messaging providers, common cloud providers; power outages).
  - Pandemics (staff absenteeism; loss of key personnel; pandemic management teams and plans).
  - Terrorism (physical attacks on RTGS, clearing houses, cross-border systems; unilateral sanctions suspending access).
  - Cyber-attacks and cyber-warfare (financial sector estimated to be "three times more at risk" of cyber-attacks than any other sector; rising cyber incidents).
  - Natural disasters (hurricane/earthquake damage to primary and secondary sites; climate change increasing uncertainties).
  - Mega disasters (high-impact, low-probability complex events; example: Great East Japan Earthquake).
- Observations:
  - Coordinated attacks on primary and secondary sites lacking distinct risk profiles could cripple systems and prolong outages, potentially necessitating a third site.
  - Cross-sectoral and cross-border crisis management arrangements are needed involving multiple authorities and stakeholders.

#### C. International principles and assessments
- CPSS-IOSCO PFMI provides an assessment framework with seven key considerations for operational risk.
- BCBS Principles for Operational Resilience (BCBS, 2021) complements PFMI for banks’ critical functions associated with payments, clearing and settlement.
- IMF and World Bank SIPS detailed assessments benchmark practices against PFMI; operational risk and cyber security supervision have been included in assessments.
- Interdependency mapping is critical—interdependencies exist at three levels: institutions, systems, and environments; mapping can identify concentration risks and single points of failure.

### III. International experiences — observed shortcomings and incidents
- High-level observation: existing electronic payment systems have experienced operational incidents in critical service providers, payments and currency services, and SIPS.
- Four major RTGS outages (October 2014 to February 2021) — transaction figures for 2019; outage hours from public incident information:
  - Australia (RITS)
    - Volume (millions): 13
    - Value of transactions (USD billions): 36,629
    - Average value per transaction (USD thousands): 2,889
    - Value of transactions as a percentage of GDP (%): 2,642
    - Outage (hours): <2
  - Euro Area (TARGET2)
    - Volume (millions): 89
    - Value of transactions (USD billions): 509,382
    - Average value per transaction (USD thousands): 5,733
    - Value of transactions as a percentage of GDP (%): 3,813
    - Outage (hours): 10
  - United Kingdom (CHAPS Sterling)
    - Volume (millions): 49
    - Value of transactions (USD billions): 106,358
    - Average value per transaction (USD thousands): 2,186
    - Value of transactions as a percentage of GDP (%): 3,766
    - Outage (hours): 9
  - United States (Fedwire)
    - Volume (millions): 168
    - Value of transactions (USD billions): 695,835
    - Average value per transaction (USD thousands): 4,149
    - Value of transactions as a percentage of GDP (%): 3,200
    - Outage (hours): 3 to 4
- Key international lessons and observations:
  - Reliability Objectives (OROs): meeting quantitative OROs such as the 2h-RTO can be difficult; operators were generally able to complete end-of-day settlements despite breaches of short RTO benchmarks.
  - Redundancies: backup systems have fallen short of commencing immediate operations; primary and backup sites in same region can fail to provide seamless transition; staff unavailability complicates BCP execution.
  - Critical Service Providers: outages often linked to critical service providers (software failures, telecommunication services, utilities); cloud concentration risk noted.
  - Endpoint Security: cyber-attacks were often ruled out in some SIPS incident investigations, but cyber resilience efforts continue internationally; endpoint security of wholesale payments remains a focus to prevent fraud.
  - Alternative Arrangements: manual paper-based procedures proved useful in at least one incident; for wide-scale prolonged outages, the role of cash in national crisis preparedness becomes a central question.

#### Empirical incidents underscoring operational reliability objectives (selected)
- Australia — Reserve Bank of Australia’s RITS (July 6, 2020):
  - Incident time: 7.30 am.
  - Outcome: Majority of RITS services re-established and fully operational from the RBA’s primary site around 9 am and within the two-hour RTO.
  - Consequence: Opening of the RITS Daily Settlement Session delayed from 9.15 am to 9.30 am; all transactions were able to settle as expected from that time.
  - System scale: 167 participants, of which 102 are exchange settlement account holders, including the RBA.
  - Cause: Operational error (power supply inadvertently shut off during fire control system maintenance; legacy switchboard design issue; contractor non-compliance).
  - Response: Upgraded switchboard; enhanced contractor induction arrangements; improved oversight of compliance; new service delivery arrangements; increased role for staff with engineering expertise.
- Europe — TARGET2 outage (October 24, 2020):
  - Duration: around 10 hours.
  - Impact: All settlement services unavailable; backup system failed to boot up as failover to secondary disaster recovery site took many hours.
  - Financial impact: Drop in deposits worth more than EUR 400 billion (USD 473 billion).
  - Cause: Software defect at a third-party network device.
  - Response: Independent review announced to investigate failures including business continuity model, recovery testing, change management, and communication.
- Mastercard outage (July 12, 2018):
  - Duration: more than one and a half hours.
  - Impact: Some card payments failed and affected many international banks.
  - Public information on cause/impact/response: No public information available beyond a statement that situation was fully resolved.
- Visa Europe Ltd (VEL) outage (June 1, 2018):
  - Duration: partial outage around 8 hours.
  - Impact: Around 5.2 million card transactions across Europe affected.
  - Consequences: Inconvenienced travel, grocery payments, cash withdrawals; some ATMs in the United Kingdom were out of cash within a couple of hours.
  - Cause: Hardware failure (authorization to chip and pin machine).
  - Response: Internal and independent third-party review; work program to oversee remedial actions.
- Bank of England CHAPS outage (October 20, 2014):
  - Duration: over 9 hours.
  - Financial impact: Effectively froze GBP 277 billion (USD 447 billion) worth of payments; put 700 time-critical housing market transactions on hold.
  - Cause: Configuration changes triggered previously undetected design defect.
  - Response: Independent review; findings presented to Bank of England’s Court; full report with response published March 23, 2015.
- Federal Reserve Fedwire Funds Service disruption (February 24, 2021):
  - Duration: around 3 to 4 hours.
  - Additional impact: Affected 14 other services including Check 21, FedCash, Account Services, Central Bank and several FedLine services; some cryptocurrency exchanges experienced delays.
  - Cause: Operational error.
  - Response: Steps taken to help ensure resilience, including recovery to point of failure; Fedwire resumed normal operations after the outage.
- Travelex cyber-attack (December 31, 2019):
  - Vulnerability period: forced to shut down and was vulnerable for 8 months.
  - Impact: Affected customers in 21 countries; attackers claimed access to a large volume of customer data and demanded a ransom of USD 6 million.
  - Response: Travelex paid equivalent of USD 2.3 million and completed a restructuring deal August 2020 delivering GBP 84 million (USD 119 million) of new money.
- Bangladesh Bank SWIFT fraud attempt (February 4, 2016):
  - Attempted theft: USD 951 million.
  - Transfers executed: USD 101 million transferred from Bangladesh Bank’s account at the Federal Reserve Bank of New York; USD 20 million traced to Sri Lanka; USD 81 million ended up in the Philippines.
  - Cause: Wholesale payments fraud related to end-point security using malware to obtain legitimate SWIFT credentials.
  - Response: SWIFT introduced mandatory security measures including daily reports and the Customer Security Program; launched SWIFT Information Sharing and Analysis Centre May 2017.
- Zimbabwe — Econet Wireless power outages (July 2019 and July 2020):
  - Duration: power outage of 8 hours.
  - Impact: Hit 6.7 million active users; 70 percent of Zimbabweans had no phones, internet, and mobile money services; detrimental effects on the economy.
  - Cause: Utility failure; generators failed to kick in after prolonged electricity disruption.
  - Response: Investment in 520 lithium-ion Tesla batteries to provide backup to 1,300 base stations; cut reliance on diesel-run generators by 75 percent.

### IV. Development and operational issues for the future — five priority areas
- Overview: guided by PFMI and business continuity management, issues to strengthen operational resilience include impact tolerance, business continuity arrangements, tandem processing, interoperability, and cost & implementation.

#### A. Impact Tolerance
- Distinguish impact tolerances from RTOs:
  - Impact tolerance = "the maximum tolerable level of disruption for an important business service, whereby further disruption could significantly threaten the transfer of payments or the safety and efficiency of the payment system."
  - RTOs (e.g., 2h-RTO) remain relevant and should be aligned with business impact analysis and enterprise risk.
- Impact tolerance may be time-based or include metrics such as number of end-users impacted, or volume/value of payments disrupted.
- Supervisory expectation example: United Kingdom expects operators to set flexible impact tolerance levels for each business service/supporting service.
- Interdependencies: when key business services interconnect with SIPS or FMIs, their impact tolerances should consider OROs and impact tolerances set for those FMIs.

#### B. Business Continuity Arrangements
- Alternative arrangements and preparedness measures to consider:
  - Common, identical, and global payment standards.
  - Common global payment systems with geographically dispersed processing sites.
  - Real-time copies of transaction and account databases in secure environments with distinct risk profiles and redundancy.
  - Parallel satellite-based communication facilities (in addition to landline) for disaster situations.
  - Comprehensive local power-generation capabilities.
- Physical security enhancements (e.g., underground sites where practicable).
- Importance of establishing payment facilities outside impacted areas with updated complete sets of payment account and transaction data to commence immediate processing.
- Note on electricity needs: "The electricity capacity needed for payment instruments are rather limited and can be supplied with high-capacity battery, solar-electricity, small even man-powered generators" — but emergency supplies require advanced preparation.

#### C. Tandem Processing
- Definition: "each transaction is processed using two separate independent parallel processing streams" with automatic comparisons for correctness; akin to an electronic "four eyes" control.
- Distinctions and benefits:
  - Eliminates single points of failure in payment system software (current state: duplicated hardware but single processing, database, and telecommunication software).
  - Provides continuous backup and rapid error detection; automated failure correction; seamless processing if at least one parallel system remains operational.
- Key elements:
  - Duplicated independent network connections using completely independent networks and communications hardware/software; continuous dual routing with automatic fallbacks to solo-mode if a network fails.
  - Duplicated independent software/hardware components: two parallel independent payment processing systems with different risk profiles.
  - Automatic synchronization checks to compare outputs of parallel processes and alert/fallback mechanisms; AI-based algorithms could support automatic analysis and error correction.
  - Decentralized transactional data recording and storage with redundancy: at least two separate transaction records and real-time copies at secondary site.
  - Settlement integrated with payment transaction processing; "Real-time payments need immediate (and non-probabilistic) finality."
  - Resilient information security: transactions encrypted using at least two different encryption systems to avoid single-algorithm exposure.
- Policy consideration: further neutral external analysis of tandem processing feasibility by ICT experts recommended.

#### D. Interoperability
- Current landscape: many payment systems are proprietary, non-standardized, and fragmented; estimated "at least 1,200 payment systems" worldwide.
- Need for common standards to achieve:
  - Domestic and cross-border resiliency and interoperability.
  - Cost efficiency, improved service levels, and availability.
  - Ability for central banks to provide mutual backup services and allow unaffected infrastructures in neighboring or distant countries to process payments during disasters.
- Interoperability would support tandem processing and enable globally standardized systems to reduce development costs through reusable programming libraries.
- International cooperative arrangements required for information sharing and crisis management as cross-border payments and new technologies evolve.

#### E. Cost and Implementation
- Investments in resilience can be undervalued in normal times but yield benefits during crises; literature suggests structural resilience investments can "increase potential economic output, lower expected losses, and improve continuity of public services."
- Building duplicated synchronized payment systems entails upfront development and ongoing maintenance/updating costs; costs vary by country and existing standards.
- Standardization reduces development costs when payment systems reuse common libraries.
- Authorities may need regulations or sanctions to stimulate investment where FMIs and PSPs underinvest relative to public expectations; stricter resiliency requirements may be warranted for FMIs and PSPs handling urgent payments to focus resources where benefits are larger.

### V. Conclusion and future research
- Recovery time objectives can be difficult to meet under existing redundancy models.
- Strengthening operational resilience requires anticipating and preparing for wide-scale or major disruptions and recognizing that redundancy alone is insufficient in the digital era.
- Five development and operational priorities reiterated:
  - Set impact tolerances.
  - Identify alternative recovery arrangements.
  - Consider tandem processing.
  - Facilitate global interoperability.
  - Examine costs and implementation.
- Future research priorities identified:
  - How central bank money (physical cash and CBDC infrastructures) fits into operational resilience frameworks and contingency planning (Appendix 4 summarizes role of cash in contingency planning).
  - How CBDC infrastructures, if implemented, could be designed to ensure operational resilience or act as alternatives to traditional payment systems.
  - Further neutral external analysis of tandem processing feasibility by ICT experts.

### Role of cash in contingency planning and resilience (detailed)
- Functional role of cash:
  - Cash acts as legal tender, a means of payment guaranteeing privacy, and is the only form of money that can be carried or kept as savings independently of a bank and as a fallback if electronic payment systems fail.
  - Cash can provide immediate security for payments without third-party involvement, is independent from central verification, and is less dependent on electronic systems (though ATM and over-the-counter withdrawals may still require electricity).
- Contingency implications:
  - Wide-scale or major disasters may produce sudden peaks in demand for cash due to rising uncertainty.
  - Absence of non-digital fallback plans, such as cash, could expose economies to significant risks from system failure or cyber-attack.
  - For emerging and developing economies with rudimentary payment and telecommunications infrastructures, identifying cash as critical infrastructure is particularly relevant.
- Operational resilience for cash infrastructures:
  - Cash infrastructures need operational resilience; central banks and stakeholders must balance cost savings with continuity.
  - Decentralization of money processing (counting, sorting, checking, recirculating) by banks or licensed operators could improve efficiency and robustness.
  - Maintaining continuity of security transports to ATMs is essential to ensure accessibility and availability of cash.
- Policies and measures taken by jurisdictions:
  - Dutch central bank sets 4,000 ATMs as the minimum number to ensure adequate access to cash services.
  - Swedish government law requires the largest banks to continue to provide adequate access to cash and requires the Swedish central bank to have sufficient locations for professional parties to obtain banknotes.
  - Central banks can use crisis management frameworks to guide deployment of cash as a fallback, including timely communication to market participants and operational measures (example: Federal Reserve actions after September 11 to reassure banks and armored carriers, extend hours, arrange special deliveries, and waive certain fees).

*Source: wpiea2021288-print-pdf - Section 4: Role of Cash in Contingency Planning*

### REFERENCES                                                                                                              

### REFERENCES

### Figures
- 1. Risk Scenarios—Scale of Disruption and Outage Time ______________________________________ 7
- 2. Major Power Outages, 1999-2019 ______________________________________________________ 8
- 3. Cyber Incidents Involving Financial Institutions, 2007–2020 _________________________________ 10
- 4. Mapping of Interdependencies for Payment and Market Infrastructures ________________________ 11
- 5. Synchronized Tandem Payment Processing Design _______________________________________ 19

### Tables
- 1. Outage of Systemically Important Payment Systems in Selected Jurisdictions ___________________ 13

### Appendixes
- 1. Key Challenges for Improving Operational Resilience in Payment Systems _____________________ 27
- 2. Principles Relevant for Assessing Operational Resilience in Payments ________________________ 29
- 3. Selected Operational Incidents in Payment and Settlement Services __________________________ 32

*Source: wpiea2021288-print-pdf - REFERENCES*

### 4. Role of Cash in Contingency Planning                            36

### 4. Role of Cash in Contingency Planning

### I. Introduction — context and motivation
- Digitalization has increased operational risks through:
  - New technologies: "application programming interfaces, big data analytics, biometric technology, cloud computing, contactless technologies, digital identification, distributed ledger technology, and the Internet of things."
  - New products and service offerings: "instant payments, central bank digital currencies (CBDC), and stablecoins."
  - New access points: "electronic wallets, open banking, and super apps."
- Rising interdependencies and outsourcing raise resiliency challenges related to third-party service providers, critical infrastructures, and cross-border links.
- Evolving risks (cyber-attacks, natural disasters, pandemics) increase oversight expectations for continuous improvement of operational resilience.
- Paper organization (high level): Section II defines operational resilience; Section III shares international observations; Section IV explores development and operational issues; Section V concludes and suggests future research.

### II. Operational resilience — definition, scenarios, and principles
- Definition:
  - "Operational resilience is defined as the ability of a payment system and its service providers to deliver critical operations in disaster situations and extreme circumstances."
  - PFMI guidance summarized: FMI business continuity plans should "incorporate the use of a secondary site" and ensure "critical information technology systems can resume operations within two hours following disruptive events" and "enable the FMI to complete settlement by the end of the day of the disruption."
- Key expected features (PFMI and related guidance):
  - Redundancy for primary and secondary sites; duplication of software and hardware; data replication consistent with recovery-point objectives.
  - Consider distinct risk profiles between primary and secondary sites, single point of failure mitigation, and alternative arrangements (e.g., manual paper-based procedures).
  - Four sound practices from international experience: (i) identify clearing and settlement activities supporting critical markets; (ii) determine appropriate recovery and resumption objectives; (iii) maintain geographically dispersed resources; (iv) routinely use and test recovery and resumption arrangements.
- Operational resilience elements (BoE, 2021a; Tiernan et al., 2019):
  - Develop an operational resilience framework subject to governance and audit.
  - Identify important services and set impact tolerance for each service.
  - Map interdependencies; test, monitor, and report.
- Regulatory harmonization example:
  - European Commission’s proposed Digital Operational Resilience Act (DORA) would apply to financial entities and ICT third-party service providers.
- Technical complications to improving resilience include: technical and organizational complexity; single software-version dependence; external infrastructure dependence; information security vulnerabilities; geographical/jurisdictional/time-zone issues; insufficient incentives for resilience.

#### B. Wide-Scale or Major Disruption Scenarios — ten scenario types
- PFMI definition of wide-scale or major disruption: event causing "a severe disruption or destruction of transportation, telecommunications, power, or other critical infrastructure components across a metropolitan or other geographic area" or resulting in "a wide-scale evacuation or inaccessibility of the population within normal commuting range of the disruption's origin."
- Ten illustrated risk scenarios (internal and external):
  - Internal: (1) inadequate arrangements; (2) internal sabotage; (3) workplace violence (active shooter).
  - External: (4) civil disturbances; (5) failure of critical service providers or utilities; (6) pandemics; (7) terrorism and warfare; (8) cyber-attacks and cyber-warfare; (9) natural disasters (bushfires, earthquakes, floods, haze, hurricanes, tornadoes, winter storms); (10) mega disasters.
- Six plausible wide-scale or major disruption scenarios (discussed in depth):
  - Critical service provider or utility failures (telecommunications, messaging providers, common cloud providers; power outages).
  - Pandemics (staff absenteeism; loss of key personnel; pandemic management teams and plans).
  - Terrorism (physical attacks on RTGS, clearing houses, cross-border systems; unilateral sanctions suspending access).
  - Cyber-attacks and cyber-warfare (financial sector estimated to be "three times more at risk" of cyber-attacks than any other sector; rising cyber incidents).
  - Natural disasters (hurricane/earthquake damage to primary and secondary sites; climate change increasing uncertainties).
  - Mega disasters (high-impact, low-probability complex events; example: Great East Japan Earthquake).
- Observations on coordinated or simultaneous incidents: coordinated attacks on primary and secondary sites lacking distinct risk profiles could cripple systems and prolong outages, potentially necessitating a third site.
- Need for cross-sectoral and cross-border crisis management arrangements involving multiple authorities and stakeholders.

#### C. International principles and assessments
- CPSS-IOSCO PFMI provides assessment framework with seven key considerations for operational risk (assessment guided by specific questions; see Appendix 2 in source).
- Basel Committee on Banking Supervision (BCBS) Principles for Operational Resilience provides complementary principles for banks’ critical functions associated with payments, clearing and settlement (BCBS, 2021).
- IMF and World Bank SIPS detailed assessments benchmark practices against PFMI; operational risk and cyber security supervision have been included in assessments.
- Interdependency mapping is critical—interdependencies exist at three levels: institutions, systems, and environments; mapping can identify concentration risks and single points of failure.

### III. International experiences — observed shortcomings and incidents
- High-level observation: existing electronic payment systems have experienced operational incidents in critical service providers, payments and currency services, and SIPS.
- Four major RTGS outages (October 2014 to February 2021) — Table 1 (transaction figures for 2019; outage hours from public incident information):
  - Australia (RITS)
    - Volume (millions): 13
    - Value of transactions (USD billions): 36,629
    - Average value per transaction (USD thousands): 2,889
    - Value of transactions as a percentage of GDP (%): 2,642
    - Outage (hours): <2
  - Euro Area (TARGET2)
    - Volume (millions): 89
    - Value of transactions (USD billions): 509,382
    - Average value per transaction (USD thousands): 5,733
    - Value of transactions as a percentage of GDP (%): 3,813
    - Outage (hours): 10
  - United Kingdom (CHAPS Sterling)
    - Volume (millions): 49
    - Value of transactions (USD billions): 106,358
    - Average value per transaction (USD thousands): 2,186
    - Value of transactions as a percentage of GDP (%): 3,766
    - Outage (hours): 9
  - United States (Fedwire)
    - Volume (millions): 168
    - Value of transactions (USD billions): 695,835
    - Average value per transaction (USD thousands): 4,149
    - Value of transactions as a percentage of GDP (%): 3,200
    - Outage (hours): 3 to 4
- Key international lessons and observations:
  - Reliability Objectives (OROs): meeting quantitative OROs such as the 2h-RTO can be difficult; operators were generally able to complete end-of-day settlements despite breaches of short RTO benchmarks.
  - Redundancies: backup systems have fallen short of commencing immediate operations; primary and backup sites in same region can fail to provide seamless transition; staff unavailability complicates BCP execution.
  - Critical Service Providers: outages often linked to critical service providers (software failures, telecommunication services, utilities); cloud concentration risk noted.
  - Endpoint Security: cyber-attacks were often ruled out in some SIPS incident investigations, but cyber resilience efforts continue internationally; endpoint security of wholesale payments remains a focus to prevent fraud.
  - Alternative Arrangements: manual paper-based procedures proved useful in at least one incident; for wide-scale prolonged outages, the role of cash in national crisis preparedness becomes a central question (summary and background in Appendix 4).

### IV. Development and operational issues for the future — five priority areas
- Overview: guided by PFMI and business continuity management, issues to strengthen operational resilience include impact tolerance, business continuity arrangements, tandem processing, interoperability, and cost & implementation.

#### A. Impact Tolerance
- Distinguish impact tolerances from RTOs:
  - Impact tolerance = "the maximum tolerable level of disruption for an important business service, whereby further disruption could significantly threaten the transfer of payments or the safety and efficiency of the payment system."
  - RTOs (e.g., 2h-RTO) remain relevant and should be aligned with business impact analysis and enterprise risk.
- Impact tolerance may be time-based or include metrics such as number of end-users impacted, or volume/value of payments disrupted.
- Supervisory expectation example: United Kingdom expects operators to set flexible impact tolerance levels for each business service/supporting service.
- Interdependencies: when key business services interconnect with SIPS or FMIs, their impact tolerances should consider OROs and impact tolerances set for those FMIs.

#### B. Business Continuity Arrangements
- Alternative arrangements and preparedness measures to consider:
  - Common, identical, and global payment standards.
  - Common global payment systems with geographically dispersed processing sites.
  - Real-time copies of transaction and account databases in secure environments with distinct risk profiles and redundancy.
  - Parallel satellite-based communication facilities (in addition to landline) for disaster situations.
  - Comprehensive local power-generation capabilities.
- Physical security enhancements (e.g., underground sites where practicable).
- Importance of establishing payment facilities outside impacted areas with updated complete sets of payment account and transaction data to commence immediate processing.
- Note on electricity needs: "The electricity capacity needed for payment instruments are rather limited and can be supplied with high-capacity battery, solar-electricity, small even man-powered generators" — but emergency supplies require advanced preparation.

#### C. Tandem Processing
- Tandem processing defined: "each transaction is processed using two separate independent parallel processing streams" with automatic comparisons for correctness; akin to an electronic "four eyes" control.
- Distinctions and benefits:
  - Eliminates single points of failure in payment system software (current state: duplicated hardware but single processing, database, and telecommunication software).
  - Provides continuous backup and rapid error detection; automated failure correction; seamless processing if at least one parallel system remains operational.
- Key elements described:
  - Duplicated independent network connections using completely independent networks and communications hardware/software; continuous dual routing with automatic fallbacks to solo-mode if a network fails.
  - Duplicated independent software/hardware components: two parallel independent payment processing systems with different risk profiles.
  - Automatic synchronization checks to compare outputs of parallel processes and alert/fallback mechanisms; AI-based algorithms could support automatic analysis and error correction.
  - Decentralized transactional data recording and storage with redundancy: at least two separate transaction records and real-time copies at secondary site.
  - Settlement integrated with payment transaction processing; "Real-time payments need immediate (and non-probabilistic) finality."
  - Resilient information security: transactions encrypted using at least two different encryption systems to avoid single-algorithm exposure.

#### D. Interoperability
- Current landscape: many payment systems are proprietary, non-standardized, and fragmented; estimated "at least 1,200 payment systems" worldwide.
- Need for common standards to achieve:
  - Domestic and cross-border resiliency and interoperability.
  - Cost efficiency, improved service levels, and availability.
  - Ability for central banks to provide mutual backup services and allow unaffected infrastructures in neighboring or distant countries to process payments during disasters.
- Interoperability would support tandem processing and enable globally standardized systems to reduce development costs through reusable programming libraries.
- International cooperative arrangements required for information sharing and crisis management as cross-border payments and new technologies evolve.

#### E. Cost and Implementation
- Investments in resilience can be undervalued in normal times but yield benefits during crises; literature suggests structural resilience investments can "increase potential economic output, lower expected losses, and improve continuity of public services."
- Building duplicated synchronized payment systems entails upfront development and ongoing maintenance/updating costs; costs vary by country and existing standards.
- Standardization reduces development costs when payment systems reuse common libraries.
- Authorities may need regulations or sanctions to stimulate investment where FMIs and PSPs underinvest relative to public expectations; stricter resiliency requirements may be warranted for FMIs and PSPs handling urgent payments to focus resources where benefits are larger.

### V. Conclusion and future research
- Recovery time objectives can be difficult to meet under existing redundancy models.
- Strengthening operational resilience requires anticipating and preparing for wide-scale or major disruptions and recognizing that redundancy alone is insufficient in the digital era.
- Five development and operational priorities reiterated: set impact tolerances; identify alternative recovery arrangements; consider tandem processing; facilitate global interoperability; examine costs and implementation.
- Future research priorities identified:
  - How central bank money (physical cash and CBDC infrastructures) fits into operational resilience frameworks and contingency planning (Appendix 4 summarizes role of cash in contingency planning).
  - How CBDC infrastructures, if implemented, could be designed to ensure operational resilience or act as alternatives to traditional payment systems.
  - Further neutral external analysis of tandem processing feasibility by ICT experts.

*IMF Working Papers — Operational Resilience in Digital Payments — Section 4: Role of Cash in Contingency Planning*

### References

### References

### Appendix 1. Key Challenges for Improving Operational Resilience in Payment Systems
- Efforts to improve operational resiliency are confronted with the following challenges: (i) technical complexity; (ii) organizational complexity; (iii) dependence on single software versions; (iv) dependence on external infrastructures; (v) vulnerability of information security solutions; (vi) geographical, jurisdictional, and time-zone issues; and (vii) insufficient incentives for resiliency.
- Technical complexity.
  - In online and real-time infrastructures, transactions are processed individually involving several applications from several service providers. Each of application and service provider utilize different components. Most of these are acquired from external sources based on outsourcing agreements, to a large extent.
  - Many of these components are duplicated in different parts of the payment processing chain e.g., a database manager software can be employed by several partners. A failure in one underlying technical component can thereby affect several parties.
  - Problem exist when failure in a party or a critical component might completely halt the processing chain. As many components involved the overall chain, the probability of failure is growing. The effects on customers are immediately apparent in online and real-time systems.
  - Footnote: This consist of server equipment (and sometimes mainframes), which contain different operating systems and middleware applications, database manager software (and sometimes separate servers), telecommunication software (and sometimes separate servers), different information security firewalls/servers and the actual payment processing software.
- Organizational complexity.
  - More parties are involved in the processing chain and most of them rely on outsourcing e.g., software developments, telecommunication services, security solutions, payment instrument production (chip-cards, e-banking user facilities etc.).
  - There seems to be a lack of overall resiliency responsibility as it depends on the resiliency level of each individual component.
  - Complexity is even larger with globalization and would not be sufficient to handle continues processing, there is a need for specialized organizational solutions to enable an immediate action against major disruption through local and international coordination.
  - Deciding on parties who is held responsible to find actual problem and fix it would be much harder as technical and organizational complexity are rising.
- Dependence on single software versions.
  - The current payment software is mostly running in one operating system environment.
  - Most PSPs’ system design is characterized with single payment software developed by single developer with single database management system.
  - Any undetected errors/bugs in software or supporting database will require a roll back to the previous version of the software (or fixing or circumventing the bug in the current version) to acquire redundancy copy.
  - This process would be time consuming and complex, especially when major software changes are made within the last update. The risks for internal bugs increase with growing system complexity.
- Dependence on external electronic infrastructures.
  - E-payments rely heavily on electricity and telecommunication networks.
  - Although, the overall quality of these services has improved over the years, idiosyncratic risk like climate change due partly to global warming can impose uncertainty in the availability and failure rates of these infrastructures.
  - Readiness for the backup solutions for these infrastructures would be important to complete the processing chain.
- Vulnerability of information security solutions.
  - Ransomware incidents have been increasing. Cybercriminals have increased their efforts and investments for bypassing different kinds of security and privacy solutions.
  - Severe security breaches will require closing affected systems and services partly and for shorter or longer periods.
  - Typical examples in the past have been replacing off-line ATM systems with online systems because cybercriminals have learnt how to forge off-line cards.
  - Well-organized terrorist/criminal attacks on electronic payment systems can potentially result in major unrecoverable losses due to fraudulent transactions.
  - The openness and anonymity set up within current internet services provide criminals and terrorists opportunities to attack with limited risks of being detected.
  - Extensive investments and restructuring in information security solutions would be needed to protect payment systems.
- Geographical, jurisdictional, and time-zone issues.
  - Geographical proximity of backup sites will increase risks.
  - Domestic PSPs will have less control on payment service resources and employment with the increasing use of cross-border payments and offshore outsourcing partners, cloud computing etc.
  - Smaller PSPs in developing countries will probably have less bargaining power compared to larger competitors in the event of failure.
  - In addition, conflicting rules and regulations (e.g., finality, prefunding etc.) may occur if payment transactions are processed in different jurisdictions.
  - The fourth era of payment systems will operate continuously 24/7, in multicurrency mode and thereby without time-zones and end-of-day timings. This will require new type of solutions and open hours for settlement systems like central banks RTGS systems.
- Insufficient incentives for resiliency improvements.
  - Private service providers are generally less risk-averse compare with their customers and society in general.
  - Aside from their profit motive, they tend to maximize transaction volumes and underinvest, for instance, in costly backup, security and similar solutions until the decision becomes necessary due to competition or regulatory requirements.
  - Consequently, different kinds of failures may occur.

### Appendix 2. Principles Relevant for Assessing Operational Resilience in Payments
- CPSS-IOSCO Principles for Financial Market Infrastructures—Operational Risk
  - Principle 17: Operational Risk
    - An FMI should identify the plausible sources of operational risk, both internal and external, and mitigate their impact through the use of appropriate systems, policies, procedures, and controls. Systems should be designed to ensure a high degree of security and operational reliability and should have adequate, scalable capacity. Business continuity management should aim for timely recovery of operations and fulfilment of the FMI’s obligations, including in the event of a wide-scale or major disruption.
  - Key Considerations
    1. An FMI should establish a robust operational risk-management framework with appropriate systems, policies, procedures, and controls to identify, monitor, and manage operational risks.
    2. An FMI’s board of directors should clearly define the roles and responsibilities for addressing operational risk and should endorse the FMI’s operational risk-management framework. Systems, operational policies, procedures, and controls should be reviewed, audited, and tested periodically and after significant changes.
    3. An FMI should have clearly defined operational reliability objectives and should have policies in place that are designed to achieve those objectives.
    4. An FMI should ensure that it has scalable capacity adequate to handle increasing stress volumes and to achieve its service-level objectives.
    5. An FMI should have comprehensive physical and information security policies that address all potential vulnerabilities and threats.
    6. An FMI should have a business continuity plan that addresses events posing a significant risk of disrupting operations, including events that could cause a wide-scale or major disruption. The plan should incorporate the use of a secondary site and should be designed to ensure that critical information technology (IT) systems can resume operations within two hours following disruptive events. The plan should be designed to enable the FMI to complete settlement by the end of the day of the disruption, even in case of extreme circumstances. The FMI should regularly test these arrangements.
    7. An FMI should identify, monitor, and manage the risks that key participants, other FMIs, and service and utility providers might pose to its operations. In addition, an FMI should identify, monitor, and manage the risks its operations might pose to other FMIs.
  - Source: CPSS-IOSCO (2012a).
- CPSS-IOSCO Principles for Financial Market Infrastructures—Questions for Assessing Operational Resilience of Payment Systems
  - Key Considerations and Questions
    - 3 An FMI should have clearly defined operational reliability objectives and should have policies in place that are designed to achieve those objectives.

*wpiea2021288-print-pdf - References*

### 1. What are the FMI’s operational reliability objectives, both qualitative and quantitative? Where and how

### 1. What are the FMI’s operational reliability objectives, both qualitative and quantitative? Where and how

### Business continuity objectives and quantitative targets
- Business continuity plan objectives:
  - Address events posing a significant risk of disrupting operations, including events that could cause a wide-scale or major disruption.
  - Incorporate the use of a secondary site.
  - Ensure that critical information technology (IT) systems can resume operations within two hours following disruptive events.
  - Be designed to enable the FMI to complete settlement by the end of the day of the disruption, even in case of extreme circumstances.
  - Be regularly tested.
- Specific assessment questions that document objectives and expectations:
  - How and to what extent does the FMI’s business continuity plan reflect objectives, policies and procedures that allow for the rapid recovery and timely resumption of critical operations following a wide-scale or major disruption?
  - How and to what extent is the FMI’s business continuity plan designed to enable critical IT systems to resume operations within two hours following disruptive events, and to enable the FMI to facilitate or complete settlement by the end of the day even in extreme circumstances?

### Design features and contingency procedures
- Contingency and transaction status:
  - The plan should ensure that the status of all transactions can be identified in a timely manner at the time of the disruption.
  - If there is a possibility of data loss, procedures should exist to deal with such loss (for example, reconciliation with participants or third parties).
- Crisis management and communications:
  - Crisis management procedures should address the need for effective communications internally and with key external stakeholders and authorities.
- Alternative arrangements:
  - Consideration of manual, paper-based procedures or other alternatives to allow processing of time-critical transactions in extreme circumstances.

### Secondary site requirements
- Secondary site expectations:
  - The business continuity plan should incorporate the use of a secondary site, ensuring the secondary site has sufficient resources, capabilities, functionalities and appropriate staffing arrangements.
  - Assessment of geographic separation: To what extent is the secondary site located a sufficient geographic distance from the primary site such that it has a distinct risk profile?

### Review, testing, and participation
- Review and testing requirements:
  - Business continuity and contingency arrangements should be reviewed and tested, including with respect to scenarios related to wide-scale and major disruptions.
  - Frequency of reviews and tests should be specified and assessed.
- Involvement of stakeholders in testing:
  - Review and testing should involve the FMI’s participants, critical service providers and linked FMIs as relevant.
  - Frequency and extent of participant, service provider, and linked FMI involvement in review and testing should be assessed.

### Monitoring and managing risks from participants, FMIs, and service providers
- Risk identification and management:
  - Identify, monitor, and manage the risks that key participants, other FMIs, and service and utility providers might pose to the FMI’s operations.
  - Identify, monitor, and manage the risks the FMI’s operations might pose to other FMIs.
- Outsourcing expectations:
  - If services critical to operations are outsourced, ensure that the operations of a critical service provider meet the same reliability and contingency requirements they would need to meet if provided internally.
- Coordination with interdependent FMIs:
  - Assess the extent to which the FMI coordinates its business continuity arrangements with those of other interdependent FMIs.

### BCBS Principles for Operational Resilience (selected considerations)
- Governance:
  - Use existing governance structure to establish, oversee and implement an effective operational resilience approach that enables response, adaptation, recovery and learning from disruptive events to minimize impact on critical operations.
- Operational risk management:
  - Leverage functions to identify external and internal threats, assess vulnerabilities of critical operations and manage resulting risks.
- Business continuity planning and testing:
  - Establish business continuity plans and conduct exercises under a range of severe but plausible scenarios to test delivery of critical operations through disruption.
- Mapping interconnections and interdependencies:
  - Map internal and external interconnections and interdependencies necessary for delivery of critical operations.
- Third-party dependency management:
  - Manage dependencies on relationships, including third parties or related entities, for delivery of critical operations.
- Incident management:
  - Develop and implement response and recovery plans in line with risk appetite and tolerance; continuously improve plans by incorporating lessons learned.
- ICT including cyber security:
  - Ensure resilient ICT and cyber security with protection, detection, response and recovery programs that are regularly tested and provide situational awareness and timely information for risk management and decision-making.

### Empirical incidents underscoring operational reliability objectives
- Australia — Reserve Bank of Australia’s RITS (July 6, 2020):
  - Incident time: 7.30 am.
  - Outcome: Majority of RITS services re-established and fully operational from the RBA’s primary site around 9 am and within the two-hour RTO.
  - Consequence: Opening of the RITS Daily Settlement Session delayed from 9.15 am to 9.30 am; all transactions were able to settle as expected from that time.
  - System scale: 167 participants, of which 102 are exchange settlement account holders, including the RBA.
  - Cause: Operational error (power supply inadvertently shut off during fire control system maintenance; legacy switchboard design issue; contractor non-compliance).
  - Response: Upgraded switchboard; enhanced contractor induction arrangements; improved oversight of compliance; new service delivery arrangements; increased role for staff with engineering expertise.
- Europe — TARGET2 outage (October 24, 2020):
  - Duration: around 10 hours.
  - Impact: All settlement services unavailable; backup system failed to boot up as failover to secondary disaster recovery site took many hours.
  - Financial impact: Drop in deposits worth more than EUR 400 billion (USD 473 billion).
  - Cause: Software defect at a third-party network device.
  - Response: Independent review announced to investigate failures including business continuity model, recovery testing, change management, and communication.
- Mastercard outage (July 12, 2018):
  - Duration: more than one and a half hours.
  - Impact: Some card payments failed and affected many international banks.
  - Public information on cause/impact/response: No public information available beyond a statement that situation was fully resolved.
- Visa Europe Ltd (VEL) outage (June 1, 2018):
  - Duration: partial outage around 8 hours.
  - Impact: Around 5.2 million card transactions across Europe affected.
  - Consequences: Inconvenienced travel, grocery payments, cash withdrawals; some ATMs in the United Kingdom were out of cash within a couple of hours.
  - Cause: Hardware failure (authorization to chip and pin machine).
  - Response: Internal and independent third-party review; work program to oversee remedial actions.
- Bank of England CHAPS outage (October 20, 2014):
  - Duration: over 9 hours.
  - Financial impact: Effectively froze GBP 277 billion (USD 447 billion) worth of payments; put 700 time-critical housing market transactions on hold.
  - Cause: Configuration changes triggered previously undetected design defect.
  - Response: Independent review; findings presented to Bank of England’s Court; full report with response published March 23, 2015.
- Federal Reserve Fedwire Funds Service disruption (February 24, 2021):
  - Duration: around 3 to 4 hours.
  - Additional impact: Affected 14 other services including Check 21, FedCash, Account Services, Central Bank and several FedLine services; some cryptocurrency exchanges experienced delays.
  - Cause: Operational error.
  - Response: Steps taken to help ensure resilience, including recovery to point of failure; Fedwire resumed normal operations after the outage.
- Travelex cyber-attack (December 31, 2019):
  - Vulnerability period: forced to shut down and was vulnerable for 8 months.
  - Impact: Affected customers in 21 countries; attackers claimed access to a large volume of customer data and demanded a ransom of USD 6 million.
  - Response: Travelex paid equivalent of USD 2.3 million and completed a restructuring deal August 2020 delivering GBP 84 million (USD 119 million) of new money.
- Bangladesh Bank SWIFT fraud attempt (February 4, 2016):
  - Attempted theft: USD 951 million.
  - Transfers executed: USD 101 million transferred from Bangladesh Bank’s account at the Federal Reserve Bank of New York; USD 20 million traced to Sri Lanka; USD 81 million ended up in the Philippines.
  - Cause: Wholesale payments fraud related to end-point security using malware to obtain legitimate SWIFT credentials.
  - Response: SWIFT introduced mandatory security measures including daily reports and the Customer Security Program; launched SWIFT Information Sharing and Analysis Centre May 2017.
- Zimbabwe — Econet Wireless power outages (July 2019 and July 2020):
  - Duration: power outage of 8 hours.
  - Impact: Hit 6.7 million active users; 70 percent of Zimbabweans had no phones, internet, and mobile money services; detrimental effects on the economy.
  - Cause: Utility failure; generators failed to kick in after prolonged electricity disruption.
  - Response: Investment in 520 lithium-ion Tesla batteries to provide backup to 1,300 base stations; cut reliance on diesel-run generators by 75 percent.

### Role of cash in contingency planning and resilience
- Functional role of cash:
  - Cash acts as legal tender, a means of payment guaranteeing privacy, and is the only form of money that can be carried or kept as savings independently of a bank and as a fallback if electronic payment systems fail.
  - Cash can provide immediate security for payments without third-party involvement, is independent from central verification, and is less dependent on electronic systems (though ATM and over-the-counter withdrawals may still require electricity).
- Contingency implications:
  - Wide-scale or major disasters may produce sudden peaks in demand for cash due to rising uncertainty.
  - Absence of non-digital fallback plans, such as cash, could expose economies to significant risks from system failure or cyber-attack.
  - For emerging and developing economies with rudimentary payment and telecommunications infrastructures, identifying cash as critical infrastructure is particularly relevant (blank spots in remote areas lacking fixed broadband or mobile data signals).
- Operational resilience for cash infrastructures:
  - Cash infrastructures need operational resilience; central banks and stakeholders must balance cost savings with continuity.
  - Decentralization of money processing (counting, sorting, checking, recirculating) by banks or licensed operators could improve efficiency and robustness.
  - Maintaining continuity of security transports to ATMs is essential to ensure accessibility and availability of cash.
- Policies and measures taken by jurisdictions:
  - Dutch central bank sets 4,000 ATMs as the minimum number to ensure adequate access to cash services.
  - Swedish government law requires the largest banks to continue to provide adequate access to cash and requires the Swedish central bank to have sufficient locations for professional parties to obtain banknotes.
  - Central banks can use crisis management frameworks to guide deployment of cash as a fallback, including timely communication to market participants and operational measures (example: Federal Reserve actions after September 11 to reassure banks and armored carriers, extend hours, arrange special deliveries, and waive certain fees).

*Source: CPSS-IOSCO (2012b); BCBS (2021); selected incident reports and central bank sources as compiled in the IMF Working Paper.*

---


_Source: https://www.imf.org/-/media/files/publications/wp/2021/english/wpiea2021288-print-pdf.pdf_
