This is a draft version! Do not share the link externally!

Article

The Payments Resilience Imperative

Speed without resilience is not a product. It is a liability.

Payments have simultaneously become an engineering challenge and a board-level imperative–not just a financial problem.

Key Takeaways

  • Stop measuring uptime. Start measuring journeys. Regulators and customers don't experience "99.99% available" — they experience a payment, a fraud check, a settlement. Journey-level SLOs are the only resilience metric that's provable, not performative.
  • ‍Resilience has shifted from a compliance requirement to a core business capability. In payments, regulation, customer expectations, and the economics of downtime are converging; making resilience a board-level commercial priority that must be engineered into platforms, not managed as a parallel risk program.
  • ‍AI can compress triage from hours to minutes but only with the right telemetry feeding it. AI is correlating signals and drafting DORA incident reports in real time, cutting root-cause detection dramatically. But that speed is only as good as the data underneath it.
No items found.

The industry now operates at a colossal scale. In the 2025 financial year, a leading global card network processed 329 billion transactions, and a leading global payment processor peaked at 199,000 transactions per minute on Black Friday. Because of this, the margin for failure is measured in seconds.

‍

A single degraded API, a missed sanctions screening window, or an unmanaged software release can breach a regulatory obligation, trigger a DORA incident report, or cost a key merchant relationship. What was once an IT concern has transformed into commercial and regulatory exposure, and it is built into every line of code.  

‍

For CIOs, the question is no longer whether to invest in operational resilience. It is time to determine whether your organization is building operational resilience with the discipline of an engineering company (or just managing it like a compliance programme).

‍

Across payments infrastructure today, three key forces are reshaping the very meaning of operational resilience, and a robust response is required.  

‍

1. Regulation is now an engineering mandate

With DORA in full force across the EU, impact tolerance settings, end-to-end ICT risk frameworks, continuous resilience testing, and direct regulatory supervision of critical third-party providers are now mandatory.

‍

In November 2025, the European Supervisory Authorities designated the first 19 critical third-party ICT providers, bringing them under direct EU oversight (backed by penalties of up to 1% of average daily worldwide turnover). DORA requires that initial notifications are submitted within four hours of classifying an incident as major, an intermediate report is provided within 72 hours, and a final report is completed within a month.

‍

When it comes to other parts of the world, the US Incident Notification Rule demands breach reporting within 36 hours, while in Singapore, 99.95% system availability is mandated. At the global level, the PCI DSS v4.0 (Payment Card Industry Data Security Standard) completed its transition in March 2025, moving from periodic validation to continuous security assurance.

‍

Regulators are going beyond auditing policies to test operational behaviour, and now they publish the results - CIOs need to view their SRE (Site Reliability Engineering) practice as their compliance posture. In June 2026, the ESA’s (European Supervisory Authorities) first annual incident report logged 3,383 major ICT incidents across EU financial entities during 2025. These were concentrated in the credit and payments spaces, with almost a third of cases traced to third-party failures, a tenth to cyberattacks, and the rest to system failures and external events.  

‍

2. Fraud has gone AI-native

According to BCG’s June 2026 analysis, agentic capability improved roughly tenfold in a year since 2024, and when agents can run scams end-to-end, they will also cost 90% less. The implication of this is a potential twofold (or greater) rise in successful scams and fraudulent activity.

‍

Scams alone already cost consumers and businesses around $442 billion a year, and now enumeration attacks, synthetic identity creation, and deepfake-enabled social engineering are automated, cheap, and globally scalable. Leading global card networks are acting on this arms race, blocking nearly twice as many fraudulent e-commerce transactions in 2025 as the year prior, while also cutting rates of fraud across its ecosystem by 8%.  This same company offers another example, having paired its network fraud signals with the threat intelligence product launched by a threat intelligence vendor in late 2025.

‍

Crucially, these industry-leading players are demonstrating the need to fight fire with fire – the race will not be won by compliance teams with spreadsheets, but with AI at industrial scale.  

‍

3. Availability is a commercial metric, not a technical one

A senior executive at another leading global payment processor frames it directly: 40% of customers who have a payment declined abandon that merchant entirely.

‍

Having carried $1.9 trillion of volume in 2025, equating to roughly 1.6% of global GDP,  That company could lose tens of millions of dollars of commerce in just five minutes of downtime. Every second of degradation now comes with immediate, measurable commercial consequences that are visible to customers, merchants, and regulators alike.  

‍

The required shift: from risk item to system property

The three forces above are distinct in nature but share a common implication: resilience can no longer be managed as a parallel workstream. It must be engineered into the system, and SRE is the discipline that makes that shift operational.

‍

SRE applies software engineering discipline to operational resilience: measurable service-level objectives instead of aspirational uptime targets, automated failure response instead of human escalation chains, and chaos testing to find failure modes before customers do.

‍

Every one of the seven principles of SRE map directly to a DORA article. 

Exhibit 1: Each of SRE's 7 core principles maps directly to a DORA regulatory article — making SRE the natural implementation vehicle for operational resilience. Source: BCG DORA Overview for Financial Institutions; BCG, Risk and Compliance 2026 (April 2026).

‍

The firms setting the benchmark are those that treat availability as a DORA compliance engine, not a back-office IT discipline. That leading global payment processor's chaos testing engine deliberately triggers multiple concurrent faults in production, because that is the only way to discover failure modes that only emerge under compound stress.

‍

A streaming technology company pioneered the same philosophy with Chaos Monkey1 in 2011. The logic: if your systems cannot survive a randomly terminated instance, you have a dependency you do not know about. Today, Chaos Kong can take down an entire data centre and this company will keep on streaming. 

ResilienTech: seven principles that turn resilience into a system property

BCG's ResilienTech data shows that SRE-mature organizations experience 10 to 30% less downtime, 10 to 15% lower IT costs, and two to five times faster release cycles. None of these benefits are soft, they directly reduce the cost of DORA compliance, reduce the frequency of material incidents that require regulatory reporting, and compress the time between a threat and your response. 

‍

Exhibit 2: BCG's ResilienTech® framework . -7 principles that operationalise resilience through pipelines, code, and automation rather than policies and controls. Source: BCG ResilienTech Roadshow, 2025.

‍

The organizations that achieve these outcomes share one characteristic; they treated reliability as a design input rather than a post-launch audit. Resilience bolted onto a system after it is built costs significantly more (and typically delivers weaker results). The failure modes are already baked into the architecture, the dependencies are already invisible, and the team has already normalised the risk.

‍

SRE by design means error budgets are defined before a service is built, failure modes are mapped during architecture review, and chaos experiments are part of the release process. The cultural implication is equally significant: engineers own reliability as a product property, product managers treat error budget consumption as a delivery constraint, and technology leadership makes the investment case before the incident - not during the post-mortem.

‍

What industry leaders do differently

The risk surface facing payment service providers has never been broader. Financial, operational, and strategic risks are converging, driven by macroeconomic pressure, AI-enabled fraud, regulatory mandates like DORA, and an accelerating competitive landscape.

‍

No single control or programme addresses this in full. What industry leaders demonstrate is that these risks can be mastered, but only when engineering discipline, AI at scale, and regulatory compliance are treated as a single integrated architecture rather than parallel workstreams.

‍

Exhibit 3: The payments risk taxonomy financial, operational, and strategic risks that CIOs must manage across a single integrated architecture. Source: BCG Global Payments Report, 2025.

‍

GenAI changes the economics of resilience and risk, if governed properly 

What will separate the next wave of leaders from today's benchmark firms is how they deploy GenAI. Frontrunners will treat the technology as a core component of their resilience and compliance architecture, rather than a productivity tool.

‍

BCG’s October 2025 know-your-customer work revealed that banks are targeting KYC cost reductions of up to 50% using predictive, generative and agentic AI. The prize is large because the baseline is poor; BCG’s compliance benchmarking found alert false-positive rates that reach or exceed 90%. 

When it comes to SRE, GenAI is compressing triage timelines from tens of minutes to mere seconds: correlating telemetry, surfacing probable root causes, and drafting the incident communications required by DORA. In compliance, large language models are enriching KYC profiles with adverse media, detecting indirect ownership links to sanctioned entities that name-matching cannot see, and drafting policy updates when regulations change.

‍

‍The governance obligation is real. AI systems operating in fraud, triage, or incident response are themselves critical ICT systems under DORA. They require audit trails, explainability, and human override at high-consequence decision points. Ungoverned GenAI in your SRE and compliance stack creates a new risk surface.

‍

Three key actions for CIOs in 2026 

Each of the three forces reshaping payments resilience demands a concrete response.

  1. ‍Treat DORA as a technology transformation, not a compliance workstream. The organizations failing to comply with DORA tend to be those that assigned it to Legal and IT Risk. The ones succeeding have a board-level resilience owner, a live ICT risk register, and resilience testing practices built into every stage of their release pipeline. Now is the time to close the gap between your paper framework and your production reality.
  2. ‍Embed GenAI into SRE and compliance by design, with governance built in from the start. Deploy AI-assisted triage, alert auto-closure, and compliance intelligence. Embed the audit trail, the explainability layer, and the human override at the same time. The cost of retrofitting governance is an order of magnitude higher than building it in.
  3. ‍Make journey-level SLIs2 your primary resilience metric. System uptime is not what regulators or customers experience. Define SLOs3 for every critical customer journey, payment initiation, fraud screening, and settlement confirmation – then instrument them end-to-end. If you cannot measure it, you cannot prove it to a regulator or to yourself. 

The defining question

The leaders in this space have built their resilience and security postures deliberately over the course of years. These leading global payment processors did not maintain 99.9999% uptime through the busiest trading weekend of 2025, and another leading global card network did not double the fraudulent e-commerce transactions it blocks, by running compliance programmes. They engineered it.

‍

‍For payments CIOs, the question is simple: are you building resilience as a system property, or managing it as a risk item? The answer will determine whether your organization leads the next phase of payments or is regulated, disrupted, or both.

‍

‍To continue the conversation, contact our BCG Platinion expert team. 


More to Explore

No items found.
No items found.
No items found.
Financial Institutions
Consumer
No items found.