What the AI cyber risk debate misses
AI Systemic risk Policy
The financial authorities are worried about AI cyber risk. Policy reports, speeches and financial stability reviews all warn that AI lowers the cost and raises the power of attacks on the financial system, and the volume has only grown as frontier models such as Mythos take on tasks that once required skilled human attackers. We are now seeing open-source models that seem to be nearly as capable as Mythos, without any safeguards.
The concern is justified. AI helps criminals find loopholes, lets attackers orchestrate synchronised strikes on financial infrastructure, and gives nation states plausible deniability.
A decade ago, only major states could mount aggressive cyber operations against the financial system. That bar is falling fast. The lag between closed and open models is now under a year, so frontier capability will soon be widely available, and extortion becomes more profitable for the likes of North Korea. The more capable attackers there are, the more likely it is that one of them times an attack for a moment of stress.
The advantage runs to the attacker. A defender has to protect the whole system, every institution, every connection, every ageing legacy component, while an attacker needs only one critical weak point. Finding weak points is exactly what AI is good at. And the weak points that matter most are not individual banks but the plumbing that connects them, the payment systems, the clearing houses and the custodian banks.
But one ingredient tends to be missing from the discussion. The double coincidence.
Suppose a serious cyber attack lands on the financial system on a calm day, as it almost always would. Liquidity is ample, trust is intact, and the private and public sectors absorb the hit. The event ends up as a costly operational incident and a case study.
Now let the same attack land in the middle of a liquidity crisis, like the Covid dash for cash in March 2020 or the days after Lehman failed in September 2008. The attack and the crisis amplify each other. Institutions that would have absorbed the loss are busy protecting themselves and everybody wants cash at the same time. Payments become uncertain, liquidity is hoarded and settlement fails, deepening the very stress that made the attack dangerous.
The authorities are stretched as their attention and resources are consumed by failing markets and institutions, so a cyber attack landing at the same time competes with everything else for scarce capacity.
And it arrives just as their credibility, the main tool for calming a panic, is draining away.
On a calm day a failed payment is an IT problem. In a crisis, when everyone is asking which counterparty goes next, a payment that does not arrive is read as evidence about the bank rather than about its systems. The operational failure and the solvency fear become indistinguishable, and a glitch can start a run.
This is why the same system that absorbs an attack on a calm day can amplify it in a crisis, what I called the double coincidence back in 2016.
All of this assumes the timing is bad luck. It need not be. Stress is easy to observe from the outside, in funding spreads, in volatility and in the use of central bank facilities. An attacker who can wait will strike when the damage is greatest. Nation states have the patience and the resources, and waiting buys deniability as well.
A state does not even have to wait for stress, because it can manufacture it. Selling into a thin market, planting a false report that a bank is failing or disabling a smaller institution first will spread doubt about counterparties before the main attack lands. The first stage looks like an ordinary market event, so the defenders do not know the attack has begun. The coincidence in the double coincidence is then no coincidence at all. The worst case can be engineered.
The severity of the potential attack is therefore not the only question to ask. The state of the system when it arrives matters as much.
Why is the double coincidence so easy to overlook? Because the data lies. Almost every cyber attack is contained. The system handles incidents all the time, and each contained attack is a data point saying that attacks of that size are manageable. The defences worked, the incident report is filed, and confidence grows. The authorities conclude they know how to handle it.
When contingency plans are informed solely by the record of incidents contained under favourable conditions, they will be over-optimistic. The result is false confidence and preparations for the wrong event.
Attacks do the most damage when they land in times of stress, and that carries two implications. First, the post mortem of an attack should ask under what more challenging circumstances it would have done the most damage, not only how it was contained. Second, the authorities should make sure the tools they plan to rely on, liquidity support, communications, and the payment and settlement systems themselves, are likely to still be available when financial stress is already high.
The state of the system matters as much as the size of the shock — the double coincidence problem.