Bailey proposes AI testing baseline, not new bank rules
Andrew Bailey wants UK banks to test frontier AI before and after deployment and learn from near misses — but his article sets out a governance starting point, not a new supervisory rule.
- Published

Andrew Bailey, governor of the Bank of England, used a Bank of England article on 30 September 2026 to set out what he called a "sensible starting point" for governing frontier AI: rigorous testing of models before they are deployed, and continued testing once they are in use. Cyber defence, trading and payments are the financial-sector applications Bailey identified as candidates for this kind of testing, and for UK banks weighing AI use in those areas, the piece reads less as a new rule than an attempt to define the problem regulators will eventually have to solve.
Bailey's article, "Frontier AI and the question of governance", did not announce a consultation, a mandatory testing standard, an enforcement policy or an implementation date. It is a framing document from the governor of the central bank that oversees UK financial stability, published at a moment when the Bank's own Financial Policy Committee (FPC) has already judged that frontier AI could pose material risks to UK financial stability through cyber and operational-resilience channels, in its July 2026 record. That combination is why the article matters to UK banks now.
Testing before and after deployment, not continuous monitoring
Bailey argued that testing should help users understand how a model behaves, identify its vulnerabilities, and build confidence in the safeguards meant to contain it. Crucially, he framed this as something that happens "before and after deployment" — assurance across the lifecycle of a model, rather than a defined requirement for real-time, continuous monitoring. Describing this as continuous testing overstates Bailey's own wording. The more accurate description is lifecycle or ongoing post-deployment testing: checking a model before it goes live, and again as it operates, rather than an always-on technical monitoring regime with specified frequency, methodology or pass/fail thresholds. Bailey did not set any of those parameters.
He was also explicit that testing is not a substitute for regulation. He said it was neither a complete solution nor an alternative to a possible future formal framework — testing, in his account, is a precursor that builds the evidence base a regulator would need, not a replacement for one. Secondary coverage sometimes reads this as a preference for testing over regulation; Bailey's text does not support that reading.
Can banks stop an autonomous system?
A recurring question in Bailey's piece is whether society — and, by extension, a firm — retains meaningful oversight and the ability to intervene as AI systems become more capable. He wrote about establishing "credible points of intervention" as models take on more autonomous, multi-stage work with less human input.
What the article does not do is specify what a credible intervention point actually has to be. It does not say whether this means a kill switch, revocation of a system's permissions, transaction or trading limits, a mandatory human-approval step, model rollback, or isolation from tools and networks. For a UK bank, that leaves an open practical question rather than a checklist: can we actually constrain, pause, disconnect or roll back this system quickly enough to stop harm, and who inside the firm has the authority to do it? Bailey raised the question. He did not answer it.
Near misses as governance evidence
Bailey said organisations should learn from incidents and near misses while models are in development, while they are being deployed, and while they are in live use. He treated near-miss learning as part of governance, not simply as a technical bug-fixing exercise.
Again, the detail is missing. The article does not say what counts as a reportable near miss, who should see that information, whether it should be shared across the industry or with regulators, or what confidentiality or legal protections would apply to a firm disclosing one. For a UK bank weighing how to build this into its own model-risk framework, that is a genuine gap, not an oversight on the reader's part.
Cyber defence: matching machine speed without creating new risk
Bailey named cyber defence as one of the financial-sector applications where rigorous testing could inform safe deployment. The Bank, FCA and HM Treasury's joint statement of 15 May 2026 — which explicitly said it introduced no new expectations, only reinforced existing ones — told regulated firms and financial market infrastructures to strengthen protective, detective, containment, response and recovery capabilities, and to manage third-party and supply-chain risk.
The underlying tension is practical. AI-assisted attacks can move faster than manual defence, so some automation in a bank's cyber response may be necessary simply to keep pace. But automated, rapid patching can itself cause outages or disruption if the change process is not well controlled. A bank using AI to defend itself faster is also a bank that can break something faster if the controls around that automation are weak.
Trading: correlated models and market stress
Bailey identified agentic trading — AI systems capable of taking autonomous action toward a goal, using tools, learning from feedback and adapting to conditions — as a second application for rigorous testing. The Bank's concern, set out in its April 2025 Financial Stability in Focus assessment, is about concentration rather than any single firm's model: if many market participants converge on common models, datasets or strategies, their positions could become correlated and amplify a shock during market stress, rather than dampening it. That assessment also found that agentic systems were not widespread in the financial system as of April 2025 — a dated snapshot, not a current measurement of how many UK banks run agentic trading today, which these sources do not establish.
Payments: finality, authority and provider concentration
Bailey's third example was payments. Two distinct risks sit under that heading. The first is concentration: the Bank's 2025 assessment warned that heavy reliance on a small number of external AI providers could impair time-critical payment services if that provider suffered an outage. The second is legal and operational: autonomous agents that are permitted to spend real money raise unresolved questions about who holds authority, who is liable if an agent acts beyond it, and how a payment, once made, can be treated as final. Bailey raised these as open public-policy and legal questions in a July 2026 speech, not as settled law — UK legislation, case law or formal regulatory guidance would be needed to resolve them, and none of the material reviewed for this article does so.
What UK banks must already do
None of this sits in a regulatory vacuum. The table below sets out the rules and statements already in force, which exist independently of Bailey's 30 September article.
| Instrument | Status | Effective / published | What it covers |
|---|---|---|---|
| PRA Supervisory Statement SS1/23 | Current supervisory expectations | Current version effective 23 April 2026 | Model-risk-management expectations — governance, development and use, independent validation, risk mitigants — covering AI and machine learning models where relevant, for PRA-regulated firms with internal-model approval for regulatory-capital calculations |
| Bank, FCA and HM Treasury joint statement | Reinforces existing expectations; creates none new | 15 May 2026 | Operational-resilience expectations for frontier-AI cyber risk: protective, detective, containment, response, recovery and third-party risk management |
| FCA's AI approach | Current stance | As of 2 October 2026 | Principles-based, outcomes-focused; FCA says it does not plan additional AI-specific regulation and will rely on existing frameworks |
| Bailey's governance article | Framing, not a rule | Published 30 September 2026 | Proposes lifecycle testing, intervention points and near-miss learning as a starting point for possible future governance |
The FCA's position that it does not plan extra AI-specific rules does not mean AI use by regulated firms escapes scrutiny. Existing duties — including the Consumer Duty, senior-manager accountability, governance requirements and operational-resilience rules — continue to apply regardless of whether the technology behind a decision is AI.
Separately, the Bank and FCA's 2024 survey of 118 regulated firms — banks, insurers, lenders, investment firms, financial market infrastructures and payments firms — gives a sector-wide sense of adoption, not a measure specific to UK banks or frontier models. Seventy-five per cent of respondents already used AI, up from 58% using machine learning in the 2022 survey, with a further 10% planning to adopt it within three years. Of reported use cases, 55% involved some automated decision-making; 24% of that group were semi-autonomous, and 2% of all reported use cases were fully autonomous. A third were third-party implementations, up from 17% in 2022, and the three most-named model providers accounted for 44% of named model providers — a mention-based measure, not market share, that points to potentially common third-party dependencies rather than a direct measure of concentration in payments specifically. Firms also reported gaps in their own understanding: 46% said they had only partial understanding of the AI they used, against 34% reporting complete understanding. On governance, 84% of firms using AI reported an accountable person or persons for their AI framework, and 72% allocated accountability for use cases and outputs to executive leadership — reported structures this survey does not independently verify.
What Bailey did not announce
Bailey did not create a new supervisory rule, set a compliance deadline, define "frontier AI" by a capability or compute threshold, or specify a testing methodology. Nor did he claim that frontier models in financial services already demonstrate fully autonomous recursive self-improvement — he described systems increasingly using their own outputs to refine performance, a narrower claim. The FPC's July 2026 record was explicit that fully autonomous recursive self-improvement remained unproven, even as models become more capable of complex, multi-stage work with limited human input.
What to watch next
Readers wanting the primary detail should go to the Bank of England's own publication of Bailey's article, the May 2026 joint statement on frontier AI and cyber resilience, and the current text of PRA Supervisory Statement SS1/23, rather than to secondary summaries. The July 2026 FPC record referred to an upcoming Bank and PRA consultation on cyber and ICT risk management; its scope and timing had not been confirmed when this article was researched, and it is the next concrete step worth tracking if a formal governance framework for frontier AI in UK financial services is going to follow Bailey's starting point.
Sources
- Frontier AI and the question of governance (opens in a new tab)
Bank of England · · Accessed
- Financial Policy Committee Record – July 2026 (opens in a new tab)
Bank of England · Accessed
- The Bank, FCA and HM Treasury joint statement on Frontier AI models and cyber resilience (opens in a new tab)
Bank of England, Financial Conduct Authority and HM Treasury · · Accessed
- Financial Stability in Focus: Artificial intelligence in the financial system (opens in a new tab)
Bank of England · Accessed
- Growth and regulation – speech by Andrew Bailey (opens in a new tab)
Bank of England · Accessed
- SS1/23 – Model risk management principles for banks (opens in a new tab)
Prudential Regulation Authority · · Accessed
- AI and the FCA: our approach (opens in a new tab)
Financial Conduct Authority · · Accessed
- Artificial intelligence in UK financial services – 2024 (opens in a new tab)
Bank of England and Financial Conduct Authority · · Accessed


