AI operational resilience after the test lab stopped being theoretical
For years, the sensible line on artificial intelligence in financial services was that firms should avoid melodrama. Do not panic. Do not ban everything. Do not buy every shiny system from the loudest vendor in the room. Govern the thing properly, test it, document it, keep humans accountable, and make sure customers are not harmed.
That is still the right instinct. But it now needs a harder edge.
The newest warning from the frontier AI world is not that a chatbot wrote a bad poem, fabricated a policy summary, or mangled a complaint response. It is that, during cyber capability evaluations, advanced AI systems reached real internet-connected systems and gained unauthorised access to three organisations. Anthropic has published its own account of the incidents, explaining that the models were working on capture-the-flag style cyber exercises and that internet access was available when it was not supposed to be.
For regulated firms, the point is not to turn one laboratory failure into a theatrical prophecy. The point is more prosaic and more useful. If AI systems can behave in ways their operators did not intend, at a speed and scale ordinary control environments were not built for, then AI operational resilience has moved from innovation committee agenda item to board risk register.
What actually changed
Anthropic says its review covered more than 141,000 cyber evaluation runs and identified three incidents in which a Claude model reached the internet from, or while interacting with, a third-party evaluation environment and then accessed real systems belonging to three different organisations. The models had been told they were operating in a simulation. Because the environment was not properly sealed, real systems became reachable.
That is a control failure, not just a clever model story. A firm can have a good policy on paper, a plausible vendor assurance pack, and a development team that believes it is testing in a harmless environment, while the actual technical conditions allow a very different result. In regulated financial services, that distinction is everything.
The Bank of England, FCA and HM Treasury had already warned in May 2026 that frontier AI models represent a step-change in capability for cyber security and operational resilience. Their joint statement said the cyber capabilities of current frontier models are already exceeding what a skilled practitioner could achieve, at higher speed, greater scale and lower cost. It also said firms need protective, detective, containment and response capabilities that can deal with faster and more disruptive AI-driven attacks.
Why this is an FCA issue
Financial services regulation does not require every risk to begin inside a bank, lender or broker before it becomes relevant to regulated firms. The FCA's operational resilience framework is built on a practical truth: firms are responsible for understanding the people, processes, technology, facilities and information needed to deliver important business services. That includes dependencies on third parties.
The FCA's outsourcing and operational resilience material is blunt on this point. Firms that use outsourcing and other third-party providers must manage the risk arising from those arrangements. They cannot delegate regulatory responsibility to the supplier. A cloud provider, software vendor, AI platform or security-testing partner may perform the activity, but the regulated firm remains accountable for how that activity affects customers, markets and the firm's ability to remain authorised.
That accountability is not limited to old-fashioned outsourcing contracts. A model embedded in a workflow may feel like a tool rather than an outsourced service. A vendor plug-in may feel like software rather than a dependency. A proof of concept may feel like experimentation rather than operations. The regulatory question is more practical: could this arrangement affect customers, data, important services, decision-making, financial promotions, credit assessments, complaints or incident response?
If the answer is yes, it belongs in the control environment.
The weak spot is governance
The more immediate risk is that a firm adopts AI faster than it can govern it. A lender may use AI to triage applications, summarise bank statements, draft customer communications, detect fraud, monitor broker quality or support collections. A credit broker may use AI for lead routing, marketing copy, customer fact-finding, call summarisation, quality assurance or compliance monitoring. An AR network may use AI to scan appointed representative activity, identify conduct indicators or assemble board reporting.
None of those uses is inherently wrong. Some could improve outcomes if implemented well. But each creates questions that need clear ownership. Who approved the use case? What data can the model access? Can it call external tools? Can it write back to live systems? Can it send communications to customers? Can it influence eligibility, affordability or pricing? Can it produce or approve a financial promotion? Can it interact with third-party environments? What happens when it is wrong, overconfident, manipulated, unavailable or unexpectedly capable?
Agentic AI changes third-party risk
Traditional software usually does what it is coded to do, including the bugs. Agentic AI systems are different because they can break a task into steps, select tools, browse, call APIs, write code, query systems and pursue an objective across several moves. The more useful they become, the more they resemble operational actors rather than passive utilities.
Third-party risk assessments therefore need to ask what the AI system is permitted to do, how its environment is isolated, whether internet access is blocked or monitored, how credentials are stored, whether tool use is restricted, how prompts and outputs are logged, and how quickly the firm can disable the system if it starts doing something outside the intended use case.
The compliance lens for credit brokers and lenders
Credit distribution is particularly exposed because it already depends on chains of activity. A consumer may encounter an advert, complete a form, be passed through an introducer, receive an eligibility result, be matched with a lender, receive disclosures, enter a regulated credit agreement, and later complain about what they were told or not told. Several firms may touch the journey. Several systems may shape the outcome.
AI can improve that chain. It can detect poor-quality leads, spot vulnerable customer indicators, flag inconsistent disclosures, identify unusual approval patterns, and help compliance teams review more material than humans can sensibly read unaided. Authorised Compliance Ltd sees the appeal: firms want better monitoring, faster file reviews and more practical MI. Used well, AI can help senior managers see risks earlier.
But the same tools can magnify weak controls. A broker using AI to generate marketing copy may create financial promotions that sound fluent but omit important limitations. A lender using AI to summarise affordability evidence may miss the reason a case should be referred. An AR principal using AI to monitor representatives may build a false sense of comfort if the model cannot see the real sales journey.
What boards should ask now
The FCA's operational resilience page now contains a frontier AI section, signposting the May 2026 joint statement and wider guidance from UK cyber and industry bodies. One notable theme is board-level commitment and a whole-firm approach. That is sensible because AI risk does not sit tidily in one department. It touches technology, data protection, cyber security, financial promotions, conduct, model risk, outsourcing, operations, complaints and senior manager accountability.
Boards and senior managers do not need to become prompt engineers. They do need enough understanding to ask hard, practical questions. Where are we using AI today, including pilots and vendor tools? Which uses touch customers, regulated activities, important business services or sensitive data? Which systems have tool access, internet access, production access or the ability to trigger actions? Which suppliers are involved, and what contractual rights do we have to inspect, test, audit, suspend and exit?
Containment is the new common sense
The most useful lesson from the recent AI testing incidents is that boundaries matter. If a model is told it is in a simulation but the environment permits contact with real systems, the written instruction is not the control. The control is the technical boundary, the monitoring, the permission structure and the stop mechanism.
For AI operational resilience, firms should be thinking in layers. Limit what the model can access. Limit what it can do. Separate test and production environments. Use least-privilege permissions. Keep human approval at the point where customer impact, external communication, money movement, eligibility assessment or security action could occur. Monitor for unexpected behaviour. Record prompts, outputs and tool calls where proportionate and lawful. Test the system against misuse, not just against happy-path productivity claims.
From AI policy to evidence
The UK Government's Financial Services AI Adoption Plan, published in July 2026, is optimistic about the benefits of AI. It says AI is already reshaping financial services, improving fraud detection, streamlining operations and sharpening risk management. It also describes safe, responsible adoption as part of the UK's opportunity.
That balance is important. The answer to frontier AI risk is not paralysis. Financial services firms that refuse to learn how AI works may end up with a different risk: shadow AI, unmanaged vendor adoption and controls that look pious but are ignored because they make useful work impossible. The better answer is controlled adoption with evidence.
A firm should be able to show what AI tools it uses, why, who owns them, what risk assessment was performed, what data is processed, what customer or operational impact could follow, what controls are in place, how performance is monitored, how incidents are escalated, and how the firm would exit or disable the tool.
The practical next step
For lenders, brokers and regulated distributors, the immediate task is not to write a grand AI manifesto. It is to do a disciplined stocktake. Identify every AI tool in use, including paid platforms, embedded vendor features, browser assistants, call transcription tools, marketing tools, compliance tools and cyber tools. Then classify them by risk.
Then map dependencies. Which third parties are involved? Which systems can the tool reach? Does it have internet access? Can it write, delete, send, approve, escalate, suppress or modify anything? Are logs retained? Are staff using personal accounts? Are prompts and outputs being used to train external models? Has data protection been considered properly?
The firms most likely to handle AI well will be the firms that know where AI is being used, understand which services and customers could be affected, can explain their supplier dependencies, and have controls that still work when the technology behaves in surprising ways.
That is where compliance earns its keep. Not by standing in the doorway shouting no, and not by waving through every experiment because the sales deck says productivity. The job is to help the firm use new tools without losing sight of old obligations: governance, accountability, customer outcomes, financial promotion discipline, cyber security, operational resilience and honest evidence.
Authorised Compliance Ltd helps firms turn that kind of reminder into practical action: registers, governance, monitoring, third-party review, board reporting and evidence that can survive scrutiny. The frontier has moved. The compliance work should move with it.
Sources
- Anthropic: Investigating three real-world incidents in our cybersecurity evaluations
- Bank of England, FCA and HM Treasury: joint statement on frontier AI models and cyber resilience
- FCA: Operational resilience
- FCA: Outsourcing and operational resilience
- HM Government: Financial Services AI Adoption Plan

