Regulatory
Banking and insurance, taken separately where they diverge. For each: what the rule requires of an agentic system, where programmes lose against it, what our method does about it, and where our own evidence stops.
Every UK authorised firm. Conduct sits with the FCA. Prudential soundness for banks, building societies, insurers and major investment firms sits with the PRA. Dual-regulated firms answer to both, and the two ask different questions of the same agent.
- What it requires of an agent
- There is no separate AI rulebook, and both regulators have said they do not intend to write one. Existing obligations apply in full: senior management accountability under SM&CR, the systems and controls rules, Consumer Duty, model risk management and operational resilience. In supervisory terms an agentic system is not a new category of thing to be permitted, it is a new way of breaching rules that already bind you, and the questions in the room are who was accountable, what testing was done, what the system actually decided, and where the record is.
- Where programmes fall down
- Firms look for the AI rule, do not find one, and treat the space as unregulated until a supervisor asks something they cannot answer from artefacts. The opposite failure is just as common and costs more: a programme so hedged that it stays stuck in pilot for two years, carrying cost, learning nothing, and leaving the firm no better able to answer the same questions.
- What our method does about it
- We treat the supervisory question as a design input. The Agent-Readiness Audit produces a decision inventory saying which decisions an agent may take and which need a human, an autonomy boundary agreed with the second line, and an evidence trail that falls out of running the system rather than being assembled afterwards. That artefact set is close to what a supervisor, an internal auditor or a section 166 skilled person would ask for, which is deliberate.
Where our evidence stops: Tenhaw is an agentic AI consultancy and delivery partner. We hold no regulatory permission, we do not give regulatory advice, and we do not sign anything off. Interpretation stays with your risk, compliance and legal functions. What we change is that they design with us rather than review after us.
FCA-regulated firms across the retail distribution chain, manufacturers and distributors alike. In force for open products since July 2023 and for closed products since July 2024.
- What it requires of an agent
- The Duty attaches to outcomes, not to mechanisms, so it draws no distinction between a decision made by a person, a rules engine or an agent. Firms must act to deliver good outcomes across products and services, price and value, consumer understanding and consumer support; must act in good faith, avoid foreseeable harm and support customers in pursuing their financial objectives; must monitor those outcomes and evidence them, including for customers with characteristics of vulnerability; and must report on that annually at board level.
- Where programmes fall down
- Three, repeatedly. An agent drafts customer communications and nobody can evidence they met the consumer understanding outcome, because testing measured accuracy rather than comprehension. Vulnerability signals get flattened by a triage step tuned for handling time. And the outcome data the Duty requires is never emitted by the agent, because monitoring was scoped as a reporting workstream for later, leaving a firm with a system shaping customer outcomes and no evidence of what it did.
- What our method does about it
- For any workflow touching a customer outcome we require the outcome measure to be defined before the build and emitted by the system as it runs, rather than reconstructed from logs a quarter later. Vulnerability handling is an explicit route to a human, not a confidence threshold, because a confidence score describes the model's certainty and not the customer's circumstances. And we sequence customer-facing decisioning after internal work, because the evidence base you will need to defend it is cheaper to build on workflows that cannot create foreseeable harm while you are learning.
Where our evidence stops: We have not delivered an agentic system into a customer-facing journey in an FCA-regulated firm. The HSBC voice-insights work was a proof of concept over enterprise voice data, and our live insurance engagement is not customer-facing decisioning. If you need someone to attest that a design satisfies the Duty, you need your compliance function or a firm that carries that liability.
The PRA's model risk management principles, effective from 17 May 2024, for UK banks, building societies and PRA-designated investment firms with internal model permissions. Insurers sit outside the formal scope, and the PRA has been clear it regards the principles as good practice more widely, which is why insurance risk functions raise it anyway.
- What it requires of an agent
- Five principles: model identification and risk classification, governance, development and implementation and use, independent validation, and mitigants where models are third-party or known to be deficient. The part that bites for agents is the breadth of the definition. A model is a quantitative method turning input into output for use in a decision, which captures a language model in a decision path without needing a special case. An agentic workflow is usually several models, plus prompts, retrieval corpora and tool permissions that all change behaviour, and almost nobody versions those the way they version models.
- Where programmes fall down
- The model inventory has no row for the agent, because the agent was procured as a tool. Prompt changes are treated as configuration, so a change that materially alters output never triggers review. Validation methods built for a fixed model cannot answer the first question an agentic system raises, which is whether this is even the same model as last month. And the third-party principle lands hardest on foundation models, where the firm cannot inspect the thing it is being asked to get comfortable with.
- What our method does about it
- The audit starts from a model and decision inventory that treats prompts, retrieval corpora, tool permissions and model versions as versioned artefacts with named owners, and we get the risk classification agreed with the second line before anything is built, because classification decides how much validation the system needs and therefore what it costs to run. Then we design for validation: reproducible evaluation sets, recorded provenance, and change control that fires on a prompt or corpus change rather than only on a model upgrade.
Where our evidence stops: Independent validation has to be independent, which by definition excludes the team that built the system. We build so your validation function or a specialist can do their job. We do not perform model validation and we claim no capability in it.
UK banks, building societies, insurers, payment and e-money institutions and designated investment firms. Firms have had to be able to remain within their impact tolerances since 31 March 2025, and the critical third parties regime, live since the start of 2025, extends the regulators' reach to designated providers underneath them.
- What it requires of an agent
- Identify your important business services. Set an impact tolerance for the maximum tolerable disruption to each. Map the people, processes, technology, facilities, information and third parties that support them. Test to tolerance under severe but plausible scenarios, and stay inside it.
- Where programmes fall down
- Agents get inside the mapped chain of an important business service without appearing on the map, because the mapping was done before them and is refreshed annually. The testing is also the wrong shape: resilience exercises are built around outage, and an agent's characteristic failure is silent degradation, where output keeps arriving and quietly gets worse. Worst of all, the manual fallback that the impact tolerance quietly assumes has been decommissioned, or has atrophied because the team that used to do the work is now half the size.
- What our method does about it
- For any workflow inside an important business service we require a fallback that is exercised rather than documented, degradation monitoring alongside availability monitoring, and the substitution question answered: if this agent stops today, who does the work, for how long can they sustain it, and does that fit inside the tolerance. An agent whose fallback is that people go back to doing it by hand is exactly as resilient as the number of people who still remember how.
Where our evidence stops: Impact tolerance setting and scenario testing belong to your resilience and risk functions, and carry a mandate we do not have. We design so an agent is testable and its degradation is visible, and we will tell you when an important business service is the wrong place to start.
An EU regulation applying since 17 January 2025 to EU financial entities. It reaches UK firms through EU subsidiaries and branches, through group functions serving them, and through supplying ICT services to EU financial entities. The UK has no DORA. It has the operational resilience regime and the critical third parties regime, which ask similar questions in different words.
- What it requires of an agent
- ICT risk management, incident classification and reporting on tight clocks, resilience testing including threat-led penetration testing for larger entities, and ICT third-party risk management: a register of information covering contractual arrangements, mandatory contract terms including audit and access rights, conditions on subcontracting, and documented exit strategies, with a direct oversight regime for designated critical ICT third-party providers. AI suppliers are in scope where they provide ICT services supporting a financial function, and that includes the model providers underneath your platform, not only the vendor whose name is on the invoice.
- Where programmes fall down
- The AI supplier was bought as a tool rather than onboarded as an ICT third-party service, so it never entered the register, the contract carries none of the required terms, and nobody has answered the exit question. Exit is the one that changes architecture: if the provider is unavailable, changes its terms, or a supervisor tells you to move, what breaks. A programme built around one provider's proprietary features has answered that by accident, and badly.
- What our method does about it
- The audit maps where the supply chain actually goes, subprocessors included, and we ask the exit question at design time because it determines how much of the system can stay provider-neutral. In practice that means abstracting the model interface, holding prompts, evaluation sets and retrieval corpora as your assets rather than a vendor's, and being explicit about which capabilities are genuinely provider-specific and what they cost you in concentration risk. The same discipline answers the EU AI Act for anyone placing systems on the EU market, and the clock there moved: the AI omnibus agreed in 2026 pushed the high-risk obligations back to 2 December 2027 for stand-alone systems and 2 August 2028 for AI embedded in regulated products. That is more time and the same evidence, which is only useful to a firm that starts producing it during the build.
Where our evidence stops: We do not draft your DORA contract terms, maintain your register of information, or perform threat-led penetration testing. Those belong to legal, vendor risk and specialist testers. On our own side of the relationship, what Tenhaw holds as a supplier is published on the security page, including what is certified and what is still in progress.
UK insurers and reinsurers under the PRA, and EU entities under Solvency II proper. The UK reforms are branded Solvency UK, and the governance and data requirements that bear on AI were not loosened by them.
- What it requires of an agent
- A system of governance with four effective key functions, risk management, compliance, internal audit and actuarial. An ORSA that reflects the firm's real risk profile rather than last year's. Data used for technical provisions that is accurate, complete and appropriate, with the actuarial function accountable for saying so. Internal model firms carry a model change policy and validation on top. And using a supplier for a critical or important operational function is outsourcing, with the notification, contractual and oversight duties that follow.
- Where programmes fall down
- An underwriting-support agent enriches submission data, the enrichment becomes an input to pricing, and nobody attested to its quality because everyone involved thought of it as a productivity tool. Provenance is the recurring gap: a number in a risk file that cannot be traced to a source is a data quality problem the actuarial function inherits long after the pilot team has moved on. Reserving is where it surfaces, because reserving is where data quality gets examined hardest and every year.
- What our method does about it
- We treat provenance as an output of the system rather than as documentation about it. On our live specialty insurance engagement the extraction pipeline scores confidence from the provenance of each enrichment source alongside model certainty and a search-based cross-check, so human review is routed by confidence and consequence, not worked as a queue, and each field traces back to the document or API it came from. That is the property an actuarial function needs, and it is far cheaper to build in than to retrofit.
Where our evidence stops: That pipeline is a proof of concept feeding business intelligence. It is not a rated pricing model and not an input to technical provisions. Tenhaw holds no actuarial capability. Anything touching technical provisions or an internal model needs your actuarial and validation functions in the design from the first week.
The London market: managing agents and syndicates under Lloyd's oversight as well as PRA and FCA regulation, and the coverholders, MGAs and brokers who underwrite or place business under delegated authority.
- What it requires of an agent
- Underwriting authority here is delegated by contract. A binding authority sets what may be written, within what limits and on whose paper, with the managing agent accountable for the performance and conduct of business written under it and Lloyd's own standards behind that. Coverholder audits ask whether what was written matched what was permitted. Delegation to a person is a well-understood arrangement. Delegation exercised partly by a machine is not, and the first question is whether the binder contemplates it at all.
- Where programmes fall down
- An agent inside an MGA's submission pipeline starts influencing risk selection, which is precisely what the binder governs, without the binder being revisited and without the managing agent knowing. Then an audit asks how a particular risk came to be accepted, and the honest answer is a prompt nobody kept and a model version nobody recorded. Broker workflows carry a quieter version of the same problem, where the record of what was disclosed to whom becomes partly machine-generated and nobody decided that it would be.
- What our method does about it
- We read the binder as a design constraint, the same way we read a peak trading freeze in retail. The autonomy boundary is drawn inside what the delegated authority actually permits, the record of why a risk was routed, flagged or deprioritised is retained as part of the workflow rather than as logs with a thirty-day retention, and where the binder does not contemplate machine involvement, the conversation with the managing agent comes before the build.
Where our evidence stops: We have run a month-one audit and a two-week proof of concept inside a London specialty insurance business, with month three standing up a team to productionise it. We have not taken an agentic system through a coverholder audit, and we would be sceptical of anyone claiming that yet.
Tenhaw builds agentic systems and the operating models around them, and works alongside the risk, compliance, legal and actuarial functions who own the interpretation of these regimes. Our security and assurance position states what we hold today and what is still in progress.