Sector ยท Financial Services

Agentic transformation in Financial Services

Agentic transformation inside the constraints that actually bind.

teams in the HSBC operating model we designed and piloted
500
hours/year of admin the voice-insights proof of concept was projected to remove
1.5M+
to a working proof of concept in specialty insurance, on work that had run a year
2 weeks
Work delivered at
HSBCLondon specialty insurance market

Live now, and unnamed at the client's request.

The short answer

Agentic transformation in UK financial services fails on governance far more often than on technology. The FCA and the PRA have written no separate AI rulebook and have said they do not intend to, so the obligations you already answer to apply to agents in full. They decide which workflows can move to agents at all, and in what order.

Banking

SM&CR keeps accountability with a named individual, and it cannot be discharged onto a system. SS1/23 model risk management captures a language model in a decision path through a deliberately broad definition of a model. Consumer Duty attaches to outcomes rather than mechanisms, so it draws no distinction between a person, a rules engine and an agent. Operational resilience asks whether an agent has appeared inside an important business service, and DORA reaches firms with EU exposure. Tenhaw's founder has worked inside HSBC across two engagements: an executive delivery governance role spanning 150+ global teams and a $102M budget, and leading the proof of concept for an AI voice-insights platform.

Insurance

Solvency II and Solvency UK govern insurers on top, through the system of governance and through the requirement that data used for technical provisions is accurate, complete and appropriate, with the actuarial function accountable for saying so. Underwriting, claims and reserving carry different data and different regulatory weight, and an answer there has to satisfy an actuarial function as well as a product owner. Where authority is delegated to brokers, MGAs or Lloyd's coverholders, a binding authority draws the boundary of a decision before any agent does. Tenhaw's current agentic engagement is with a London specialty insurance business, unnamed at the client's request.

Tenhaw designs agentic operating models where the audit trail and the human-in-the-loop points are part of the design rather than a retrofit after an incident. Tenhaw is an AI delivery partner, not a compliance consultancy: the method produces the evidence your second line needs, it does not replace them.

Regulation in financial services

The constraints that decide the sequence

Not the ones that sound good in a deck. These determine which workflows can move to agents at all, and in what order.

01

Accountability has to resolve to a person

Under SM&CR a named individual is accountable for outcomes, and 'the agent decided' is not a defence a regulator accepts. The operating model has to map accountability for agent decisions onto real people with real authority, before anything is deployed.

02

Model risk governance was not built for this

Existing model risk frameworks assume a model that is validated, versioned and periodically reviewed. Agentic systems compose models at runtime and behave differently week to week. SS1/23 already captures them through a deliberately broad definition of a model, so the framework needs extending on purpose rather than pretending the old one covers it.

03

The data that matters is the data you cannot move

The highest-value agentic workflows sit on customer and transaction data with the strictest residency and access constraints. Sequencing matters enormously: the workflows that are easiest to automate are frequently the least valuable, and vice versa.

04

Insurance is not one problem with one answer

Underwriting submission triage, claims handling and reserving have different data, different regulatory weight and different appetites for autonomy. Where authority is delegated to brokers, MGAs or Lloyd's coverholders, a binding authority agreement already defines what may be decided and by whom, and it was not written with a machine in mind.

05

The evidence has to satisfy people who audit for a living

Second line, internal audit, external auditors, the actuarial function, and occasionally a skilled person appointed under section 166. A programme that cannot produce its own evidence as it runs turns every one of those reviews into an archaeology exercise, months after the people who made the decisions have moved on.

06

Change fatigue is real and earned

Most large banks and insurers have run continuous transformation for a decade. Staff have watched initiatives arrive and evaporate. Adoption planning has to account for justified scepticism rather than treating it as a comms problem.

Data residency, model risk and auditability are questions about us as much as about your estate. Our security and assurance position says what we hold today and what we do not.

Sequencing

Where agents land first, and where they should not

Both halves matter. A supplier who only shows you the left-hand column is selling you the second year of the programme as though it were the first.

Start here

Where the value is real and the risk is contained.

  1. 1Operational process with high volume and cheaply reversible decisions
  2. 2Internal knowledge retrieval across fragmented policy and procedure estates
  3. 3Call and case summarisation: the workflow behind our HSBC voice-insights proof of concept
  4. 4Underwriting submission triage, turning broker documents into structured, traceable data before an underwriter sees them
  5. 5Claims document handling and triage, with settlement authority left with people
  6. 6Control testing and evidence gathering, where the audit trail is the product
  7. 7Developer and delivery workflow, where the risk surface is contained

Real constraints

The things that will bite, and worth pricing in before you sign.

  • Anything touching customer outcomes needs human-in-the-loop until the evidence base exists
  • Data residency frequently rules out the strongest available models
  • Second-line risk functions must be designers of the governance, not reviewers of it
  • Model risk classification has to be agreed before the build, because it decides how much validation the system will need and therefore what it costs
  • Anything feeding pricing, technical provisions or an internal model pulls in the actuarial function on day one, not at year end
  • Regulatory interpretation stays with your risk, compliance and legal functions: we are not a compliance consultancy
  • Procurement and supplier onboarding realistically add 8โ€“12 weeks before work starts
Regulatory

The regimes that actually gate this, and what each one asks of an agent

Banking and insurance, taken separately where they diverge. For each: what the rule requires of an agentic system, where programmes lose against it, what our method does about it, and where our own evidence stops.

The FCA and the PRA

Every UK authorised firm. Conduct sits with the FCA. Prudential soundness for banks, building societies, insurers and major investment firms sits with the PRA. Dual-regulated firms answer to both, and the two ask different questions of the same agent.

What it requires of an agent
There is no separate AI rulebook, and both regulators have said they do not intend to write one. Existing obligations apply in full: senior management accountability under SM&CR, the systems and controls rules, Consumer Duty, model risk management and operational resilience. In supervisory terms an agentic system is not a new category of thing to be permitted, it is a new way of breaching rules that already bind you, and the questions in the room are who was accountable, what testing was done, what the system actually decided, and where the record is.
Where programmes fall down
Firms look for the AI rule, do not find one, and treat the space as unregulated until a supervisor asks something they cannot answer from artefacts. The opposite failure is just as common and costs more: a programme so hedged that it stays stuck in pilot for two years, carrying cost, learning nothing, and leaving the firm no better able to answer the same questions.
What our method does about it
We treat the supervisory question as a design input. The Agent-Readiness Audit produces a decision inventory saying which decisions an agent may take and which need a human, an autonomy boundary agreed with the second line, and an evidence trail that falls out of running the system rather than being assembled afterwards. That artefact set is close to what a supervisor, an internal auditor or a section 166 skilled person would ask for, which is deliberate.

Where our evidence stops: Tenhaw is an agentic AI consultancy and delivery partner. We hold no regulatory permission, we do not give regulatory advice, and we do not sign anything off. Interpretation stays with your risk, compliance and legal functions. What we change is that they design with us rather than review after us.

Consumer Duty, where agents touch customer outcomes

FCA-regulated firms across the retail distribution chain, manufacturers and distributors alike. In force for open products since July 2023 and for closed products since July 2024.

What it requires of an agent
The Duty attaches to outcomes, not to mechanisms, so it draws no distinction between a decision made by a person, a rules engine or an agent. Firms must act to deliver good outcomes across products and services, price and value, consumer understanding and consumer support; must act in good faith, avoid foreseeable harm and support customers in pursuing their financial objectives; must monitor those outcomes and evidence them, including for customers with characteristics of vulnerability; and must report on that annually at board level.
Where programmes fall down
Three, repeatedly. An agent drafts customer communications and nobody can evidence they met the consumer understanding outcome, because testing measured accuracy rather than comprehension. Vulnerability signals get flattened by a triage step tuned for handling time. And the outcome data the Duty requires is never emitted by the agent, because monitoring was scoped as a reporting workstream for later, leaving a firm with a system shaping customer outcomes and no evidence of what it did.
What our method does about it
For any workflow touching a customer outcome we require the outcome measure to be defined before the build and emitted by the system as it runs, rather than reconstructed from logs a quarter later. Vulnerability handling is an explicit route to a human, not a confidence threshold, because a confidence score describes the model's certainty and not the customer's circumstances. And we sequence customer-facing decisioning after internal work, because the evidence base you will need to defend it is cheaper to build on workflows that cannot create foreseeable harm while you are learning.

Where our evidence stops: We have not delivered an agentic system into a customer-facing journey in an FCA-regulated firm. The HSBC voice-insights work was a proof of concept over enterprise voice data, and our live insurance engagement is not customer-facing decisioning. If you need someone to attest that a design satisfies the Duty, you need your compliance function or a firm that carries that liability.

SS1/23 model risk management

The PRA's model risk management principles, effective from 17 May 2024, for UK banks, building societies and PRA-designated investment firms with internal model permissions. Insurers sit outside the formal scope, and the PRA has been clear it regards the principles as good practice more widely, which is why insurance risk functions raise it anyway.

What it requires of an agent
Five principles: model identification and risk classification, governance, development and implementation and use, independent validation, and mitigants where models are third-party or known to be deficient. The part that bites for agents is the breadth of the definition. A model is a quantitative method turning input into output for use in a decision, which captures a language model in a decision path without needing a special case. An agentic workflow is usually several models, plus prompts, retrieval corpora and tool permissions that all change behaviour, and almost nobody versions those the way they version models.
Where programmes fall down
The model inventory has no row for the agent, because the agent was procured as a tool. Prompt changes are treated as configuration, so a change that materially alters output never triggers review. Validation methods built for a fixed model cannot answer the first question an agentic system raises, which is whether this is even the same model as last month. And the third-party principle lands hardest on foundation models, where the firm cannot inspect the thing it is being asked to get comfortable with.
What our method does about it
The audit starts from a model and decision inventory that treats prompts, retrieval corpora, tool permissions and model versions as versioned artefacts with named owners, and we get the risk classification agreed with the second line before anything is built, because classification decides how much validation the system needs and therefore what it costs to run. Then we design for validation: reproducible evaluation sets, recorded provenance, and change control that fires on a prompt or corpus change rather than only on a model upgrade.

Where our evidence stops: Independent validation has to be independent, which by definition excludes the team that built the system. We build so your validation function or a specialist can do their job. We do not perform model validation and we claim no capability in it.

Operational resilience and important business services

UK banks, building societies, insurers, payment and e-money institutions and designated investment firms. Firms have had to be able to remain within their impact tolerances since 31 March 2025, and the critical third parties regime, live since the start of 2025, extends the regulators' reach to designated providers underneath them.

What it requires of an agent
Identify your important business services. Set an impact tolerance for the maximum tolerable disruption to each. Map the people, processes, technology, facilities, information and third parties that support them. Test to tolerance under severe but plausible scenarios, and stay inside it.
Where programmes fall down
Agents get inside the mapped chain of an important business service without appearing on the map, because the mapping was done before them and is refreshed annually. The testing is also the wrong shape: resilience exercises are built around outage, and an agent's characteristic failure is silent degradation, where output keeps arriving and quietly gets worse. Worst of all, the manual fallback that the impact tolerance quietly assumes has been decommissioned, or has atrophied because the team that used to do the work is now half the size.
What our method does about it
For any workflow inside an important business service we require a fallback that is exercised rather than documented, degradation monitoring alongside availability monitoring, and the substitution question answered: if this agent stops today, who does the work, for how long can they sustain it, and does that fit inside the tolerance. An agent whose fallback is that people go back to doing it by hand is exactly as resilient as the number of people who still remember how.

Where our evidence stops: Impact tolerance setting and scenario testing belong to your resilience and risk functions, and carry a mandate we do not have. We design so an agent is testable and its degradation is visible, and we will tell you when an important business service is the wrong place to start.

DORA, for firms in scope

An EU regulation applying since 17 January 2025 to EU financial entities. It reaches UK firms through EU subsidiaries and branches, through group functions serving them, and through supplying ICT services to EU financial entities. The UK has no DORA. It has the operational resilience regime and the critical third parties regime, which ask similar questions in different words.

What it requires of an agent
ICT risk management, incident classification and reporting on tight clocks, resilience testing including threat-led penetration testing for larger entities, and ICT third-party risk management: a register of information covering contractual arrangements, mandatory contract terms including audit and access rights, conditions on subcontracting, and documented exit strategies, with a direct oversight regime for designated critical ICT third-party providers. AI suppliers are in scope where they provide ICT services supporting a financial function, and that includes the model providers underneath your platform, not only the vendor whose name is on the invoice.
Where programmes fall down
The AI supplier was bought as a tool rather than onboarded as an ICT third-party service, so it never entered the register, the contract carries none of the required terms, and nobody has answered the exit question. Exit is the one that changes architecture: if the provider is unavailable, changes its terms, or a supervisor tells you to move, what breaks. A programme built around one provider's proprietary features has answered that by accident, and badly.
What our method does about it
The audit maps where the supply chain actually goes, subprocessors included, and we ask the exit question at design time because it determines how much of the system can stay provider-neutral. In practice that means abstracting the model interface, holding prompts, evaluation sets and retrieval corpora as your assets rather than a vendor's, and being explicit about which capabilities are genuinely provider-specific and what they cost you in concentration risk. The same discipline answers the EU AI Act for anyone placing systems on the EU market, and the clock there moved: the AI omnibus agreed in 2026 pushed the high-risk obligations back to 2 December 2027 for stand-alone systems and 2 August 2028 for AI embedded in regulated products. That is more time and the same evidence, which is only useful to a firm that starts producing it during the build.

Where our evidence stops: We do not draft your DORA contract terms, maintain your register of information, or perform threat-led penetration testing. Those belong to legal, vendor risk and specialist testers. On our own side of the relationship, what Tenhaw holds as a supplier is published on the security page, including what is certified and what is still in progress.

Solvency II and Solvency UK, for insurers

UK insurers and reinsurers under the PRA, and EU entities under Solvency II proper. The UK reforms are branded Solvency UK, and the governance and data requirements that bear on AI were not loosened by them.

What it requires of an agent
A system of governance with four effective key functions, risk management, compliance, internal audit and actuarial. An ORSA that reflects the firm's real risk profile rather than last year's. Data used for technical provisions that is accurate, complete and appropriate, with the actuarial function accountable for saying so. Internal model firms carry a model change policy and validation on top. And using a supplier for a critical or important operational function is outsourcing, with the notification, contractual and oversight duties that follow.
Where programmes fall down
An underwriting-support agent enriches submission data, the enrichment becomes an input to pricing, and nobody attested to its quality because everyone involved thought of it as a productivity tool. Provenance is the recurring gap: a number in a risk file that cannot be traced to a source is a data quality problem the actuarial function inherits long after the pilot team has moved on. Reserving is where it surfaces, because reserving is where data quality gets examined hardest and every year.
What our method does about it
We treat provenance as an output of the system rather than as documentation about it. On our live specialty insurance engagement the extraction pipeline scores confidence from the provenance of each enrichment source alongside model certainty and a search-based cross-check, so human review is routed by confidence and consequence, not worked as a queue, and each field traces back to the document or API it came from. That is the property an actuarial function needs, and it is far cheaper to build in than to retrofit.

Where our evidence stops: That pipeline is a proof of concept feeding business intelligence. It is not a rated pricing model and not an input to technical provisions. Tenhaw holds no actuarial capability. Anything touching technical provisions or an internal model needs your actuarial and validation functions in the design from the first week.

Lloyd's, delegated authority, brokers and MGAs

The London market: managing agents and syndicates under Lloyd's oversight as well as PRA and FCA regulation, and the coverholders, MGAs and brokers who underwrite or place business under delegated authority.

What it requires of an agent
Underwriting authority here is delegated by contract. A binding authority sets what may be written, within what limits and on whose paper, with the managing agent accountable for the performance and conduct of business written under it and Lloyd's own standards behind that. Coverholder audits ask whether what was written matched what was permitted. Delegation to a person is a well-understood arrangement. Delegation exercised partly by a machine is not, and the first question is whether the binder contemplates it at all.
Where programmes fall down
An agent inside an MGA's submission pipeline starts influencing risk selection, which is precisely what the binder governs, without the binder being revisited and without the managing agent knowing. Then an audit asks how a particular risk came to be accepted, and the honest answer is a prompt nobody kept and a model version nobody recorded. Broker workflows carry a quieter version of the same problem, where the record of what was disclosed to whom becomes partly machine-generated and nobody decided that it would be.
What our method does about it
We read the binder as a design constraint, the same way we read a peak trading freeze in retail. The autonomy boundary is drawn inside what the delegated authority actually permits, the record of why a risk was routed, flagged or deprioritised is retained as part of the workflow rather than as logs with a thirty-day retention, and where the binder does not contemplate machine involvement, the conversation with the managing agent comes before the build.

Where our evidence stops: We have run a month-one audit and a two-week proof of concept inside a London specialty insurance business, with month three standing up a team to productionise it. We have not taken an agentic system through a coverholder audit, and we would be sceptical of anyone claiming that yet.

Tenhaw builds agentic systems and the operating models around them, and works alongside the risk, compliance, legal and actuarial functions who own the interpretation of these regimes. Our security and assurance position states what we hold today and what is still in progress.

Evidence

What we have done here

Named clients where we have permission to name them, and the evidence basis stated on every one.

Live now ยท specialty insurance3 months, ongoing

Roughly a year of stalled work, rebuilt as a working proof of concept in two weeks

If you are an insurer rather than a bank, this is the closest evidence we own, and it is running now.

Read the engagement
12 months โ†’ 2 weeks
prior build effort rebuilt as a working proof of concept

What our financial services evidence is, and what it is not

Our named work is in banking, at HSBC: an operating model designed, piloted and validated with global rollout scheduled for 2026, and a proof of concept for an AI voice-insights platform. Neither is a production agentic rollout in a bank. Our live agentic engagement is with a London specialty insurance business, confidential at the client's request. If you are an insurer, a broker or an MGA, ask on the call which of the banking work transfers to your regime and which does not.

Financial Services: your questions

How do banks govern AI agent decisions?

By mapping accountability for each agent decision onto a named individual with matching authority, defining explicit human-in-the-loop points for consequential or irreversible decisions, and extending model risk governance to cover systems that compose models at runtime. Under SM&CR the accountability cannot rest with the system, so the operating model has to resolve it to people before deployment.

Does Consumer Duty apply to decisions made by AI agents?

Yes, if your firm is in scope of the Duty. It attaches to outcomes, not to mechanisms, so it makes no distinction between a decision made by a person, a rules engine or an agent. The four outcomes still have to be delivered and evidenced, products and services, price and value, consumer understanding and consumer support, alongside the cross-cutting obligations to act in good faith, avoid foreseeable harm and support customers in pursuing their financial objectives. The design consequences are concrete: the outcome measure has to be emitted by the system as it runs rather than reconstructed later, and vulnerability handling has to be an explicit route to a human, because a confidence score describes the model's certainty and not the customer's circumstances. Tenhaw has not delivered an agentic system into a customer-facing journey in an FCA-regulated firm.

What does the FCA expect when an AI agent makes a customer-facing decision?

The FCA has not published an AI rulebook and has said it does not intend to, so what it expects is what it already expects. A named senior manager accountable under SM&CR. Governance and controls proportionate to the risk. Evidence that the Consumer Duty outcomes are being delivered and monitored, including for customers with characteristics of vulnerability. And the ability to explain a decision to the customer who received it and to a supervisor afterwards. The practical test is whether you can answer who was accountable, what testing was done, what the system actually decided and where the record is, from artefacts the programme produced anyway rather than from an archaeology exercise months later.

How does SS1/23 apply to agentic systems?

SS1/23 took effect on 17 May 2024 for UK banks, building societies and PRA-designated investment firms with internal model permissions, and it uses a deliberately broad definition of a model: a quantitative method turning input into output for use in a decision. A language model inside a decision path meets that definition without needing a special case. An agentic workflow is usually several models plus prompts, retrieval corpora and tool permissions, all of which change behaviour, so the practical work is treating those as versioned artefacts with named owners, entering the system on the model inventory, agreeing its risk classification with the second line before you build, and designing so independent validation is possible: reproducible evaluation sets, recorded provenance, and change control that fires on a prompt change rather than only on a model upgrade. Insurers are outside the formal scope and raise it anyway, because the PRA treats the principles as good practice more widely.

Does DORA cover AI suppliers?

Yes, where the supplier provides ICT services supporting a financial function of an entity in scope, and DORA has applied since 17 January 2025. That brings the AI vendor, and usually the model providers underneath it, into ICT third-party risk management: an entry on the register of information, contract terms covering audit and access rights, conditions on subcontracting, and a documented exit strategy, with a direct oversight regime for designated critical providers on top. UK-only firms are not in scope of DORA itself, and meet similar questions through the operational resilience regime and the critical third parties regime. The question that actually changes an architecture is exit: if the provider is unavailable, changes its terms, or a supervisor tells you to move, what breaks. Programmes built tightly around one provider's proprietary features have answered that by accident, and badly.

What does Solvency II require of AI in underwriting?

It does not name AI, and it binds it anyway, through governance and through data. The system of governance requires the four key functions to be effective, the ORSA has to reflect the firm's real risk profile, and data used for technical provisions must be accurate, complete and appropriate, with the actuarial function accountable for saying so. If an agent enriches submission data and that enrichment reaches pricing or reserving, someone has to be able to trace every field back to its source, which makes provenance an engineering requirement rather than a documentation exercise. Internal model firms add a model change policy and validation. And using a supplier for a critical or important operational function is outsourcing, with the notification and contractual duties that follow. Tenhaw holds no actuarial capability: our specialty insurance work is a proof of concept feeding business intelligence, not a rated pricing model.

Do you work with insurers, or only banks?

Both. Our named financial services work is banking, at HSBC. Our current agentic engagement is with a London specialty insurance business, confidential at the client's request: a month-one audit, then a two-week proof of concept turning PDFs into business intelligence on Azure, covering ground the business had circled for roughly a year, with month three standing up a team to productionise it. Insurance differs from banking in ways that matter to the design. Underwriting, claims and reserving carry different data and different regulatory weight, Solvency II and the actuarial function govern where SS1/23 would in a bank, and delegated authority through brokers, MGAs and Lloyd's coverholders means a binding authority draws the boundary of a decision before any agent does. Ask on the call which of the banking work transfers to your regime and which does not.

Where should a bank start with agentic AI?

With high-volume workflows where decisions are observable and cheaply reversible, and where the audit trail is naturally part of the output: control testing, internal knowledge retrieval, and case and call summarisation. Customer-facing decisioning should come later, once the governance evidence base exists.

Has Tenhaw delivered AI in a regulated bank?

Yes, at HSBC, as pilot and proof of concept rather than production. Tenhaw's founder led the proof of concept for an AI Voice Insights platform, applying natural language processing, sentiment analysis and entity recognition to enterprise voice data, projected to remove 1.5M+ hours of manual administration annually. That figure was a projection from a proof of concept, not a measured result from a production rollout. Separately, the operating-model work across HSBC's Global Payment Solutions division was designed, piloted and validated, with global rollout scheduled for 2026. Both are written up in full on the case studies page.

How does SM&CR affect AI agent deployment?

It requires a named senior manager to be accountable for the outcomes of the function, including those produced by agents. In practice this means the operating model must specify which decisions agents may take autonomously, which require human approval, and who holds accountability at each point, documented before deployment rather than reconstructed after an incident.

Can AI agents do KYC and customer onboarding?

They can do the document and evidence layer, which is where the elapsed time sits, and they should not take the decision. What an agent handles well is reading incorporation documents, structure charts and identity evidence into structured fields with provenance recorded per field, resolving entities across registries and third-party sources, assembling the file, and saying what is missing. What it must not do is set the risk rating, clear a politically exposed person or close an alert, because customer due diligence under the Money Laundering Regulations 2017 is a decision the firm has to defend to its supervisor, and enhanced due diligence exists precisely for the cases where the machine-legible answer is least reliable. The design question is not accuracy, it is what the exception path costs: route by confidence and by consequence, score confidence from where each value came from rather than from the model's own certainty, and the accuracy target falls out of it. Tenhaw has built this pipeline shape on a live insurance engagement as a proof of concept, and has never taken one into production.

Where do AI agents help with anti-money-laundering and transaction monitoring?

In alert triage and narrative assembly, which is the part everyone under-resources, and not in the disposition itself. An agent can pull the customer history, the prior alerts, the counterparty context and the relevant documents into one place, draft the investigation narrative with every claim linked to its source, and order the queue by consequence rather than by arrival time. That is analyst minutes per alert on a volume where minutes are the entire budget. What an agent must not do is close an alert, suppress one, or decide not to file, and it must never be allowed to tune the alert threshold: a system optimised to reduce alert volume has learned exactly the wrong objective and will be extremely good at it. For a bank, SS1/23 already reaches the monitoring models, so an agent in the disposition path joins the model inventory. Fraud disputes follow the same split: an agent assembles the evidence pack and drafts the communication, and a person declines the payment, because a false positive there is a customer locked out of their money, which is a Consumer Duty question before it is an accuracy question.

Can an AI agent handle an insurance claim?

It can run intake, triage and the document work that spans the file, and it must not settle anything. Reading the notification and the evidence into structured fields, checking completeness against what the policy actually requires, routing by complexity and consequence, drafting the chronology and keeping the customer communication current are all document and coordination work. Deciding coverage, setting or moving a reserve, declining a claim and authorising a payment are not. Three things decide whether it works. Cycle time is made of waiting rather than of handling, so assisting each step leaves the end-to-end number almost unchanged and you have to map the waits first. Vulnerability has to be an explicit route to a person defined by circumstance, not a confidence threshold, because a confidence score describes the model's certainty and not the customer's situation. And reserving feeds technical provisions, so any field an agent extracted that reaches a reserve is now data the actuarial function is accountable for, which makes provenance an engineering output, not a documentation exercise. Tenhaw has not delivered an agentic system into claims handling.

Can AI agents make credit decisions?

Not the decision, and this is the workflow where the law is most direct about it. Section 80 of the Data (Use and Access) Act 2025 replaced Article 22 of the UK GDPR with Articles 22A to 22D, which turn on whether there is meaningful human involvement in a significant decision about a person, and a credit refusal is the textbook example of one. Consumer Duty adds price and value and consumer understanding on top, and a decision you cannot give the customer a reason for is not one you can defend. Where agents do earn their place is everything around the decision: assembling the application file, reading bank statements and accounts into structured data with provenance, drafting the credit paper and surfacing the inconsistencies a human would want to ask about. In commercial lending that file assembly is most of the elapsed time and almost none of the judgement. For a bank, an agent in that path also meets SS1/23, because a quantitative method turning input into output for use in a decision is a model under its definition whether or not anyone called it one.

Can agents handle complaints, and what does Consumer Duty require?

Agents belong in investigation support and root cause, not in the outcome. An agent can assemble the full relationship history into a chronology, retrieve the terms in force on the relevant date, not the current version, draft the response for a person to own, and cluster complaints across the book so the same root cause is not rediscovered five times by five handlers. That clustering is the output most worth showing your second line early, because it is evidence of outcome monitoring, not a productivity claim. What an agent must not do is decide the outcome or send a final response unreviewed: a complaint is a customer disputing the firm's judgement, and having the machine decide it is marking your own homework at the moment it matters most. The FCA's complaints rules in DISP govern the handling and the Financial Ombudsman Service sits behind it, and the Duty expects the outcome, and not only the handling time, to be monitored and evidenced, which means the outcome measure has to be emitted by the system as it runs rather than reconstructed from logs a quarter later.

How does AI help with underwriting submission triage?

By turning the submission into structured, traceable data before an underwriter opens it, and by ordering the queue by appetite fit rather than by arrival. It is the one workflow on this page where our evidence is a build: on a live engagement in the London specialty insurance market we produced a working proof of concept in two weeks, PDFs in and business intelligence out on Azure, from a blank repository, over ground the business had circled for roughly a year, pair-programmed throughout with the client's own engineer. The pipeline extracts each document to a readable markdown form first so it is inspectable and re-runnable when the field list changes, then narrows to the fields that matter, normalises, enriches against third-party APIs and adds semantic grouping. Confidence comes from the provenance of each value combined with model certainty and an independent cross-check, not from the model's self-reported score, which is what makes routing explainable to the underwriter. What it does not do is decline a risk, set a price or bind, and extracted data should not reach a rating model without an explicit quality attestation, because the moment enrichment becomes an input to pricing the actuarial function inherits it. Where authority is delegated, the binding authority draws the boundary of a decision before any agent does, and that conversation with the managing agent belongs before the build. This is a proof of concept feeding business intelligence, not a production deployment.

Thirty minutes on Financial Services, with James Rooney

We'll be specific about what applies in your sector and what does not, and you'll leave with a rough scope whether you engage us or not.

30 minutesWith James personallyNo obligation

Most organisations start with a fixed-price Agent-Readiness Audit ยท ยฃ30kโ€“ยฃ90k ยท 6โ€“8 weeks