Document and voice intelligence, answered in full.
- questions in this group, each answered in full
- 36
- pages the answers are written on, every one linked
- 2
- questions across the whole FAQ
- 1424
36 questions on document and voice intelligence, answered by Tenhaw, a UK AI consultancy and AI delivery partner based in London. Nothing here is a summary: each answer is the exact text from the page that owns it, and every group links back to that page for the context around it.
On this page
Document intelligence to business intelligence
Answered on Document intelligence to business intelligence, and rendered here in the same words.
Read the page these answers live on →
How long does a document intelligence proof of concept take?
Two to four weeks for a working proof of concept against one document type and one real workflow. On a live engagement in the London insurance market, Tenhaw delivered a working proof of concept extracting information from PDFs into business intelligence on Azure in two weeks, ground the business had been circling for roughly a year. Productionising it is a separate phase, scoped at four to six weeks with a dedicated team.
Why do document extraction projects stall after the demo?
Usually for one of four reasons: accuracy on a curated sample does not survive the long tail and no exception path was designed; nobody established what a correct answer is, so every accuracy discussion becomes an argument about the benchmark; the extraction works but the BI layer has no semantic layer or governance, so the business gets confident answers from the wrong table; or human review was built as a queue rather than routed by confidence and consequence, so expert time becomes the bottleneck.
Our document AI proof of concept worked and never shipped. What now?
Start by working out which of the three gaps you are in, because they need different money and different people. Tenhaw does that diagnosis before scoping a build. If accuracy collapsed on the long tail, the missing piece is exception routing rather than a better model. If nobody can agree what a correct answer is, the next work is a ground-truth set built with the people who own the decision, and it is a fortnight rather than a phase. If the extraction is fine and the numbers are not trusted, you are waiting on a semantic layer and lineage, a data programme with a different sponsor and a different budget line. A proof of concept that never shipped is rarely blocked on the model.
What accuracy is good enough for document intelligence?
There is no universal number, and quoting one is a warning sign. The right question is what the exception path costs. A process that tolerates review can run at accuracy that would be unacceptable for straight-through processing. Design the routing first, by extraction confidence and business consequence, and the accuracy target falls out of it rather than being asserted up front.
Do we need to fix our data platform before doing this?
Not before a proof of concept, and yes before production. Whether the extraction is viable at all is the cheaper question to answer first, so run it against the estate as it stands. Tenhaw did exactly that for a London insurance market business, taking PDFs to business intelligence on Azure in a two-week proof of concept, ahead of any data platform work. But structured output with no semantic layer, agreed definitions or lineage produces confident answers from the wrong table, so data foundations belong in the productionisation scope rather than being discovered during it. That scope usually sits with a different sponsor and a different budget line, which is where programmes with working extraction and untrusted numbers get stuck.
What is intelligent document processing?
Intelligent document processing, or IDP, also called document intelligence, turns unstructured source material such as PDFs, scans, email and forms into structured data that lands in a warehouse and drives business intelligence. It is where most enterprises meet agentic AI first, because the source material already exists, nothing is customer-facing, and the value is easy to describe to a board. It is also the pattern most likely to stall between a convincing extraction demo and a number the business will actually act on, which is why exception handling, ground truth and the BI layer matter as much as the model.
How do you establish ground truth for document extraction?
Build a ground-truth set with the people who own the decision, before anyone argues about models. Two experienced underwriters will disagree about what a document says, so if the decision owners have not agreed what a correct answer is, every accuracy conversation becomes an argument about the benchmark rather than the system. Done properly it is a fortnight of work rather than a phase, built from real documents, including the long tail of scanned faxes, amended schedules and important numbers buried in footnotes, with the experts' disagreements resolved into agreed answers. Skip it and the pilot never fails a test, it simply never gets one, which is how document programmes sit unresolved for quarters.
What does a good document extraction pipeline look like?
Separable stages, in a deliberate order. First use a language model to turn each document into a markdown representation of itself, so the extraction is inspectable by a human and re-runnable when the field list changes, then narrow to the fields that matter. Normalise those fields into consistent formats, enrich them against third-party APIs, then add semantic enhancement, context and thematic grouping, so downstream logic works with meaning rather than strings. Finally apply business logic and surface the result in a dashboard, not just a populated table. Keeping the stages separable matters because enrichment and semantics are where accuracy problems are usually diagnosable.
How do you stop human review becoming the bottleneck in document AI?
Route work by confidence and consequence instead of building a first-in, first-out queue. Every serious implementation keeps humans in the loop, but most implement it as a queue worked in order, which spends scarce expert time uniformly across easy and hard cases until review becomes the bottleneck and the business case shows a saving the operation cannot feel. Score confidence using the provenance of the data, which enrichment source it came from, alongside model certainty and a search-based cross-check, because a model's own confidence is a weak signal on its own. Then send experts only the cases where low confidence meets real consequence.
Does it matter which cloud we use for document intelligence?
Less than the vendors suggest. Tenhaw has delivered this pattern end to end on Azure, including Azure OpenAI. A London specialty insurance proof of concept took PDFs to business intelligence in two weeks, ground the business had circled for roughly a year, with month three productionising it against the client's security standards. The parts where programmes actually fail, ground truth, exception design, human routing, the semantic layer and lineage, are provider-independent and transfer intact. What genuinely differs is the document-understanding service and how it handles layout and tables, how permissions propagate from the source repository through to the warehouse, and the residency position. On AWS or Google Cloud those service choices are scoped with your own platform engineers and priced into the engagement.
We already have an OCR tool. What would document intelligence add?
Everything that happens after the text comes off the page. Optical character recognition and template extraction solve the easy half, which is getting the characters out. Tenhaw builds the other half. Document intelligence carries those fields on: normalised into consistent formats, enriched against third-party sources, grouped semantically so downstream logic works with meaning rather than strings, then interpreted by business logic and surfaced somewhere people actually decide with. That last step is what separates document intelligence from document extraction, and it is the one most often descoped when a programme runs late, which is how organisations end up with a populated table nobody uses. If your current tool reads a stable template accurately and somebody already acts on the output, keep it.
How much does an intelligent document processing project cost?
In two pieces, and it is the second one budgets tend to miss. Tenhaw's Agentic Proof of Concept against one document type and one real workflow is a fixed £20,000 to £55,000 over two to four weeks, set by the complexity of the workflow and the state of the data underneath it. Productionising it is a separate phase, four to six weeks with a dedicated team, and an Agentic Build Team is £70,000 to £85,000 a month. Then there are the data foundations. A governed semantic layer, agreed definitions and lineage back to the source document usually sit with a different sponsor and a different budget line, and programmes that meet that late stall with the extraction working and the numbers untrusted.
Which document type should we start with?
One document type, one real workflow, and a workflow whose output somebody actually decides with. Document work is usually where an organisation meets agentic AI first because the source material already exists, nothing is customer-facing, and the value is easy to describe to a board, so the temptation is to pick the tidiest document type in the building. Resist it. Accuracy on a curated sample tells you almost nothing about the scanned fax, the amended schedule, or the important number sitting in a footnote. Choose something with real exceptions and a named destination for the data, or the proof of concept proves the wrong thing.
Who from the business needs to be involved in a document AI project?
Four groups, and missing any one of them is the usual cause of a stall. The people who own the decision, because ground truth has to be agreed with them and two experienced experts will disagree about what a document says. Your own engineers, because Tenhaw pairs with them throughout the build and the point is that your team can run it afterwards. Whoever owns the warehouse and the semantic layer, since agreed definitions and lineage are usually a different sponsor with a different budget line. And the operations people whose review time you are about to route by confidence and consequence instead of a first-come queue.
Will it handle tables and awkward layouts in our PDFs?
Layout and table handling is one of the genuinely provider-specific decisions in this pattern, so test it on your own documents rather than taking a vendor demo on trust. The pipeline shape helps, and Tenhaw ran this one on Azure for a London insurance market client. Each source document is first turned into a markdown representation of itself, which keeps what the model read inspectable by a human, and only then is it narrowed to the fields that matter. When a layout defeats it, you want that arriving as low confidence and routing to a person rather than as a confident wrong value in a dashboard. Design that exception path before anyone argues about which document-understanding service wins.
What happens when a document format changes after go-live?
Less than you would fear, if the pipeline was built in separable stages. Each document is first turned into a markdown representation of itself and only then narrowed to the fields that matter, so extraction stays inspectable by a human and re-runnable when the field list changes rather than being a black box from PDF to database column. A shape the pipeline has not seen before tends to arrive as low confidence, which is exactly what routing by confidence and consequence exists for. Validating a system like this is only feasible during operation, the published engineering literature says, and its robustness evolves rather than being designed in at the start.
Do our documents leave our own cloud during processing?
Not if the pipeline is deployed the way Tenhaw deploys it, on your infrastructure and under your policies. The delivered version ran end to end on Azure inside a London insurance market client's own regulated estate, with UK data residency the default and the EU available where that is the requirement. The residency question worth settling early is narrower than most people expect, coming down to which region the document-understanding service processes a file in and how identity and permissions propagate from the source repository through to the warehouse. Both are genuinely provider-specific, and both are far cheaper to answer before any documents move.
How do we trace a dashboard number back to the source document?
Design for it, because a sceptical reviewer asks for it first and it is not free. Three mechanisms do the work. Extracting each document into a readable markdown form before narrowing to fields keeps the extraction inspectable by a human rather than a black box. Confidence is scored using the provenance of the data, meaning which third-party enrichment source a value came from, alongside model certainty and a search-based cross-check, so you can say why a number was trusted. And lineage back to the source document belongs in the build alongside the semantic layer, because without it the programme has delivered extraction and not intelligence.
If the sources do not answer it, a call will.
Talk it throughVoice agents and conversation intelligence
Answered on Voice agents and conversation intelligence, and rendered here in the same words.
Read the page these answers live on →
What is conversation intelligence?
Turning recorded voice (calls, meetings) into structured, searchable intelligence: summaries, recurring themes, entities, and compliance or risk signals. It is distinct from real-time voice agents, and it is usually the lower-risk starting point because the data already exists, nothing is customer-facing, and it tests how well models handle your actual accents, jargon and line quality before anything goes live.
Can we use our existing call recordings to train or run AI analysis?
Often yes, but not automatically. Recordings captured for quality monitoring or regulatory purposes were collected under a specific processing purpose, and analysing them with AI is generally a different one. The consent basis, retention position and residency need establishing before the build rather than during it, because it is answerable, and programmes that leave it to month four lose months.
What makes real-time voice agents hard?
Latency and interruption. Transcription, reasoning, tool calls and speech synthesis all have to complete inside the window where a human would have started speaking, and users interrupt constantly. Architectures that work asynchronously fail immediately under those conditions. The handoff to a human is the other hard part, and it is usually designed as an edge case when it is the main event.
Has Tenhaw delivered voice AI?
Yes, on both halves of it. At HSBC Tenhaw's founder led an AI Voice Insights proof of concept that integrated with the contact-centre system, transcribed inbound handler calls into a vector database, classified each call for recurring themes, and drove automated agent notes and business intelligence on what customers were actually calling about, with 1.5M+ hours of annual manual administration projected rather than realised. Real-time agentic voice is shipped work at Velocity84, Tenhaw's build lab, which has produced more than twenty agentic products across voice, video, document reading, mobile and go-to-market, at startup rather than enterprise scale. Taking that into a regulated contact centre is scoped as a staged build with your telephony and compliance people in the room, and priced a stage at a time.
Should we start with call analytics or a live voice agent?
Start with the recorded estate. Analysing calls you already hold is lower risk because the data already exists, nothing is customer-facing, and it produces business value without touching a live customer interaction. Tenhaw's founder began there at HSBC, integrating with the existing contact-centre system and processing inbound handler calls. It also tells you how well the models handle your real accents, jargon and line quality before anything goes live, which is knowledge you will need either way. Latency, interruption handling and the handoff to a human make a live voice agent a different problem, and one unforgiving of architectures that only worked asynchronously. Prove value on the recordings first, then decide whether the live channel is worth the harder build.
Is sentiment analysis on customer calls reliable?
Not reliable enough to be a headline metric. Sentiment scoring on real calls is noisy, culturally variable, and frequently wrong on exactly the calls that matter most. Treat it as one weak signal among several rather than the number the programme reports upwards, because using it as the headline is a reliable way to lose the confidence of the operational team that has to act on it. Thematic classification, showing what customers were actually calling about, is usually the more dependable output, and it answers a question most contact centres can currently only answer anecdotally.
How should a voice agent hand off to a human?
Design it as the main event, not an edge case, because the handoff decides whether customers experience the agent as helpful or as an obstacle. Three things need specifying: how the agent decides to escalate, what the human receives when it does (the context, the transcript, the reason the agent gave up), and what the customer experiences at the boundary. Each should be specified and tested as carefully as the automation itself, and the escalation decision works on the same confidence-and-consequence routing logic used for document review. Most implementations design the happy path thoroughly and the handoff barely at all, which is the wrong way round.
How do you choose a speech-to-text provider for a contact centre?
Run a bake-off against your own recordings. It is the one place Tenhaw would insist on one. Speech recognition quality on your specific accents, jargon and line conditions varies meaningfully between providers, and it is worth testing rather than assuming from published benchmarks. For real-time work the orchestration primitives and telephony integration also differ substantially, so a transcription score alone does not settle the choice. In regulated environments, establish data residency for voice before any provider selection, because it is frequently the binding constraint and can rule providers out before quality is even measured.
How do you analyse years of call recordings with AI?
Transcribe the calls into a retrievable store first, then classify each call for recurring themes, account issues and similar categories. The AI Voice Insights proof of concept Tenhaw's founder led at HSBC used exactly that architecture, integrating with the existing contact-centre system, processing inbound handler calls and landing transcripts in a vector database before thematic classification. The ordering matters, because with a searchable corpus in place, each new analytical question can be asked of the same store without reprocessing, rather than every question requiring a new pipeline. Outputs then need to land somewhere governed, with lineage back to the source conversation, or you have replaced one unsearchable estate with another. Settle the data-protection position in month one, before the build.
What are the benefits of AI call analytics for a contact centre?
Two concrete outputs, serving different people. The AI Voice Insights proof of concept Tenhaw's founder led at HSBC produced both, and was projected to remove more than 1.5 million hours of manual administration a year, a projection rather than a realised figure. The operational output is automated agent notes, which remove the manual write-up from a call handler's day. The analytical one is thematic classification, a structured view of what customers were actually calling about, which most contact centres can only answer anecdotally. Aim at both, because programmes that pursue only the analytical output tend to struggle for operational buy-in; nothing changes for the people whose calls are being analysed.
Do call recordings have to stay in the UK to be analysed?
They can, and on regulated work they usually should. Tenhaw's default is UK data residency with the EU available, and the pipeline runs on your infrastructure under your policies, not Tenhaw's. Residency for voice is frequently the binding constraint, so it is established before any provider is selected rather than after, because it can rule an otherwise strong speech provider out before quality is even measured. It also sits alongside processing purpose, consent basis and retention as a position settled with your DPO before the build starts, since recordings of customers are personal data and moving them is a decision rather than a deployment detail.
Will an off-the-shelf call analytics tool do, or do we need a build?
Often it will. A packaged tool answers packaged questions perfectly well, and if yours are generic, Tenhaw will say so. The build case turns on three things. Whether the categories you need are your own recurring themes and account issues rather than a fixed set someone else chose. Whether the summaries and themes have to land somewhere governed, with lineage back to the source conversation, because outputs nobody can trace have replaced one unsearchable estate with another. And whether the pipeline has to sit inside the contact-centre system you already run, which is how the AI Voice Insights proof of concept Tenhaw's founder led at HSBC was built.
How much does a conversation intelligence pilot cost?
Between £20,000 and £55,000, fixed, over two to four weeks. That is Tenhaw's published Agentic Proof of Concept price, and voice work is engaged through it. For that money the scope is usually the recorded estate rather than a live agent: transcribe a slice of your calls into a retrievable store, classify them for recurring themes and account issues, and find out how the models cope with your real accents, jargon and line quality. It ends in something running rather than a slide pack, and it commits you to nothing afterwards. Settle the processing purpose for those recordings inside the same weeks, not later.
Who needs to be involved before we analyse call recordings?
Your DPO, as a co-author rather than a reviewer at the end. Tenhaw settles purpose, consent basis, retention and residency with them inside the two to four weeks of an Agentic Proof of Concept, because Article 5(1)(b) of the UK GDPR limits further processing incompatible with the purpose recordings were collected for. Then whoever owns the contact-centre system the pipeline has to integrate with, since that integration is most of the engineering. Then the operational leader whose handlers are on the calls, because their people have to act on whatever comes back. Getting those three in early is what decides whether the work ships or stalls in month four.
Will call handlers see AI call analysis as surveillance?
They will if the programme gives them nothing back, and the answer is in the design rather than the communications. The AI Voice Insights proof of concept Tenhaw's founder led at HSBC aimed at an operational output as well as an analytical one, so the same pipeline that tells the business what customers were actually calling about also produces automated agent notes that take the manual write-up out of the handler's day. Do the same, then keep sentiment scoring away from anybody's performance review, because it is noisy and culturally variable and people can tell when they are being measured by a number that is wrong about their hardest calls.
How do you set a latency budget for a voice agent?
Give each stage an explicit number and treat the total as an architecture requirement rather than something discovered in production. Transcription, reasoning, tool calls and speech synthesis all have to finish inside the window where a human would already have started speaking, so writing the budget down early is what keeps the design honest once real systems are attached to it. Then test against interruption from the first week, because users interrupt constantly and designs that behave beautifully asynchronously fall apart the moment somebody talks over them. That is the most transferable lesson from shipping consumer voice products, where users are far less forgiving than internal testers.
Won't customers hate talking to an AI voice agent?
Mostly they hate being trapped by one. Two design choices decide the experience. The first is interruption, because people talk over an agent constantly and an architecture that only ever worked asynchronously falls apart the moment they do, which makes the whole thing feel like a machine however good the voice sounds. The second is the boundary, the moment the agent gives up and a person takes over. Treat that as the main event and callers experience the agent as helpful. Most of the effort goes into the happy path and almost none into the boundary, which is why so many callers come away feeling obstructed.
Can a voice agent actually do things, or only answer questions?
Taking an action is part of the definition. A real-time voice agent holds a live conversation, does something in a system and hands off to a human, so answering questions is only the first part of it. Time is the constraint rather than capability. The tool call has to complete inside the pause a caller will tolerate, which means a slow back-end system quietly decides what the agent can finish while someone is still on the line. Work that out in the architecture instead of discovering it in production, and specify where the agent stops and a person picks up.
If the sources do not answer it, a call will.
Talk it through1424 questions, grouped by subject
Every question answered anywhere on tenhaw.com sits in one of 51 groups. This is one of them.
- Using the pattern guides15
- End-to-end agentic workflow18
- Retrieval and knowledge access18
- Retrieval, fine-tuning or prompting18
- Tools and system integration18
- Agent identity and access18
- Guardrails and accuracy18
- Agent evaluation and assurance18
- The business case18
- Governance and regulatory evidence18
All 1424questions, and every group →
Or ask the question directly and skip the categories.
Talk it throughStill have a question?
A 30-minute discovery call with James Rooney. Bring the question this page did not answer. You'll leave with a rough scope whether you engage us or not.
most start with a fixed-price AI Readiness Audit · £44,000 · 4 weeks · working prototypes
Calendar not loading? Open it on cal.com or email hello@tenhaw.com.