The programmes people are running, in the FAQ

Retrieval, fine-tuning or prompting, answered in full.

Which of the three a problem actually needs, what each costs to run, and the cases where the cheapest option is the right one.

questions in this group, each answered in full
6
pages the answers are written on, every one linked
1
questions across the whole FAQ
316

6 questions on retrieval, fine-tuning or prompting, answered by Tenhaw, a UK AI consultancy and AI delivery partner based in London. Nothing here is a summary: each answer is the exact text from the page that owns it, and every group links back to that page for the context around it.

6 questions

RAG, fine-tuning or prompting: how to choose

Answered on RAG, fine-tuning or prompting: how to choose, and rendered here in the same words.

Read the page these answers live on →

Do we need to fine-tune or use RAG?

For getting your own knowledge into the system, retrieval, and the evidence is fairly direct: a controlled comparison of knowledge injection found retrieval consistently outperformed unsupervised fine-tuning, both for knowledge the model had seen in training and for knowledge that was entirely new, and that models struggle to learn new facts through unsupervised fine-tuning at all. Retrieval also gives you three things fine-tuning cannot: content that is current without retraining, answers that cite a source the reader can open, and access control applied per user at query time. Fine-tuning earns its place for behaviour, format and cost, not for facts. Most enterprise systems that work are prompting plus retrieval, and a fine-tune later if the economics call for one.

When is fine-tuning actually the right answer?

Four cases, and they are all about form or money rather than knowledge. When you need an output shape or house style that prompting cannot hold reliably across the long tail. When you have a narrow classification or extraction task with plenty of labelled examples and a stable definition of correct. When volume is high enough that running a smaller, cheaper, fine-tuned model beats a large general one on cost at the same measured quality. And when latency matters enough that a smaller model is the only way to meet the budget. In all four, the case is measurable before you commit, which is the test: if you cannot state the number that would prove it worked, it is not the right answer yet.

Is prompting on its own enough?

More often than the market implies, and you should find out before spending anything else. Careful prompting against a strong model, with structured output and a well-designed retrieval step, covers a large share of enterprise use cases, and it is the only option with no training lifecycle attached. The reason to establish it first is not economy for its own sake: it is that without a prompting baseline you have nothing to measure the expensive options against, so you cannot demonstrate that they were worth their cost even when they were.

What does fine-tuning commit us to after the first run?

A lifecycle rather than a deliverable. Training data has to be curated, kept current and governed, because it is now part of how your system behaves. Every base model deprecation forces a retrain on the provider's timetable rather than yours. Every retrain forces a full re-evaluation, so you need the evaluation set anyway. And you carry a version history, because when someone asks why a decision came out as it did in March, the answer involves which model version was live. None of that is a reason not to fine-tune. It is a reason to make sure the case is about form or cost, where the benefit is durable, rather than about facts, where it is not.

How do we choose without running a three-way bake-off?

By sorting the requirement first, which is an afternoon rather than a quarter. Split what you need into knowledge, behaviour, format and cost. Anything in the knowledge bucket that has to be current, attributable or permission-bound is retrieval, and that is settled. Everything else starts at prompting with structured output, because that is the cheapest thing that can work and it establishes the baseline. Only what is left after that, and only where the measured gap is large enough to justify a training lifecycle, is a fine-tuning candidate. A full bake-off is worth running for one decision at most, and usually the sorting has already made it unnecessary.

Can we combine them?

Yes, and most systems that work in production do. A typical shape is careful prompting for the reasoning and the house conventions, retrieval for anything that has to be current or permission-bound, structured output for the shape, and possibly a smaller fine-tuned model handling one high-volume narrow step inside the workflow where the economics justify it. The discipline that makes combining safe is holding the evaluation set constant across every configuration, so you can attribute a change in the score to the thing you changed rather than to the general direction of travel.

All pattern guides

Still have a question?

A 30-minute discovery call with James Rooney. Bring the question this page did not answer. You'll leave with a rough scope whether you engage us or not.

30 minutesWith James personallyNo obligation

Most organisations start with a fixed-price Agent-Readiness Audit · £30k–£90k · 6–8 weeks