The programme is trying to eliminate hallucination rather than bound it
Every week spent hunting a model that does not make things up is a week not spent designing what happens when one does. The incentives that produce a confident guess sit in how models are trained and scored, so the mitigation available to you is architectural: constrain what the system can say, make it cite where each claim came from, give it a supported way to decline, and route by consequence. Programmes that treat this as a procurement question, solvable by the next model, are still having the same conversation two quarters later with a larger bill.