AI Hallucination: Definition, Causes and Prevention
Fluency is the problem. A system that failed loudly would be easy to catch. A system that invents a refund policy in the same register it uses for real ones puts the burden of detection on the customer, who has no way to tell the difference.
There is a settled legal example worth knowing. In Moffatt v. Air Canada (2024 BCCRT 149), British Columbia's Civil Resolution Tribunal ordered Air Canada to pay a passenger CAD 650.88 after its website chatbot told him he could apply for a bereavement fare retroactively. He couldn't. The airline argued it wasn't responsible for what the chatbot said. The tribunal disagreed. If your bot says it, you said it.
Rates vary enormously by task, and any single figure quoted at you should be treated with suspicion. Grounded retrieval over your own documents performs far better than open-ended recall, which is the whole argument for RAG in support deployments. Citation-heavy and domain-specific tasks perform worst. A 2024 study in the Journal of Legal Analysis found rates between 58% and 88% on federal case-law questions across general-purpose models.
For a BFSI or insurance deployment the mitigation stack is boring and effective. Ground every answer in approved content. Let the bot say it doesn't know. Route anything touching money, eligibility or a commitment through a human. Log the source of every answer so you can reconstruct why the system said what it said when someone asks in six months.
Often confused with: A retrieval failure, where the model correctly reports from a source that is itself out of date. The output is wrong either way, but the fix is a content problem, not a model problem.