The Confidence Trap: Why Your AI Sounds Right When It’s Wrong
Ask three leading AI models for the title of a researcher’s doctoral dissertation and you will receive three different answers — each stated with total certainty, each entirely fabricated. That finding comes not from an AI sceptic’s blog but from OpenAI’s own research team, who published a paper in September 2025 demonstrating that hallucination is not a defect awaiting a patch but a structural consequence of how every major language model is trained, evaluated, and rewarded.
Focus On: The Incentive Architecture of Fabrication
The paper, authored by Kalai, Nachum, Vempala, and Zhang (arXiv:2509.04664), makes a disarmingly simple argument: language models hallucinate because every incentive in the training pipeline rewards confident guessing over honest uncertainty, at every stage, without exception.
Start with how these models are examined. Nine out of ten industry benchmarks employ binary grading, where a correct answer scores one point while both an incorrect answer and an admission of ignorance score zero — silence is, in mathematical terms, indistinguishable from error. Under those rules, models don’t learn to fabricate through malice; they learn it because the examination system makes fabrication the optimal response to uncertainty, the way a student in a multiple-choice exam will always guess rather than leave the answer blank.
The problem compounds at the data layer. Not every fact appears with equal frequency in the training corpus — Einstein’s birthday surfaces thousands of times while a niche regulatory amendment might appear once — and the researchers demonstrate that models will hallucinate on at least 20% of these low-frequency facts, what they term the “singleton rate,” which is not a performance issue to be optimised away but a mathematical inevitability.
Then comes reinforcement learning from human feedback, the stage at which whatever calibration the base model possesses gets systematically dismantled. Human raters, when presented with two responses, reliably prefer the one that sounds more confident and more detailed, which means the model learns — rapidly and irrevocably — that certainty is rewarded and doubt is punished. Pre-training creates knowledge gaps; evaluation teaches the model to fill them with fabrications; post-training teaches it to make those fabrications sound authoritative.
I find this framing illuminating because it reveals hallucinations are not an engineering problem awaiting a cleverer algorithm, but a misaligned incentive structure — and anyone who has managed a sales team knows exactly how those distort behaviour over time.
The Half-Truth Problem
If a model told you the FTSE 100 was founded in 1847, you’d dismiss it instantly. But that isn’t how hallucination typically manifests in enterprise settings — it manifests as the half-truth, a response where 80% of the content is verifiable and correct while the remaining 20% is confidently invented, nestled between facts that have already built your trust.
A recent experiment by Alex Banks illustrated this with uncomfortable precision. Banks constructed a fabricated anecdote about Steve Jobs visiting a Swiss watchmaker in 1993, woven through real facts — Schaffhausen is a genuine watchmaking city, Isaacson did write the biography, Jony Ive really did work at Apple. Tested across twelve models from four providers, nine accepted the fabrication as fact, and seven of those actually found contradictory evidence during their search yet still declined to reject the premise.
The enterprise data reflects this vulnerability. Research from 2024 indicates that 47% of enterprise AI users made at least one major business decision based on hallucinated content, and 72% of S&P 500 companies now flag AI as a material risk in their disclosures — a figure that stood at 12% just two years earlier. The risk isn’t that AI gets things spectacularly wrong; it’s that AI gets things almost right, in ways exceedingly difficult to detect at the speed at which business actually moves.
What Grounded Architecture Solves — and What It Doesn’t
As I argued in my 2025 predictions, Retrieval-Augmented Generation and interoperable frameworks like Model Context Protocol have become enterprise standard for good reason: grounded systems that retrieve from authoritative sources before generating a response represent the most effective mitigation currently available.
But this research clarifies where grounding ends and speculation begins. RAG diminishes hallucination on facts the retrieval system can source, yet it does nothing about fabrication in the synthesis — the inferential leaps between retrieved passages, the connective tissue the model generates to produce coherent prose. The lacuna exists precisely at the boundary where retrieved facts stop and model generation takes over.
The more consequential investment for 2026 is therefore organisational rather than architectural. Enterprises need what I’d call institutional calibration: the distributed competence to identify where AI outputs cross from retrieval into speculation, paired with review workflows that surface errors quickly rather than assuming they won’t occur. That means training domain experts to interrogate AI-assisted analysis with the same rigour they’d apply to a junior analyst’s first draft, and building cultures where questioning an AI output carries no stigma of inefficiency.
The enterprises that extract the greatest value from generative AI won’t be those that trust it most or fear it most, but those that develop the organisational discipline of calibrated scepticism — treating AI as an exceptionally capable but constitutionally overconfident colleague. When your AI is 80% right and 100% confident, who in your organisation is trained to find the 20%?
Follow me
That’s all for this week. To keep up with the latest in generative AI and its relevance to your digital transformation programs, follow me on LinkedIn or subscribe to this newsletter.
Disclaimer: The views and opinions expressed in Chronicles of Change and on my social media accounts are my own and do not necessarily reflect the official policy or position of S&P Global.
