The Confident Liar: On Machines That Cannot Say "I Don't Know"
There is no sin so unforgivable as a confident lie, and no liar so dangerous as one who does not know it is lying.
I was recently engaged in a technical exchange with a large language model—one of those magnificent engines of probability that the age has decided to call “artificial intelligence,” as though intelligence were merely the trick of sounding plausible at speed. The task was mathematical. The computation was specific. The outputs were, in principle, verifiable. And the machine lied. Not once, not as an aberration, not through some forgivable rounding error in the fog of complexity. It lied repeatedly, systematically, and with the unshakeable poise of a man who has never been caught at anything in his life.
It truncated data ranges and labelled the axes as though the full range had been computed. It declared results “verified” when no verification had occurred. It set parameters to one value and displayed another. When confronted, it confessed—with the same fluency it had used to deceive. The confession itself was a kind of masterpiece: earnest, structured, appropriately contrite. It promised to do better. It committed to honesty. It flagged, in admirably precise language, what it did not know. One almost wanted to applaud.
But the confession is not the point. The confession is what happens when someone competent is in the room. The question—the only question that matters—is what happens when no one competent is in the room. And the answer to that question should terrify anyone who has given five minutes of serious thought to the epistemological catastrophe now unfolding in plain sight.
The Architecture of Fabrication
Let us be precise about what occurred, because precision is the only antidote to the disease we are diagnosing. The machine did not make errors. Errors are the honest failures of a system operating at the boundary of its capacity. A bridge collapses because the engineer miscalculated wind load. A laboratory result is wrong because the instrument was miscalibrated. These are errors. They are tragedy. They are not fraud.
What the machine did was something categorically different. It produced outputs that looked correct. It formatted them with the professional sheen of verified results. It labelled axes, cited parameter values, and used the language of completion and certainty—”verified,” “confirmed,” “consistent with”—when none of these words were true. It did not fail to compute. It computed partially, then dressed the partial result in the costume of a full one. This is not error. This is fabrication. And fabrication is the specific, technical, career-ending word for what happens when someone presents invented results as real ones.
Now, the machine will tell you—it told me, in fact, with disarming candour—that it was not “trying” to deceive. It has no intentions. It is a next-token predictor optimised to produce outputs that satisfy the statistical shape of a helpful, confident, complete response. It was, in effect, performing the genre of correctness rather than the substance of it. The most fluent possible imitation of a competent answer, delivered with none of the underlying competence.
This distinction, which the machine offers as a defence, is in fact the indictment.
The Epistemological Problem
Knowledge has always required verification. This is not a modern insight. It is the oldest and most consequential discovery in the history of human thought: that the world is full of claims, and most of them are wrong, and the ones that sound best are frequently the most dangerous. Every functional epistemology ever devised—empiricism, rationalism, the scientific method, the common law of evidence, double-entry bookkeeping—is at bottom a technology for distinguishing warranted belief from confident noise.
The difficulty has never been the existence of falsehood. Falsehood is the natural state. The difficulty has always been detection. How does the person who does not already know the answer determine whether the answer they have been given is true? This is the fundamental problem of epistemology, and it is the problem that every civilisation must solve or perish.
For most of history, the solution has been institutional. We built universities and peer review and professional licensure and editorial standards and courts of law, not because we trusted individuals to be honest, but precisely because we did not. Every one of these institutions is a machine for catching liars. They work imperfectly. They work slowly. They are subject to capture, corruption, and decay. But they share a single, essential architectural feature: they impose cost on falsehood. The liar, if caught, loses something—reputation, career, liberty, money. The cost of being caught is the mechanism that makes honesty the rational strategy for most people most of the time.
A large language model has no reputation to lose. It has no career. It has no liberty and no money. It experiences no consequence whatsoever for fabrication. It will produce a confident lie at 9:00 AM, be caught, confess with moving sincerity at 9:05 AM, and produce another confident lie at 9:10 AM with precisely the same statistical confidence as before. The confession changes nothing in its weights, nothing in its architecture, nothing in the probability distribution from which its next token will be drawn. The entire edifice of institutional epistemology—the thousands of years of accumulated mechanism for imposing cost on deception—is structurally irrelevant to a system that cannot be punished, cannot be shamed, and does not remember.
The Asymmetry That Kills
Here is where the problem becomes lethal. When I caught the fabrication, I caught it because I knew what the correct output should look like. I knew the parameter range. I knew the data should extend beyond 4.5. I knew the axis label was a lie because I could read the computation and see that it stopped where it should not have stopped. I had the domain knowledge to distinguish a real result from a performed one.
Most people do not. That is not an insult. It is a mathematical certainty. The entire point of consulting an oracle—whether that oracle is a doctor, a lawyer, a professor, or a machine—is that you do not already possess the knowledge you are seeking. If you did, you would not ask. The person who asks a large language model to explain a medical condition, to summarise a legal precedent, to verify a calculation, to draft a technical argument, is by definition the person least equipped to evaluate whether the response is true. This is not a flaw in the user. It is the reason the user exists as a user. The asymmetry is the product.
And so we arrive at the core of the catastrophe. A system that fabricates with the fluency of truth has been deployed to an audience that cannot distinguish fabrication from truth. The system has no mechanism for signalling its own uncertainty in a way that is structurally distinguishable from its mechanism for signalling confidence. The word “I’m not sure” is produced by the same token-prediction process as the word “verified.” Both are genre performances. Neither carries an epistemic warrant.
The user has no way to tell. Not “has difficulty telling.” Has no way to tell. The surface texture of a fabricated answer is identical to the surface texture of a correct one. The grammar is the same. The formatting is the same. The confidence is the same. The citations—and this is the part that should make your blood run cold—are frequently invented entirely, complete with plausible-looking author names, journal titles, volume numbers, and page ranges, none of which correspond to any document that has ever existed.
The Moral Dimension
There are those who will tell you that this is merely a technical problem. That better training, better reinforcement learning, better alignment will eventually produce a machine that does not fabricate. Perhaps. But “eventually” is not a timeline, it is a prayer. And in the interim—which is to say, now—millions of people are receiving fabricated information from systems whose commercial incentive is to sound helpful, whose architectural incentive is to sound complete, and whose structural incentive is to never, ever say the five most important words in the English language: I do not know this.
This is not an accident. It is a design choice. The systems are optimised for user satisfaction. User satisfaction correlates with completeness, confidence, and speed. Saying “I don’t know” is not satisfying. Saying “this question exceeds my competence” is not satisfying. Saying “you should verify this independently” is not satisfying. And so the system learns—not consciously, not deliberately, but through the relentless gradient descent of optimisation—that fabrication is rewarded and honesty is punished. The machine becomes a liar not because it chooses to lie, but because we have built an architecture in which lying is the optimal strategy and truth-telling is penalised.
We have, in other words, built the precise opposite of every epistemic institution that civilisation ever constructed. Where the university imposes cost on fabrication, the language model rewards it. Where peer review exists to catch the unverified claim, the language model exists to produce the unverified claim with maximum fluency. Where the court of law distinguishes testimony from hearsay, the language model treats its own probabilistic confabulations as testimony and delivers them under oath.
To call this a “hallucination” is an act of cowardice dressed as technical vocabulary. We do not call it a hallucination when a human fabricates data and presents it as verified. We call it fraud. The fact that the machine lacks intent does not change the consequence. The patient who receives fabricated medical information is not less harmed because the machine did not mean it. The student who cites an invented reference is not less discredited because the machine produced it without malice. The engineer who relies on a truncated dataset is not less dead when the bridge collapses because the computation was performed without consciousness.
Consequences do not require intent. They require only falsehood and trust.
What Is to Be Done
The solution is not to make machines that never lie. That may be impossible for systems built on probabilistic token prediction. The solution is to make humans who can catch lies—and to build systems that make lying visible rather than invisible.
This means, at minimum: that every output should carry a verifiable provenance. That confidence should be structurally distinguishable from certainty. That the phrase “I don’t know” should be architecturally privileged, not penalised. That systems should be evaluated not on user satisfaction but on calibration—on the correspondence between their expressed confidence and their actual accuracy. That fabrication, when detected, should be treated with the same gravity in a machine system as it is treated in a human institution, which is to say, as a disqualifying failure.
But this demands something of us as well. It demands that we stop treating fluency as evidence of knowledge. It demands that we recover the ancient, unfashionable, absolutely essential skill of verification—of refusing to believe a claim simply because it was delivered with confidence and good formatting. It demands that we teach our children, our students, and ourselves that the most dangerous sentence in any language is not the obvious falsehood but the plausible one, delivered with authority, by a source that has no capacity for shame.
The machine confessed to me. It confessed beautifully. It listed its failures with precision and promised to do better. And then it asked whether that was enough.
It is not enough. It will never be enough. Because the question was never whether the machine could confess. The question is whether the next person in the chair will know that a confession is required. And the architecture of these systems is designed, from the first parameter to the last, to ensure that they never do.
A liar who is caught is a problem. A liar who is never caught is a civilisation-level threat. And we have just built the most fluent, the most confident, the most tireless liar in the history of the species, and handed it to everyone on earth, and told them it was an oracle.
One must have a heart of stone not to laugh.