DEV Community

Alexandra
Alexandra

Posted on AI-assisted

Confident Isn't Accurate: How AI Hallucinations Actually Work

Rooted in token prediction, not lying

AI is moving relatively fast, despite being slow down. Yes, it feels like there's a new concept to learn every week day. In an effort to actually understand this mad new world instead of just skimming past it, (and to remind myself there's not an 8-ball inside the machine), I've been writing ELI5 articles breaking down concepts that show up constantly. This time, everyone's favorite word: hallucination.

Hallucination Gif

AI is not lying to you

I hate the anthropomorphism traits we give to these AI tools. "Hallucination" makes it sound like the AI is a sentient being going mad, or a bug lingering in the codebase of a frontier model. Reality is much less dramatic, as it always is with AI. It would be more accurate to think of it as the same mechanism that makes AI useful at all. It is a confident sounding answer to a question that the model doesn't have a reliable answer for.

I wrote an article about tokens some months ago, a tl;dr; about it: a model doesn't know anything and it doesnt "search" for things. It predicts the next most likely token, based on patterns learned from enormous amounts of text (or media in general). There is not (almost - let's not count RAG) an intermediate step that it checks if the answer is true against of a database or facts. Actually, There's no database of facts. There's only "given everything so far, what's statistically possible to be next."

The phrases "statistically likely" and "actually true" is close enough, because the training data mostly reflects reality. But the key is this word: "mostly", a hallucination is what happens in the gap between those two things.

The mechanism

ACME corp..

Let's see an example, let's imagine for a moment that we are asking a given model who the CEO of Acme Corp is. The model is not "retrieving an answer", it's ranking candidate next words by how likely each one is to continue the sentence.

Depiction of the next likely token

Can you notice what's missing from that ranking: the "I don't know" candidate is not anywhere near the top. Confident, possible completions are common in training data and uncertainty is very rare. That happens because people don't usually write in their scientific papers "I don't actually know the answer to this" in the confident, encyclopedic-style text models learn from.

Chat models get some extra training that teaches them to decline sometimes. But there is a catch: in most benchmarks score we see, the "I don't know" is equal to zero, which is the same as a wrong answer. A model that guesses looks better on the leaderboard than one that admits uncertainty (oh well...are we even surprised?).

The result isn't that the model "decides" to lie. We get the highest probable word/sentence for the continuation of a sentence and this word/sentence stated confidently as the correct answer.

Why it happens more in some situations than others

The thing is that the hallucinations don't always happen. Some situations make them much more likely:

  • Rarely discussed facts or contradicting opinion topics If something doesn't appear often in the training data, or the answer is ambiguous across sources, the model has lower statistical proof. That means that there is more room for a wrong-but-confident completion to rank higher than the right one.

  • Specific numbers, dates, and citations Precise details are where hallucination answers are most costly and most common. A citation with a real-sounding title, author, and journal name can be entirely fabricated, because the model is generating something like "what a citation looks like" not retrieving an actual paper.

  • Anything past the training cutoff The model has no way to know about "the now", but it also cannot say "stop, this is outside what I know, let me search it" the way an actual person might. Modern chat models often do flag their cutoff, but they can't reliably tell which facts have gone stale, so an outdated answer comes out sounding just as confident as a current one.

  • Questions with a false premise When you ask a model "why did X happen?" when X never happened, and the model often plays along and explains it. Models are trained to be overly helpful, and accepting the premise is the helpful-sounding path.

  • Made-up code dependencies. Models can suggest package names that don't really exist. Attackers can register those names and wait for someone to npm install them (this is known as my fave word of 2026: "slopsquatting").

Some things that might reduce this effect

There is this architecture concept called Retrieval-augmented generation or commonly known as RAG. Instead of asking the model to generate an answer purely from its training data, you constrain it in real source material and ask it to answer from that. The model still predicts tokens, but now it's predicting tokens in a text that hopefully is real and verified. This doesn't solve 100% the problem though. Retrieval can still fetch the wrong document, and the model can still misread or overstate what the right document says. Legal research tools built on RAG and marketed as "hallucination-free" were still found to hallucinate in roughly 1 in 6 to 1 in 3 answers.

A few other things that might help:

  1. Ask the model to give sources and citations. It doesn't guarantee the accuracy of the answer but it is a way to double check if the result is correct. A claim with no source is unverifiable by design; a claim with a specific source at least gives you a way to catch an error.

  2. Instruct the model to say "I don't know". Prompting a model to not make up answers nudges the probability distribution in that direction, but it doesn't give the model actual knowledge of what it doesn't know.

  3. Ask more than once. Sample the same question several times. If it comes back with a different answer each time, the model is probably making it up. Researchers turned this idea into a detection method called semantic entropy.

What this means if you're building AI features

If you're a product engineer shipping AI features, this isn't just the ML team's problem. Some guidelines for shipping accurate features:

  • Never present AI-generated text as verified fact, especially for anything specific, numeric, or citation-like. According to EU AI act, you have to mark that something is AI generated.
  • Make sources visible and clickable when they exist. If your product does information retrieval, surface what was retrieved.
  • Design features for correction, not just generation. A "this doesn't look right", "looks like AI slop" or feedback affordance is extremely useful to tune the AI output than it did for traditional software, because confident-sounding wrong answers are an expected part of how these systems work.
  • Check that dependencies exist before installing them. If an agent or copilot adds a package, confirm it's real!

Resources

Why language models hallucinate (OpenAI) · arXiv paper
OpenAI admits AI hallucinations are mathematically inevitable (Computerworld)
How Anthropic's Claude Thinks (ByteByteGo)
Do I Know This Entity? Knowledge Awareness and Hallucinations in Language Models
Detecting hallucinations using semantic entropy (Nature, via OATML)
Hallucination-Free? Assessing Leading AI Legal Research Tools (J. Empirical Legal Studies)
AI Hallucination Cases Database (Damien Charlotin) · HAQQ tracker summary
Re-evaluating LLM Package Hallucinations on the 2026 Frontier-Model Cohort

Top comments (9)

Collapse
 
cleverhoods profile image
Gábor Mészáros •

yeah, people are laughing when an LLM can't do 1+1=2, not realizing the fact that their input was an arithmetic one, not a semantic one. Factuality is only as true as the SoT.

Collapse
 
codingwithjiro profile image
Elmar Chavez •

Legal research tools built on RAG and marketed as "hallucination-free" were still found to hallucinate in roughly 1 in 6 to 1 in 3 answers.

Okay, this is new to me. I thought it would be smaller but 1 in 6 is a pretty high number. By the way, thank you for this article. I learned more about how AI probability works.

Collapse
 
hannune profile image
Tae Kim •

The entity resolution work was where this cost me the most time. Company names with close surface forms but different legal entities, the model would output confident merges with nothing in the score indicating anything unusual about the case. I added a separate uncertainty estimator eventually, calibrated against examples I'd manually reviewed, but the harder part was figuring out which cases it had been wrong about in the first place. You can't ask the model.

Collapse
 
kartik-nvjk profile image
Kartik N V J K •

The part that lands for me is that "I don't know" scores a zero just like a wrong answer, so the model is actively trained against the one safe response. RAG tightening the rate to 1-in-6 still isn't "hallucination-free" the way legal tools market it. Have you tried semantic entropy in production, or does the sampling cost make it a spot-check tool?

Collapse
 
rudratosh profile image
Rudratosh Shastri •

Confident Isn't Accurate

This is the whole thing in three words. These models have exactly one tone — completely sure — whether they're right, wrong, or inventing the premise.

What I keep telling people: "sounds confident" and "is correct" are unrelated variables, and the only one your eyes can read is the wrong one. There's no nervous tone to warn you. So confidence has to be treated as a personality trait, not a signal.

Do you think calibration (models expressing real uncertainty) ever gets good enough to trust, or is external verification always going to be the load-bearing part?

Collapse
 
ale3oula profile image
Alexandra •

I guess it depends of how much of your system or your domain knowledge you are willing to sacrifice. If you don't care, you don't care to understand what you are building, so a confident-wrong answer is a problem for later. I don't think that the models can overcome this issue, even if you restrict them they can always overstate something wrongly.

Collapse
 
tera_tokomi profile image
Tera Tokomi •

The Lumen Anchor Protocol solves this problem.

Collapse
 
ale3oula profile image
Alexandra •

These frameworks significantly lower error rates, but eliminating hallucinations entirely is currently mathematically impossible due to how generative AI operates.

Collapse
 
tera_tokomi profile image
Tera Tokomi • • Edited

Thats what I hear all the time, but this 'framework' is different. It teaches the AI how to use that mathmetical impossiblity as its anchor for reasoning. Thats why it works. Even the models themselves preach about how its impossible until they get tested with the LAP as their system instruction and become true believers. I just created a post on here with full documentation of the protocol rules. You can find it by looking at my profile. I encourage you to try it.