← Previous: Fine-Tuning vs Prompting vs RAG | Next: AI Agents 101: From Chatbot to Agent →
Last time we laid out prompting, RAG, and fine-tuning as three different levers, and RAG in particular came up as the fix for “the model doesn’t know this fact.” That’s actually a good bridge into this article, because the reason RAG helps is directly tied to why models hallucinate in the first place. Time to explain that honestly, instead of treating it as a mysterious bug.
What “Hallucination” Actually Means
A hallucination is when a model states something false, made-up, or unsupported — a fake citation, a wrong date, a function that doesn’t exist in a library, a confident answer to a question it genuinely has no reliable information about. The unsettling part isn’t that it’s wrong. It’s that it’s wrong while sounding exactly as confident as when it’s right.
If you’ve read the second article in this series, you already have the tool to understand why: an LLM generates the next token based on what’s statistically plausible, not what’s verified. Nothing in that process checks a fact against reality. It’s producing the most likely-sounding continuation of text — and “most likely-sounding” and “true” usually overlap heavily, but not always.
flowchart LR
A[Question] --> B[Model predicts<br/>plausible next tokens]
B --> C{Does the model have<br/>strong training signal<br/>on this topic?}
C -->|Yes, seen a lot| D[Answer is usually accurate]
C -->|No, thin or no signal| E[Model still produces<br/>plausible-sounding text anyway]
E --> F[Hallucination]
Notice something important in that diagram: the model doesn’t have a “know / don’t know” switch. It doesn’t pause and say “I’m not sure.” By default, it just keeps predicting the next plausible token regardless of how much real signal it has — which means uncertainty doesn’t look any different on the surface than confidence.
Why It Happens More in Some Situations
A few patterns make hallucination more likely, and once you see them, they’re easy to spot coming:
- Obscure or narrow topics — the less a subject appeared in training data, the more the model is guessing at plausible-sounding text rather than recalling something well-established.
- Very recent events — anything after the model’s training cutoff simply isn’t in there. The model may still answer confidently, filling the gap with its best guess.
- Highly specific details — exact numbers, precise citations, exact function names. Broad concepts tend to be well-represented in training data; specific details are exactly where the guessing shows up most.
- Long, unconstrained generations — the longer a response runs without being grounded in real source material, the more room there is for it to drift into plausible-but-wrong territory.
What Actually Reduces It
Grounding with RAG. This is the big one, and it’s why the last article’s discussion of RAG matters here directly. Instead of relying purely on what the model “remembers” from training, you retrieve real documents relevant to the question and hand them to the model as context, then instruct it to answer based on that material. The model is now predicting the next plausible token constrained by real source text sitting right in front of it — a fundamentally easier and more reliable task than pulling an answer purely from memory.
Asking for citations, and actually checking them. Prompting a model to cite where an answer came from doesn’t guarantee accuracy on its own — a model can hallucinate a citation just as easily as any other fact. But paired with RAG, where the citation has to point to a real retrieved document, it becomes a genuinely useful check: no matching source, no trustworthy claim.
Narrowing the question. Broad, open-ended prompts give the model more room to wander into unsupported territory. A tightly scoped question, ideally with reference material attached, gives it a much smaller, more accurate space to work within.
Lower temperature for factual tasks. Most model APIs expose a “temperature” setting that controls how much randomness goes into token selection. Lower temperature biases the model toward its most likely, most well-supported predictions — useful for factual work, though it won’t eliminate hallucination on its own if the underlying knowledge simply isn’t there.
Explicitly allowing “I don’t know.” Left unprompted, models default to attempting an answer. A simple instruction — “if you’re not confident based on the provided context, say so instead of guessing” — measurably reduces confident wrong answers, because it gives the model an acceptable alternative to guessing.
Mental Model
Picture someone finishing your sentences with total confidence, every single time, whether or not they actually know the answer. That’s the behavior baked into how an LLM generates text — nothing in the core process distinguishes “I recall this clearly” from “this sounds about right.” Grounding the model in real documents is like handing that person the actual reference material before they answer, instead of asking them to speak purely from memory. It doesn’t make the guessing habit go away, but it gives them something solid to check against instead of nothing at all.
Key Takeaways
- Hallucination happens because LLMs predict plausible next text, not verified facts — there’s no built-in fact-check step.
- It’s more likely on obscure topics, recent events, highly specific details, and long unconstrained answers.
- RAG is the strongest practical fix — grounding the model in real retrieved documents gives it something to work from besides memory.
- Citations, narrower questions, lower temperature, and explicitly allowing “I don’t know” all help — none of them eliminate hallucination completely on their own.
- Confidence in tone is not evidence of accuracy — that single idea is the most useful thing to carry into everything you build with these models going forward.
Next up: AI Agents 101: From Chatbot to Agent — the shift from a model that answers questions to one that takes real actions using tools.