← Previous: What Is Generative AI (and What It Isn’t) | Next: Tokens, Context Windows, and Why They Matter →
Last time, we said generative AI builds new content instead of picking from a fixed list. Text is the clearest example of that, and the model doing most of the building is the LLM, the Large Language Model. Fair question at this point: how does something “build” a sentence out of nothing?
Here’s the part that surprised me the first time I really dug into it. The honest answer is almost anticlimactic. An LLM is, at its core, a really, really good autocomplete.
I know that sounds like it’s underselling something that can write code and explain physics. It isn’t. Let’s see why.
You Already Know This Pattern
Your phone’s keyboard does next-word prediction. You type “I’ll see you” and it suggests “tomorrow,” “later,” “soon.” It’s guessing the most likely next word based on patterns it’s seen.
An LLM does the exact same job: just at a scale and quality that makes it feel like something entirely different. Instead of learning from a few years of your texting history, it’s learned from a huge slice of the internet, books, articles, and code. And instead of predicting one word at a time in a fairly crude way, it predicts the next token with startling nuance, then does it again, and again, and again, each time reading everything generated so far.
One Token at a Time
Here’s the loop, visually:
flowchart LR
A["Prompt: 'The capital of France is'"] --> B[Model predicts next token]
B --> C["' Paris'"]
C --> D[Add it to the text]
D --> E[Model predicts next token again]
E --> F["'.'"]
F --> G[Repeat until done]
In slow motion, that’s:
- You give the model a starting point: your prompt.
- It looks at everything so far and asks: “what’s the single most likely next chunk of text?”
- It picks one (in our example, " Paris").
- That chunk gets glued onto the text.
- The model looks at the new, slightly longer text and does step 2 again.
- This repeats, one small piece at a time, until the response is complete.
There’s no paragraph being planned in advance. There’s no outline sitting in the model’s head. It’s genuinely one step at a time: predict, append, repeat.
Where Does the “Knowing” Come From?
This is the part people find hard to believe: the model isn’t looking anything up. It doesn’t have a database of facts it queries. Everything it “knows” was compressed into its parameters during training, a process where it read enormous amounts of text and adjusted itself, over and over, to get better at predicting what word comes next.
flowchart TB
A[Massive amount of text] --> B[Training: adjust the model<br/>to predict next words better]
B --> C[Model gets good at this one task]
C --> D[That single skill turns out to be enough<br/>to write, explain, summarize, and reason]
The surprising discovery behind the entire LLM boom is this: if you get a model good enough at predicting the next word, that one narrow skill generalizes into something that looks a lot like writing, explaining, summarizing, translating, and even reasoning. Nobody explicitly programmed “explain this like I’m five”; it emerged from being extremely good at next-token prediction over a huge, diverse dataset.
Why It Sometimes Gets Things Wrong
This mental model also explains something important for later: an LLM is optimized to produce plausible-sounding next text, not verified facts. Most of the time those overlap heavily, because plausible text usually is correct; the model has seen the real answer thousands of times in training. But when it hasn’t seen enough about something, it will still confidently produce the next most-likely-sounding token anyway. That’s the seed of what we’ll later call hallucination, a topic that deserves (and gets) its own article.
Mental Model
Picture someone finishing your sentences, but they’ve read practically everything ever written first. You say “The sky is…” and they say “blue,” not because they looked outside, but because in everything they’ve ever read, “blue” was overwhelmingly what came next after “the sky is.” Stack that trick a few hundred times in a row, one word after another, and you get a full paragraph, a poem, or working code.
That’s an LLM: a next-word guesser, running in a loop, that got so good at guessing it can write.
Key Takeaways
- An LLM generates text one small piece (token) at a time, not all at once.
- Each new token is predicted based on everything written so far: the prompt plus its own output.
- The model isn’t retrieving facts from a database; its “knowledge” is patterns compressed into it during training.
- Being extremely good at “predict the next word” turns out to generalize into writing, summarizing, explaining, and more.
- Because it’s predicting plausible text rather than verified text, LLMs can sound confident while being wrong: this is the root of hallucinations, covered later in this series.
Next up: Tokens, Context Windows, and Why They Matter, what a “token” really is, and why every model has a limit on how much it can remember at once.