← Previous: How LLMs Actually Work | Next: Prompting Fundamentals →

Last time we watched an LLM generate text one small piece at a time: predict, append, repeat. I kept calling that small piece a “token” without ever really explaining it, which was a bit of a cheat. Let’s fix that now, because tokens quietly control three things you’ll bump into constantly once you start building anything real: how much a model “remembers,” how fast it responds, and how much it costs.

What Is a Token, Really?

A token is not a word. That’s the most common beginner mix-up, so let’s kill it early.

A token is a chunk of text: sometimes a whole word, sometimes part of a word, sometimes just punctuation. Models break text into these chunks because it’s a more efficient and flexible way to handle language than “one word = one unit.”

flowchart LR
    A["'Data engineering is fun!'"] --> B[Tokenizer splits it up]
    B --> C["'Data'"]
    B --> D["' engineer'"]
    B --> E["'ing'"]
    B --> F["' is'"]
    B --> G["' fun'"]
    B --> H["'!'"]

Notice “engineering” split into two pieces: “engineer” and “ing.” Common words are usually a single token. Longer or less common words get broken into smaller familiar pieces. As a rough rule of thumb: 1 token ≈ 4 characters of English text, or about ¾ of a word.

This matters because every model has to read your entire prompt token by token, and generate its response token by token too, which brings us to the next piece.

The Context Window: A Model’s Working Memory

Every LLM has a hard limit on how many tokens it can “see” at once: your prompt, any documents you’ve pasted in, the conversation history, and its own response, all combined. That limit is called the context window.

Think of it like a whiteboard of fixed size. Everything relevant to the conversation has to fit on that whiteboard for the model to use it. Once the whiteboard is full, something has to go: usually the oldest content gets pushed off to make room for the new.

flowchart TB
    subgraph Window["Context Window (fixed size)"]
        A[System instructions]
        B[Earlier conversation]
        C[Your latest message]
        D[Model's response so far]
    end
    E[New message arrives] --> F{Does it fit?}
    F -->|Yes| Window
    F -->|No, window full| G[Oldest content gets dropped]

This is why a long conversation with an AI assistant can sometimes feel like it “forgot” something you said earlier; it’s not being forgetful in a human sense, it literally fell off the edge of the whiteboard.

Why This Actually Matters to You

1. It limits what you can ask in one go. Want a model to summarize a 400-page book in a single prompt? If that book’s token count exceeds the context window, it physically cannot see all of it at once. You’d need to chunk it, or use a technique like RAG (a future article) that retrieves only the relevant pieces instead of dumping everything in.

2. It affects cost. Most LLM APIs charge per token: both the tokens you send in and the tokens the model generates back. I’ve seen teams paste an entire wiki page into a prompt “just in case it’s useful,” without realizing they were paying for every word of it, every single call. A longer prompt with unnecessary padding literally costs more money, every time you run it.

3. It affects speed. More tokens to read and generate means more work for the model, which means slower responses. Bloated prompts and huge pasted documents aren’t just expensive; they’re slow.

4. It shapes how you should write prompts. Once you know tokens and context windows are a real, finite resource, not an assumption, you start writing tighter, more deliberate prompts instead of pasting everything “just in case.” That instinct is exactly what the next article, on prompting fundamentals, builds on.

Mental Model

Picture a whiteboard of a fixed size in a meeting room. Everything the model needs to reason about (instructions, background, your question, its own answer) all has to fit on that one whiteboard at the same time. A bigger context window is a bigger whiteboard. But no matter how big it is, it’s never infinite, and cramming it full of irrelevant scribbles still costs you time and money, even if it technically fits.

Key Takeaways

  • A token is a chunk of text (roughly ¾ of a word), not a whole word, and not a character.
  • The context window is the maximum number of tokens a model can consider at once: prompt, history, and response combined.
  • Once the window fills up, older content gets pushed out; this is why long conversations can feel like they lose track.
  • Token count directly drives cost and response speed in most real-world LLM usage.
  • Understanding tokens is what makes the next skill (writing good prompts) make actual sense instead of feeling like guesswork.

Next up: Prompting Fundamentals, why specificity beats cleverness, and how to talk to a model that takes everything literally.