← Previous: Prompting Fundamentals | Next: RAG (Retrieval-Augmented Generation) Explained Simply, coming soon

Last time we treated the model like a literal intern who just needs clear instructions. Fair enough, but there’s a limit to what instructions alone can do: an intern, however well-briefed, still can’t answer questions about a document they’ve never actually seen. To fix that, the model needs a way to search through your data and pull out what’s relevant. That’s where a concept called an embedding comes in, and honestly it’s one of the more elegant ideas in this whole field.

The Problem: Computers Don’t Understand “Meaning”

I’ve seen this exact scenario play out in a support queue: a ticket comes in saying “my order hasn’t arrived,” and sitting right there in the knowledge base is an article titled “tracking a late package.” Any human agent connects those instantly, even though the wording barely overlaps.

A computer comparing raw text has no such instinct. Word-for-word matching would completely miss the connection, because “order hasn’t arrived” and “tracking a late package” share zero overlapping words. We need a way to represent meaning, not just spelling, and that’s exactly what embeddings do.

What an Embedding Actually Is

An embedding is a list of numbers (a vector) that represents the meaning of a piece of text. A model reads your text and outputs something like:

"my order hasn't arrived"  →  [0.12, -0.48, 0.91, ... ] (hundreds of numbers)

That list of numbers on its own looks meaningless. But here’s the key property that makes it useful: text with similar meaning produces vectors that sit close together, and text with unrelated meaning produces vectors that sit far apart.

flowchart LR
    A["'my order hasn't arrived'"] --> C[Embedding Model]
    B["'tracking a late package'"] --> C
    D["'best pizza recipe'"] --> C
    C --> E[Vector space]
    E --> F["Close together<br/>(similar meaning)"]
    E --> G["Far apart<br/>(unrelated meaning)"]

The two package-related sentences land near each other in this “vector space,” even with no shared words. The pizza sentence lands somewhere completely different. The model learned this sense of closeness from being trained on huge amounts of text where related concepts kept appearing in similar contexts.

A Simpler Way to Picture It

Imagine a giant map where every possible sentence has a spot. Sentences about similar topics cluster into neighborhoods: a “shipping problems” neighborhood, a “cooking” neighborhood, a “sports” neighborhood. Embeddings are just the coordinates that place a piece of text on this map. Once everything has coordinates, finding “what’s related to this” becomes a simple, fast geometry problem: find the nearest neighbors.

flowchart TB
    A[New question comes in] --> B[Turn it into a vector]
    B --> C[Compare against stored vectors]
    C --> D[Return the closest matches]

Simple as that sounds, it’s the entire trick behind semantic search, recommendation engines, and (critically for where this series is headed) RAG.

Why This Matters for Everything After This Article

This is the piece that connects prompting to real, private data. On its own, an LLM only knows what it learned during training: it has never seen your company’s internal documents, your product catalog, or last week’s support tickets. Embeddings are what let a system search through your data for the pieces relevant to a question, so those pieces can be handed to the model as context.

That’s the exact idea the next article (RAG) is built entirely around: use embeddings to find the right information, then give that information to the model so it can answer accurately instead of guessing.

Mental Model

Picture a massive library where every book has been placed on the shelf not alphabetically, but by meaning: books about similar topics end up physically near each other, even if their titles share no words. Embeddings are the process of figuring out exactly where on that shelf a new piece of text belongs. Once it’s placed correctly, finding “what else is related to this” is as simple as looking at what’s sitting on the same shelf.

Key Takeaways

  • An embedding turns text into a list of numbers (a vector) that captures its meaning.
  • Text with similar meaning produces vectors that land close together, even with completely different wording.
  • This “closeness” is what powers semantic search: finding relevant results by meaning, not by matching exact words.
  • Embeddings are the mechanism that lets an LLM work with your data instead of only what it learned during training.
  • This sets up the next article directly: RAG uses embeddings to retrieve the right information before the model ever generates an answer.

Next up: RAG (Retrieval-Augmented Generation) Explained Simply, coming soon. It’ll cover why models hallucinate, and how grounding answers in real data fixes it.