← Previous: Getting an LLM to Follow the Rules | Next: Vector Databases 101 →

Last time we made sure the model’s output is reliable and structured. This time, let’s fix something on the other side of the interaction: the input. If you’ve built more than one or two things with LLMs, you’ve probably already lived this — a prompt string hardcoded directly into your application code, tweaked in place every time something breaks, with no history of what changed or why.

That works for a weekend project. It falls apart fast in anything real.

The Problem With Prompts Living in Code

Picture a prompt buried in a function like this:

1
2
3
def summarize(text):
    prompt = f"Summarize this in 3 bullet points: {text}"
    return call_model(prompt)

Looks harmless. Then six months later:

  • Someone tweaks the wording directly in the function, with no record of what the old version said or why it changed.
  • The same “summarize this” logic gets copy-pasted into two other files, and now three slightly different versions exist, drifting apart silently.
  • Nobody can answer “did this prompt actually get better when we changed it?” because there was never a before/after to compare.

None of this is a model problem. It’s a plain old software engineering problem — the same one version control and modularity solved for code decades ago, just not yet applied to prompts.

The Fix: Treat Prompts Like Code

flowchart LR
    A[Prompt as a<br/>template file] --> B[Fill in variables<br/>at runtime]
    B --> C[Send to model]
    A --> D[Version controlled,<br/>reviewed, tested]

A prompt template separates the fixed instructional wording from the variable content plugged in at runtime:

Summarize the following {content_type} in {bullet_count} bullet points,
written for a {audience} audience:

{text}

Now content_type, bullet_count, audience, and text are parameters — the same template serves many different calls, and the instructional wording lives in exactly one place. Change the wording once, and every caller benefits immediately, instead of hunting down three copy-pasted versions.

What “Reusable” Actually Buys You

Version control. Prompt templates as files (not inline strings) mean you get the same history, diffs, and rollback ability you already rely on for code. When a prompt change makes things worse, you can see exactly what changed and revert it — something nearly impossible with prompts scattered inline.

Testability. A separated template can be run against a fixed set of test inputs and checked against expected outputs — directly connecting to the evaluation practices from Foundations. Without separation, there’s nothing consistent to test against; every call is a slightly different one-off.

Consistency across a team. If five people are calling an LLM for related tasks, a shared template library means they’re not each independently reinventing (and subtly breaking) the same prompt. One well-tested “summarize this” template beats five slightly different homegrown versions.

Safer iteration. Because changes happen in one place with history attached, trying a new phrasing becomes a low-risk experiment instead of a live edit with no way back.

A Simple Structure to Start With

You don’t need a fancy framework to get most of the benefit. A folder of template files, organized by task, is often enough:

prompts/
  summarize.txt
  classify_sentiment.txt
  extract_entities.txt
  generate_email_reply.txt

Each file holds the instructional template with {placeholders}, your application code loads the relevant file and fills in the variables, and version control tracks every change the same way it tracks your source code. As things mature, dedicated prompt-management tools add versioning UIs, A/B testing, and team collaboration features on top of this same core idea — but the underlying principle doesn’t change.

Mental Model

Think of the difference between a chef who scribbles a recipe from memory slightly differently every time they cook, versus a kitchen with a written, tested recipe card that any cook on staff can follow and produce the same dish. Neither cook is doing anything wrong in the moment — but only one of them can improve the recipe deliberately, hand it to someone else reliably, or explain why last Tuesday’s version turned out better. Prompt templates are the recipe card: separating “what varies” from “what’s fixed,” written down once, and improvable on purpose instead of by accident.

Key Takeaways

  • Prompts hardcoded inline in application code drift, duplicate, and lose their history — the same problems software engineering already solved for regular code.
  • A prompt template separates fixed instructional wording from variable inputs filled in at runtime.
  • Templating unlocks version control, testability, team consistency, and safer iteration — none of which are possible with scattered inline prompts.
  • A simple folder of template files is often enough to start; dedicated prompt-management tooling adds versioning UIs and testing on top of the same core idea later.
  • This connects directly to evaluation from Foundations — you can’t meaningfully test a prompt’s quality over time if it isn’t stable and separated enough to test in the first place.

Next up: Vector Databases 101 — what they actually add on top of the embeddings concept from Foundations, and when you need one.