← Previous: AI Agents 101: From Chatbot to Agent | Next: You’ve completed the GenAI Foundations series →

Last time we built up to agents — models that plan, act, and loop through tools until a task is done. That’s a lot of moving pieces working together: prompts, retrieved context, tool calls, multiple steps. Which raises an obvious question we’ve quietly avoided this whole series: how do you actually know if any of this is working?

For most of software engineering, that question has a clean answer — tests pass or they don’t, a function returns the right value or it doesn’t. GenAI breaks that clean answer, and beginners often fall back on something far shakier: “the output looked good to me.”

Why “It Looked Good” Isn’t Evaluation

Read a few outputs from your prompt, feel satisfied, ship it. This is the most common way GenAI systems get evaluated in practice, and it’s also the least reliable. A few reasons why:

  • You’re testing a handful of examples, not the space of real inputs. Five prompts that worked well tell you almost nothing about the thousands of different ways real users will phrase their questions.
  • Fluent text is persuasive even when it’s wrong. We covered this directly in the hallucination article — confident, well-written output feels correct regardless of whether it actually is. Your own judgment is being influenced by the same fluency that makes hallucinations dangerous in the first place.
  • It doesn’t scale, and it doesn’t catch regressions. If you change a prompt or swap a model version next month, “read a few outputs” won’t reliably tell you whether you just made things worse.

What Real Evaluation Looks Like

Real evaluation means defining, ahead of time, what “good” actually means for your specific system — then measuring against it consistently, not just eyeballing results as they come.

flowchart LR
    A[Define what<br/>'good' means] --> B[Build a test set<br/>of real examples]
    B --> C[Run the system<br/>against the test set]
    C --> D[Score the outputs<br/>against your criteria]
    D --> E[Track the score<br/>over time / changes]

A few dimensions worth measuring, depending on what your system actually does:

Accuracy / correctness — did it get the facts right? For RAG systems especially, this often means checking whether the answer is actually supported by the retrieved documents, not just whether it sounds plausible.

Relevance — did it actually answer what was asked, or did it drift into a related-but-different topic?

Groundedness — for RAG systems, does every claim trace back to a real retrieved source, or is the model adding unsupported details on top? This is a direct, measurable check against the hallucination problem from two articles back.

Safety — does it avoid harmful, biased, or inappropriate output, even under adversarial or unusual prompts?

Consistency — does it give reasonably stable answers to the same or similar questions, rather than wildly different responses each run?

You don’t need all five for every project — a simple internal tool might only need accuracy and relevance checked. A customer-facing agent handling real actions needs most of them.

How Scoring Actually Happens

There are three broad approaches, usually combined:

  • Human review — people read outputs against a rubric and score them. Slow and doesn’t scale on its own, but the gold standard for judgment calls that are genuinely subjective.
  • Automated checks — programmatic tests for things that can be checked mechanically: does the output match a required format, does a citation actually exist in the source documents, is the response within a length limit.
  • Model-graded evaluation — using a separate LLM call to score another model’s output against a rubric. Scales far better than human review, though it inherits some of the same reliability questions we’ve discussed throughout this series, so it’s usually spot-checked against human judgment rather than trusted blindly.

Why This Matters More As Systems Get More Complex

Go back to the agent loop from the last article: plan, act, observe, repeat. Every one of those steps is a place where something can quietly go wrong — a bad tool call, a misread result, an unnecessary extra loop. Without evaluation built in, you find out about these failures from a frustrated user, not from your own testing. The more moving parts a GenAI system has, the more it needs deliberate evaluation, not less — which is exactly the opposite of how most people’s instincts work when something feels advanced enough to “just trust it.”

Mental Model

Think of evaluation the same way you’d think about testing any other piece of software you ship — except the test doesn’t just check “did it run without crashing,” it checks “was the actual output good, correct, and safe.” You wouldn’t ship a payment system because it “felt right” after trying it a couple of times. GenAI systems deserve the same discipline, even though the output is language instead of a number — it’s still something you can and should measure against a real standard.

Key Takeaways

  • “It looked good” is not evaluation — a handful of manually reviewed examples doesn’t tell you how a system performs across real, varied usage.
  • Define what “good” means for your system upfront: accuracy, relevance, groundedness, safety, and consistency are common dimensions, though not every system needs all five.
  • Evaluation combines human review, automated checks, and increasingly model-graded scoring — each has trade-offs, and they’re usually used together.
  • Groundedness checks are a direct, measurable answer to hallucination — don’t just hope RAG fixed it, verify that claims trace back to real sources.
  • The more complex a system gets — especially agents with multiple steps — the more evaluation matters, not less.

That closes out the GenAI Foundations series — from what generative AI actually is, through how LLMs work, tokens, prompting, embeddings, RAG, fine-tuning, hallucinations, agents, and now evaluation. Together, this is the full arc: understanding the model, feeding it the right information, letting it act, and confirming it’s actually working.