Evaluating GenAI Systems (Why 'It Feels Good' Isn't Enough)
GenAI systems need real evaluation — accuracy, relevance, safety — just like any other engineering system, not vibes-based judgment.
GenAI systems need real evaluation — accuracy, relevance, safety — just like any other engineering system, not vibes-based judgment.
The shift from a model that answers questions to one that takes real actions using tools — the plan, act, observe, repeat loop.
The honest explanation for why LLMs make things up, and practical ways to reduce it — grounding, citations, and a few other levers.
How text gets turned into numbers that capture meaning, and why ‘similar meaning = nearby numbers’ is the foundation of search, recommendations, and RAG.
Why specificity beats cleverness, and how to talk to a model that takes everything you say literally.
What a token actually is, why every model has a memory limit, and why this quietly affects cost, speed, and what you can even ask a model to do.
No math-heavy transformer internals, just the core idea: an LLM is a very sophisticated autocomplete engine trained on huge amounts of text.
The simple difference between predictive AI and generative AI, explained without the jargon.