Evaluating GenAI Systems (Why 'It Feels Good' Isn't Enough)

GenAI systems need real evaluation — accuracy, relevance, safety — just like any other engineering system, not vibes-based judgment.

August 19, 2026 · 5 min read

AI Agents 101: From Chatbot to Agent

The shift from a model that answers questions to one that takes real actions using tools — the plan, act, observe, repeat loop.

August 18, 2026 · 5 min read

Hallucinations: Why They Happen and How to Reduce Them

The honest explanation for why LLMs make things up, and practical ways to reduce it — grounding, citations, and a few other levers.

August 17, 2026 · 5 min read

The Modern Data Stack, Mapped Out

The Modern Data Stack, Mapped Out You’ve learned about data pipelines. You understand data types. You know about data models and data quality. You can explain ETL vs ELT, warehouses vs lakes, and batch vs streaming. But here’s the thing: all of these are just components. The real power comes from understanding how they fit together. That’s what the modern data stack is: it’s the complete architecture that takes raw data from everywhere and turns it into insights that drive decisions. Think of it like a city’s water system Imagine every building in a city needs water. Source: Water comes from rivers, lakes, wells (raw data from apps, databases, APIs) Treatment plant: Water gets cleaned, filtered, made safe to drink (your pipeline with quality checks and transformations) Storage: Cleaned water goes into reservoirs and distribution centers (your warehouse or lake) Plumbing: Pipes carry water to different buildings (data access layer: dashboards, APIs, reports) Monitoring: Engineers check water quality constantly, alert if something’s wrong (data quality monitoring) The city doesn’t work if you skip any step. Water isn’t useful if it’s not clean. A big reservoir doesn’t help if there’s no plumbing to deliver it. And none of it matters if nobody monitors for contamination. ...

August 10, 2026 · 6 min

The Modern Data Stack, Mapped Out

The Modern Data Stack, Mapped Out You’ve learned about data pipelines. You understand data types. You know about data models and data quality. You can explain ETL vs ELT, warehouses vs lakes, and batch vs streaming. But here’s the thing: all of these are just components. The real power comes from understanding how they fit together. That’s what the modern data stack is — it’s the complete architecture that takes raw data from everywhere and turns it into insights that drive decisions. Think of it like a city’s water system Imagine every building in a city needs water. Source: Water comes from rivers, lakes, wells (raw data from apps, databases, APIs) Treatment plant: Water gets cleaned, filtered, made safe to drink (your pipeline with quality checks and transformations) Storage: Cleaned water goes into reservoirs and distribution centers (your warehouse or lake) Plumbing: Pipes carry water to different buildings (data access layer — dashboards, APIs, reports) Monitoring: Engineers check water quality constantly, alert if something’s wrong (data quality monitoring) The city doesn’t work if you skip any step. Water isn’t useful if it’s not clean. A big reservoir doesn’t help if there’s no plumbing to deliver it. And none of it matters if nobody monitors for contamination. ...

August 10, 2026 · 8 min

Data Quality: What Can Actually Go Wrong

Data Quality: What Can Actually Go Wrong You’ve built an elegant data pipeline. Your data model is clean. Queries are fast. Your dashboard shows revenue up 50% this quarter. The Business controller makes decisions based on it. And then, three weeks later, you discover: the revenue calculation was wrong the whole time because someone entered “500,000” as “500,00” in a legacy system, and nobody caught it. That’s a data quality problem. And here’s the scary part: nobody noticed until you dug deeper. Your pipeline ran perfectly. Your warehouse loaded perfectly. Your SQL was correct. But the data itself was broken, silently, invisibly, for weeks. Think of it like a supply chain Imagine you’re running a restaurant and ordering ingredients. Good supply chain: You order 100 tomatoes. They arrive. You inspect them: most look good, but you spot 3 that are rotten. You remove them. You use 97 good tomatoes. Dinner is great. Bad supply chain: You order 100 tomatoes. They arrive in unmarked boxes. You dump them in the fridge without checking. You don’t know if any are rotten until a customer bites into a bad one during dinner. Now you have an angry customer and a problem. Your data pipeline is the supply chain. If you don’t check quality before it reaches your customers (dashboards, reports, decisions), you’ll serve them rotten data without knowing it. ...

August 9, 2026 · 6 min

Data Quality: What Can Actually Go Wrong

Data Quality: What Can Actually Go Wrong You’ve built an elegant data pipeline. Your data model is clean. Queries are fast. Your dashboard shows revenue up 50% this quarter. The Business controller makes decisions based on it. And then, three weeks later, you discover: the revenue calculation was wrong the whole time because someone entered “500,000” as “500,00” in a legacy system, and nobody caught it. That’s a data quality problem. And here’s the scary part: nobody noticed until you dug deeper. Your pipeline ran perfectly. Your warehouse loaded perfectly. Your SQL was correct. But the data itself was broken — silently, invisibly, for weeks. Think of it like a supply chain Imagine you’re running a restaurant and ordering ingredients. Good supply chain: You order 100 tomatoes. They arrive. You inspect them — most look good, but you spot 3 that are rotten. You remove them. You use 97 good tomatoes. Dinner is great. Bad supply chain: You order 100 tomatoes. They arrive in unmarked boxes. You dump them in the fridge without checking. You don’t know if any are rotten until a customer bites into a bad one during dinner. Now you have an angry customer and a problem. Your data pipeline is the supply chain. If you don’t check quality before it reaches your customers (dashboards, reports, decisions), you’ll serve them rotten data without knowing it. ...

August 9, 2026 · 8 min

What Is Data Modeling (and Why It Matters)

What Is Data Modeling (and Why It Matters) You’ve learned about data pipelines, data types, and how data flows from source to warehouse. But once data lands in your warehouse, you face a new question: how should you organize it? That’s where data modeling comes in. A data model is like an architectural blueprint for your data: it decides which tables exist, how they connect to each other, and what information goes where. Get it right, and queries are fast and answers are clear. Get it wrong, and you’re stuck rewriting everything later. Think of it like designing a restaurant Imagine you’re opening a restaurant and need to keep track of everything. Bad design: You throw all information into one giant notebook. Every page has customer names, orders, menu items, and prices all mixed together. When your manager asks “How much revenue did we make yesterday?” you have to flip through hundreds of pages, manually collecting and calculating. Slow, error-prone, and painful. Good design: You have separate notebooks organized like a filing system. One notebook for customers (name, phone, email), one for orders (date, customer ID, total), one for menu items (name, price, ingredients). When you need to answer a question, you know exactly where to look. Customer name? Check the customer notebook. Find their orders? Match the customer ID in the orders notebook. Quick, organized, reliable. That’s the difference between having no data model and having a good data model. ...

August 8, 2026 · 5 min

What Is Data Modeling (and Why It Matters)

What Is Data Modeling (and Why It Matters) You’ve learned about data pipelines, data types, and how data flows from source to warehouse. But once data lands in your warehouse, you face a new question: how should you organize it? That’s where data modeling comes in. A data model is like an architectural blueprint for your data — it decides which tables exist, how they connect to each other, and what information goes where. Get it right, and queries are fast and answers are clear. Get it wrong, and you’re stuck rewriting everything later. Think of it like designing a restaurant Imagine you’re opening a restaurant and need to keep track of everything. Bad design: You throw all information into one giant notebook. Every page has customer names, orders, menu items, and prices all mixed together. When your manager asks “How much revenue did we make yesterday?” you have to flip through hundreds of pages, manually collecting and calculating. Slow, error-prone, and painful. Good design: You have separate notebooks organized like a filing system. One notebook for customers (name, phone, email), one for orders (date, customer ID, total), one for menu items (name, price, ingredients). When you need to answer a question, you know exactly where to look. Customer name? Check the customer notebook. Find their orders? Match the customer ID in the orders notebook. Quick, organized, reliable. That’s the difference between having no data model and having a good data model. ...

August 8, 2026 · 7 min

Structured vs Semi-Structured vs Unstructured Data

Structured vs Semi-Structured vs Unstructured Data In the last post, we talked about data pipelines and how data flows from source to storage. But here’s something important: not all data looks the same. Some data is clean and organized, some is messy but has a pattern, and some is just… chaos. That’s where structured, semi-structured, and unstructured data come in. And this distinction is the entire reason why data warehouses and data lakes are built differently. Think of it like organizing a filing system Imagine you’re in charge of filing documents for a company. Structured data is like a perfectly organized file cabinet. Every file has the same format: name, date, department, and amount. You know exactly where everything goes, and you can quickly find information. “How much did we spend in sales last month?” You can answer that in seconds by looking at your organized files. Semi-structured data is like a folder of emails. Emails have a subject, sender, date, and content, but some emails have attachments, some don’t. Some have multiple recipients, some don’t. There’s a structure, but it’s flexible. Unstructured data is like a box of printed photographs and handwritten notes. Sure, they all came from your company, but there’s no consistent format. Some photos have dates written on the back, some don’t. Some notes are one line, others are pages long. You can read them, but you can’t instantly summarize them. Same company, same filing system, three very different types of information. ...

August 7, 2026 · 6 min