The Modern Data Stack, Mapped Out
You’ve learned about data pipelines. You understand data types. You know about data models and data quality. You can explain ETL vs ELT, warehouses vs lakes, and batch vs streaming.
But here’s the thing: all of these are just components. The real power comes from understanding how they fit together. That’s what the modern data stack is: it’s the complete architecture that takes raw data from everywhere and turns it into insights that drive decisions.
Think of it like a city’s water system
Imagine every building in a city needs water.
- Source: Water comes from rivers, lakes, wells (raw data from apps, databases, APIs)
- Treatment plant: Water gets cleaned, filtered, made safe to drink (your pipeline with quality checks and transformations)
- Storage: Cleaned water goes into reservoirs and distribution centers (your warehouse or lake)
- Plumbing: Pipes carry water to different buildings (data access layer: dashboards, APIs, reports)
- Monitoring: Engineers check water quality constantly, alert if something’s wrong (data quality monitoring)
The city doesn’t work if you skip any step. Water isn’t useful if it’s not clean. A big reservoir doesn’t help if there’s no plumbing to deliver it. And none of it matters if nobody monitors for contamination.
That’s the modern data stack.
The Seven Layers of the Modern Data Stack
The modern data stack flows through these seven layers:
Layer 1: Data Sources
- Applications (e-commerce platforms, mobile apps)
- Databases (production systems, customer records)
- APIs and Services (payment processors, third-party tools)
- Sensors and Events (IoT devices, clickstream data)
Layer 2: Ingestion
- Batch Schedulers (Airflow, dbt Cloud)
- Message Queues (Kafka, Pub/Sub)
- API Connectors (Fivetran, Airbyte)
Layer 3: Transformation
- ETL/ELT tools (Spark, dbt, Dataflow)
- Quality Checks (validation, testing)
Layer 4: Storage
- Data Warehouse (Snowflake, BigQuery, Redshift)
- Data Lake (S3, ADLS)
- Lakehouse (Delta Lake, Apache Iceberg)
Layer 5: Semantic Layer
- Data Modeling (Star Schema, dimensional tables)
- Metric definitions and business logic
Layer 6: Analytics & Compute
- BI Tools (Tableau, Looker, Mode)
- SQL Analysis and Python notebooks
- Machine Learning models
Layer 7: Governance & Monitoring
- Access Control and security
- Documentation and data lineage
- Monitoring and alerting
Data flows downstream from Layer 1 through Layer 7, with each layer adding structure, reliability, and accessibility.
Let’s walk through each layer and see how the concepts you’ve learned fit in.
Layer 1: Data Sources
Where it all begins. Data lives everywhere:
- Applications: Your e-commerce platform, mobile app, SaaS tool
- Databases: Production systems tracking transactions, customer records
- APIs and Services: Third-party data (payment processors, marketing tools, weather services)
- Sensors and Events: IoT devices, clickstream data, server logs
This is where your structured, semi-structured, and unstructured data comes from. A database gives you clean structured data. An API gives you JSON (semi-structured). A sensor gives you a raw event stream.
Example: An e-commerce company has orders in a database (structured), clickstream events from the website (semi-structured JSON), and customer photos (unstructured images).
Layer 2: Ingestion
Getting data out of source systems and into your pipeline.
This is where extraction happens. Tools here decide when and how data moves:
- Batch schedulers (Airflow): “Every night at 2 AM, grab yesterday’s data from the database”
- Message queues (Kafka): “Stream events continuously as they happen”
- API connectors: “Poll this API every hour for new data”
This is where your batch vs streaming choice matters. Batch ingestion is simpler but has latency. Streaming ingestion is complex but gives fresh data.
Example: Orders are batch-ingested nightly (simple, sufficient for daily reports). Customer clicks are streamed continuously (needed for real-time fraud detection).
Layer 3: Transformation
Making raw data usable. This is where transformation happens and data quality gets enforced.
- Transformation logic: Convert date formats, join customer data with orders, aggregate daily sales
- Quality checks: “If this amount is negative, something’s wrong”
- Validation: “If we expect 50K rows and only got 100, alert the team”
This is critical because bad data here flows downstream to your warehouse and dashboards.
Example: Raw orders might have a customer ID of 0 (invalid), dates in mixed formats, and duplicate rows. The transformation layer cleans it all, validates it, and only loads good data.
Layer 4: Storage
Where processed data lives. Three options, each with trade-offs:
- Data Warehouse (Snowflake, BigQuery): Structured, optimized for fast queries, expects clean data
- Data Lake (S3, ADLS): Stores anything, flexible, cheap, but you have to organize it yourself
- Lakehouse (Delta, Iceberg): Tries to give you both: data lake flexibility with warehouse performance
Your choice depends on your data type:
- Structured data → Warehouse (ready to query immediately)
- Semi/Unstructured data → Lake (store everything, figure it out later)
Example: Sales orders (structured) go to the warehouse. Customer photos (unstructured) go to the lake.
Layer 5: Semantic Layer
Organizing warehouse data so queries make sense. This is data modeling.
The semantic layer creates the bridge between raw tables and what analysts actually need:
- Fact tables: Numbers you measure (sales amount, click count)
- Dimension tables: Descriptions (who, what, when, where)
- Pre-calculated metrics: Revenue, customer count (instead of calculating from scratch each time)
A good star schema means:
- Queries are fast (the work is already done)
- Analysts don’t need to know how tables connect (the model handles it)
- Definitions are consistent (everyone agrees on what “revenue” means)
Example: Instead of analysts manually joining ORDERS, CUSTOMERS, and PRODUCTS tables, you provide a pre-built SALES_FACT table ready to query.
Layer 6: Compute & Analytics
Actually using the data. Three main approaches:
- BI Tools (Tableau, Looker, Mode): Visual dashboards and reports. Business controllers check these every morning.
- SQL Analytics: Analysts write queries to explore data and answer specific questions
- Machine Learning: Models that predict future behavior (churn prediction, recommendation engines)
All three pull from the semantic layer or warehouse, trusting that data quality checks have already caught problems.
Example: A dashboard shows daily revenue (BI tool). An analyst queries “what’s our customer retention rate?” (SQL). A model predicts which customers will churn (ML).
Layer 7: Governance
Keeping everything trustworthy and secure.
- Access Control: Who can see what data? (Finance sees revenue, not individual purchases)
- Documentation: What does this field mean? How is it calculated? When was it last updated?
- Monitoring & Alerts: Is the pipeline running? Is data arriving on time? Did something break?
This is where data quality monitoring lives. If the pipeline fails or data quality drops, alerts fire and the team investigates.
Example: An alert fires if yesterday’s orders never loaded. The team checks the pipeline, finds the source database was down, and waits for it to come back online.
A Complete End-to-End Flow
Let’s trace data through all seven layers for an e-commerce company:
flowchart TD
A["📱 Customer places order<br/>in mobile app"] -->|Layer 1: Source| B["🗄️ Order in production DB"]
B -->|Layer 2: Ingestion<br/>Batch @ 2 AM| C["Raw order files<br/>from extraction"]
C -->|Layer 3: Transformation| D["✅ Clean & validate<br/>Check for duplicates<br/>Join with customer data"]
D -->|Layer 4: Storage| E["📊 Data Warehouse<br/>SALES_FACT table"]
E -->|Layer 5: Modeling| F["🎯 Star Schema<br/>Pre-built joins<br/>Pre-calculated metrics"]
F -->|Layer 6: Analytics| G["📈 Dashboard<br/>Business controller sees<br/>revenue by category"]
G -->|Layer 7: Governance| H["⚠️ Monitoring alerts<br/>if data doesn't arrive<br/>or quality drops"]
How These Concepts Connect
Here’s how everything you’ve learned actually works together:
| Concept | Layer | Purpose |
|---|---|---|
| Data Pipelines | 2-3 | Move and transform data automatically |
| Data Types | 1, 4 | Determine which storage and tools you need |
| ETL vs ELT | 3 | Choose when and where to transform |
| Warehouse vs Lake | 4 | Choose where to store based on data type and use case |
| Batch vs Streaming | 2 | Choose timing based on how fresh data needs to be |
| Data Modeling | 5 | Organize warehouse so queries are fast and clear |