Data Engineering & Analytics
Operational databases are built to run the product. Analytics needs something different: years of history, joins across every system the company owns, and queries that scan billions of rows. Data engineering is the discipline that bridges the two — extracting data from sources, transforming it into trustworthy models, and serving it to dashboards, analysts, and machine learning.
This hub covers the modern data stack end to end. Individual database engines live in Databases; streaming transport lives in Messaging & Event Streaming; modeling and ML live in Data Science & ML.
TL;DR
- ELT has mostly replaced ETL. Load raw data into cheap cloud storage first, then transform it with SQL.
- Separate storage from compute. Warehouses and lakehouses scale each independently.
- Orchestrate, don't cron. Tools like Airflow give you dependencies, retries, and backfills.
- Version your transformations with dbt so models are tested and reviewed like code.
- Open table formats like Iceberg let many engines share one copy of data.
- Data quality is a product feature. Test it, own it, and fix it at the source.
The Modern Data Stack
Featured Topics
Foundations
- Data Engineering — Pipelines, ETL vs ELT, data contracts, and the role itself
- Data Quality Management — Rules, ownership, and fixing defects at the source
Storage & Query
- Data Warehousing — Columnar storage and dimensional modeling
- Data Lakehouse — Lake economics with warehouse guarantees
- Apache Iceberg — The open table format for lakehouses
- Snowflake — The warehouse that separated storage and compute
- Trino & Presto — Federated SQL over data where it lives
Processing & Orchestration
- dbt — SQL transformations with tests, refs, and lineage
- Apache Spark — Distributed processing for data too big for one machine
- Apache Airflow — DAG-based pipeline orchestration
Warehouse or Lakehouse?
Common Mistakes
🚫 Transforming in the ingestion tool — Business logic hidden in connectors is untested and unreviewable. Load raw, transform in dbt.
🚫 No ownership of datasets — When a dashboard breaks, nobody knows who fixes the source. Assign owners.
🚫 Full reloads forever — Fine at 10k rows, ruinous at 10 billion. Learn incremental models early.
🚫 Spark for small data — A single DuckDB or warehouse query often beats a cluster. Distribute only when you must.
🚫 Skipping data tests — Nulls, duplicates, and broken joins silently corrupt every downstream number.
Learning Path
Beginner
Learn SQL deeply, load a public dataset into a warehouse, and build a star schema following Data Warehousing.
Intermediate
Model it with dbt including tests, then schedule it with Airflow and add incremental loads.
Advanced
Design a lakehouse on Iceberg, query it from Trino and Spark, and run a data quality program with SLAs.
Related Topics
- Databases — Engines including DuckDB, ClickHouse, and PostgreSQL
- Change Data Capture — Streaming operational changes into the platform
- Data Science & ML — Consumers of well-modeled data
- Compliance & Privacy — Retention, PII, and governance rules