dbt (data build tool)

A transformation framework that lets data teams write SQL to define, test, and document data models in the warehouse.

dbt has become the standard transformation layer in modern data stacks, sitting between raw data ingestion (Fivetran, Airbyte) and BI tools (Looker, Tableau). It brings software engineering practices — version control, testing, documentation — to SQL-based data transformation. dbt proficiency is now table stakes for data engineering and analytics engineering roles.

Typical time to job-readiness: ~4 weeks.

Learning dbt (data build tool)

Beginner

Understand what dbt does — it transforms data in-warehouse using SQL and version control. Write your first model, run it, and understand refs and sources.

Intermediate

Build modular model layers (staging → intermediate → marts), write data tests (not_null, unique, accepted_values, relationships), and use macros and packages.

Advanced

Incremental models for large tables, multi-project dbt environments (dbt Mesh), and designing warehouse architecture with a team. Data engineer interviews frequently include dbt schema design questions.

Key concepts

  • Models — SQL SELECT statements that define a transformed table or view in the warehouse
  • ref() function — declares dependencies between models so dbt builds in order
  • Sources — declarations of raw tables that dbt reads from but doesn't own
  • Tests — schema tests (not_null, unique) and custom SQL tests on model output
  • Materializations: view, table, incremental, ephemeral — control how models are stored
  • Lineage graph — dbt generates a DAG showing data flow from source to mart

Common interview topics

  • What is dbt and what problem does it solve in a data stack
  • Explain the difference between a view and a table materialization in dbt
  • How does dbt's ref() function work and why does it matter
  • How would you test data quality in dbt
  • What is an incremental model and when would you use one

Browse dbt (data build tool) jobs