Databricks Training
Modules

← All modules

R2 — Databricks DE Field Manual

Why this module exists

The other modules teach the platform from the documentation. This one is different: it's synthesized from ten Data + AI Summit sessions (~6.5 hours) given by the engineers and PMs who build the products — Lakeflow, Auto Loader, Unity Catalog managed tables, declarative pipelines, Spark 4.1, real-time mode — plus two production war stories (a trading-dashboard build and a 10–20-billion-events-per-day architecture).

What conference sessions give you that docs don't: the numbers ("listing 10M files takes 40 minutes; file events do it in one"), the honest caveats ("real-time mode is at-least-once — only adopt it where duplicates are acceptable"), and the decision framing ("stop asking what source am I reading; ask how much management vs. customization do I want"). This module is that material, organized, with a consolidated gotcha list in the cheat sheet.

Status caveat: GA / Preview / roadmap labels here are as stated at the conference — verify current status in the docs before designing around anything not marked GA. Source transcripts live in the shared workspace (/Volumes/hd2_shared_docs/aide-shared/db-training-shared/de/transcripts/).

Lessons

  1. The spectrum: streaming, batch, and the two engines — the Lakeflow mental model, Structured Streaming vs Enzyme, the streaming-vs-batch decision per medallion layer.
  2. Auto Loader, next generation — file events, trigger selection, the small-files fix, cleanSource, schema evolution vs VARIANT, and the observability recipe.
  3. Connectors and managed tables — Lakeflow Connect's catalog and billing gotchas, why managed tables win, and the SET MANAGED migration runbook with its DBR version matrix.
  4. Declarative pipelines, real-time mode, and Spark 4.1 — SDP lands in open-source Spark, how it compares to dbt and Airflow, the sub-second trigger's at-least-once catch, and what ships in 4.1.
  5. Serving and scale: two production stories — sub-second dashboards via the OLTP fast path, Swiggy's L0–L4 layered architecture, and the expert-DE roadmap.

How to use it

Listen straight through once (about 34 minutes), then keep the cheat sheet — especially the 27-item gotcha table — as the pre-design-review checklist. The lab converts a real table with SET MANAGED and proves the file-events claim on your own bucket.