Databricks interview questions in 2026 test whether you can run a governed data and AI platform, not whether you can recite what a cluster is. Interviewers want to hear how you choose between serverless and classic compute, design Unity Catalog permissions, build Lakeflow pipelines that survive schema changes, control DBU spend, and ship a RAG or agent application on the same governed data. This guide collects 60 high-value questions with model answers, from fundamentals to 13 production scenarios.
How to use this guide
Databricks renamed a lot between 2025 and 2026. Using the current names, and knowing the old ones, tells an interviewer that your knowledge is recent. Each answer below says "previously X" where a product was renamed. Generic Spark and lakehouse theory (shuffles, skew, join strategies, the Delta transaction log, Iceberg versus Hudi) is covered in our data engineering interview questions. This page stays on what is specific to Databricks.
- Freshers and early-career engineers: the workspace and compute types, DBUs, the Unity Catalog namespace, managed versus external tables, Auto Loader basics and the three pipeline dataset types.
- Mid-level Databricks data engineers: liquid clustering, predictive optimization, expectations, AUTO CDC, Lakeflow Jobs triggers, SQL warehouse sizing and Declarative Automation Bundles.
- Senior engineers and architects: ABAC and PII governance, migrating from the Hive metastore, cost attribution with system tables, network security, and GenAI on Databricks: AI Search, Model Serving, Agent Bricks, MLflow tracing and Unity Gateway.
Feature status (Beta, Public Preview, GA) changes often. In an interview, say "check the current documentation for status" rather than guessing.
- Platform fundamentals (Q1βQ9)
- Unity Catalog and governance (Q10βQ17)
- Delta Lake on Databricks (Q18βQ22)
- Lakeflow: ingestion, pipelines and jobs (Q23βQ31)
- Databricks SQL, AI/BI and Genie (Q32βQ34)
- MLflow, Model Serving and GenAI (Q35βQ42)
- CI/CD, cost and security (Q43βQ47)
- Real-world scenario questions (Q48βQ60)
- Key takeaways
- Interview preparation checklist
- FAQ
Platform fundamentals
1. Explain the Databricks architecture: what are the control plane and the compute plane?
Answer: Databricks splits into a control plane and a compute plane. The control plane is run by Databricks in its own cloud account. It hosts the web application, notebooks and workspace metadata, the job scheduler, cluster management and Unity Catalog's metadata services. The compute plane is where your data is processed. For classic compute, it runs in your own cloud account (your VPC or VNet), so clusters read your storage directly. For serverless compute, it runs in a compute plane that Databricks manages, isolated per workspace and connected to your storage. Your data at rest normally stays in your own object storage (S3, ADLS Gen2, GCS).
2. What is a workspace, and how does it relate to the Databricks account?
Answer: The account is the top-level container for billing, identity (users, groups, service principals), Unity Catalog metastores and workspace creation. A workspace is the environment where teams work: notebooks, Git folders (previously Repos), jobs, pipelines, dashboards, compute and serving endpoints. Large enterprises commonly run several workspaces, such as dev, test and prod per business unit, attached to one regional Unity Catalog metastore. Data governance then lives at account and metastore level, and the workspace becomes mostly a unit of isolation for compute, code and people.
3. Compare the compute types: all-purpose (classic) compute, job compute, serverless compute and SQL warehouses.
Answer: Each type matches a workload pattern:
| Compute | Used for | Key trait |
|---|---|---|
| All-purpose (classic) compute | Interactive notebooks, exploration | You configure node types, autoscaling and auto-termination. Shared by users and billed while running. |
| Job compute | Scheduled production jobs | Created for a job run and terminated after it. Billed at a lower rate than all-purpose compute. |
| Serverless compute | Notebooks, Lakeflow Jobs, Lakeflow pipelines | Databricks provisions and scales it. Fast start, no cluster tuning, fewer knobs. |
| SQL warehouses | SQL, BI, dashboards, Genie | Serverless, pro or classic. Built for concurrent SQL with Photon. |
4. What are the access modes on classic compute, and why do they matter for Unity Catalog?
Answer: Classic compute has two Unity Catalog access modes. Standard (previously "shared") is multi-user compute where Lakeguard isolates users from each other, so it can safely enforce row filters, column masks and per-user permissions. Dedicated (previously "single user") is assigned to one user or group. It supports more low-level workloads, such as some ML libraries and RDD APIs, and Databricks now supports fine-grained access control on dedicated compute too. The access mode decides which code can run and how governance is enforced.
5. What is a DBU, and how is Databricks billed?
Answer: A Databricks Unit (DBU) is a normalised unit of processing capability per hour. The DBUs a workload consumes depend on the compute type and size, and the DBU rate depends on the product (all-purpose, jobs, SQL, serverless, Model Serving and so on). For classic compute you pay Databricks for DBUs and pay your cloud provider separately for the VMs. Serverless prices bundle the infrastructure into the Databricks charge. Usage is recorded in the system.billing.usage system table, which you can join with list prices to attribute cost.
Interview tip: Do not quote rates from memory. They differ by cloud, region, tier and contract.
6. What is the Databricks Runtime, and when do you pick an LTS or ML runtime?
Answer: The Databricks Runtime is the versioned software image on classic compute. It bundles Apache Spark, Delta Lake, Photon support, Python and Java libraries, and Databricks optimizations. LTS (long-term support) versions get maintenance updates for longer, so production jobs normally pin an LTS version and upgrade on a planned cycle. ML runtimes add common ML libraries and GPU support. Serverless compute uses Databricks-managed versions and environment versions, so you choose fewer things but give up some control over the exact library set.
7. What is Photon, and when does it help?
Answer: Photon is the Databricks-native vectorized query engine, written in C++. Spark's Catalyst optimizer still plans the query, and Photon takes over execution for supported operations, processing data in columnar batches. Unsupported operations fall back to the Spark runtime. Photon helps most with longer SQL and DataFrame workloads: large scans, joins, aggregations and Delta writes such as MERGE. Queries that finish in a second or two see little gain, because planning and scheduling dominate. Python UDF-heavy code also benefits less, since the UDF itself is not executed by Photon.
8. What are Git folders, and how do notebooks fit into engineering practice?
Answer: Git folders (previously Repos) sync a workspace folder with a remote Git repository, so notebooks, Python modules and configuration files are version-controlled and reviewed like normal code. Mature teams keep reusable logic in Python files or packages with unit tests, use notebooks as thin entry points, and deploy with Declarative Automation Bundles (see Q43).
9. What does "Delta is the default" mean, and what are managed versus external tables?
Answer: On Databricks, a table is a Delta Lake table unless you specify otherwise. In Unity Catalog, a managed table stores its data in managed storage that Unity Catalog controls. Unity Catalog manages the lifecycle and file layout, the table is eligible for predictive optimization, and dropping it removes the data after a retention period. An external table registers data at a path you control through an external location. Unity Catalog governs access, but you manage the files, and dropping the table leaves the data in place. Databricks recommends managed tables by default. Use external tables when non-Databricks engines must write the files directly, or as a migration step.
Unity Catalog volumes are the governed equivalent for non-tabular files, such as PDFs, images and raw CSV drops. They replace the old pattern of storing files in DBFS root or mounts, which Databricks now describes as deprecated.
Unity Catalog and governance
10. Explain the Unity Catalog object model.
Answer: Unity Catalog is the governance layer for data and AI on Databricks. At the top is the metastore, normally one per region, attached to workspaces. Below it, data and AI assets use a three-level namespace, catalog.schema.object. Objects include tables, views, volumes, functions, registered models, and model and MCP services. Other securables sit directly under the metastore: storage credentials, external locations, connections (for federation and external APIs) and shares. Every object is a securable that you can grant privileges on to users, groups or service principals.
Account
ββ Metastore (per region)
ββ Catalog (e.g. prod_finance)
β ββ Schema (e.g. gold)
β ββ Tables / Views
β ββ Volumes
β ββ Functions
β ββ Models
ββ Storage credentials, External locations
ββ Connections
ββ Shares, Recipients
Interview tip: Explain how you would lay out catalogs. Common patterns are catalog per environment (dev, prod), catalog per business domain, or both (prod_finance). Say why: workspace bindings and ownership are easier at catalog level.
11. How do privileges and inheritance work in Unity Catalog?
Answer: Privileges are granted with SQL (GRANT SELECT ON SCHEMA prod.gold TO `analysts`) or in Catalog Explorer, and they inherit downward. A grant on a catalog applies to every current and future schema and table inside it. To read a table, a principal needs USE CATALOG on the catalog, USE SCHEMA on the schema, and SELECT on the table (or on a parent). Each object has an owner who can manage it, and MANAGE lets non-owners administer grants. Practical rules:
- Grant to groups, never to individuals, and sync groups from your identity provider (for example Microsoft Entra ID through SCIM or automatic identity management).
- Have service principals, not people, own and run production pipelines.
- Grant at schema level for domain teams and at table level only for exceptions.
12. What are storage credentials and external locations?
Answer: A storage credential wraps a cloud identity that can reach storage, such as an IAM role on AWS or a managed identity on Azure. An external location pairs a storage path with a storage credential. Privileges on the external location (READ FILES, WRITE FILES, CREATE EXTERNAL TABLE) control who can use that path. Users never handle cloud keys, and all path access is governed and audited in one place.
13. How do you implement row-level and column-level security?
Answer: There are two layers. Row filters and column masks are SQL UDFs attached to a table. A filter returns a Boolean per row, and a mask returns either the real value or a redacted one depending on, for example, is_account_group_member('pii_readers'). Attribute-based access control (ABAC) scales this: you tag data with governed tags (for example pii=aadhaar) and attach policies at catalog, schema or table level (metastore level is in Beta). Any column carrying the tag is masked automatically, including tables created later. ABAC also supports GRANT policies and, in Beta, DENY policies.
Real-world example: An insurer tags every column holding phone numbers, PAN or Aadhaar. One ABAC policy on the claims catalog masks them for everyone outside the claims_investigators group, so analysts who create new tables do not have to remember to mask.
14. How does lineage work in Unity Catalog, and what is it used for?
Answer: Unity Catalog captures lineage automatically for queries run on Databricks, down to column level, across all workspaces attached to a metastore. It records which jobs, notebooks, pipelines and dashboards read and write each table. Lineage is visible in Catalog Explorer and queryable through system tables in the system.access schema. Uses include impact analysis before changing a column, root-cause analysis when a report is wrong, and compliance evidence that shows where regulated data flows. Lineage now also covers AI dependencies such as model services governed through Unity Gateway, and external lineage lets you register upstream systems like Salesforce. Know the limitations: lineage is not preserved across renames of catalogs, schemas, tables or columns, and code paths that Databricks cannot observe are not captured.
15. What are system tables, and which ones should a platform engineer know?
Answer: System tables are Databricks-managed tables in the system catalog that expose operational data with SQL. The ones that come up most:
system.billing.usageandsystem.billing.list_pricesfor cost.system.access.auditfor the audit log.- Lineage tables in
system.access. - Query history for SQL warehouses.
- Jobs and pipeline tables in
system.lakeflow. - Compute tables for clusters, node types and warehouses.
Access is governed like any other table, so grant it deliberately: the audit log is sensitive.
16. What is Lakehouse Federation, and when would you use it instead of ingestion?
Answer: Lakehouse Federation lets Unity Catalog register an external database, such as PostgreSQL, MySQL, SQL Server, Snowflake or another warehouse, as a foreign catalog through a connection. You can then query it in place with Unity Catalog permissions and lineage. Use it for occasional queries, exploration, or joining a small reference table during a migration. Use ingestion (Lakeflow Connect or a pipeline) when data is queried heavily, needs history, or must not load the source system. Hive metastore federation applies the same idea to a legacy Hive metastore (see Q52).
17. What is OpenSharing (previously Delta Sharing), and how does it differ from giving a partner a login?
Answer: OpenSharing, previously called Delta Sharing, is Databricks' open protocol for sharing data and AI assets across organisations. A provider creates a share (a read-only collection of tables, partitions, views and, for Databricks recipients, volumes, notebooks and models) and grants it to a recipient. With Databricks-to-Databricks sharing, the recipient mounts the share as a catalog in their own Unity Catalog and queries it on their own compute, and both sides get auditing. With Databricks-to-Open sharing, a recipient on any platform reads through a token- or OIDC-based protocol. No data is copied into the partner's tenancy by default, and you never create accounts for them in your workspace. OpenSharing also underpins Databricks Marketplace and Clean Rooms.
Delta Lake on Databricks
18. What is liquid clustering, and why does it replace partitioning and Z-ordering for most tables?
Answer: Liquid clustering (CLUSTER BY (col1, col2)) organises data files by the clustering keys so that queries filtering on those keys skip most files. Unlike Hive-style partitioning, it does not create a directory per value, so you avoid the small-file problem on high-cardinality columns. You can change the keys later without rewriting the table up front, and clustering is applied incrementally by OPTIMIZE. Z-ordering required full rewrites and could not be combined cleanly with partitions. CLUSTER BY AUTO lets predictive optimization choose keys from observed query patterns. After you change keys, OPTIMIZE ... FULL reclusters existing data, which can take hours on a large table.
19. What does predictive optimization do?
Answer: Predictive optimization automatically runs OPTIMIZE (including incremental clustering), VACUUM and ANALYZE on Unity Catalog managed tables, when Databricks judges that they will pay off. It also collects statistics as data is written. This removes most hand-scheduled maintenance jobs, which teams often forget or run too often. It does not apply to external tables, which is one more reason to prefer managed tables.
20. What are deletion vectors, and what trade-off do they introduce?
Answer: Without deletion vectors, a DELETE, UPDATE or MERGE that touches one row rewrites the whole Parquet file containing it. With deletion vectors enabled, Delta writes a small file marking which rows are logically removed, and readers skip those rows. Writes become much faster, especially for frequent small updates and GDPR or DPDP erasure requests. The trade-off is slightly more work at read time until the files are physically rewritten by OPTIMIZE or REORG TABLE ... APPLY (PURGE). For erasure, remember that the data is gone physically only after the rewrite and a VACUUM past the retention window.
21. How do change data feed, time travel and VACUUM interact?
Answer: Change data feed (CDF), when enabled, records row-level inserts, updates and deletes per table version, so downstream jobs can read only the changes (table_changes() or readChangeFeed). Time travel lets you query or restore an earlier version (VERSION AS OF, TIMESTAMP AS OF, RESTORE). VACUUM deletes data files that are no longer referenced and are older than the retention threshold. Once vacuumed, older versions can no longer be read. Retention settings are therefore a business decision: long enough for recovery and downstream CDF consumers, short enough to control storage and to honour deletion obligations.
22. How do Iceberg clients read Delta tables on Databricks?
Answer: Enabling Iceberg reads on a Delta table makes Databricks generate Iceberg metadata asynchronously alongside the Delta log. This feature uses the Universal Format, known as UniForm. A single copy of the Parquet files then serves both Delta and Iceberg clients. Unity Catalog can act as an Iceberg REST catalog so engines such as Trino, Spark or Snowflake read the tables with governance. Databricks also supports managed Iceberg tables natively. The interview point is interoperability without copies, after checking which table features each client supports.
Lakeflow: ingestion, pipelines and jobs
23. What is Lakeflow, and what are its parts?
Answer: Lakeflow is Databricks' umbrella for data engineering, and its parts carry newer names for older products:
- Lakeflow Connect: ingestion connectors. Managed connectors for databases (CDC), SaaS apps, file sources and streaming sources, plus standard connectors such as Auto Loader.
- Lakeflow Spark Declarative Pipelines, which the docs now mostly call Lakeflow pipelines. Previously Delta Live Tables (DLT), and briefly "Lakeflow Declarative Pipelines". This is declarative batch and streaming ETL that extends the open-source Apache Spark Declarative Pipelines.
- Lakeflow Jobs (previously Databricks Jobs, often called Workflows): orchestration of tasks with dependencies, triggers and control flow.
- Lakeflow Designer: a visual, low-code way to build pipelines.
24. How does Auto Loader work, and what are its key options?
Answer: Auto Loader is a Structured Streaming source (cloudFiles) that incrementally processes new files arriving in cloud storage or a Unity Catalog volume. It tracks discovered files in a RocksDB store in the checkpoint location, so each file is processed exactly once and a restart resumes where it stopped. Key points:
- Discovery: directory listing, or file notification and file events, which scale better for large directories.
- Schema:
cloudFiles.schemaLocationstores the inferred schema. Schema evolution modes such asaddNewColumnsandrescuedecide what happens when new columns appear. Unexpected data lands in the_rescued_datacolumn instead of being lost. - Formats: JSON, CSV, XML, Parquet, Avro, ORC, text and binary files.
- Triggers: run continuously or with
availableNowfor incremental batch runs on a schedule.
(spark.readStream.format("cloudFiles")
.option("cloudFiles.format", "json")
.option("cloudFiles.schemaLocation", chk)
.load("/Volumes/raw/pos/landing/")
.writeStream.option("checkpointLocation", chk)
.trigger(availableNow=True)
.toTable("bronze.pos_events"))
25. When would you use a Lakeflow Connect managed connector instead of building your own ingestion?
Answer: Use a managed connector when one exists for your source and your needs are standard: Salesforce, Workday, SharePoint or Google Drive files, or a SQL Server, MySQL or PostgreSQL database through CDC. The connector handles authentication, incremental reads, CDC, schema evolution, retries and API changes, and writes to Unity Catalog streaming tables on serverless compute. Database connectors add an ingestion gateway and staging storage for continuous change capture. Drop to a standard connector (Auto Loader, Kafka, custom Spark code) when you need transformations during ingestion, an unsupported source or unusual control. Databricks' own guidance is to start with the most managed layer and go lower only when it does not meet the need.
26. Explain streaming tables, materialized views and views in a Lakeflow pipeline.
Answer: A pipeline declares datasets, and the engine works out the dependency graph and orchestration.
| Dataset | Processing | Typical use |
|---|---|---|
| Streaming table | Each record processed once, from an append-only source | Ingestion (bronze), incremental appends, AUTO CDC targets |
| Materialized view | Results kept current, recomputed incrementally where possible | Joins, aggregations, gold tables |
| View | Evaluated on demand, not persisted | Intermediate logic and checks |
Flows are the units that write into these datasets, and sinks write to external targets such as Kafka. Streaming tables and materialized views can also be created standalone in Databricks SQL.
Interview tip: A classic mistake is using a streaming table over a source that gets updated or deleted. A streaming table expects append-only input. Use a materialized view, CDF or AUTO CDC instead.
27. What are expectations, and how do you choose between warn, drop and fail?
Answer: Expectations are named data quality constraints on pipeline datasets, written as SQL Boolean expressions, for example CONSTRAINT valid_amount EXPECT (amount >= 0) ON VIOLATION DROP ROW. There are three actions. Warn (the default) keeps the record and records a metric. Drop removes invalid records. Fail stops the update. Choose by business impact. Fail for invariants where wrong data is worse than late data, such as negative ledger balances or a missing primary key in finance. Drop for noise you can safely discard. Warn while you are learning the data. Metrics appear in the pipeline event log, which you can query and alert on.
Production consideration: Dropped rows disappear silently unless you route them somewhere. Many teams write a parallel quarantine table with the inverse condition so the source system's owners can fix the data.
28. How do the AUTO CDC APIs work?
Answer: AUTO CDC (which replaces APPLY CHANGES with the same syntax; the old name still works) applies a change feed to a target streaming table. You give it the keys, a sequencing column to order events, and how to treat deletes. It then produces SCD Type 1 (latest state) or SCD Type 2 (full history with validity ranges). It handles out-of-order events from the sequence column without you writing watermark or MERGE logic. AUTO CDC FROM SNAPSHOT derives changes by comparing full snapshots when the source has no CDC feed.
CREATE FLOW customers_cdc AS AUTO CDC INTO silver.customers
FROM STREAM(bronze.customers_cdc)
KEYS (customer_id)
APPLY AS DELETE WHEN op = 'D'
SEQUENCE BY event_ts
STORED AS SCD TYPE 2;
Interview tip: Name a good sequencing column. A source commit sequence number or log position is better than wall-clock ingestion time, because ingestion order is not change order.
29. What do triggered versus continuous mode and development versus production mode mean for pipelines?
Answer: Triggered pipelines process available data and stop, usually run from a job on a schedule. Continuous pipelines keep running for low latency and cost more because compute stays up. Real-time mode in Spark Declarative Pipelines on Lakeflow serves the very low latency cases (check its current status). Development mode reuses compute between runs and does not retry, which speeds up iteration. Production mode restarts compute on failure and retries progressively, from the Spark task to the flow to the whole pipeline.
30. What can Lakeflow Jobs do beyond a cron schedule?
Answer: Lakeflow Jobs orchestrates tasks: notebooks, Python scripts and wheels, SQL, dbt, pipelines, other jobs and more, with dependencies between them. Features interviewers probe:
- Triggers: scheduled, file arrival in a Unity Catalog location, table update (run when upstream tables change), model update, continuous and manual.
- Control flow: if/else condition tasks, run-if conditions (for example run on failure) and for-each loops over a parameter list.
- Parameters at job and task level, and task values passed between tasks.
- Repair runs that rerun only failed tasks and their dependents. Backfill runs for date ranges.
- Notifications, duration thresholds, and retries per task.
Interview tip: Table-update triggers are a strong answer to "how do you avoid running a job before its upstream data is ready?" Event-driven dependencies replace fragile "run at 6:30 and hope" schedules.
31. When would you use a Lakeflow pipeline versus a job running plain Spark code?
Answer: Use a pipeline when the work is mainly a graph of tables: ingestion, cleaning, CDC and aggregations. You get dependency ordering, incremental materialized views, expectations, lineage and retries without writing them. Use jobs with plain Spark or Python when you need imperative control: calling external APIs, complex per-batch logic in foreachBatch, ML training, or libraries pipelines do not support. Most platforms use both, with a job orchestrating pipelines, SQL tasks and model training.
Databricks SQL, AI/BI and Genie
32. How do you size and scale a SQL warehouse?
Answer: Warehouse size (2X-Small upwards) sets the compute behind one cluster and mainly affects how fast a single heavy query runs. Scaling (minimum and maximum clusters) adds clusters to handle concurrency, since queries queue when all clusters are busy. Serverless warehouses start quickly and scale with intelligent workload management. Pro and classic warehouses run in your account with slower start-up. A working method:
- Start small and read the query profile. Spill to disk or long scans suggest a larger size.
- Rising queue time with many users suggests a higher maximum cluster count, not a larger size.
- Set auto-stop for idle warehouses, and separate ETL and BI warehouses so a heavy refresh does not slow dashboards.
- Use query tags and query history to attribute cost and find the worst queries.
33. What are Genie One, Genie Agents and Genie Code?
Answer: Genie is the family of Databricks AI experiences for users, all grounded in Unity Catalog permissions:
- Genie Agents (previously Genie spaces): domain-specific natural-language interfaces over curated datasets. Analysts configure the tables, example SQL, metrics, business rules and instructions, and users get SQL, results and charts back.
- Genie One: a single interface for business users to ask questions and use dashboards, Genie Agents and Databricks Apps.
- Genie Code (previously Databricks Assistant): the AI coding and data assistant inside notebooks, the SQL editor, pipelines and dashboards.
Genie works within the asking user's permissions, so it cannot reveal a table the user cannot query. Answer quality depends on curation: clean gold tables, column descriptions, Unity Catalog metric views, and verified example queries.
Interview tip: Connect this to text-to-SQL fundamentals. The hard part is semantics, meaning what "active customer" means, not SQL syntax. See our text-to-SQL agent project for the general pattern.
34. What are Unity Catalog metric views, and why do they matter for AI/BI?
Answer: Metric views define business metrics, such as revenue, active users or claim ratio, once in Unity Catalog as a semantic layer: measures, dimensions and the joins behind them. Dashboards, SQL users and Genie Agents then reuse the same definition instead of each re-implementing it slightly differently. For AI/BI this is the single biggest quality lever. A Genie Agent answering from a governed metric definition is far more trustworthy than one inferring "revenue" from raw columns.
MLflow, Model Serving and GenAI
35. How is MLflow used on Databricks, and what changed with MLflow 3?
Answer: MLflow is the open-source platform for the ML and GenAI lifecycle, and Databricks runs a managed version. For classic ML it provides experiment tracking, model logging, evaluation and a model registry. On Databricks the registry lives in Unity Catalog: models are catalog.schema.model securables with versions, aliases (for example @champion), permissions and lineage to training data. MLflow 3 adds LoggedModels, which track a model or agent version across runs and environments, and deployment jobs, which gate promotion with evaluation and approval steps. For GenAI it adds tracing, evaluation with LLM judges and custom scorers, human feedback and prompt management. Traces can be stored and governed in Unity Catalog.
Interview tip: Stages ("Staging", "Production") belong to the old workspace registry. In Unity Catalog you use aliases. Generic MLOps lifecycle questions are covered in our MLOps interview questions.
36. What does Model Serving offer, and how do pay-per-token and provisioned throughput differ?
Answer: Model Serving deploys models as REST endpoints on serverless infrastructure that autoscales. It serves three kinds of model:
- Custom models: MLflow-packaged models such as scikit-learn, XGBoost, PyTorch or agent code.
- Databricks-hosted foundation models through Foundation Model APIs.
- External models from providers, routed through Unity Gateway.
For foundation models there are two pricing modes. Pay-per-token is shared capacity, ideal for development and moderate traffic. Provisioned throughput reserves capacity for production workloads that need predictable performance or fine-tuned variants. Endpoints can log requests and responses to inference tables in Unity Catalog for monitoring and evaluation.
37. What are AI Functions, and when would you use them instead of an endpoint call from Python?
Answer: AI Functions are built-in SQL functions that apply models to data in tables. Task-specific functions include ai_parse_document, ai_extract, ai_classify and others for summarisation, translation and sentiment. ai_query is the general-purpose function for calling a chosen Foundation Model API or serving endpoint with your own prompt. They run from Databricks SQL, notebooks, pipelines and jobs, and they scale batch inference without you managing concurrency or retries. Use them for set-based enrichment, for example classifying a million support tickets or extracting fields from invoices into a silver table. Use an endpoint call from application code for interactive, per-request work. Note that AI Functions are not available on classic SQL warehouses, and model inference may be billed separately from the compute that runs the query.
38. What is Databricks AI Search (previously Vector Search), and what index types does it support?
Answer: AI Search, previously called Databricks Vector Search, is the built-in retrieval service for RAG and similarity search, governed by Unity Catalog. It uses HNSW approximate nearest neighbour search with L2 distance (normalise embeddings if you want cosine ranking), plus hybrid keyword and vector search (BM25 for the keyword part), filtering and reranking. Index choices:
- Delta Sync index with managed embeddings: you point it at a source table and a text column, and Databricks computes embeddings and keeps the index in sync as the table changes.
- Delta Sync index with self-managed embeddings: your pipeline writes the vectors, and the index syncs them.
- Direct Vector Access index: you write and delete vectors through the API, with no source table.
Endpoints come in standard and storage-optimized types, which trade latency against scale and cost. Generic indexing theory is covered in our vector database interview questions.
Interview tip: Mention hybrid search for enterprise corpora full of IDs, SKUs and policy numbers, where pure embeddings miss exact matches. Our guide to hybrid search and reranking explains why.
39. What is Agent Bricks, and how does it relate to the older Mosaic AI Agent Framework?
Answer: Agent Bricks is now the Databricks agent developer platform for code-first agents. Earlier documentation presented this tooling as the Mosaic AI Agent Framework, and the Mosaic AI brand has largely gone from the current docs. You bring an agent built with any framework or harness, such as LangGraph or the OpenAI Agents SDK, and Agent Bricks supplies the production pieces:
- Agent Runtime: managed hosting.
- Agent server:
DurableAgentServerfrom the AgentKit library, for durable runs that survive restarts. - Databricks Sandbox for running code the agent writes.
- Managed memory backed by Lakebase.
- MCP servers and tools, model access through Unity Gateway, and MLflow tracing and evaluation.
The Agent Bricks CLI scaffolds and deploys agents from an agent.toml definition. For no-code agents over tables and documents, Databricks points to Genie Agents, which a custom agent can also call as tools.
Interview tip: This area changes month to month. Describe the components (runtime, memory, tools, tracing, governance) and say you would confirm current names and release status in the docs. Our agentic AI interview questions cover framework-neutral agent design.
40. What is Unity Gateway, and why put model and MCP access behind it?
Answer: Unity Gateway is Databricks' governance layer for AI traffic. Databricks documentation called this capability AI Gateway until mid-2026; the current docs use Unity Gateway. Built on Unity Catalog, it registers foundation models served in system.ai, external providers (bring your own key), MCP servers, Unity Catalog functions used as tools, and HTTP connections as governed securables. It applies permissions, rate limits, traffic splitting and fallbacks, guardrail service policies, and usage tracking with cost attribution by user, team or project. The value is the same as an API gateway for microservices: one place to enforce who may call which model or tool, with what limits, and to see the spend. See our explainer on the LLM gateway pattern for the vendor-neutral version.
41. How do you evaluate and monitor a GenAI application on Databricks?
Answer: Use MLflow on three levels:
- Tracing: every request records spans for retrieval, model calls and tool calls, with latency, tokens and inputs and outputs.
- Offline evaluation: run the application over an evaluation dataset of real questions with expected facts, scoring with built-in LLM judges (groundedness, relevance, safety), custom scorers and code-based checks. Compare versions before promotion.
- Production monitoring: sample live traces and score them with the same judges. Collect user and expert feedback through review apps and turn failures into new evaluation cases.
Store traces in Unity Catalog so quality, cost and incidents can be analysed with SQL. The method is the same as in our AI agent evaluation guide, and Databricks supplies the plumbing.
42. What are Databricks Apps, and how do they handle identity?
Answer: Databricks Apps hosts data and AI applications (Python with frameworks such as Streamlit, Dash or Gradio, or Node.js with React and similar) on serverless compute inside the workspace, so you do not run separate infrastructure. Typical apps are RAG chat front ends, data-entry forms and operational tools. An app declares resources such as a SQL warehouse, serving endpoint, AI Search index, secret or Lakebase database. It authenticates either as its own app service principal or on behalf of the signed-in user, so Unity Catalog permissions apply per user. For sensitive data, prefer on-behalf-of-user so row filters and masks apply to the real user. Apps are billed per hour while running, and they are deployed through bundles and CI/CD like other code.
CI/CD, cost and security
43. What are Declarative Automation Bundles (previously Databricks Asset Bundles), and how do you use them for CI/CD?
Answer: Declarative Automation Bundles, renamed from Databricks Asset Bundles in 2026, are Databricks' infrastructure-as-code for projects. A databricks.yml file and included YAML define resources: jobs, pipelines, dashboards, serving endpoints, MLflow experiments, registered models and apps, along with source files and per-environment targets (dev, staging, prod) holding workspace URLs, variables and permissions. The Databricks CLI validates, deploys and runs a bundle (databricks bundle validate, deploy -t prod, run). In development mode, resources are prefixed per user so engineers do not overwrite each other.
PR opened ββ> unit tests + bundle validate
merge main ββ> deploy -t staging ββ> integration run
tag/approve ββ> deploy -t prod (service principal)
Production consideration: Production deploys should run from CI (GitHub Actions, Azure DevOps) as a service principal using OAuth, with no personal access tokens and no manual edits in the prod workspace. Use Terraform for workspace, network and metastore infrastructure, and bundles for the project resources inside the workspace.
44. How do you attribute and control Databricks cost?
Answer: Attribution first, then guardrails, then optimisation:
- Attribute: tag compute, jobs, pipelines and warehouses with cost centre and owner. Apply serverless usage policies (previously budget policies) so serverless usage carries tags. Build a dashboard on
system.billing.usagejoined with list prices, by product, workspace, tag and job. - Guardrails: compute policies that cap node types, worker counts and auto-termination. Budgets and alerts. Restrict who can create all-purpose compute.
- Optimise: move scheduled work from all-purpose to job compute or serverless, right-size warehouses, enable auto-stop, use incremental processing instead of full refreshes, and let predictive optimization handle maintenance.
For the general FinOps method, see our guide on cloud cost optimization for AI.
45. Which security controls would you configure for a regulated Databricks deployment?
Answer: Work through identity, network, data and monitoring:
- Identity: SSO with your IdP (Microsoft Entra ID, Okta and others), MFA, SCIM or automatic identity management, and service principals with OAuth for automation.
- Network: a customer-managed VPC or VNet for classic compute, private connectivity for front-end and back-end traffic, IP access lists or context-based ingress control, and serverless egress control to limit where serverless workloads can connect.
- Data: Unity Catalog everywhere, customer-managed keys where required, secrets in secret scopes, and credential redaction.
- Compliance and monitoring: the compliance security profile for standards such as HIPAA or PCI-DSS where applicable, plus audit log monitoring through system tables.
46. How should service principals and tokens be handled?
Answer: Production jobs, pipelines, apps and CI/CD should run as service principals, not as employees, so they survive people changing roles and have narrowly scoped permissions. Prefer OAuth (machine-to-machine client credentials, or workload identity federation from your CI system) over personal access tokens, which Databricks now labels legacy. Where PATs still exist, give them short lifetimes and monitor them. Never store secrets in notebooks. Use secret scopes or Unity Catalog service credentials for cloud services.
47. What is Lakebase, and when does an OLTP database belong in a lakehouse platform?
Answer: Lakebase is Databricks' managed PostgreSQL service for operational, low-latency reads and writes. It can sync with lakehouse tables and is integrated with Unity Catalog and Databricks Apps. Use it where an application needs transactional, row-level access: app state, feature or lookup serving, or agent memory (which Agent Bricks' managed memory uses). Do not use it as the analytical store. Analytical queries still belong on Delta tables and SQL warehouses.
If you want to practise lakehouse pipelines, governance and data-for-AI work hands-on rather than only reading about them, Cloudsoft's HORIZON Data Engineering & AI program covers modern data engineering together with the data foundations that AI systems depend on.
Real-world scenario questions
48. A nightly Databricks job that used to take 40 minutes now takes three hours. How do you investigate?
Answer: Find out what changed (data, code, compute or contention) before tuning anything. Compare the slow run with a good run task by task in the Lakeflow Jobs run history, and then use the Spark UI or query profile for the slow task.
What I would check:
- Which task got slower, and did input volume change? A source that started sending full snapshots instead of increments is a common cause.
- Recent deployments: a new join, a Python UDF that blocks Photon, a
collect(), or an accidental cross join. - In the Spark UI: one stage dominating, a few long tasks (skew), spill to disk, or very many small tasks (small files).
- Table health:
DESCRIBE DETAILfor file count and size. Check whether maintenance stopped (an external table that predictive optimization does not cover) or whether clustering keys no longer match the filters. - Compute: a runtime upgrade, a changed node type, autoscaling capped by a policy, or spot instance loss.
- Contention: concurrent writers causing
MERGEconflicts and retries, or a shared all-purpose cluster with other workloads.
Production consideration: Add a duration threshold notification to the job so you hear about drift before the business does.
49. The monthly Databricks bill jumped sharply. Finance wants an explanation by Friday. What do you do?
Answer: Query system.billing.usage to break the increase down by product (all-purpose, jobs, SQL, serverless, Model Serving, AI Search, Apps), by workspace and by tag or job. Then explain the top contributors with evidence.
What I would check:
- Day-by-day usage to find when the jump started, and match it to deployments or new projects.
- All-purpose clusters with no or long auto-termination, and clusters running jobs that should be on job compute.
- SQL warehouses that never stop because a dashboard refreshes every few minutes, or whose maximum cluster count was raised.
- Pipelines switched from triggered to continuous, or full refreshes replacing incremental runs.
- GenAI spend: a batch
ai_queryjob over a whole table, provisioned throughput endpoints left running, or AI Search endpoints nobody uses. - Untagged usage, which is itself a finding.
Production consideration: Fix the specific cause, then add structural controls: compute policies, tags enforced by policy, serverless usage policies, budgets with alerts, and a weekly cost review per team.
50. A hospital group wants analysts to use patient data on Databricks without exposing PII. Design the governance.
Answer: Classify, tag, enforce with policies, separate duties and audit. Under India's DPDP Act, apply purpose limitation and minimisation to the design as well as access control. Our DPDP Act guide maps the obligations.
What I would check:
- Data classification in Unity Catalog to find columns holding names, phone numbers, Aadhaar, addresses and diagnoses, followed by human review of the results.
- Governed tags on those columns, and ABAC column mask policies at catalog level, so new tables inherit protection.
- Row filters by hospital or region where analysts should see only their facility.
- Separate catalogs for identified and de-identified data. Most analysts get only the de-identified gold catalog with tokenised patient IDs.
- Standard or serverless compute only, so Lakeguard enforces fine-grained policies. No legacy clusters.
- Audit log queries for access to sensitive tables, with alerts on unusual volume.
- Genie Agents, apps and AI features see the same masked data, because they run with user permissions.
Production consideration: Test the policies as you would test code: a CI step that queries masked columns as a test principal and fails if the real values come back.
51. Design a RAG assistant over an insurer's policy documents, built on Databricks.
Answer: Keep ingestion, retrieval, serving and evaluation on governed components, and design for freshness and permissions from day one.
SharePoint/PDFs
β Lakeflow Connect / Auto Loader
βΌ
UC volume (raw files) β> ai_parse_document
βΌ
silver.chunks (text, metadata, ACL tags)
βΌ Delta Sync
AI Search index (hybrid + filters)
βΌ
Agent (Agent Bricks / Model Serving)
β models via Unity Gateway
βΌ
Databricks App (on-behalf-of-user)
βΌ
MLflow traces β> judges β> eval set
What I would check:
- Ingestion: a managed file connector or Auto Loader into a Unity Catalog volume, parsing with
ai_parse_document, and chunking that respects headings and tables. - Permissions: carry document ACLs or business-line tags into chunk metadata and filter retrieval by the user's groups. Retrieval must not return chunks the user could not open.
- Freshness and deletes: a Delta Sync index on the chunks table so updated or withdrawn policies leave the index automatically.
- Retrieval quality: hybrid search for policy numbers and clause IDs, reranking, and citations in every answer.
- Governance: model access, rate limits and guardrails through Unity Gateway. PII handling in prompts and logs.
- Evaluation: a set of real underwriter and claims questions, scored offline before each release and on sampled production traces.
Production consideration: Most RAG failures are data failures: stale documents, bad parsing, missing permissions. The model is rarely the problem. See data pipelines for RAG and our RAG interview questions for the retrieval side.
52. A GCC team must migrate a five-year-old workspace from the Hive metastore to Unity Catalog. Plan it.
Answer: Assess first, migrate in waves by domain, and keep the old and new running side by side until consumers have moved.
What I would check:
- Assessment: run UCX (databrickslabs/ucx) to inventory tables, mounts, cluster types, grants and code incompatible with Unity Catalog, such as RDD use on shared compute, DBFS paths and init scripts.
- Foundations: metastore, catalog design, storage credentials and external locations replacing mounts, account-level groups replacing workspace-local groups.
- Table strategy per table: register external tables in place with
SYNCfor a fast start, deepCLONEor CTAS into managed tables where you want predictive optimization (history and time travel do not carry over), or Hive metastore federation to govern the old metastore while you migrate. - Permissions: translate legacy table ACLs into Unity Catalog grants, ideally simplified to group and schema level.
- Code: replace
hive_metastore.db.tableand mount paths with three-level names and volumes, move jobs to standard, dedicated or serverless compute, and use Genie Code to help find and rewrite references. - Cutover: mark old tables deprecated with the documented comment format, test by revoking access to the old tables, then drop them after a quiet period.
53. A pipeline with expectations suddenly drops a large share of records. What do you do?
Answer: Treat it as a data incident. Check the pipeline event log for which expectation is failing, since when and on which flow, then sample the dropped records from the quarantine table or by re-running the logic in a notebook.
What I would check:
- Did the source change format, units, a date pattern, or start sending nulls for a mandatory field?
- Did someone change the expectation itself in the last deployment?
- Did Auto Loader push unexpected values into
_rescued_dataafter a schema change?
Production consideration: Alert on the drop rate, not only on failures. A DROP ROW expectation that quietly removes a large share of data for a week is worse than a pipeline that fails loudly.
54. An upstream system added and renamed columns in its JSON files, and the Auto Loader stream stopped. How do you handle it?
Answer: With the default addNewColumns mode, Auto Loader stops the stream when it detects new columns, updates the schema in the schema location, and succeeds on restart. That is by design. A rename looks like a new column plus a column that went null, so it needs a decision.
What I would check:
- Whether the job has retries configured. With production pipelines or job retries, the restart picks up the evolved schema automatically.
- Whether the renamed field should map to the old column in silver (coalesce old and new names) to keep downstream contracts stable.
- Whether
rescuemode suits this source better: keep the schema fixed and capture everything unexpected in_rescued_data. - Schema hints for fields that must have a specific type.
Production consideration: Bronze should absorb change and silver should enforce the contract. Agree a data contract with the source team so renames are announced. Our data engineering guide covers contracts in depth.
55. Two jobs write to the same Delta table and one keeps failing with concurrent modification errors. How do you fix it?
Answer: Delta uses optimistic concurrency. Two MERGE, UPDATE or DELETE operations that read and rewrite overlapping files conflict, and one fails with an exception such as ConcurrentAppendException or ConcurrentDeleteReadException.
What I would check:
- Whether the writers really need the same table at the same time. Often they should be serialised in one job, or one should write to a staging table.
- Whether the
MERGEconditions can be made disjoint, for example by explicit region or date predicates, so the operations touch different files. - Whether liquid clustering on the merge key and deletion vectors (which reduce rewritten files) narrow the conflict window. Where supported, row-level concurrency resolves many of these conflicts automatically; check the current requirements in the docs.
- Retry with backoff for the remaining, rare conflicts.
56. Business users say the sales Genie Agent gives wrong numbers for "last quarter's revenue". How do you fix it?
Answer: This is a semantics problem, not a model problem. Review the generated SQL in the conversation history and find where the interpretation diverges: fiscal versus calendar quarter, gross versus net revenue, returns, currency, or the wrong table.
What I would check:
- Whether the agent's datasets are curated gold tables or raw tables with ambiguous columns.
- Column and table comments in Unity Catalog explaining what each field means.
- A Unity Catalog metric view for revenue, so the definition is governed and reused.
- Instructions stating the fiscal calendar, plus verified example queries for common questions.
- A benchmark set of business questions with expected answers, run after every change.
57. Several teams call external LLMs directly from notebooks with personal API keys. How do you bring this under control?
Answer: Centralise model access through Unity Gateway and remove the direct path.
What I would check:
- An inventory of current use from audit logs, code search and egress logs.
- Registering the approved providers once, with keys held centrally, and granting access to model services by group.
- Rate limits and usage tracking per team, with cost attributed by tags.
- Guardrail service policies for PII and unsafe content where the use case needs them.
- Serverless egress control and network rules so notebooks cannot reach provider APIs directly.
- A migration window, then revoking the personal keys.
58. A developer edited a production job in the UI to "fix it quickly", and the next bundle deploy overwrote the fix. How do you prevent this?
Answer: Make Git the only source of truth for production. Deploy production only from CI with a service principal. Remove edit permissions on production jobs for humans, who get view and run access, plus a break-glass procedure for emergencies. Require any hotfix to go through a pull request, even a fast-tracked one. Run databricks bundle validate in CI and review the bundle diff in the pull request, so reviewers see every production change before it is deployed.
59. A retailer must share daily sales data with a supplier who does not use Databricks. How do you do it?
Answer: Use OpenSharing's Databricks-to-Open protocol. Create a share with a view that exposes only that supplier's products and the agreed columns, create a recipient for the supplier, and send the activation link. The supplier reads with an OpenSharing client (pandas, Spark, BI connectors) using a bearer token or OIDC federation.
What I would check:
- Row-level scoping by supplier ID in the shared view, and no customer PII.
- Token lifetime and rotation, or OIDC where the supplier supports it.
- IP access lists on the recipient if their network is known.
- Audit of recipient access through system tables.
- A contract for schema changes, since the share becomes an external interface.
60. Dashboards on a SQL warehouse are slow every Monday morning. What do you investigate?
Answer: Monday morning means concurrency: everyone opens dashboards at once, often right after weekend batch loads. Check the warehouse's monitoring for queued queries, running clusters and peak concurrency, and look in query history for the slowest and most frequent queries.
What I would check:
- Queueing: raise the maximum cluster count (scale out) rather than the size.
- Whether the weekend ETL on the same warehouse is still running. Separate ETL from BI.
- Expensive dashboard queries that scan raw silver tables. Pre-aggregate into materialized views or metric views.
- Gold table health: clustering on the dashboard filter columns, file counts, and statistics.
Key takeaways
- Use current names and know the old ones: Lakeflow pipelines (previously DLT), Lakeflow Jobs, Declarative Automation Bundles (previously Asset Bundles), AI Search (previously Vector Search), OpenSharing (previously Delta Sharing), Genie Agents (previously Genie spaces), Genie Code (previously Databricks Assistant).
- Unity Catalog is the centre of every senior answer: namespace design, group-based grants, ABAC for PII, lineage and system tables.
- Prefer managed tables, liquid clustering and predictive optimization over hand-tuned partitions and maintenance jobs.
- Use the most managed layer that fits: Lakeflow Connect, then pipelines, then custom Spark code.
- Cost questions are answered with
system.billing.usage, tags, policies and the right compute type for each workload. - GenAI on Databricks reuses the same governance: AI Search on Delta tables, models and tools behind Unity Gateway, and quality measured with MLflow traces and judges.
- Scenario answers win interviews: find what changed, show evidence, fix the cause, then add a control so it cannot recur.
Interview preparation checklist
- Create a Databricks Free Edition workspace and build one end-to-end project: Auto Loader into bronze, a Lakeflow pipeline with expectations and AUTO CDC, gold materialized views, and a dashboard.
- Design a Unity Catalog layout for a fictional bank and write the grants, a row filter and an ABAC mask policy.
- Practise reading a Spark UI and a SQL query profile. Be able to explain spill, skew and Photon fallback from a screenshot.
- Write a cost query against
system.billing.usagethat groups by product and tag. - Package the project as a Declarative Automation Bundle with dev and prod targets, deployed from GitHub Actions.
- Build a small RAG app: parse documents, index them with AI Search, serve an agent, and evaluate it with MLflow.
- Read the latest Databricks release notes the week before your interview. Names and feature status change often.
- Prepare two stories from your own work, a performance fix and a governance or cost fix, told as problem, evidence, action and outcome.
- Revise SQL window functions and joins with our SQL interview questions for data and AI.
FAQ
What skills are required for a Databricks data engineer interview?
Strong SQL and Python, PySpark fundamentals, Delta Lake, Unity Catalog permissions, Lakeflow pipelines and jobs, Auto Loader, and performance tuning with the Spark UI. Senior roles add cost management, security, CI/CD with bundles and increasingly data pipelines for GenAI.
How should I prepare for a Databricks interview?
Build one end-to-end project in a free workspace, from ingestion to governed gold tables and a dashboard, and practise explaining every design choice. Then rehearse scenario answers about slow jobs, cost spikes and access control, because those separate candidates more than definitions do.
Is Databricks certification necessary to get a job?
No. Certifications can help your resume pass initial screening, but interviewers judge hands-on reasoning. A well-documented project and clear answers to scenario questions usually carry more weight than a certificate alone.
Are DLT and Workflows still asked about in interviews?
Yes, many interviewers still use the old names. Delta Live Tables is now Lakeflow Spark Declarative Pipelines, usually called Lakeflow pipelines, and Databricks Jobs or Workflows is now Lakeflow Jobs. Answer using the concepts and mention the current names once.
What Unity Catalog topics come up most in interviews?
The three-level namespace, privilege inheritance, managed versus external tables, storage credentials and external locations, row filters and column masks, ABAC with governed tags, lineage, system tables, and migrating from the Hive metastore.
Do Databricks interviews include GenAI questions now?
Increasingly, yes, especially for platform and ML roles. Expect questions on AI Search, Model Serving, AI Functions, Agent Bricks, MLflow tracing and evaluation, and governing model access with Unity Gateway.
Can a fresher get a Databricks role?
Yes, if they have solid SQL, Python and Spark basics and one project built on Databricks that they can explain in depth. Freshers are usually assessed on fundamentals and problem solving rather than years with the platform.
Is Databricks a good skill for Indian engineers?
Databricks is widely used by GCCs, product companies and services firms in Hyderabad, Bengaluru and other cities for data engineering, analytics and AI platforms. Pairing it with strong SQL, a major cloud and data-for-AI skills keeps your options broad.
How long does it take to prepare for a Databricks interview?
It depends on your background. An engineer who already knows Spark and SQL can cover the Databricks-specific topics and build a project in a few focused weeks. A fresher should plan for longer and spend most of the time building and debugging pipelines.
Ready to build production data pipelines and the governed data foundations behind enterprise AI? The HORIZON data engineering and AI course combines hands-on pipeline engineering with AI data work, in classroom sessions in Ameerpet or live online. If you want a broader path across AI, ML, cloud and security, look at the APEX AI, ML, Cloud and Cyber Security program. Call +91 96660 19191 to book a free demo.



