Ketch and Databricks

Find where personal data lives inside a lakehouse platform, and fulfill deletion requests against it directly, managed through the same role-based access control already governing everything else in the workspace.

Databricks

About Databricks

Databricks runs analytics and machine learning workloads across data lakes and warehouses at once, holding data that can span both raw and heavily transformed states. Managing access to that data through the platform's own RBAC model, rather than a workaround, keeps a privacy integration consistent with how every other access decision in the workspace already gets made.

Ketch connects to Databricks through a Ketch Transponder with that consistency in mind.

Capabilities

How Ketch works with Databricks

The Ketch Databricks integration covers Discovery and Rights Orchestration for both Right to Access and Right to Delete.

Discovery

Scans databases, tables, and columns using read-only `SELECT` access.

Rights Orchestration

Requires `UPDATE` and `DELETE` access granted on top of that, scoped to specific tables and columns. Access is controlled through Databricks's own role-based access control, governing permissions on SQL warehouses, catalogs, and tables the same way every other access decision in the workspace is made.

With Ketch, teams can

  • Discover which Databricks tables hold personal data using read-only access
  • Grant rights-execution privileges separately from Discovery, deciding independently whether and when to authorize direct deletion and update
  • Manage the Ketch service account through Databricks's own RBAC model, consistent with how every other access decision in the workspace is made

The gap

The problem this integration solves

A lakehouse platform spanning both raw and transformed data still needs personal data actually located and reachable through the platform's own access model:

01. Manually reviewing tables and columns for personal data across a lakehouse doesn't scale and goes stale as data pipelines evolve

02. A privacy integration managed outside the platform's own RBAC model creates an inconsistent access pattern worth avoiding

03. Fulfilling a deletion request against a lakehouse has historically meant a manually written query, slow and hard to prove complete

Ketch resolves this by connecting through Databricks's own RBAC and access token model, and by separating discovery from execution as two distinct privilege grants.

Why Ketch

Why teams choose Ketch for Databricks privacy compliance

Permissioning infrastructure that governs Databricks the same way it governs every other system in your stack — not a one-off connector bolted onto a banner.

  • Managed through Databricks's own RBAC

    The Ketch service account's permissions on SQL warehouses, catalogs, and tables are governed the same way every other access decision in the workspace is.

  • Least-privilege by design

    Discovery requires only `SELECT`; rights execution is a separate, explicit grant.

  • Backed by enforcement precedent

    Regulators increasingly expect businesses to prove technical enforcement, not just describe it on paper, the same underlying expectation that applies to knowing what personal data a lakehouse holds.

Questions about the Databricks integration

Integrations

Pre-built APIs with 1,000+ systems, apps, and models

Ketch ships connectors and SDKs so consent, rights, and policy flow into your CDPs, warehouses, ad platforms, and AI stack — without a custom data pipeline.

Browse All Integrations

See Databricks permissioning running end to end

Book a demo to walk through rights, consent, and preference orchestration on your stack — or start free and connect Databricks yourself.

Get Started Free

Get started in less than 5 min