Ketch and Trino

Discover personal data across a distributed SQL query engine, with an authentication path that matches however your specific Trino cluster is actually deployed.

Trino

About Trino

Trino queries across multiple, often heterogeneous data sources through a single distributed SQL engine, which means personal data discovered "in Trino" may actually be sourced from any number of underlying systems it federates. How Ketch authenticates to Trino also depends on the deployment: a standalone cluster typically uses ordinary username and password credentials, while a cluster running within Google Dataproc can authenticate through a Google service account instead.

Ketch connects to Trino through a Ketch Transponder, supporting both paths.

Capabilities

How Ketch works with Trino

Discovery

Automatically finding catalogs, schemas, tables, and columns, using `SELECT` access against the `information_schema` of the target catalogs and schemas. Two authentication options are supported. For standard deployments, username and password authentication works against the Trino coordinator directly, scoped to a specific catalog and, optionally, a specific schema within it. For clusters running within Google Dataproc specifically, a Google service account with a JSON key, added to the Dataproc Service Agent role, authenticates instead.

With Ketch, teams can

  • Discover personal data across Trino catalogs and schemas, whichever specific data sources Trino happens to federate underneath
  • Authenticate using standard username and password credentials, or a Google service account for Dataproc-hosted clusters
  • Scope discovery to a specific catalog, and optionally a specific schema within it, rather than an entire cluster by default

The gap

The problem this integration solves

A distributed query engine spanning multiple underlying data sources, deployed in different ways across different organizations, creates real authentication and scoping questions:

01. A single, fixed authentication method doesn't fit every way Trino actually gets deployed, standalone versus within managed infrastructure like Dataproc

02. Personal data surfaced through Trino may live in any number of federated underlying systems, and knowing it appears in a Trino query result is a starting point, not the full picture of where it's actually stored

03. Manually reviewing Trino-accessible tables and views for personal data doesn't scale

Ketch resolves the authentication problem by supporting both a standard username/password path and a Google service account path for Dataproc specifically, and addresses the third by connecting continuously through the Transponder rather than a one-time manual review.

Why Ketch

Why teams choose Ketch for Trino privacy compliance

Permissioning infrastructure that governs Trino the same way it governs every other system in your stack — not a one-off connector bolted onto a banner.

  • Two authentication paths, matched to real deployment patterns

    Standard username and password for typical clusters, Google service account credentials for Trino running within Dataproc.

  • Backed by enforcement precedent

    Regulators increasingly expect businesses to prove technical enforcement, not just describe it on paper, the same underlying expectation that applies to knowing what personal data a query engine surfaces.

Questions about the Trino integration

Integrations

Pre-built APIs with 1,000+ systems, apps, and models

Ketch ships connectors and SDKs so consent, rights, and policy flow into your CDPs, warehouses, ad platforms, and AI stack — without a custom data pipeline.

Browse All Integrations

See Trino permissioning running end to end

Book a demo to walk through rights, consent, and preference orchestration on your stack — or start free and connect Trino yourself.

Get Started Free

Get started in less than 5 min