So I think the security barrier / bypass path avoidance can be handled in a few different ways (I've got one proposal I'm working on a draft off but it's maybe a little early and shouldn't block this work).
On 2026/09/23 16:48:33 Dongjoon Hyun wrote: > Hi, Huaxin. > > Thank you for proposing this SPIP. I support the direction. Moving row > filtering and column masking from private Catalyst hacks to a public DSv2 > contract would be a clear improvement for the ecosystem, and the proposed > semantics are clear. > > Since the value of this SPIP rests on its "fail-closed" and "optimizer-safe" > guarantees, I'd like the SPIP to address the following, grouped by priority. > > ### 1. Security guarantees (should be addressed before a vote) > > - **Security barrier**: The SPIP protects masks from the optimizer, but not > hidden rows from user expressions. Leaky user predicates (e.g. ANSI errors, > Python UDFs, predicates pushed to the connector) can be evaluated before the > policy filter and expose rows that should be hidden. The SPIP should define > the policy filter as a barrier, similar to PostgreSQL's `security_barrier` > and `LEAKPROOF`. > - **DML read path**: DELETE/UPDATE/MERGE also read the target table. Applying > the policy on that read can silently delete hidden rows or write masked > values back. These operations should fail closed. > - **Older Spark versions**: An engine that does not know the new interface > will silently ignore it, so the table is fail-open there. We need a handshake > between the engine and the connector. > - **Bypass paths**: The policy is resolved once at analysis time and then > stays in the plan. The SPIP should define enforcement for paths that reuse an > analyzed plan or bypass the analyzer rewrite: global temp views, streaming, > catalog-side `Table` caching, `TableProvider` loads, V1 fallback, aggregate > pushdown, and statistics. > > ### 2. Design > > - Restate "non-removable Project" as an invariant rather than a special plan > node. The invariant: a mask expression is never replaced by the original > attribute, while unused mask outputs can still be pruned. > - Clarify the trust model when the policy filter is fully pushed down to the > connector. > - Resolve mask functions safely. Temporary and session functions must not be > able to shadow a mask. Reuse existing built-ins where possible, avoid the > `hash` name collision, and drop weak hashes such as MD5. > - Clean up the API: consistent null vs. empty semantics, type rules for each > mask function, a more specific name than `AccessControl`, and consider V2 > connector expressions instead of String function names. > > ### 3. Clarifications > > - Semantics for views (invoker vs. definer), time travel, nested and metadata > columns, and schema visibility of non-readable columns. > - Add test cases for group 1 to Q8, and revisit the Q7 timeline accordingly. > > I'm looking forward to the next revision. Thank you again for driving this, > Huaxin. > > Dongjoon. > > On 2026/09/23 08:48:42 Yang Jie wrote: > > Thanks for putting this together. I'm supportive of the direction. > > > > This closes a real gap and I think putting enforcement in the engine, > > fail-closed and optimizer-aware, is the right call, and keeping it a > > vendor-neutral DSv2 capability lets Iceberg, Ranger, Polaris, and Delta > > retire their private-API workarounds instead of each carrying its own. > > > > I have a few questions on the enforcement semantics and the standard mask > > set that I will raise in the SPIP document. None of them change my view on > > the direction. > > > > Thanks, > > Jie Yang > > > > On 2026/09/23 03:39:25 huaxin gao wrote: > > > Hi all, > > > > > > I would like to start a discussion on a SPIP that adds a public DataSource > > > V2 > > > API for a table or catalog to declare a read access policy (readable > > > columns, > > > a row filter, and column masks) that Spark enforces in the query plan. > > > Here > > > are > > > the jira <https://issues.apache.org/jira/browse/SPARK-59726> and SPIP doc > > > <https://docs.google.com/document/d/1hYEHORjHUnFzSoBY2DeOPILVJyu6iWXBusYwpF9FQJM/edit?tab=t.0> > > > . > > > > > > Today Spark has no supported API for this. Integrators inject Catalyst > > > rules > > > through SparkSessionExtensions and build on internal Catalyst APIs, as > > > Apache > > > Ranger (via the Kyuubi Spark AuthZ plugin) and a recent Apache Iceberg POC > > > both do. That is brittle across versions and unsafe: a masking projection > > > added naively to a plan can be removed or collapsed by the optimizer, > > > silently > > > returning unmasked data. > > > > > > The proposal adds SupportsAccessControl, enforced during analysis, with > > > three > > > guarantees: the masks cannot be optimized away, the row filter always sees > > > original pre-mask values, and anything Spark cannot resolve fails the read > > > rather than returning unprotected data. > > > > > > Details and proposed interfaces are in the doc. Feedback is very welcome. > > > > > > Thanks, > > > Huaxin > > > > > > > --------------------------------------------------------------------- > > To unsubscribe e-mail: [email protected] > > > > > > --------------------------------------------------------------------- > To unsubscribe e-mail: [email protected] > > --------------------------------------------------------------------- To unsubscribe e-mail: [email protected]
