Thanks for taking the time to write this proposal. It's an exciting direction.
The most important gap in the current proposal is to outline the security boundaries of the execution. The current proposal focuses too much on a rather vague definition of how the filters or masks are not elided, but that is not enough from a security perspective. As part of the SPIP, we should outline both what the guarantees are and how we plan on enforcing them. I have seen a lot of weirdness in the past when it comes to row filters and column masks. Just to give some examples: - How do we handle leaking expressions and what is the way to prevent information disclosure? (see https://rhaas.blogspot.com/2012/03/security-barrier-views.html) - How do we handle type mismatches between masks and colum types? - How do we handle time travel, clones, etc? Generally, I would recommend decoupling built-in masking functions from the overall proposal. Masking functions are just standard Spark built-ins, so once the SPIP is on its way, we can always add them to Spark without defining them in this proposal. I don't think having a fixed list of masking functions is useful. From my experience observing user behavior, there is a lot of complexity and variation in these functions. Giving users the ability to reference any UDF or built-in function is a much more convenient approach. Looking forward to seeing the updated proposal. Martin On Wed, Sep 23, 2026 at 7:28 PM Holden Karau <[email protected]> wrote: > So I think the security barrier / bypass path avoidance can be handled in > a few different ways (I've got one proposal I'm working on a draft off but > it's maybe a little early and shouldn't block this work). > > On 2026/09/23 16:48:33 Dongjoon Hyun wrote: > > Hi, Huaxin. > > > > Thank you for proposing this SPIP. I support the direction. Moving row > filtering and column masking from private Catalyst hacks to a public DSv2 > contract would be a clear improvement for the ecosystem, and the proposed > semantics are clear. > > > > Since the value of this SPIP rests on its "fail-closed" and > "optimizer-safe" guarantees, I'd like the SPIP to address the following, > grouped by priority. > > > > ### 1. Security guarantees (should be addressed before a vote) > > > > - **Security barrier**: The SPIP protects masks from the optimizer, but > not hidden rows from user expressions. Leaky user predicates (e.g. ANSI > errors, Python UDFs, predicates pushed to the connector) can be evaluated > before the policy filter and expose rows that should be hidden. The SPIP > should define the policy filter as a barrier, similar to PostgreSQL's > `security_barrier` and `LEAKPROOF`. > > - **DML read path**: DELETE/UPDATE/MERGE also read the target table. > Applying the policy on that read can silently delete hidden rows or write > masked values back. These operations should fail closed. > > - **Older Spark versions**: An engine that does not know the new > interface will silently ignore it, so the table is fail-open there. We need > a handshake between the engine and the connector. > > - **Bypass paths**: The policy is resolved once at analysis time and > then stays in the plan. The SPIP should define enforcement for paths that > reuse an analyzed plan or bypass the analyzer rewrite: global temp views, > streaming, catalog-side `Table` caching, `TableProvider` loads, V1 > fallback, aggregate pushdown, and statistics. > > > > ### 2. Design > > > > - Restate "non-removable Project" as an invariant rather than a special > plan node. The invariant: a mask expression is never replaced by the > original attribute, while unused mask outputs can still be pruned. > > - Clarify the trust model when the policy filter is fully pushed down to > the connector. > > - Resolve mask functions safely. Temporary and session functions must > not be able to shadow a mask. Reuse existing built-ins where possible, > avoid the `hash` name collision, and drop weak hashes such as MD5. > > - Clean up the API: consistent null vs. empty semantics, type rules for > each mask function, a more specific name than `AccessControl`, and consider > V2 connector expressions instead of String function names. > > > > ### 3. Clarifications > > > > - Semantics for views (invoker vs. definer), time travel, nested and > metadata columns, and schema visibility of non-readable columns. > > - Add test cases for group 1 to Q8, and revisit the Q7 timeline > accordingly. > > > > I'm looking forward to the next revision. Thank you again for driving > this, Huaxin. > > > > Dongjoon. > > > > On 2026/09/23 08:48:42 Yang Jie wrote: > > > Thanks for putting this together. I'm supportive of the direction. > > > > > > This closes a real gap and I think putting enforcement in the engine, > fail-closed and optimizer-aware, is the right call, and keeping it a > vendor-neutral DSv2 capability lets Iceberg, Ranger, Polaris, and Delta > retire their private-API workarounds instead of each carrying its own. > > > > > > I have a few questions on the enforcement semantics and the standard > mask set that I will raise in the SPIP document. None of them change my > view on the direction. > > > > > > Thanks, > > > Jie Yang > > > > > > On 2026/09/23 03:39:25 huaxin gao wrote: > > > > Hi all, > > > > > > > > I would like to start a discussion on a SPIP that adds a public > DataSource > > > > V2 > > > > API for a table or catalog to declare a read access policy (readable > > > > columns, > > > > a row filter, and column masks) that Spark enforces in the query > plan. Here > > > > are > > > > the jira <https://issues.apache.org/jira/browse/SPARK-59726> and > SPIP doc > > > > < > https://docs.google.com/document/d/1hYEHORjHUnFzSoBY2DeOPILVJyu6iWXBusYwpF9FQJM/edit?tab=t.0 > > > > > > . > > > > > > > > Today Spark has no supported API for this. Integrators inject > Catalyst rules > > > > through SparkSessionExtensions and build on internal Catalyst APIs, > as > > > > Apache > > > > Ranger (via the Kyuubi Spark AuthZ plugin) and a recent Apache > Iceberg POC > > > > both do. That is brittle across versions and unsafe: a masking > projection > > > > added naively to a plan can be removed or collapsed by the optimizer, > > > > silently > > > > returning unmasked data. > > > > > > > > The proposal adds SupportsAccessControl, enforced during analysis, > with > > > > three > > > > guarantees: the masks cannot be optimized away, the row filter > always sees > > > > original pre-mask values, and anything Spark cannot resolve fails > the read > > > > rather than returning unprotected data. > > > > > > > > Details and proposed interfaces are in the doc. Feedback is very > welcome. > > > > > > > > Thanks, > > > > Huaxin > > > > > > > > > > --------------------------------------------------------------------- > > > To unsubscribe e-mail: [email protected] > > > > > > > > > > --------------------------------------------------------------------- > > To unsubscribe e-mail: [email protected] > > > > > > --------------------------------------------------------------------- > To unsubscribe e-mail: [email protected] > >
