Hi all, Thank you for the detailed feedback, both here and in the doc. There has been a lot of it and it is all useful.
I am working through everything now and will address the comments in the doc and the comments on this thread, then post a revised doc. Thanks, Huaxin On Thu, Sep 24, 2026 at 9:45 PM Yuming Wang <[email protected]> wrote: > Huaxin, thank you for proposing this SPIP. I think this feature is > necessary, and Spark is the right place to enforce it. I left a comment in > the doc. > > On Thu, Sep 24, 2026 at 5:48 PM Peter Toth <[email protected]> wrote: > >> Thanks for the proposal Huaxin. I like the direction. >> I agree with the points earlier reviewers raised and I left a few >> comments in the doc. >> >> Best, >> Peter >> >> On Wed, Sep 23, 2026 at 11:09 PM Martin Grund via dev < >> [email protected]> wrote: >> >>> Thanks for taking the time to write this proposal. It's an exciting >>> direction. >>> >>> The most important gap in the current proposal is to outline the >>> security boundaries of the execution. The current proposal focuses too much >>> on a rather vague definition of how the filters or masks are not elided, >>> but that is not enough from a security perspective. >>> >>> As part of the SPIP, we should outline both what the guarantees are and >>> how we plan on enforcing them. I have seen a lot of weirdness in the past >>> when it comes to row filters and column masks. >>> >>> Just to give some examples: >>> >>> - How do we handle leaking expressions and what is the way to prevent >>> information disclosure? (see >>> https://rhaas.blogspot.com/2012/03/security-barrier-views.html) >>> - How do we handle type mismatches between masks and colum types? >>> - How do we handle time travel, clones, etc? >>> >>> Generally, I would recommend decoupling built-in masking functions from >>> the overall proposal. Masking functions are just standard Spark built-ins, >>> so once the SPIP is on its way, we can always add them to Spark without >>> defining them in this proposal. I don't think having a fixed list of >>> masking functions is useful. From my experience observing user behavior, >>> there is a lot of complexity and variation in these functions. Giving users >>> the ability to reference any UDF or built-in function is a much more >>> convenient approach. >>> >>> Looking forward to seeing the updated proposal. >>> >>> Martin >>> >>> On Wed, Sep 23, 2026 at 7:28 PM Holden Karau <[email protected]> wrote: >>> >>>> So I think the security barrier / bypass path avoidance can be handled >>>> in a few different ways (I've got one proposal I'm working on a draft off >>>> but it's maybe a little early and shouldn't block this work). >>>> >>>> On 2026/09/23 16:48:33 Dongjoon Hyun wrote: >>>> > Hi, Huaxin. >>>> > >>>> > Thank you for proposing this SPIP. I support the direction. Moving >>>> row filtering and column masking from private Catalyst hacks to a public >>>> DSv2 contract would be a clear improvement for the ecosystem, and the >>>> proposed semantics are clear. >>>> > >>>> > Since the value of this SPIP rests on its "fail-closed" and >>>> "optimizer-safe" guarantees, I'd like the SPIP to address the following, >>>> grouped by priority. >>>> > >>>> > ### 1. Security guarantees (should be addressed before a vote) >>>> > >>>> > - **Security barrier**: The SPIP protects masks from the optimizer, >>>> but not hidden rows from user expressions. Leaky user predicates (e.g. ANSI >>>> errors, Python UDFs, predicates pushed to the connector) can be evaluated >>>> before the policy filter and expose rows that should be hidden. The SPIP >>>> should define the policy filter as a barrier, similar to PostgreSQL's >>>> `security_barrier` and `LEAKPROOF`. >>>> > - **DML read path**: DELETE/UPDATE/MERGE also read the target table. >>>> Applying the policy on that read can silently delete hidden rows or write >>>> masked values back. These operations should fail closed. >>>> > - **Older Spark versions**: An engine that does not know the new >>>> interface will silently ignore it, so the table is fail-open there. We need >>>> a handshake between the engine and the connector. >>>> > - **Bypass paths**: The policy is resolved once at analysis time and >>>> then stays in the plan. The SPIP should define enforcement for paths that >>>> reuse an analyzed plan or bypass the analyzer rewrite: global temp views, >>>> streaming, catalog-side `Table` caching, `TableProvider` loads, V1 >>>> fallback, aggregate pushdown, and statistics. >>>> > >>>> > ### 2. Design >>>> > >>>> > - Restate "non-removable Project" as an invariant rather than a >>>> special plan node. The invariant: a mask expression is never replaced by >>>> the original attribute, while unused mask outputs can still be pruned. >>>> > - Clarify the trust model when the policy filter is fully pushed down >>>> to the connector. >>>> > - Resolve mask functions safely. Temporary and session functions must >>>> not be able to shadow a mask. Reuse existing built-ins where possible, >>>> avoid the `hash` name collision, and drop weak hashes such as MD5. >>>> > - Clean up the API: consistent null vs. empty semantics, type rules >>>> for each mask function, a more specific name than `AccessControl`, and >>>> consider V2 connector expressions instead of String function names. >>>> > >>>> > ### 3. Clarifications >>>> > >>>> > - Semantics for views (invoker vs. definer), time travel, nested and >>>> metadata columns, and schema visibility of non-readable columns. >>>> > - Add test cases for group 1 to Q8, and revisit the Q7 timeline >>>> accordingly. >>>> > >>>> > I'm looking forward to the next revision. Thank you again for driving >>>> this, Huaxin. >>>> > >>>> > Dongjoon. >>>> > >>>> > On 2026/09/23 08:48:42 Yang Jie wrote: >>>> > > Thanks for putting this together. I'm supportive of the direction. >>>> > > >>>> > > This closes a real gap and I think putting enforcement in the >>>> engine, fail-closed and optimizer-aware, is the right call, and keeping it >>>> a vendor-neutral DSv2 capability lets Iceberg, Ranger, Polaris, and Delta >>>> retire their private-API workarounds instead of each carrying its own. >>>> > > >>>> > > I have a few questions on the enforcement semantics and the >>>> standard mask set that I will raise in the SPIP document. None of them >>>> change my view on the direction. >>>> > > >>>> > > Thanks, >>>> > > Jie Yang >>>> > > >>>> > > On 2026/09/23 03:39:25 huaxin gao wrote: >>>> > > > Hi all, >>>> > > > >>>> > > > I would like to start a discussion on a SPIP that adds a public >>>> DataSource >>>> > > > V2 >>>> > > > API for a table or catalog to declare a read access policy >>>> (readable >>>> > > > columns, >>>> > > > a row filter, and column masks) that Spark enforces in the query >>>> plan. Here >>>> > > > are >>>> > > > the jira <https://issues.apache.org/jira/browse/SPARK-59726> and >>>> SPIP doc >>>> > > > < >>>> https://docs.google.com/document/d/1hYEHORjHUnFzSoBY2DeOPILVJyu6iWXBusYwpF9FQJM/edit?tab=t.0 >>>> > >>>> > > > . >>>> > > > >>>> > > > Today Spark has no supported API for this. Integrators inject >>>> Catalyst rules >>>> > > > through SparkSessionExtensions and build on internal Catalyst >>>> APIs, as >>>> > > > Apache >>>> > > > Ranger (via the Kyuubi Spark AuthZ plugin) and a recent Apache >>>> Iceberg POC >>>> > > > both do. That is brittle across versions and unsafe: a masking >>>> projection >>>> > > > added naively to a plan can be removed or collapsed by the >>>> optimizer, >>>> > > > silently >>>> > > > returning unmasked data. >>>> > > > >>>> > > > The proposal adds SupportsAccessControl, enforced during >>>> analysis, with >>>> > > > three >>>> > > > guarantees: the masks cannot be optimized away, the row filter >>>> always sees >>>> > > > original pre-mask values, and anything Spark cannot resolve fails >>>> the read >>>> > > > rather than returning unprotected data. >>>> > > > >>>> > > > Details and proposed interfaces are in the doc. Feedback is very >>>> welcome. >>>> > > > >>>> > > > Thanks, >>>> > > > Huaxin >>>> > > > >>>> > > >>>> > > >>>> --------------------------------------------------------------------- >>>> > > To unsubscribe e-mail: [email protected] >>>> > > >>>> > > >>>> > >>>> > --------------------------------------------------------------------- >>>> > To unsubscribe e-mail: [email protected] >>>> > >>>> > >>>> >>>> --------------------------------------------------------------------- >>>> To unsubscribe e-mail: [email protected] >>>> >>>>
