Huaxin, thank you for proposing this SPIP. I think this feature is necessary, and Spark is the right place to enforce it. I left a comment in the doc.
On Thu, Sep 24, 2026 at 5:48 PM Peter Toth <[email protected]> wrote: > Thanks for the proposal Huaxin. I like the direction. > I agree with the points earlier reviewers raised and I left a few comments > in the doc. > > Best, > Peter > > On Wed, Sep 23, 2026 at 11:09 PM Martin Grund via dev < > [email protected]> wrote: > >> Thanks for taking the time to write this proposal. It's an exciting >> direction. >> >> The most important gap in the current proposal is to outline the security >> boundaries of the execution. The current proposal focuses too much on a >> rather vague definition of how the filters or masks are not elided, but >> that is not enough from a security perspective. >> >> As part of the SPIP, we should outline both what the guarantees are and >> how we plan on enforcing them. I have seen a lot of weirdness in the past >> when it comes to row filters and column masks. >> >> Just to give some examples: >> >> - How do we handle leaking expressions and what is the way to prevent >> information disclosure? (see >> https://rhaas.blogspot.com/2012/03/security-barrier-views.html) >> - How do we handle type mismatches between masks and colum types? >> - How do we handle time travel, clones, etc? >> >> Generally, I would recommend decoupling built-in masking functions from >> the overall proposal. Masking functions are just standard Spark built-ins, >> so once the SPIP is on its way, we can always add them to Spark without >> defining them in this proposal. I don't think having a fixed list of >> masking functions is useful. From my experience observing user behavior, >> there is a lot of complexity and variation in these functions. Giving users >> the ability to reference any UDF or built-in function is a much more >> convenient approach. >> >> Looking forward to seeing the updated proposal. >> >> Martin >> >> On Wed, Sep 23, 2026 at 7:28 PM Holden Karau <[email protected]> wrote: >> >>> So I think the security barrier / bypass path avoidance can be handled >>> in a few different ways (I've got one proposal I'm working on a draft off >>> but it's maybe a little early and shouldn't block this work). >>> >>> On 2026/09/23 16:48:33 Dongjoon Hyun wrote: >>> > Hi, Huaxin. >>> > >>> > Thank you for proposing this SPIP. I support the direction. Moving row >>> filtering and column masking from private Catalyst hacks to a public DSv2 >>> contract would be a clear improvement for the ecosystem, and the proposed >>> semantics are clear. >>> > >>> > Since the value of this SPIP rests on its "fail-closed" and >>> "optimizer-safe" guarantees, I'd like the SPIP to address the following, >>> grouped by priority. >>> > >>> > ### 1. Security guarantees (should be addressed before a vote) >>> > >>> > - **Security barrier**: The SPIP protects masks from the optimizer, >>> but not hidden rows from user expressions. Leaky user predicates (e.g. ANSI >>> errors, Python UDFs, predicates pushed to the connector) can be evaluated >>> before the policy filter and expose rows that should be hidden. The SPIP >>> should define the policy filter as a barrier, similar to PostgreSQL's >>> `security_barrier` and `LEAKPROOF`. >>> > - **DML read path**: DELETE/UPDATE/MERGE also read the target table. >>> Applying the policy on that read can silently delete hidden rows or write >>> masked values back. These operations should fail closed. >>> > - **Older Spark versions**: An engine that does not know the new >>> interface will silently ignore it, so the table is fail-open there. We need >>> a handshake between the engine and the connector. >>> > - **Bypass paths**: The policy is resolved once at analysis time and >>> then stays in the plan. The SPIP should define enforcement for paths that >>> reuse an analyzed plan or bypass the analyzer rewrite: global temp views, >>> streaming, catalog-side `Table` caching, `TableProvider` loads, V1 >>> fallback, aggregate pushdown, and statistics. >>> > >>> > ### 2. Design >>> > >>> > - Restate "non-removable Project" as an invariant rather than a >>> special plan node. The invariant: a mask expression is never replaced by >>> the original attribute, while unused mask outputs can still be pruned. >>> > - Clarify the trust model when the policy filter is fully pushed down >>> to the connector. >>> > - Resolve mask functions safely. Temporary and session functions must >>> not be able to shadow a mask. Reuse existing built-ins where possible, >>> avoid the `hash` name collision, and drop weak hashes such as MD5. >>> > - Clean up the API: consistent null vs. empty semantics, type rules >>> for each mask function, a more specific name than `AccessControl`, and >>> consider V2 connector expressions instead of String function names. >>> > >>> > ### 3. Clarifications >>> > >>> > - Semantics for views (invoker vs. definer), time travel, nested and >>> metadata columns, and schema visibility of non-readable columns. >>> > - Add test cases for group 1 to Q8, and revisit the Q7 timeline >>> accordingly. >>> > >>> > I'm looking forward to the next revision. Thank you again for driving >>> this, Huaxin. >>> > >>> > Dongjoon. >>> > >>> > On 2026/09/23 08:48:42 Yang Jie wrote: >>> > > Thanks for putting this together. I'm supportive of the direction. >>> > > >>> > > This closes a real gap and I think putting enforcement in the >>> engine, fail-closed and optimizer-aware, is the right call, and keeping it >>> a vendor-neutral DSv2 capability lets Iceberg, Ranger, Polaris, and Delta >>> retire their private-API workarounds instead of each carrying its own. >>> > > >>> > > I have a few questions on the enforcement semantics and the standard >>> mask set that I will raise in the SPIP document. None of them change my >>> view on the direction. >>> > > >>> > > Thanks, >>> > > Jie Yang >>> > > >>> > > On 2026/09/23 03:39:25 huaxin gao wrote: >>> > > > Hi all, >>> > > > >>> > > > I would like to start a discussion on a SPIP that adds a public >>> DataSource >>> > > > V2 >>> > > > API for a table or catalog to declare a read access policy >>> (readable >>> > > > columns, >>> > > > a row filter, and column masks) that Spark enforces in the query >>> plan. Here >>> > > > are >>> > > > the jira <https://issues.apache.org/jira/browse/SPARK-59726> and >>> SPIP doc >>> > > > < >>> https://docs.google.com/document/d/1hYEHORjHUnFzSoBY2DeOPILVJyu6iWXBusYwpF9FQJM/edit?tab=t.0 >>> > >>> > > > . >>> > > > >>> > > > Today Spark has no supported API for this. Integrators inject >>> Catalyst rules >>> > > > through SparkSessionExtensions and build on internal Catalyst >>> APIs, as >>> > > > Apache >>> > > > Ranger (via the Kyuubi Spark AuthZ plugin) and a recent Apache >>> Iceberg POC >>> > > > both do. That is brittle across versions and unsafe: a masking >>> projection >>> > > > added naively to a plan can be removed or collapsed by the >>> optimizer, >>> > > > silently >>> > > > returning unmasked data. >>> > > > >>> > > > The proposal adds SupportsAccessControl, enforced during analysis, >>> with >>> > > > three >>> > > > guarantees: the masks cannot be optimized away, the row filter >>> always sees >>> > > > original pre-mask values, and anything Spark cannot resolve fails >>> the read >>> > > > rather than returning unprotected data. >>> > > > >>> > > > Details and proposed interfaces are in the doc. Feedback is very >>> welcome. >>> > > > >>> > > > Thanks, >>> > > > Huaxin >>> > > > >>> > > >>> > > --------------------------------------------------------------------- >>> > > To unsubscribe e-mail: [email protected] >>> > > >>> > > >>> > >>> > --------------------------------------------------------------------- >>> > To unsubscribe e-mail: [email protected] >>> > >>> > >>> >>> --------------------------------------------------------------------- >>> To unsubscribe e-mail: [email protected] >>> >>>
