documenting small conversation thread , I had with Huaxin: Hi , Long back I had implemented the row level filter integrated, with ddl enhancement and access control using ldap policy. Had also created an inbuilt UDF which could check if the user executing the query belonged to the privileged user group. Basically a ddl which could create a row level policy specifying the criteria , and access applied using user or user group as defined in LDAP . Later I think integrated it with ranger.. that work might be lying some where... Plz let me know if you want to reuse any part. Regards Asif
Actually skimmed through your proposal.. Looks like you have all sorted out and thought through well, so pls ignore my mail... Regards Asif Hi Asif, Thanks for the note, and good to hear you built something similar before. Would you mind sharing that on the dev@ thread? A short note that you've done this work and think a supported Spark API would help goes a long way. Thanks, Huaxin I had ranger integrated code tucked away in a machine in previous company , and left it there so that is lost..but the non ranger code developed on some spark 2 or 3 version must be lying with me, will find out in a day or two and let you know. I will also read the spip in detail, I just took cursory look. But from what I just skimmed through, on mobile .. it's my understanding that filters would need to be applied before masking..in case masked cols are in filter... Then there are aggregate functions , join keys on masked cols... I suppose you have thought about it... So just to brainstorm what thoughts came to me, ..why not mask only at the top level of tree?.. say MaskedProject?.. rest all the push down filter rules and all remain unchanged... Only mask at the top level?.. If this idea , I am sure , must have come to your mind, so what was the reason to not go by that route?.. The purpose of masking...row level filtering...is to control what the final result is seen by the user , right? Regards Asif On Wed, Sep 23, 2026 at 9:48 AM Dongjoon Hyun <[email protected]> wrote: > Hi, Huaxin. > > Thank you for proposing this SPIP. I support the direction. Moving row > filtering and column masking from private Catalyst hacks to a public DSv2 > contract would be a clear improvement for the ecosystem, and the proposed > semantics are clear. > > Since the value of this SPIP rests on its "fail-closed" and > "optimizer-safe" guarantees, I'd like the SPIP to address the following, > grouped by priority. > > ### 1. Security guarantees (should be addressed before a vote) > > - **Security barrier**: The SPIP protects masks from the optimizer, but > not hidden rows from user expressions. Leaky user predicates (e.g. ANSI > errors, Python UDFs, predicates pushed to the connector) can be evaluated > before the policy filter and expose rows that should be hidden. The SPIP > should define the policy filter as a barrier, similar to PostgreSQL's > `security_barrier` and `LEAKPROOF`. > - **DML read path**: DELETE/UPDATE/MERGE also read the target table. > Applying the policy on that read can silently delete hidden rows or write > masked values back. These operations should fail closed. > - **Older Spark versions**: An engine that does not know the new interface > will silently ignore it, so the table is fail-open there. We need a > handshake between the engine and the connector. > - **Bypass paths**: The policy is resolved once at analysis time and then > stays in the plan. The SPIP should define enforcement for paths that reuse > an analyzed plan or bypass the analyzer rewrite: global temp views, > streaming, catalog-side `Table` caching, `TableProvider` loads, V1 > fallback, aggregate pushdown, and statistics. > > ### 2. Design > > - Restate "non-removable Project" as an invariant rather than a special > plan node. The invariant: a mask expression is never replaced by the > original attribute, while unused mask outputs can still be pruned. > - Clarify the trust model when the policy filter is fully pushed down to > the connector. > - Resolve mask functions safely. Temporary and session functions must not > be able to shadow a mask. Reuse existing built-ins where possible, avoid > the `hash` name collision, and drop weak hashes such as MD5. > - Clean up the API: consistent null vs. empty semantics, type rules for > each mask function, a more specific name than `AccessControl`, and consider > V2 connector expressions instead of String function names. > > ### 3. Clarifications > > - Semantics for views (invoker vs. definer), time travel, nested and > metadata columns, and schema visibility of non-readable columns. > - Add test cases for group 1 to Q8, and revisit the Q7 timeline > accordingly. > > I'm looking forward to the next revision. Thank you again for driving > this, Huaxin. > > Dongjoon. > > On 2026/09/23 08:48:42 Yang Jie wrote: > > Thanks for putting this together. I'm supportive of the direction. > > > > This closes a real gap and I think putting enforcement in the engine, > fail-closed and optimizer-aware, is the right call, and keeping it a > vendor-neutral DSv2 capability lets Iceberg, Ranger, Polaris, and Delta > retire their private-API workarounds instead of each carrying its own. > > > > I have a few questions on the enforcement semantics and the standard > mask set that I will raise in the SPIP document. None of them change my > view on the direction. > > > > Thanks, > > Jie Yang > > > > On 2026/09/23 03:39:25 huaxin gao wrote: > > > Hi all, > > > > > > I would like to start a discussion on a SPIP that adds a public > DataSource > > > V2 > > > API for a table or catalog to declare a read access policy (readable > > > columns, > > > a row filter, and column masks) that Spark enforces in the query plan. > Here > > > are > > > the jira <https://issues.apache.org/jira/browse/SPARK-59726> and SPIP > doc > > > < > https://docs.google.com/document/d/1hYEHORjHUnFzSoBY2DeOPILVJyu6iWXBusYwpF9FQJM/edit?tab=t.0 > > > > > . > > > > > > Today Spark has no supported API for this. Integrators inject Catalyst > rules > > > through SparkSessionExtensions and build on internal Catalyst APIs, as > > > Apache > > > Ranger (via the Kyuubi Spark AuthZ plugin) and a recent Apache Iceberg > POC > > > both do. That is brittle across versions and unsafe: a masking > projection > > > added naively to a plan can be removed or collapsed by the optimizer, > > > silently > > > returning unmasked data. > > > > > > The proposal adds SupportsAccessControl, enforced during analysis, with > > > three > > > guarantees: the masks cannot be optimized away, the row filter always > sees > > > original pre-mask values, and anything Spark cannot resolve fails the > read > > > rather than returning unprotected data. > > > > > > Details and proposed interfaces are in the doc. Feedback is very > welcome. > > > > > > Thanks, > > > Huaxin > > > > > > > --------------------------------------------------------------------- > > To unsubscribe e-mail: [email protected] > > > > > > --------------------------------------------------------------------- > To unsubscribe e-mail: [email protected] > >
