Thanks for putting this together. I'm supportive of the direction. This closes a real gap and I think putting enforcement in the engine, fail-closed and optimizer-aware, is the right call, and keeping it a vendor-neutral DSv2 capability lets Iceberg, Ranger, Polaris, and Delta retire their private-API workarounds instead of each carrying its own.
I have a few questions on the enforcement semantics and the standard mask set that I will raise in the SPIP document. None of them change my view on the direction. Thanks, Jie Yang On 2026/09/23 03:39:25 huaxin gao wrote: > Hi all, > > I would like to start a discussion on a SPIP that adds a public DataSource > V2 > API for a table or catalog to declare a read access policy (readable > columns, > a row filter, and column masks) that Spark enforces in the query plan. Here > are > the jira <https://issues.apache.org/jira/browse/SPARK-59726> and SPIP doc > <https://docs.google.com/document/d/1hYEHORjHUnFzSoBY2DeOPILVJyu6iWXBusYwpF9FQJM/edit?tab=t.0> > . > > Today Spark has no supported API for this. Integrators inject Catalyst rules > through SparkSessionExtensions and build on internal Catalyst APIs, as > Apache > Ranger (via the Kyuubi Spark AuthZ plugin) and a recent Apache Iceberg POC > both do. That is brittle across versions and unsafe: a masking projection > added naively to a plan can be removed or collapsed by the optimizer, > silently > returning unmasked data. > > The proposal adds SupportsAccessControl, enforced during analysis, with > three > guarantees: the masks cannot be optimized away, the row filter always sees > original pre-mask values, and anything Spark cannot resolve fails the read > rather than returning unprotected data. > > Details and proposed interfaces are in the doc. Feedback is very welcome. > > Thanks, > Huaxin > --------------------------------------------------------------------- To unsubscribe e-mail: [email protected]
