Hi all, I would like to start a discussion on a SPIP that adds a public DataSource V2 API for a table or catalog to declare a read access policy (readable columns, a row filter, and column masks) that Spark enforces in the query plan. Here are the jira <https://issues.apache.org/jira/browse/SPARK-59726> and SPIP doc <https://docs.google.com/document/d/1hYEHORjHUnFzSoBY2DeOPILVJyu6iWXBusYwpF9FQJM/edit?tab=t.0> .
Today Spark has no supported API for this. Integrators inject Catalyst rules through SparkSessionExtensions and build on internal Catalyst APIs, as Apache Ranger (via the Kyuubi Spark AuthZ plugin) and a recent Apache Iceberg POC both do. That is brittle across versions and unsafe: a masking projection added naively to a plan can be removed or collapsed by the optimizer, silently returning unmasked data. The proposal adds SupportsAccessControl, enforced during analysis, with three guarantees: the masks cannot be optimized away, the row filter always sees original pre-mask values, and anything Spark cannot resolve fails the read rather than returning unprotected data. Details and proposed interfaces are in the doc. Feedback is very welcome. Thanks, Huaxin
