Hi all,

I would like to start a discussion on a SPIP that adds a public DataSource
V2
API for a table or catalog to declare a read access policy (readable
columns,
a row filter, and column masks) that Spark enforces in the query plan. Here
are
the jira <https://issues.apache.org/jira/browse/SPARK-59726> and SPIP doc
<https://docs.google.com/document/d/1hYEHORjHUnFzSoBY2DeOPILVJyu6iWXBusYwpF9FQJM/edit?tab=t.0>
.

Today Spark has no supported API for this. Integrators inject Catalyst rules
through SparkSessionExtensions and build on internal Catalyst APIs, as
Apache
Ranger (via the Kyuubi Spark AuthZ plugin) and a recent Apache Iceberg POC
both do. That is brittle across versions and unsafe: a masking projection
added naively to a plan can be removed or collapsed by the optimizer,
silently
returning unmasked data.

The proposal adds SupportsAccessControl, enforced during analysis, with
three
guarantees: the masks cannot be optimized away, the row filter always sees
original pre-mask values, and anything Spark cannot resolve fails the read
rather than returning unprotected data.

Details and proposed interfaces are in the doc. Feedback is very welcome.

Thanks,
Huaxin

Reply via email to