Hi, Huaxin. Thank you for proposing this SPIP. I support the direction. Moving row filtering and column masking from private Catalyst hacks to a public DSv2 contract would be a clear improvement for the ecosystem, and the proposed semantics are clear.
Since the value of this SPIP rests on its "fail-closed" and "optimizer-safe" guarantees, I'd like the SPIP to address the following, grouped by priority. ### 1. Security guarantees (should be addressed before a vote) - **Security barrier**: The SPIP protects masks from the optimizer, but not hidden rows from user expressions. Leaky user predicates (e.g. ANSI errors, Python UDFs, predicates pushed to the connector) can be evaluated before the policy filter and expose rows that should be hidden. The SPIP should define the policy filter as a barrier, similar to PostgreSQL's `security_barrier` and `LEAKPROOF`. - **DML read path**: DELETE/UPDATE/MERGE also read the target table. Applying the policy on that read can silently delete hidden rows or write masked values back. These operations should fail closed. - **Older Spark versions**: An engine that does not know the new interface will silently ignore it, so the table is fail-open there. We need a handshake between the engine and the connector. - **Bypass paths**: The policy is resolved once at analysis time and then stays in the plan. The SPIP should define enforcement for paths that reuse an analyzed plan or bypass the analyzer rewrite: global temp views, streaming, catalog-side `Table` caching, `TableProvider` loads, V1 fallback, aggregate pushdown, and statistics. ### 2. Design - Restate "non-removable Project" as an invariant rather than a special plan node. The invariant: a mask expression is never replaced by the original attribute, while unused mask outputs can still be pruned. - Clarify the trust model when the policy filter is fully pushed down to the connector. - Resolve mask functions safely. Temporary and session functions must not be able to shadow a mask. Reuse existing built-ins where possible, avoid the `hash` name collision, and drop weak hashes such as MD5. - Clean up the API: consistent null vs. empty semantics, type rules for each mask function, a more specific name than `AccessControl`, and consider V2 connector expressions instead of String function names. ### 3. Clarifications - Semantics for views (invoker vs. definer), time travel, nested and metadata columns, and schema visibility of non-readable columns. - Add test cases for group 1 to Q8, and revisit the Q7 timeline accordingly. I'm looking forward to the next revision. Thank you again for driving this, Huaxin. Dongjoon. On 2026/09/23 08:48:42 Yang Jie wrote: > Thanks for putting this together. I'm supportive of the direction. > > This closes a real gap and I think putting enforcement in the engine, > fail-closed and optimizer-aware, is the right call, and keeping it a > vendor-neutral DSv2 capability lets Iceberg, Ranger, Polaris, and Delta > retire their private-API workarounds instead of each carrying its own. > > I have a few questions on the enforcement semantics and the standard mask set > that I will raise in the SPIP document. None of them change my view on the > direction. > > Thanks, > Jie Yang > > On 2026/09/23 03:39:25 huaxin gao wrote: > > Hi all, > > > > I would like to start a discussion on a SPIP that adds a public DataSource > > V2 > > API for a table or catalog to declare a read access policy (readable > > columns, > > a row filter, and column masks) that Spark enforces in the query plan. Here > > are > > the jira <https://issues.apache.org/jira/browse/SPARK-59726> and SPIP doc > > <https://docs.google.com/document/d/1hYEHORjHUnFzSoBY2DeOPILVJyu6iWXBusYwpF9FQJM/edit?tab=t.0> > > . > > > > Today Spark has no supported API for this. Integrators inject Catalyst rules > > through SparkSessionExtensions and build on internal Catalyst APIs, as > > Apache > > Ranger (via the Kyuubi Spark AuthZ plugin) and a recent Apache Iceberg POC > > both do. That is brittle across versions and unsafe: a masking projection > > added naively to a plan can be removed or collapsed by the optimizer, > > silently > > returning unmasked data. > > > > The proposal adds SupportsAccessControl, enforced during analysis, with > > three > > guarantees: the masks cannot be optimized away, the row filter always sees > > original pre-mask values, and anything Spark cannot resolve fails the read > > rather than returning unprotected data. > > > > Details and proposed interfaces are in the doc. Feedback is very welcome. > > > > Thanks, > > Huaxin > > > > --------------------------------------------------------------------- > To unsubscribe e-mail: [email protected] > > --------------------------------------------------------------------- To unsubscribe e-mail: [email protected]
