So I think the security barrier / bypass path avoidance can be handled in a few 
different ways (I've got one proposal I'm working on a draft off but it's maybe 
a little early and shouldn't block this work).

On 2026/09/23 16:48:33 Dongjoon Hyun wrote:
> Hi, Huaxin.
> 
> Thank you for proposing this SPIP. I support the direction. Moving row 
> filtering and column masking from private Catalyst hacks to a public DSv2 
> contract would be a clear improvement for the ecosystem, and the proposed 
> semantics are clear.
> 
> Since the value of this SPIP rests on its "fail-closed" and "optimizer-safe" 
> guarantees, I'd like the SPIP to address the following, grouped by priority.
> 
> ### 1. Security guarantees (should be addressed before a vote)
> 
> - **Security barrier**: The SPIP protects masks from the optimizer, but not 
> hidden rows from user expressions. Leaky user predicates (e.g. ANSI errors, 
> Python UDFs, predicates pushed to the connector) can be evaluated before the 
> policy filter and expose rows that should be hidden. The SPIP should define 
> the policy filter as a barrier, similar to PostgreSQL's `security_barrier` 
> and `LEAKPROOF`.
> - **DML read path**: DELETE/UPDATE/MERGE also read the target table. Applying 
> the policy on that read can silently delete hidden rows or write masked 
> values back. These operations should fail closed.
> - **Older Spark versions**: An engine that does not know the new interface 
> will silently ignore it, so the table is fail-open there. We need a handshake 
> between the engine and the connector.
> - **Bypass paths**: The policy is resolved once at analysis time and then 
> stays in the plan. The SPIP should define enforcement for paths that reuse an 
> analyzed plan or bypass the analyzer rewrite: global temp views, streaming, 
> catalog-side `Table` caching, `TableProvider` loads, V1 fallback, aggregate 
> pushdown, and statistics.
> 
> ### 2. Design
> 
> - Restate "non-removable Project" as an invariant rather than a special plan 
> node. The invariant: a mask expression is never replaced by the original 
> attribute, while unused mask outputs can still be pruned.
> - Clarify the trust model when the policy filter is fully pushed down to the 
> connector.
> - Resolve mask functions safely. Temporary and session functions must not be 
> able to shadow a mask. Reuse existing built-ins where possible, avoid the 
> `hash` name collision, and drop weak hashes such as MD5.
> - Clean up the API: consistent null vs. empty semantics, type rules for each 
> mask function, a more specific name than `AccessControl`, and consider V2 
> connector expressions instead of String function names.
> 
> ### 3. Clarifications
> 
> - Semantics for views (invoker vs. definer), time travel, nested and metadata 
> columns, and schema visibility of non-readable columns.
> - Add test cases for group 1 to Q8, and revisit the Q7 timeline accordingly.
> 
> I'm looking forward to the next revision. Thank you again for driving this, 
> Huaxin.
> 
> Dongjoon.
> 
> On 2026/09/23 08:48:42 Yang Jie wrote:
> > Thanks for putting this together. I'm supportive of the direction.
> > 
> > This closes a real gap and I think putting enforcement in the engine, 
> > fail-closed and optimizer-aware, is the right call, and keeping it a 
> > vendor-neutral DSv2 capability lets Iceberg, Ranger, Polaris, and Delta 
> > retire their private-API workarounds instead of each carrying its own.
> > 
> > I have a few questions on the enforcement semantics and the standard mask 
> > set that I will raise in the SPIP document. None of them change my view on 
> > the direction.
> > 
> > Thanks,
> > Jie Yang
> > 
> > On 2026/09/23 03:39:25 huaxin gao wrote:
> > > Hi all,
> > > 
> > > I would like to start a discussion on a SPIP that adds a public DataSource
> > > V2
> > > API for a table or catalog to declare a read access policy (readable
> > > columns,
> > > a row filter, and column masks) that Spark enforces in the query plan. 
> > > Here
> > > are
> > > the jira <https://issues.apache.org/jira/browse/SPARK-59726> and SPIP doc
> > > <https://docs.google.com/document/d/1hYEHORjHUnFzSoBY2DeOPILVJyu6iWXBusYwpF9FQJM/edit?tab=t.0>
> > > .
> > > 
> > > Today Spark has no supported API for this. Integrators inject Catalyst 
> > > rules
> > > through SparkSessionExtensions and build on internal Catalyst APIs, as
> > > Apache
> > > Ranger (via the Kyuubi Spark AuthZ plugin) and a recent Apache Iceberg POC
> > > both do. That is brittle across versions and unsafe: a masking projection
> > > added naively to a plan can be removed or collapsed by the optimizer,
> > > silently
> > > returning unmasked data.
> > > 
> > > The proposal adds SupportsAccessControl, enforced during analysis, with
> > > three
> > > guarantees: the masks cannot be optimized away, the row filter always sees
> > > original pre-mask values, and anything Spark cannot resolve fails the read
> > > rather than returning unprotected data.
> > > 
> > > Details and proposed interfaces are in the doc. Feedback is very welcome.
> > > 
> > > Thanks,
> > > Huaxin
> > > 
> > 
> > ---------------------------------------------------------------------
> > To unsubscribe e-mail: [email protected]
> > 
> > 
> 
> ---------------------------------------------------------------------
> To unsubscribe e-mail: [email protected]
> 
> 

---------------------------------------------------------------------
To unsubscribe e-mail: [email protected]

Reply via email to