Thanks for the proposal Huaxin. I like the direction. I agree with the points earlier reviewers raised and I left a few comments in the doc.
Best, Peter On Wed, Sep 23, 2026 at 11:09 PM Martin Grund via dev <[email protected]> wrote: > Thanks for taking the time to write this proposal. It's an exciting > direction. > > The most important gap in the current proposal is to outline the security > boundaries of the execution. The current proposal focuses too much on a > rather vague definition of how the filters or masks are not elided, but > that is not enough from a security perspective. > > As part of the SPIP, we should outline both what the guarantees are and > how we plan on enforcing them. I have seen a lot of weirdness in the past > when it comes to row filters and column masks. > > Just to give some examples: > > - How do we handle leaking expressions and what is the way to prevent > information disclosure? (see > https://rhaas.blogspot.com/2012/03/security-barrier-views.html) > - How do we handle type mismatches between masks and colum types? > - How do we handle time travel, clones, etc? > > Generally, I would recommend decoupling built-in masking functions from > the overall proposal. Masking functions are just standard Spark built-ins, > so once the SPIP is on its way, we can always add them to Spark without > defining them in this proposal. I don't think having a fixed list of > masking functions is useful. From my experience observing user behavior, > there is a lot of complexity and variation in these functions. Giving users > the ability to reference any UDF or built-in function is a much more > convenient approach. > > Looking forward to seeing the updated proposal. > > Martin > > On Wed, Sep 23, 2026 at 7:28 PM Holden Karau <[email protected]> wrote: > >> So I think the security barrier / bypass path avoidance can be handled in >> a few different ways (I've got one proposal I'm working on a draft off but >> it's maybe a little early and shouldn't block this work). >> >> On 2026/09/23 16:48:33 Dongjoon Hyun wrote: >> > Hi, Huaxin. >> > >> > Thank you for proposing this SPIP. I support the direction. Moving row >> filtering and column masking from private Catalyst hacks to a public DSv2 >> contract would be a clear improvement for the ecosystem, and the proposed >> semantics are clear. >> > >> > Since the value of this SPIP rests on its "fail-closed" and >> "optimizer-safe" guarantees, I'd like the SPIP to address the following, >> grouped by priority. >> > >> > ### 1. Security guarantees (should be addressed before a vote) >> > >> > - **Security barrier**: The SPIP protects masks from the optimizer, but >> not hidden rows from user expressions. Leaky user predicates (e.g. ANSI >> errors, Python UDFs, predicates pushed to the connector) can be evaluated >> before the policy filter and expose rows that should be hidden. The SPIP >> should define the policy filter as a barrier, similar to PostgreSQL's >> `security_barrier` and `LEAKPROOF`. >> > - **DML read path**: DELETE/UPDATE/MERGE also read the target table. >> Applying the policy on that read can silently delete hidden rows or write >> masked values back. These operations should fail closed. >> > - **Older Spark versions**: An engine that does not know the new >> interface will silently ignore it, so the table is fail-open there. We need >> a handshake between the engine and the connector. >> > - **Bypass paths**: The policy is resolved once at analysis time and >> then stays in the plan. The SPIP should define enforcement for paths that >> reuse an analyzed plan or bypass the analyzer rewrite: global temp views, >> streaming, catalog-side `Table` caching, `TableProvider` loads, V1 >> fallback, aggregate pushdown, and statistics. >> > >> > ### 2. Design >> > >> > - Restate "non-removable Project" as an invariant rather than a special >> plan node. The invariant: a mask expression is never replaced by the >> original attribute, while unused mask outputs can still be pruned. >> > - Clarify the trust model when the policy filter is fully pushed down >> to the connector. >> > - Resolve mask functions safely. Temporary and session functions must >> not be able to shadow a mask. Reuse existing built-ins where possible, >> avoid the `hash` name collision, and drop weak hashes such as MD5. >> > - Clean up the API: consistent null vs. empty semantics, type rules for >> each mask function, a more specific name than `AccessControl`, and consider >> V2 connector expressions instead of String function names. >> > >> > ### 3. Clarifications >> > >> > - Semantics for views (invoker vs. definer), time travel, nested and >> metadata columns, and schema visibility of non-readable columns. >> > - Add test cases for group 1 to Q8, and revisit the Q7 timeline >> accordingly. >> > >> > I'm looking forward to the next revision. Thank you again for driving >> this, Huaxin. >> > >> > Dongjoon. >> > >> > On 2026/09/23 08:48:42 Yang Jie wrote: >> > > Thanks for putting this together. I'm supportive of the direction. >> > > >> > > This closes a real gap and I think putting enforcement in the engine, >> fail-closed and optimizer-aware, is the right call, and keeping it a >> vendor-neutral DSv2 capability lets Iceberg, Ranger, Polaris, and Delta >> retire their private-API workarounds instead of each carrying its own. >> > > >> > > I have a few questions on the enforcement semantics and the standard >> mask set that I will raise in the SPIP document. None of them change my >> view on the direction. >> > > >> > > Thanks, >> > > Jie Yang >> > > >> > > On 2026/09/23 03:39:25 huaxin gao wrote: >> > > > Hi all, >> > > > >> > > > I would like to start a discussion on a SPIP that adds a public >> DataSource >> > > > V2 >> > > > API for a table or catalog to declare a read access policy (readable >> > > > columns, >> > > > a row filter, and column masks) that Spark enforces in the query >> plan. Here >> > > > are >> > > > the jira <https://issues.apache.org/jira/browse/SPARK-59726> and >> SPIP doc >> > > > < >> https://docs.google.com/document/d/1hYEHORjHUnFzSoBY2DeOPILVJyu6iWXBusYwpF9FQJM/edit?tab=t.0 >> > >> > > > . >> > > > >> > > > Today Spark has no supported API for this. Integrators inject >> Catalyst rules >> > > > through SparkSessionExtensions and build on internal Catalyst APIs, >> as >> > > > Apache >> > > > Ranger (via the Kyuubi Spark AuthZ plugin) and a recent Apache >> Iceberg POC >> > > > both do. That is brittle across versions and unsafe: a masking >> projection >> > > > added naively to a plan can be removed or collapsed by the >> optimizer, >> > > > silently >> > > > returning unmasked data. >> > > > >> > > > The proposal adds SupportsAccessControl, enforced during analysis, >> with >> > > > three >> > > > guarantees: the masks cannot be optimized away, the row filter >> always sees >> > > > original pre-mask values, and anything Spark cannot resolve fails >> the read >> > > > rather than returning unprotected data. >> > > > >> > > > Details and proposed interfaces are in the doc. Feedback is very >> welcome. >> > > > >> > > > Thanks, >> > > > Huaxin >> > > > >> > > >> > > --------------------------------------------------------------------- >> > > To unsubscribe e-mail: [email protected] >> > > >> > > >> > >> > --------------------------------------------------------------------- >> > To unsubscribe e-mail: [email protected] >> > >> > >> >> --------------------------------------------------------------------- >> To unsubscribe e-mail: [email protected] >> >>
