Thanks for the proposal Huaxin. I like the direction.
I agree with the points earlier reviewers raised and I left a few comments
in the doc.

Best,
Peter

On Wed, Sep 23, 2026 at 11:09 PM Martin Grund via dev <[email protected]>
wrote:

> Thanks for taking the time to write this proposal. It's an exciting
> direction.
>
> The most important gap in the current proposal is to outline the security
> boundaries of the execution. The current proposal focuses too much on a
> rather vague definition of how the filters or masks are not elided, but
> that is not enough from a security perspective.
>
> As part of the SPIP, we should outline both what the guarantees are and
> how we plan on enforcing them. I have seen a lot of weirdness in the past
> when it comes to row filters and column masks.
>
> Just to give some examples:
>
> - How do we handle leaking expressions and what is the way to prevent
> information disclosure? (see
> https://rhaas.blogspot.com/2012/03/security-barrier-views.html)
> - How do we handle type mismatches between masks and colum types?
> - How do we handle time travel, clones, etc?
>
> Generally, I would recommend decoupling built-in masking functions from
> the overall proposal. Masking functions are just standard Spark built-ins,
> so once the SPIP is on its way, we can always add them to Spark without
> defining them in this proposal. I don't think having a fixed list of
> masking functions is useful. From my experience observing user behavior,
> there is a lot of complexity and variation in these functions. Giving users
> the ability to reference any UDF or built-in function is a much more
> convenient approach.
>
> Looking forward to seeing the updated proposal.
>
> Martin
>
> On Wed, Sep 23, 2026 at 7:28 PM Holden Karau <[email protected]> wrote:
>
>> So I think the security barrier / bypass path avoidance can be handled in
>> a few different ways (I've got one proposal I'm working on a draft off but
>> it's maybe a little early and shouldn't block this work).
>>
>> On 2026/09/23 16:48:33 Dongjoon Hyun wrote:
>> > Hi, Huaxin.
>> >
>> > Thank you for proposing this SPIP. I support the direction. Moving row
>> filtering and column masking from private Catalyst hacks to a public DSv2
>> contract would be a clear improvement for the ecosystem, and the proposed
>> semantics are clear.
>> >
>> > Since the value of this SPIP rests on its "fail-closed" and
>> "optimizer-safe" guarantees, I'd like the SPIP to address the following,
>> grouped by priority.
>> >
>> > ### 1. Security guarantees (should be addressed before a vote)
>> >
>> > - **Security barrier**: The SPIP protects masks from the optimizer, but
>> not hidden rows from user expressions. Leaky user predicates (e.g. ANSI
>> errors, Python UDFs, predicates pushed to the connector) can be evaluated
>> before the policy filter and expose rows that should be hidden. The SPIP
>> should define the policy filter as a barrier, similar to PostgreSQL's
>> `security_barrier` and `LEAKPROOF`.
>> > - **DML read path**: DELETE/UPDATE/MERGE also read the target table.
>> Applying the policy on that read can silently delete hidden rows or write
>> masked values back. These operations should fail closed.
>> > - **Older Spark versions**: An engine that does not know the new
>> interface will silently ignore it, so the table is fail-open there. We need
>> a handshake between the engine and the connector.
>> > - **Bypass paths**: The policy is resolved once at analysis time and
>> then stays in the plan. The SPIP should define enforcement for paths that
>> reuse an analyzed plan or bypass the analyzer rewrite: global temp views,
>> streaming, catalog-side `Table` caching, `TableProvider` loads, V1
>> fallback, aggregate pushdown, and statistics.
>> >
>> > ### 2. Design
>> >
>> > - Restate "non-removable Project" as an invariant rather than a special
>> plan node. The invariant: a mask expression is never replaced by the
>> original attribute, while unused mask outputs can still be pruned.
>> > - Clarify the trust model when the policy filter is fully pushed down
>> to the connector.
>> > - Resolve mask functions safely. Temporary and session functions must
>> not be able to shadow a mask. Reuse existing built-ins where possible,
>> avoid the `hash` name collision, and drop weak hashes such as MD5.
>> > - Clean up the API: consistent null vs. empty semantics, type rules for
>> each mask function, a more specific name than `AccessControl`, and consider
>> V2 connector expressions instead of String function names.
>> >
>> > ### 3. Clarifications
>> >
>> > - Semantics for views (invoker vs. definer), time travel, nested and
>> metadata columns, and schema visibility of non-readable columns.
>> > - Add test cases for group 1 to Q8, and revisit the Q7 timeline
>> accordingly.
>> >
>> > I'm looking forward to the next revision. Thank you again for driving
>> this, Huaxin.
>> >
>> > Dongjoon.
>> >
>> > On 2026/09/23 08:48:42 Yang Jie wrote:
>> > > Thanks for putting this together. I'm supportive of the direction.
>> > >
>> > > This closes a real gap and I think putting enforcement in the engine,
>> fail-closed and optimizer-aware, is the right call, and keeping it a
>> vendor-neutral DSv2 capability lets Iceberg, Ranger, Polaris, and Delta
>> retire their private-API workarounds instead of each carrying its own.
>> > >
>> > > I have a few questions on the enforcement semantics and the standard
>> mask set that I will raise in the SPIP document. None of them change my
>> view on the direction.
>> > >
>> > > Thanks,
>> > > Jie Yang
>> > >
>> > > On 2026/09/23 03:39:25 huaxin gao wrote:
>> > > > Hi all,
>> > > >
>> > > > I would like to start a discussion on a SPIP that adds a public
>> DataSource
>> > > > V2
>> > > > API for a table or catalog to declare a read access policy (readable
>> > > > columns,
>> > > > a row filter, and column masks) that Spark enforces in the query
>> plan. Here
>> > > > are
>> > > > the jira <https://issues.apache.org/jira/browse/SPARK-59726> and
>> SPIP doc
>> > > > <
>> https://docs.google.com/document/d/1hYEHORjHUnFzSoBY2DeOPILVJyu6iWXBusYwpF9FQJM/edit?tab=t.0
>> >
>> > > > .
>> > > >
>> > > > Today Spark has no supported API for this. Integrators inject
>> Catalyst rules
>> > > > through SparkSessionExtensions and build on internal Catalyst APIs,
>> as
>> > > > Apache
>> > > > Ranger (via the Kyuubi Spark AuthZ plugin) and a recent Apache
>> Iceberg POC
>> > > > both do. That is brittle across versions and unsafe: a masking
>> projection
>> > > > added naively to a plan can be removed or collapsed by the
>> optimizer,
>> > > > silently
>> > > > returning unmasked data.
>> > > >
>> > > > The proposal adds SupportsAccessControl, enforced during analysis,
>> with
>> > > > three
>> > > > guarantees: the masks cannot be optimized away, the row filter
>> always sees
>> > > > original pre-mask values, and anything Spark cannot resolve fails
>> the read
>> > > > rather than returning unprotected data.
>> > > >
>> > > > Details and proposed interfaces are in the doc. Feedback is very
>> welcome.
>> > > >
>> > > > Thanks,
>> > > > Huaxin
>> > > >
>> > >
>> > > ---------------------------------------------------------------------
>> > > To unsubscribe e-mail: [email protected]
>> > >
>> > >
>> >
>> > ---------------------------------------------------------------------
>> > To unsubscribe e-mail: [email protected]
>> >
>> >
>>
>> ---------------------------------------------------------------------
>> To unsubscribe e-mail: [email protected]
>>
>>

Reply via email to