Hi all,

Thank you for the detailed feedback, both here and in the doc. There has
been a lot of it and it is all useful.

I am working through everything now and will address the comments in the
doc and the comments on this thread, then post a revised doc.

Thanks,
Huaxin

On Thu, Sep 24, 2026 at 9:45 PM Yuming Wang <[email protected]> wrote:

> Huaxin, thank you for proposing this SPIP. I think this feature is
> necessary, and Spark is the right place to enforce it. I left a comment in
> the doc.
>
> On Thu, Sep 24, 2026 at 5:48 PM Peter Toth <[email protected]> wrote:
>
>> Thanks for the proposal Huaxin. I like the direction.
>> I agree with the points earlier reviewers raised and I left a few
>> comments in the doc.
>>
>> Best,
>> Peter
>>
>> On Wed, Sep 23, 2026 at 11:09 PM Martin Grund via dev <
>> [email protected]> wrote:
>>
>>> Thanks for taking the time to write this proposal. It's an exciting
>>> direction.
>>>
>>> The most important gap in the current proposal is to outline the
>>> security boundaries of the execution. The current proposal focuses too much
>>> on a rather vague definition of how the filters or masks are not elided,
>>> but that is not enough from a security perspective.
>>>
>>> As part of the SPIP, we should outline both what the guarantees are and
>>> how we plan on enforcing them. I have seen a lot of weirdness in the past
>>> when it comes to row filters and column masks.
>>>
>>> Just to give some examples:
>>>
>>> - How do we handle leaking expressions and what is the way to prevent
>>> information disclosure? (see
>>> https://rhaas.blogspot.com/2012/03/security-barrier-views.html)
>>> - How do we handle type mismatches between masks and colum types?
>>> - How do we handle time travel, clones, etc?
>>>
>>> Generally, I would recommend decoupling built-in masking functions from
>>> the overall proposal. Masking functions are just standard Spark built-ins,
>>> so once the SPIP is on its way, we can always add them to Spark without
>>> defining them in this proposal. I don't think having a fixed list of
>>> masking functions is useful. From my experience observing user behavior,
>>> there is a lot of complexity and variation in these functions. Giving users
>>> the ability to reference any UDF or built-in function is a much more
>>> convenient approach.
>>>
>>> Looking forward to seeing the updated proposal.
>>>
>>> Martin
>>>
>>> On Wed, Sep 23, 2026 at 7:28 PM Holden Karau <[email protected]> wrote:
>>>
>>>> So I think the security barrier / bypass path avoidance can be handled
>>>> in a few different ways (I've got one proposal I'm working on a draft off
>>>> but it's maybe a little early and shouldn't block this work).
>>>>
>>>> On 2026/09/23 16:48:33 Dongjoon Hyun wrote:
>>>> > Hi, Huaxin.
>>>> >
>>>> > Thank you for proposing this SPIP. I support the direction. Moving
>>>> row filtering and column masking from private Catalyst hacks to a public
>>>> DSv2 contract would be a clear improvement for the ecosystem, and the
>>>> proposed semantics are clear.
>>>> >
>>>> > Since the value of this SPIP rests on its "fail-closed" and
>>>> "optimizer-safe" guarantees, I'd like the SPIP to address the following,
>>>> grouped by priority.
>>>> >
>>>> > ### 1. Security guarantees (should be addressed before a vote)
>>>> >
>>>> > - **Security barrier**: The SPIP protects masks from the optimizer,
>>>> but not hidden rows from user expressions. Leaky user predicates (e.g. ANSI
>>>> errors, Python UDFs, predicates pushed to the connector) can be evaluated
>>>> before the policy filter and expose rows that should be hidden. The SPIP
>>>> should define the policy filter as a barrier, similar to PostgreSQL's
>>>> `security_barrier` and `LEAKPROOF`.
>>>> > - **DML read path**: DELETE/UPDATE/MERGE also read the target table.
>>>> Applying the policy on that read can silently delete hidden rows or write
>>>> masked values back. These operations should fail closed.
>>>> > - **Older Spark versions**: An engine that does not know the new
>>>> interface will silently ignore it, so the table is fail-open there. We need
>>>> a handshake between the engine and the connector.
>>>> > - **Bypass paths**: The policy is resolved once at analysis time and
>>>> then stays in the plan. The SPIP should define enforcement for paths that
>>>> reuse an analyzed plan or bypass the analyzer rewrite: global temp views,
>>>> streaming, catalog-side `Table` caching, `TableProvider` loads, V1
>>>> fallback, aggregate pushdown, and statistics.
>>>> >
>>>> > ### 2. Design
>>>> >
>>>> > - Restate "non-removable Project" as an invariant rather than a
>>>> special plan node. The invariant: a mask expression is never replaced by
>>>> the original attribute, while unused mask outputs can still be pruned.
>>>> > - Clarify the trust model when the policy filter is fully pushed down
>>>> to the connector.
>>>> > - Resolve mask functions safely. Temporary and session functions must
>>>> not be able to shadow a mask. Reuse existing built-ins where possible,
>>>> avoid the `hash` name collision, and drop weak hashes such as MD5.
>>>> > - Clean up the API: consistent null vs. empty semantics, type rules
>>>> for each mask function, a more specific name than `AccessControl`, and
>>>> consider V2 connector expressions instead of String function names.
>>>> >
>>>> > ### 3. Clarifications
>>>> >
>>>> > - Semantics for views (invoker vs. definer), time travel, nested and
>>>> metadata columns, and schema visibility of non-readable columns.
>>>> > - Add test cases for group 1 to Q8, and revisit the Q7 timeline
>>>> accordingly.
>>>> >
>>>> > I'm looking forward to the next revision. Thank you again for driving
>>>> this, Huaxin.
>>>> >
>>>> > Dongjoon.
>>>> >
>>>> > On 2026/09/23 08:48:42 Yang Jie wrote:
>>>> > > Thanks for putting this together. I'm supportive of the direction.
>>>> > >
>>>> > > This closes a real gap and I think putting enforcement in the
>>>> engine, fail-closed and optimizer-aware, is the right call, and keeping it
>>>> a vendor-neutral DSv2 capability lets Iceberg, Ranger, Polaris, and Delta
>>>> retire their private-API workarounds instead of each carrying its own.
>>>> > >
>>>> > > I have a few questions on the enforcement semantics and the
>>>> standard mask set that I will raise in the SPIP document. None of them
>>>> change my view on the direction.
>>>> > >
>>>> > > Thanks,
>>>> > > Jie Yang
>>>> > >
>>>> > > On 2026/09/23 03:39:25 huaxin gao wrote:
>>>> > > > Hi all,
>>>> > > >
>>>> > > > I would like to start a discussion on a SPIP that adds a public
>>>> DataSource
>>>> > > > V2
>>>> > > > API for a table or catalog to declare a read access policy
>>>> (readable
>>>> > > > columns,
>>>> > > > a row filter, and column masks) that Spark enforces in the query
>>>> plan. Here
>>>> > > > are
>>>> > > > the jira <https://issues.apache.org/jira/browse/SPARK-59726> and
>>>> SPIP doc
>>>> > > > <
>>>> https://docs.google.com/document/d/1hYEHORjHUnFzSoBY2DeOPILVJyu6iWXBusYwpF9FQJM/edit?tab=t.0
>>>> >
>>>> > > > .
>>>> > > >
>>>> > > > Today Spark has no supported API for this. Integrators inject
>>>> Catalyst rules
>>>> > > > through SparkSessionExtensions and build on internal Catalyst
>>>> APIs, as
>>>> > > > Apache
>>>> > > > Ranger (via the Kyuubi Spark AuthZ plugin) and a recent Apache
>>>> Iceberg POC
>>>> > > > both do. That is brittle across versions and unsafe: a masking
>>>> projection
>>>> > > > added naively to a plan can be removed or collapsed by the
>>>> optimizer,
>>>> > > > silently
>>>> > > > returning unmasked data.
>>>> > > >
>>>> > > > The proposal adds SupportsAccessControl, enforced during
>>>> analysis, with
>>>> > > > three
>>>> > > > guarantees: the masks cannot be optimized away, the row filter
>>>> always sees
>>>> > > > original pre-mask values, and anything Spark cannot resolve fails
>>>> the read
>>>> > > > rather than returning unprotected data.
>>>> > > >
>>>> > > > Details and proposed interfaces are in the doc. Feedback is very
>>>> welcome.
>>>> > > >
>>>> > > > Thanks,
>>>> > > > Huaxin
>>>> > > >
>>>> > >
>>>> > >
>>>> ---------------------------------------------------------------------
>>>> > > To unsubscribe e-mail: [email protected]
>>>> > >
>>>> > >
>>>> >
>>>> > ---------------------------------------------------------------------
>>>> > To unsubscribe e-mail: [email protected]
>>>> >
>>>> >
>>>>
>>>> ---------------------------------------------------------------------
>>>> To unsubscribe e-mail: [email protected]
>>>>
>>>>

Reply via email to