Huaxin, thank you for proposing this SPIP. I think this feature is
necessary, and Spark is the right place to enforce it. I left a comment in
the doc.

On Thu, Sep 24, 2026 at 5:48 PM Peter Toth <[email protected]> wrote:

> Thanks for the proposal Huaxin. I like the direction.
> I agree with the points earlier reviewers raised and I left a few comments
> in the doc.
>
> Best,
> Peter
>
> On Wed, Sep 23, 2026 at 11:09 PM Martin Grund via dev <
> [email protected]> wrote:
>
>> Thanks for taking the time to write this proposal. It's an exciting
>> direction.
>>
>> The most important gap in the current proposal is to outline the security
>> boundaries of the execution. The current proposal focuses too much on a
>> rather vague definition of how the filters or masks are not elided, but
>> that is not enough from a security perspective.
>>
>> As part of the SPIP, we should outline both what the guarantees are and
>> how we plan on enforcing them. I have seen a lot of weirdness in the past
>> when it comes to row filters and column masks.
>>
>> Just to give some examples:
>>
>> - How do we handle leaking expressions and what is the way to prevent
>> information disclosure? (see
>> https://rhaas.blogspot.com/2012/03/security-barrier-views.html)
>> - How do we handle type mismatches between masks and colum types?
>> - How do we handle time travel, clones, etc?
>>
>> Generally, I would recommend decoupling built-in masking functions from
>> the overall proposal. Masking functions are just standard Spark built-ins,
>> so once the SPIP is on its way, we can always add them to Spark without
>> defining them in this proposal. I don't think having a fixed list of
>> masking functions is useful. From my experience observing user behavior,
>> there is a lot of complexity and variation in these functions. Giving users
>> the ability to reference any UDF or built-in function is a much more
>> convenient approach.
>>
>> Looking forward to seeing the updated proposal.
>>
>> Martin
>>
>> On Wed, Sep 23, 2026 at 7:28 PM Holden Karau <[email protected]> wrote:
>>
>>> So I think the security barrier / bypass path avoidance can be handled
>>> in a few different ways (I've got one proposal I'm working on a draft off
>>> but it's maybe a little early and shouldn't block this work).
>>>
>>> On 2026/09/23 16:48:33 Dongjoon Hyun wrote:
>>> > Hi, Huaxin.
>>> >
>>> > Thank you for proposing this SPIP. I support the direction. Moving row
>>> filtering and column masking from private Catalyst hacks to a public DSv2
>>> contract would be a clear improvement for the ecosystem, and the proposed
>>> semantics are clear.
>>> >
>>> > Since the value of this SPIP rests on its "fail-closed" and
>>> "optimizer-safe" guarantees, I'd like the SPIP to address the following,
>>> grouped by priority.
>>> >
>>> > ### 1. Security guarantees (should be addressed before a vote)
>>> >
>>> > - **Security barrier**: The SPIP protects masks from the optimizer,
>>> but not hidden rows from user expressions. Leaky user predicates (e.g. ANSI
>>> errors, Python UDFs, predicates pushed to the connector) can be evaluated
>>> before the policy filter and expose rows that should be hidden. The SPIP
>>> should define the policy filter as a barrier, similar to PostgreSQL's
>>> `security_barrier` and `LEAKPROOF`.
>>> > - **DML read path**: DELETE/UPDATE/MERGE also read the target table.
>>> Applying the policy on that read can silently delete hidden rows or write
>>> masked values back. These operations should fail closed.
>>> > - **Older Spark versions**: An engine that does not know the new
>>> interface will silently ignore it, so the table is fail-open there. We need
>>> a handshake between the engine and the connector.
>>> > - **Bypass paths**: The policy is resolved once at analysis time and
>>> then stays in the plan. The SPIP should define enforcement for paths that
>>> reuse an analyzed plan or bypass the analyzer rewrite: global temp views,
>>> streaming, catalog-side `Table` caching, `TableProvider` loads, V1
>>> fallback, aggregate pushdown, and statistics.
>>> >
>>> > ### 2. Design
>>> >
>>> > - Restate "non-removable Project" as an invariant rather than a
>>> special plan node. The invariant: a mask expression is never replaced by
>>> the original attribute, while unused mask outputs can still be pruned.
>>> > - Clarify the trust model when the policy filter is fully pushed down
>>> to the connector.
>>> > - Resolve mask functions safely. Temporary and session functions must
>>> not be able to shadow a mask. Reuse existing built-ins where possible,
>>> avoid the `hash` name collision, and drop weak hashes such as MD5.
>>> > - Clean up the API: consistent null vs. empty semantics, type rules
>>> for each mask function, a more specific name than `AccessControl`, and
>>> consider V2 connector expressions instead of String function names.
>>> >
>>> > ### 3. Clarifications
>>> >
>>> > - Semantics for views (invoker vs. definer), time travel, nested and
>>> metadata columns, and schema visibility of non-readable columns.
>>> > - Add test cases for group 1 to Q8, and revisit the Q7 timeline
>>> accordingly.
>>> >
>>> > I'm looking forward to the next revision. Thank you again for driving
>>> this, Huaxin.
>>> >
>>> > Dongjoon.
>>> >
>>> > On 2026/09/23 08:48:42 Yang Jie wrote:
>>> > > Thanks for putting this together. I'm supportive of the direction.
>>> > >
>>> > > This closes a real gap and I think putting enforcement in the
>>> engine, fail-closed and optimizer-aware, is the right call, and keeping it
>>> a vendor-neutral DSv2 capability lets Iceberg, Ranger, Polaris, and Delta
>>> retire their private-API workarounds instead of each carrying its own.
>>> > >
>>> > > I have a few questions on the enforcement semantics and the standard
>>> mask set that I will raise in the SPIP document. None of them change my
>>> view on the direction.
>>> > >
>>> > > Thanks,
>>> > > Jie Yang
>>> > >
>>> > > On 2026/09/23 03:39:25 huaxin gao wrote:
>>> > > > Hi all,
>>> > > >
>>> > > > I would like to start a discussion on a SPIP that adds a public
>>> DataSource
>>> > > > V2
>>> > > > API for a table or catalog to declare a read access policy
>>> (readable
>>> > > > columns,
>>> > > > a row filter, and column masks) that Spark enforces in the query
>>> plan. Here
>>> > > > are
>>> > > > the jira <https://issues.apache.org/jira/browse/SPARK-59726> and
>>> SPIP doc
>>> > > > <
>>> https://docs.google.com/document/d/1hYEHORjHUnFzSoBY2DeOPILVJyu6iWXBusYwpF9FQJM/edit?tab=t.0
>>> >
>>> > > > .
>>> > > >
>>> > > > Today Spark has no supported API for this. Integrators inject
>>> Catalyst rules
>>> > > > through SparkSessionExtensions and build on internal Catalyst
>>> APIs, as
>>> > > > Apache
>>> > > > Ranger (via the Kyuubi Spark AuthZ plugin) and a recent Apache
>>> Iceberg POC
>>> > > > both do. That is brittle across versions and unsafe: a masking
>>> projection
>>> > > > added naively to a plan can be removed or collapsed by the
>>> optimizer,
>>> > > > silently
>>> > > > returning unmasked data.
>>> > > >
>>> > > > The proposal adds SupportsAccessControl, enforced during analysis,
>>> with
>>> > > > three
>>> > > > guarantees: the masks cannot be optimized away, the row filter
>>> always sees
>>> > > > original pre-mask values, and anything Spark cannot resolve fails
>>> the read
>>> > > > rather than returning unprotected data.
>>> > > >
>>> > > > Details and proposed interfaces are in the doc. Feedback is very
>>> welcome.
>>> > > >
>>> > > > Thanks,
>>> > > > Huaxin
>>> > > >
>>> > >
>>> > > ---------------------------------------------------------------------
>>> > > To unsubscribe e-mail: [email protected]
>>> > >
>>> > >
>>> >
>>> > ---------------------------------------------------------------------
>>> > To unsubscribe e-mail: [email protected]
>>> >
>>> >
>>>
>>> ---------------------------------------------------------------------
>>> To unsubscribe e-mail: [email protected]
>>>
>>>

Reply via email to