Hi, Huaxin.

Thank you for proposing this SPIP. I support the direction. Moving row 
filtering and column masking from private Catalyst hacks to a public DSv2 
contract would be a clear improvement for the ecosystem, and the proposed 
semantics are clear.

Since the value of this SPIP rests on its "fail-closed" and "optimizer-safe" 
guarantees, I'd like the SPIP to address the following, grouped by priority.

### 1. Security guarantees (should be addressed before a vote)

- **Security barrier**: The SPIP protects masks from the optimizer, but not 
hidden rows from user expressions. Leaky user predicates (e.g. ANSI errors, 
Python UDFs, predicates pushed to the connector) can be evaluated before the 
policy filter and expose rows that should be hidden. The SPIP should define the 
policy filter as a barrier, similar to PostgreSQL's `security_barrier` and 
`LEAKPROOF`.
- **DML read path**: DELETE/UPDATE/MERGE also read the target table. Applying 
the policy on that read can silently delete hidden rows or write masked values 
back. These operations should fail closed.
- **Older Spark versions**: An engine that does not know the new interface will 
silently ignore it, so the table is fail-open there. We need a handshake 
between the engine and the connector.
- **Bypass paths**: The policy is resolved once at analysis time and then stays 
in the plan. The SPIP should define enforcement for paths that reuse an 
analyzed plan or bypass the analyzer rewrite: global temp views, streaming, 
catalog-side `Table` caching, `TableProvider` loads, V1 fallback, aggregate 
pushdown, and statistics.

### 2. Design

- Restate "non-removable Project" as an invariant rather than a special plan 
node. The invariant: a mask expression is never replaced by the original 
attribute, while unused mask outputs can still be pruned.
- Clarify the trust model when the policy filter is fully pushed down to the 
connector.
- Resolve mask functions safely. Temporary and session functions must not be 
able to shadow a mask. Reuse existing built-ins where possible, avoid the 
`hash` name collision, and drop weak hashes such as MD5.
- Clean up the API: consistent null vs. empty semantics, type rules for each 
mask function, a more specific name than `AccessControl`, and consider V2 
connector expressions instead of String function names.

### 3. Clarifications

- Semantics for views (invoker vs. definer), time travel, nested and metadata 
columns, and schema visibility of non-readable columns.
- Add test cases for group 1 to Q8, and revisit the Q7 timeline accordingly.

I'm looking forward to the next revision. Thank you again for driving this, 
Huaxin.

Dongjoon.

On 2026/09/23 08:48:42 Yang Jie wrote:
> Thanks for putting this together. I'm supportive of the direction.
> 
> This closes a real gap and I think putting enforcement in the engine, 
> fail-closed and optimizer-aware, is the right call, and keeping it a 
> vendor-neutral DSv2 capability lets Iceberg, Ranger, Polaris, and Delta 
> retire their private-API workarounds instead of each carrying its own.
> 
> I have a few questions on the enforcement semantics and the standard mask set 
> that I will raise in the SPIP document. None of them change my view on the 
> direction.
> 
> Thanks,
> Jie Yang
> 
> On 2026/09/23 03:39:25 huaxin gao wrote:
> > Hi all,
> > 
> > I would like to start a discussion on a SPIP that adds a public DataSource
> > V2
> > API for a table or catalog to declare a read access policy (readable
> > columns,
> > a row filter, and column masks) that Spark enforces in the query plan. Here
> > are
> > the jira <https://issues.apache.org/jira/browse/SPARK-59726> and SPIP doc
> > <https://docs.google.com/document/d/1hYEHORjHUnFzSoBY2DeOPILVJyu6iWXBusYwpF9FQJM/edit?tab=t.0>
> > .
> > 
> > Today Spark has no supported API for this. Integrators inject Catalyst rules
> > through SparkSessionExtensions and build on internal Catalyst APIs, as
> > Apache
> > Ranger (via the Kyuubi Spark AuthZ plugin) and a recent Apache Iceberg POC
> > both do. That is brittle across versions and unsafe: a masking projection
> > added naively to a plan can be removed or collapsed by the optimizer,
> > silently
> > returning unmasked data.
> > 
> > The proposal adds SupportsAccessControl, enforced during analysis, with
> > three
> > guarantees: the masks cannot be optimized away, the row filter always sees
> > original pre-mask values, and anything Spark cannot resolve fails the read
> > rather than returning unprotected data.
> > 
> > Details and proposed interfaces are in the doc. Feedback is very welcome.
> > 
> > Thanks,
> > Huaxin
> > 
> 
> ---------------------------------------------------------------------
> To unsubscribe e-mail: [email protected]
> 
> 

---------------------------------------------------------------------
To unsubscribe e-mail: [email protected]

Reply via email to