bharos commented on code in PR #12757:
URL: https://github.com/apache/gravitino/pull/12757#discussion_r4149947766


##########
design-docs/tag-based-access-control.md:
##########
@@ -0,0 +1,767 @@
+<!--
+  Licensed to the Apache Software Foundation (ASF) under one
+  or more contributor license agreements.  See the NOTICE file
+  distributed with this work for additional information
+  regarding copyright ownership.  The ASF licenses this file
+  to you under the Apache License, Version 2.0 (the
+  "License"); you may not use this file except in compliance
+  with the License.  You may obtain a copy of the License at
+
+   http://www.apache.org/licenses/LICENSE-2.0
+
+  Unless required by applicable law or agreed to in writing,
+  software distributed under the License is distributed on an
+  "AS IS" BASIS, WITHOUT WARRANTIES OR CONDITIONS OF ANY
+  KIND, either express or implied.  See the License for the
+  specific language governing permissions and limitations
+  under the License.
+-->
+
+# Design of Tag-Based Access Control in Gravitino
+
+**Status:** draft for discussion. The [open questions](#open-questions) are 
deliberately left
+undecided in this revision, each presented with its options; decisions will be 
folded in after
+review.
+
+Discussion: [#12619](https://github.com/apache/gravitino/discussions/12619)
+
+---
+
+## Summary
+
+An access rule is a `Policy` of type `system_access_control` whose `content` 
carries a set of
+privileges and a role condition. The policy is bound to a tag. Any object 
carrying that tag
+becomes subject to the rule.
+
+```json
+POST /api/metalakes/prod/policies
+{
+  "name": "certified_access",
+  "policyType": "system_access_control",
+  "enabled": true,
+  "content": {
+    "privileges": ["SELECT_TABLE", "MODIFY_TABLE"],
+    "applicableRoles": ["analyst", "data_engineer"]
+  }
+}
+```
+
+```
+PUT  /api/metalakes/prod/tags/certified/policies/certified_access
+     { "selector": { "type": "ALL_VALUES" } }
+
+POST /api/metalakes/prod/objects/TABLE/lakehouse.finance.orders/tags
+     { "tagsToAdd": [{ "name": "certified" }] }
+```
+
+Read together: *a caller holding either `analyst` or `data_engineer` may 
select from and modify
+any table that carries the tag `certified`.* Any listed role satisfies the 
condition, and every
+listed privilege is conferred to it.
+
+Only one thing is attached: the policy to the tag. The roles are values inside 
`content`, not a
+second link. No new user-facing entity, REST resource or client API is 
introduced.
+
+---
+
+## Background
+
+Gravitino authorizes metadata operations through RBAC. A grant names a 
securable object and a
+privilege and binds them to a role; the authorization expression on each REST 
endpoint evaluates
+those grants over the object's ancestor chain.
+
+Tags are a separate subsystem. They apply to catalogs, schemas, tables, views, 
topics, filesets,
+models, columns and functions, carry assignment values (see
+[tag-assignment-values.md](tag-assignment-values.md)), and inherit down the 
object hierarchy.
+Policy-on-tag ([policy-on-tag.md](policy-on-tag.md)) lets governance policies 
be selected by those
+tags. Authorization does not read tags at all.
+
+So a label cannot drive access. An organization that already tags tables 
`certified`, `pii` or
+`data_domain=finance` must still issue grants object by object to act on those 
tags. New objects
+need new grants, dropped objects leave stale ones, and the rule itself is 
written down nowhere — it
+exists only as the pile of grants someone remembered to issue.
+
+---
+
+## Scope
+
+### In this version
+
+- A rule of the form *(action, role condition)* bound to a tag.
+- `ALLOW` only.
+- Roles as the matched condition.
+- Evaluation inside the existing authorization-expression path, composing with 
RBAC.
+- Reuse of the `Policy` entity, the policy-to-tag relation and 
`PolicySelector`, so tag conditions
+  are written identically for governance and for authorization.
+- No new REST resource or client API. The only storage addition is an internal 
derived index, not
+  written or read by any endpoint — see [Lifecycle](#lifecycle).
+
+### Not in this version
+
+| Excluded                                                      | Reason       
                                                                                
                                                                                
                                                                                
                                                                                
                                                                                
                                                       |
+| ------------------------------------------------------------- | 
-----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------
 |
+| `DENY`                                                        | A deny that 
a descendant tag cannot undo is a separate problem. Allow-only means every rule 
set has an answer and the order rules are applied in never matters.             
                                                                                
                                                                                
                                                                                
                                                        |
+| Users and groups as the matched condition                     | The 
condition schema can gain them later without changing the model.                
                                                                                
                                                                                
                                                                                
                                                                                
                                                                |
+| Row filtering and column masking                              | Distinct 
policy types; this design governs whole-object decisions.                       
                                                                                
                                                                                
                                                                                
                                                                                
                                                           |
+| Cross-tag conditions                                          | A rule 
matches one tag. Conditions spanning several tags await the `EXPRESSION` 
selector type in [policy-on-tag.md](policy-on-tag.md).                          
                                                                                
                                                                                
                                                                                
                                                                    |
+| Column-level decisions                                        | A tag on a 
column does not affect decisions about its table.                               
                                                                                
                                                                                
                                                                                
                                                                                
                                                         |
+| A `scope` field in `content`, restricting a rule to a subtree | Not a 
security boundary: creating a policy already needs metalake-wide 
`CREATE_POLICY`, so whoever writes the rule chooses its reach anyway. It is 
also not checked when a tag is applied, so it limits where a rule takes effect 
rather than stopping a wrong tag. One tag meaning different things in different 
subtrees is already covered by tag assignment values with a value-sensitive 
selector. Can be added later, since an absent `scope` has always meant 
metalake-wide. |
+| Replacing RBAC                                                | Baseline 
privileges, ownership and traversal are unchanged. See [Composition with 
RBAC](#composition-with-rbac).                                                  
                                                                                
                                                                                
                                                                                
                                                                  |
+
+---
+
+## Alternatives considered
+
+| Option                                                                 | 
Pros                                                                            
                                 | Cons                                         
                                                                                
| Status       |
+| ---------------------------------------------------------------------- | 
----------------------------------------------------------------------------------------------------------------
 | 
----------------------------------------------------------------------------------------------------------------------------
 | ------------ |
+| **A `system_access_control` policy type bound to a tag**               | 
Reuses the entity, relation, selector and resolver; no new REST or client 
surface; one governance model to learn | The role condition lives in `content` 
JSON, so lookup by role needs a derived index rather than a foreign key         
       | **Proposed** |
+| A dedicated `tag_access_policy` entity with action and role as columns | 
Foreign key on role; indexed lookup; cascade on role deletion falls out of the 
schema                            | New table across three dialects, new REST 
resource, new client and CLI surface, a second governance model alongside 
policies | Rejected     |
+| Extend RBAC grants with a tag predicate                                | No 
new concepts                                                                    
                              | The grant table is object-identified; a 
predicate has no object, and every grant read path would change                 
     | Rejected     |
+| Evaluate tags in an external engine (OPA and similar)                  | 
Arbitrary policy language                                                       
                                 | Moves the decision out of Gravitino, 
duplicates the tag hierarchy, and cannot use the existing expression path       
        | Rejected     |
+
+That one con means the server keeps the role reference consistent, rather than 
the database schema
+doing it. [Lifecycle](#lifecycle) covers how.
+
+---
+
+## Model
+
+### Content
+
+`PolicyContent` is an interface, and each built-in policy type has a concrete 
implementation with
+typed fields and a `validate()` that runs at write time. 
`IcebergDataCompactionContent` is the
+existing example. `system_access_control` follows the same pattern with a new
+`AccessControlContent`:
+
+| Field              | Type                     | Meaning                      
                                                                                
       |
+| ------------------ | ------------------------ | 
-------------------------------------------------------------------------------------------------------------------
 |
+| `privileges`       | list of `Privilege.Name` | The privileges the rule 
confers. Each must be a permitted name — see [Permitted 
privileges](#permitted-privileges). |
+| `applicableRoles`  | list of role names       | The **condition**. Satisfied 
when any listed role is among the caller's expanded roles.                      
       |
+
+`validate()` rejects at creation rather than at evaluation:
+
+- `privileges` is non-empty and every entry parses to a permitted 
`Privilege.Name`.
+- `applicableRoles` is non-empty and every name is non-blank.
+
+Rejecting at write time matters because the alternative failure is silent: a 
policy naming a
+privilege that does not parse simply grants nothing, and nothing surfaces 
until someone notices
+the access they expected is missing.
+
+Whether `validate()` also requires the named role to *exist* is part of
+[OQ-3](#oq-3--deleting-a-referenced-role), not a separate decision.
+
+### Permitted privileges
+
+Parsing to a `Privilege.Name` is a syntax check, not a safety one. A rule may 
confer only
+privileges that grant access to the tagged object itself, so `validate()` 
checks each name against
+a fixed allowlist:
+
+| Object   | Permitted                                          |
+| -------- | -------------------------------------------------- |
+| Table    | `SELECT_TABLE`, `MODIFY_TABLE`, `PROBE_TABLE_LIKE` |
+| View     | `SELECT_VIEW`                                      |
+| Fileset  | `READ_FILESET`, `WRITE_FILESET`                    |
+| Topic    | `CONSUME_TOPIC`, `PRODUCE_TOPIC`                   |
+| Model    | `USE_MODEL`                                        |
+| Function | `EXECUTE_FUNCTION`                                 |
+
+An allowlist rather than a denylist so the boundary fails closed: a privilege 
added to
+`Privilege.Name` later confers nothing through a tag until someone adds it 
here deliberately.
+
+Everything else is rejected. Two classes are worth naming because the reasons 
differ:
+
+- **Authority over the authorization system** — `MANAGE_USERS`, 
`MANAGE_GROUPS`, `MANAGE_GRANTS`,
+  `CREATE_ROLE`, `CREATE_TAG`, `APPLY_TAG`, `CREATE_POLICY`, `APPLY_POLICY`. 
These turn one tagging
+  operation into a standing ability to widen access. `MANAGE_GRANTS` binds to 
every taggable type
+  and covers all children of whatever it binds to, so a tag carrying it on one 
catalog would let
+  every role in `applicableRoles` grant anything beneath that catalog — and 
would keep doing so
+  after the applier's own authority was revoked. `APPLY_TAG` and 
`APPLY_POLICY` close the loop
+  further, letting a conferred role extend the tag system's own reach.
+- **Traversal** — `USE_CATALOG` and `USE_SCHEMA`, for the reasons in
+  [Traversal stays RBAC](#traversal-stays-rbac).
+
+Both exclusions are the same argument: a tag confers access to data, never the 
ability to hand out
+access or to reach new territory. 
[OQ-4](#oq-4--authority-to-confer-access-through-a-tag) governs
+who may apply a rule and does not substitute for this, because the applier 
holds the authority
+privilege by construction — the check they pass is exactly the one the rule 
would make permanent.
+
+### `applicableRoles` is a condition, not a principal
+
+The rule does not grant anything to `analyst`. It states that *if* the caller 
holds `analyst`
+among their expanded roles *and* the object carries `certified`, then 
`SELECT_TABLE` is satisfied
+for this request.
+
+The distinction matters for two reasons. The rule is not a grant, so it does 
not appear in the
+role's securable objects and does not participate in grant listing. And a role 
that is never
+assigned to anyone confers nothing, exactly as an unassigned role does today.
+
+### The tag bind
+
+The policy is attached to the tag through the existing policy-to-tag relation 
and its selector,
+exactly as governance policies are. `ALL_VALUES` in the example above matches 
the tag regardless of
+assignment value; value-sensitive selectors work as they do for governance 
policies, and nothing in
+this design is specific to `ALL_VALUES`.
+
+---
+
+## Evaluation
+
+An authorization decision needs to know, for the object being accessed and the 
caller's roles,
+whether any access rule is satisfied. That requires three things:
+
+1. the tags effective at the object after nearest-wins resolution
+   ([tag-assignment-values.md](tag-assignment-values.md)), including those 
inherited from ancestors;
+2. the `system_access_control` policies bound to those tags;
+3. for each, whether any of `applicableRoles` is among the caller's expanded 
roles.
+
+Access rules only allow — `content` has a role condition but no deny effect — 
so a tag cannot
+restrict or deny. An RBAC `DENY` is unaffected; see [Allow and 
deny](#allow-and-deny).
+
+The question is *where* steps 1 and 2 happen.
+
+### Proposed: check tags at the privilege leaf
+
+Every privilege check bottoms out in `GravitinoAuthorizer.authorize`. The 
expression converter
+expands each `ANY_*` macro mechanically —
+
+```
+ANY_USE_CATALOG → ANY(USE_CATALOG, METALAKE, CATALOG) && 
!ANY(DENY_USE_CATALOG, METALAKE, CATALOG)
+```
+
+— and `hasAuthorizeWithoutDeny` walks the object's ancestor chain calling 
`authorize` and `deny` at
+each level. Tag evaluation goes inside `authorize`: when the RBAC rows do not 
allow, resolve the
+effective tags of the object being decided, load the access policies bound to 
them, and test those
+against the caller's roles.
+
+Three properties follow from the surrounding code rather than from a rule this 
design has to write:
+
+- **RBAC deny still wins.** The `!ANY(DENY_…)` conjunct is built from `deny`, 
which the tag path
+  never touches, so an allow-only tag cannot reach it. See
+  [OQ-2](#oq-2--composition-when-a-tag-allows-and-rbac-denies).
+- **Traversal stays RBAC.** Each conjunct consults tags independently, so a 
tag granting
+  `SELECT_TABLE` still cannot bypass `USE_CATALOG`.
+- **A denial stays attributable.** The tag check is a distinct step, so the 
information needed to
+  explain a decision stays separable from the grant that would otherwise have 
produced it.
+
+One rule does not come for free. Role assumption narrows a request to the 
roles the caller
+activated, but that narrowing lives inside `enforceNarrowed`, on the jCasbin 
path the tag check
+does not take. The tag check therefore applies it itself: `applicableRoles` is 
tested against the
+caller's *active* roles, and an `ActiveRoles.none()` request grants nothing. 
Otherwise a caller who
+narrowed would silently keep tag-derived access they had asked to drop.
+
+Inheritance is the one thing the walk does not supply. `authorize` resolves 
the object's effective
+tags once ([tag-assignment-values.md](tag-assignment-values.md)) rather than 
asking each level in
+turn. Different tag names still union down the chain; nearest-wins settles 
only the same name
+assigned at two levels, where the nearer assignment wins and the farther one 
is dropped:
+
+```
+catalog lakehouse        certified = gold      pii = true
+table   finance.orders   certified = bronze
+
+effective on the table   certified = bronze    pii = true
+```
+
+`pii` is inherited; `certified=gold` is gone because the table overrode it. So 
a rule bound with
+`TAG_VALUE("gold")` does not match, while asking level by level would still 
find `gold` on the
+catalog and grant. The two readings agree under `ALL_VALUES`, where only the 
presence of the name
+matters, and diverge as soon as a rule reads the value.
+
+One constraint on where the check hooks in. `hasAuthorizeWithoutDeny` walks 
the ancestor chain
+calling `authorize` at each level, so a check placed inside that per-level 
call would resolve
+effective tags once per level, each resolution walking its own chain — 
quadratic in chain depth,
+and for nothing, since the leaf's effective tags already subsume every 
ancestor's. The check runs
+once for the object under decision, not once per level of the RBAC walk.
+
+The cost lands on the request path, and is set out in [Cost](#cost).
+
+### Decision flow
+
+The following flow describes an ordinary tag-conferrable privilege check. 
Existing ownership
+branches and the enclosing authorization expression remain responsible for the 
complete endpoint
+decision, including RBAC-only traversal checks.
+
+```mermaid
+flowchart TD
+  A[Privilege check] --> B{Tag access enabled?}
+  B -->|No| C[Existing RBAC decision]
+  B -->|Yes| D{RBAC allows without an applicable deny?}
+  D -->|Yes| E[Privilege satisfied]
+  D -->|No| F[Resolve effective tags and enabled access policies once per 
object]
+  F --> G{Selector, privilege and active role match?}
+  G -->|No| H[Privilege not satisfied]
+  G -->|Yes| I{Explicit RBAC deny applies?}
+  I -->|No| E
+  I -->|Yes| J[Privilege not satisfied; record suppressed tag allow]
+```
+
+This combines both sources of authority without requiring both to run on every 
request. Tags
+cannot revoke an RBAC allow, so skipping them after a successful RBAC decision 
preserves the
+result. An RBAC miss is not an explicit deny: it gives tags an opportunity to 
grant access. An
+explicit deny remains effective even when a tag matches. This fast path 
depends on v1 being
+allow-only; adding restrictive tag policies would require revisiting it.
+
+### Independently testable evaluation
+
+M3 separates the evaluation into three components with explicit inputs:
+
+- **Resolution:** load tag assignments and policy bindings through an 
injectable reader, and
+  resolve nearest-wins inheritance. Point checks and list preloads produce the 
same resolved
+  representation. Tests can supply an in-memory hierarchy and bindings without 
a catalog or
+  database.
+- **Matching:** a deterministic evaluator takes the requested privilege, 
active role set and
+  resolved tags and enabled policies. It returns a match or no match, with the 
matching tag and
+  policy identifiers. It performs no storage reads, role expansion or cache 
mutation and has no
+  dependency on REST or jCasbin.
+- **Composition:** the authorization adapter combines the match with the 
existing RBAC allow,
+  deny and traversal checks. It also records decision provenance. Storage 
failures are errors,
+  not empty policy sets or successful matches; an evaluation error cannot 
grant tag-derived
+  access.
+
+The matching result is internal evidence, not a new public authorization 
decision or API. Keeping
+it separate lets tests assert both the decision and its explanation without 
starting the server.
+
+M3 is complete only with unit tests for the resolver, matcher and composition 
adapter, plus
+integration tests through the authorization-expression and list-filtering 
paths. Cover:
+
+- multiple privileges and any matching active role; no match for inactive 
roles or
+  `ActiveRoles.none()`;
+- same-name nearest-wins overrides, different-name inheritance, `ALL_VALUES` 
and value-sensitive
+  selectors, disabled policies and absent bindings;
+- RBAC allow, RBAC miss, explicit deny with a matching tag, missing traversal 
privileges, and the
+  feature disabled; a skipped tag path must perform no tag-resolution reads;
+- resolution failures, attribution of a suppressed tag allow, and identical 
point/list decisions
+  for the same caller, object and privilege;
+- one resolution per object per request despite repeated privilege checks, 
including an ancestor
+  value overridden at the target; the RBAC ancestor walk must not restore the 
overridden value.
+
+M4 adds multi-node revocation and freshness tests. M5 adds query-count tests 
and list benchmarks
+across candidate counts, hierarchy depths and tag counts. Record the RBAC-only 
baseline and
+feature-enabled latency and query counts so later changes to the model can be 
checked against
+both correctness and cost.
+
+### Rejected: expand tag rows when roles load
+
+The alternative writes one permission row per (role, object) when a role's 
policies load, so tag
+permissions are indistinguishable from RBAC ones at decision time.
+
+It fails on cardinality. The jCasbin matcher compares `metadataId` for 
equality with no prefix
+form, so a tag on a catalog grants on a table only if a row exists for that 
table. Tagging one
+catalog materialises a row per descendant per affected role, and every later 
`CREATE TABLE` beneath
+it has to add rows — a write-path dependency on the authorizer that does not 
exist today. Miss one
+and stale rows keep granting, which errs towards more access rather than less.
+
+### Freshness
+
+A node that has already loaded the affected roles still has to learn that tag 
state changed.
+
+Today the authorizer keeps role policies fresh by version-checking on read: 
`loadedRoles` maps role
+id to `updated_at`, and a newer `role_meta.updated_at` in the database evicts 
and reloads that
+role's policies. `groupRoleCache` is validated the same way against 
`group_meta.updated_at`. Write
+paths additionally call `handleRolePrivilegeChange`, `handleUserRoleRelChange` 
and
+`handleGroupRoleRelChange` in-process on the node that performed the write, 
and TTL bounds the rest.
+`JcasbinChangeListener` covers two further surfaces: entity changes through 
`onEntityChange`, and a
+poll of `owner_meta`.
+
+Tag state reaches none of that, and the transport differs by what changed:
+
+| Change                                         | Reaches other nodes today   
                                                                                
               |
+| ---------------------------------------------- | 
--------------------------------------------------------------------------------------------------------------------------
 |
+| Tag or policy entity created, altered, dropped | Yes — `entity_change_log` 
carries `TAG` and `POLICY`, but `JcasbinChangeListener` discards both as 
virtual-namespace types |
+| Tag applied to or removed from an object       | No — relation changes emit 
no change-log rows                                                              
                |
+| Policy bound to or unbound from a tag          | No — same                   
                                                                                
               |
+
+The first needs the existing filter relaxed. The second and third need a 
transport that does not
+exist yet: relation changes emitted into `entity_change_log`, a poll of the 
relation tables, or a
+TTL accepted as the bound. The rejected option needs the same three signals, 
and reacts to each by
+rewriting rows rather than by dropping a cache entry.
+
+See [OQ-1](#oq-1--where-tags-are-evaluated).
+
+### List filtering
+
+List endpoints filter their results through the same authorizer, so 
tag-derived permissions must be
+visible to filtering as well as to single-object checks.
+
+Filtering has a fast path today: `allVisibleViaParentScope` skips the 
per-object loop entirely when
+a parent-scope grant makes every candidate visible and no object-level deny 
exists. Tags only add
+access, so that short-circuit stays correct untouched — if RBAC already shows 
everything, no tag
+can change the answer.
+
+When it misses, the per-object loop runs, and that loop is deliberately free 
of per-object queries:
+`preloadToCache` batch-gets the entities and `preloadOwner` batch-gets their 
owners, leaving the
+loop as in-memory evaluation. Resolving tags per candidate breaks that 
property and turns one
+listing into N walks of the ancestor chain. Filtering needs a batch preload of 
tag and policy state
+for the candidate set, alongside those two paths — not as an optimisation, but 
to keep an invariant
+the loop already has.
+
+### Cost
+
+The feature is off by default, and while off the authorizer does not consult 
tags at all, so a
+deployment that leaves it off pays nothing and the rest of this section does 
not apply to it. See
+[Enabling the feature](#enabling-the-feature).
+
+With it on, tags are consulted only when the RBAC rows do not already allow, 
so a check RBAC grants
+costs nothing. The cost falls on the not-allowed branch — which is the common 
branch for list
+filtering, where most candidates are objects the caller cannot see. The 
feature therefore makes
+"no" more expensive than "yes", inverting the shape RBAC has today.
+
+On that branch, one check resolves the object's effective tags, then loads the 
policies bound to
+them:
+
+| Step                                | Cost today                             
                                                                                
                        |
+| ----------------------------------- | 
----------------------------------------------------------------------------------------------------------------------------------------------
 |
+| Resolve the object's effective tags | One relation query per level of the 
ancestor chain. Nothing caches the result — `RelationalEntityStore` caches 
entities, not relation queries. |
+| Load the policies bound to each tag | One `POLICY_METADATA_OBJECT_REL` query 
per distinct effective tag.                                                     
                        |
+| Test `applicableRoles`              | In memory, against roles the request 
has already loaded.                                                             
                          |
+
+The chain is not bounded by a constant. `getParentMetadataObjects` expands a 
hierarchical schema
+one level at a time, so a column under `catalog.a:b.table` walks five levels, 
and a deeper schema
+walks more:
+
+```
+COLUMN:catalog.a:b.table.col -> TABLE:catalog.a:b.table -> SCHEMA:catalog.a:b
+                             -> SCHEMA:catalog.a -> CATALOG:catalog
+```
+
+Two costs the query count hides. `TagManager.listTagsInfoForMetadataObject` and
+`PolicyManager.listPolicyInfosForMetadataObject` each take a tree read lock 
and call
+`checkMetadataObject`, whose existence check can reach the underlying catalog 
— so the evaluator
+has to query the relation directly rather than reuse them. And 
`batchListEntitiesByRelation`
+implements only `OWNER_REL` today, so the batch preload below is a new query, 
not a reuse.
+
+There is also a write-side cost that predates this design: applying or 
removing a tag invalidates
+the entity cache for both endpoints, and that invalidation cascades down the 
identifier hierarchy,
+so tagging a catalog drops every cached schema and table beneath it.
+
+Three things get worse when the feature is on. The second is a trade-off the 
feature asks for
+rather than a defect, but it should be a deliberate choice rather than a 
surprise:
+
+- **Listings that miss the parent-scope short-circuit.** The per-object loop 
issues no per-object
+  queries today, provided the preload runs — `preloadToCache` returns early 
when the entity cache
+  is disabled or the type is not preloadable. Each not-allowed candidate adds 
a chain walk plus a
+  policy query per tag, so a listing of N objects goes from a batch preload 
and in-memory
+  evaluation to O(N × d) queries against the relational store. Latency on a 
single call is not the
+  concern; connection-pool pressure under concurrency is.
+- **Granting through tags instead of coarse RBAC makes that miss more often.** 
The short-circuit
+  fires on a parent-scope RBAC grant. A deployment that replaces those grants 
with tag-derived
+  access removes the condition the short-circuit tests, so listings that are 
free today take the
+  slow path.
+- **Tag churn degrades RBAC.** The invalidation above is not scoped to the tag 
path — the entity
+  cache is shared, so retagging a catalog costs every request that reads 
entities beneath it,
+  including requests that never touch a tag.
+
+Two guards keep the not-allowed branch cheap where no tag could grant anyway: 
no enabled
+`system_access_control` policy in the metalake, resolved once per request; and 
no active roles on
+the request, since `applicableRoles` is tested against active roles and an 
empty set matches
+nothing. The first has no list-by-type query today, so it lists the metalake's 
policies.
+
+Three mitigations, in increasing order of what they cost to build:
+
+- **Per-request memoisation.** `AuthorizationRequestContext` already memoises 
the final decision,
+  but keys it on principal, metalake, object and privilege, while the 
expensive part — effective
+  tags and the policies bound to them — depends on neither the principal nor 
the privilege.
+  Memoising that per object keeps a request checking several privileges on one 
object to a single
+  walk.
+- **Batch preload for lists.** Resolve the candidate set's tag and policy 
state in one round trip
+  rather than N walks, as above.
+- **A cache across requests.** Not designed here, and not yet designable: 
caching tag state beyond
+  one request needs an invalidation signal, and two of the three signals in
+  [Freshness](#freshness) have no carrier today. Performance and freshness are 
the same problem.
+
+The first two land in M3 and M5. The third becomes possible once M4 does, and 
until then the
+uncached cost above is what the feature costs.
+
+---
+
+## Composition with RBAC
+
+### Traversal stays RBAC
+
+`USE_CATALOG` and `USE_SCHEMA` are not conferrable by a tag. Reaching
+`lakehouse.finance.orders` requires:
+
+```
+USE_CATALOG   on lakehouse            from RBAC
+USE_SCHEMA    on lakehouse.finance    from RBAC
+SELECT_TABLE  on the table            from RBAC or from a tag rule
+```
+
+**Tag-based access is additive within territory a role already has, not a way 
to hand out new
+territory.** For `analyst` to read `lakehouse.finance.orders` through 
`certified`, the role must
+already hold `USE_CATALOG` on `lakehouse` and `USE_SCHEMA` on `finance`. A tag 
applied to a table in
+a schema the role cannot enter has no effect. `validate()` rejects both names 
in `privileges`.
+
+That containment is what makes the worst case analyzable. The most a 
misapplied tag can do is
+expose an object the role could already traverse to, which scopes the blast 
radius to the
+territory its RBAC grants already describe rather than to the whole metalake.
+
+### Allow and deny
+
+With `ALLOW` only, two access rules cannot conflict — they union. The 
interaction that remains is
+between a tag rule that allows and an RBAC grant that denies; the proposal is 
that the deny wins. See
+[OQ-2](#oq-2--composition-when-a-tag-allows-and-rbac-denies).
+
+---
+
+## Administration
+
+Four write paths can change who has access. Their current authority 
requirements:
+
+| Operation                | Expression today                                  
                              | Scoping available                               
                 |
+| ------------------------ | 
------------------------------------------------------------------------------- 
| ---------------------------------------------------------------- |
+| Create a policy          | `METALAKE::OWNER \|\| METALAKE::CREATE_POLICY`    
                              | metalake only                                   
                 |
+| Create a tag             | `METALAKE::OWNER \|\| METALAKE::CREATE_TAG`       
                              | metalake only                                   
                 |
+| Bind a policy to a tag   | tag-scoped                                        
                              | metalake, or the specific tag                   
                 |
+| Apply a tag to an object | `METALAKE::OWNER \|\| ((TAG::OWNER \|\| 
ANY_APPLY_TAG) && CAN_ACCESS_METADATA)` | the specific tag, and only objects 
the caller can already access |
+
+Creating a policy and creating a tag are both metalake-wide, with no way to 
scope either to a
+catalog. Authoring access rules is therefore a central function today; 
delegating it per-catalog
+would need new privileges.
+
+`ANY_APPLY_TAG` above expands to
+`(METALAKE::APPLY_TAG || TAG::APPLY_TAG) && !(METALAKE::DENY_APPLY_TAG || 
TAG::DENY_APPLY_TAG)`.
+
+Applying a tag is the one operation that is already bounded on the object side.
+`CAN_ACCESS_METADATA` resolves per entity type to that type's load expression; 
for a table:
+
+```
+ANY(OWNER, METALAKE, CATALOG) ||
+SCHEMA_OWNER_WITH_USE_CATALOG ||
+ANY_USE_CATALOG && ANY_USE_SCHEMA && (TABLE::OWNER || ANY_SELECT_TABLE || 
ANY_MODIFY_TABLE)
+```
+
+So an applier must already own the object or an ancestor, or hold traversal 
plus read or write on
+the object itself. That bounds which objects they can tag, not what the tag 
may confer on them.
+
+### What binding a policy to a tag delegates
+
+Binding an access policy to a tag is a deliberate delegation: it says that 
whoever can apply this
+tag may confer this access on the named role. That is the feature, not a 
defect.
+
+Two properties of that delegation are worth recording:
+
+- `CAN_ACCESS_METADATA` establishes that the applier can *access* the object. 
It does not establish
+  that they may *confer* access on a role they do not control. These are 
different authorities.
+- `ApplyTag.canBindTo` accepts only `METALAKE` and `TAG`, so the delegation 
cannot be scoped to a
+  subtree — "may apply `certified` within `lakehouse.finance`" is not 
expressible.
+
+See [OQ-4](#oq-4--authority-to-confer-access-through-a-tag).
+
+### Enabling the feature
+
+One server-level configuration turns tag-based access on or off. It defaults 
to off. No dry-run or
+audit-only mode is proposed for v1.
+
+**Off.** Policies and tags behave as they do today: they can be created, bound 
to each other and
+applied to objects, each still requiring the authority in the table above. The 
authorizer never
+reads them, so no tag grants anyone anything.
+
+**On.** The authorizer consults tags — both when deciding access to a single 
object and when
+filtering a list. It is both or neither: enforcing only one would either grant 
a caller access to
+objects that never appear in their listings, or list objects they are then 
denied.
+
+A metalake with no `system_access_control` policy already confers nothing, so 
the flag is not what
+makes the feature opt-in. It is a kill switch. This adds a new path to the 
authorization hot path,
+and an operator who needs it gone — a wrong decision, or list filtering 
degrading under
+[Cost](#cost) — should not have to unbind policies one at a time to get there.
+
+Flipping it either way takes effect immediately:
+
+- **On to off revokes.** Access held only through a tag disappears at once. 
Nothing gains access,

Review Comment:
   You're right, "disappears at once" was wrong. Updated to explain four cases:
   
   New request, any node -> no lag.
   
   The flag itself -> read at startup, so a rolling restart is the bound. Same 
as any server config.
   
   An operation already admitted -> runs to completion. The check happens once, 
at authorization.
   
   A vended storage credential -> until it expires, and nothing recalls one. A 
static secret-key credential never expires, so only rotating the catalog's key 
ends access.
   
   Only the last outlives the decision that granted it.
   
   On a maximum stale-allow interval for tag state: there isn't one. The 
evaluator reads effective tags and bound policies through 
listEntitiesByRelation, and RelationalEntityStore caches entities, not relation 
query results. Every request reads current state. The only bound left is role 
membership which is jcasbin.cacheExpirationSecs, one hour , which is the 
existing RBAC bound, unchanged here.
   
   Both are now a table in Freshness, and M4 tests them across nodes.



-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to