Hi Ryan, I like the proposal; it is concise and clear. I would suggest making a clear distinction between two classes of expressions:
1. Expressions that produce deterministic values intrinsically tied to the input value, e.g. to_lower_case(s).This class is relatively straightforward to reason about, since the result depends only on the individual row. 2. Expressions such as Hilbert that produce values whose meaning depends on the input value + the whole distribution. Here the main issue is that the underlying distribution matters. If the distribution changes over time, hilbert(c11,c12) computed at time t may no longer be valid at t+1. This makes the second class harder to handle, because we need to establish for how long a computed value remains valid (=if in the current snapshot values computed on a previous snapshot are still valid). For this second class, I don't think persistence of the computed value alone is sufficient. We may also need to persist the distribution (or its CDF) against which the value was computed (I am not confident that min/max can be good enough in this case). Before I comment: 1. Are you targeting to manage these two classes in the same way? It seems that they have potentially different validity scope. 2. What about adding a snapshot level CDF struct per column? Cheers, Gianluca Il giorno mar 1 set 2026 alle ore 00:25 Ryan Blue <[email protected]> ha scritto: > Hi everyone, > > One of the remaining open questions for v4 metadata is how we will assign > table field IDs for values that are not written into the table. I want to > propose a solution that I think is going to be flexible, while not > introducing a lot of churn in the table or REST specs. > > Columnar field stats are written into metadata using a simple function > from table field ID to metadata field ID. We want to reuse what we already > have working for table fields and keep the spec simple. That means we need > a way to assign a table field ID to a non-materialized column so that we > can track its stats. There are a few cases we’ve identified: > > - Partition field output for non-monotonic functions, like bucket(1024, > id) > - Clustering expressions, like to_lower_case(last_name) > - Collation sequence lower and upper bounds > > We also discussed a new case this morning in the index sync: we need a > field ID for a derived value used to organize an index, like hilbert(col_a, > col_b), because we intend to use table field IDs in index schemas. > > Initially, I suggested that we keep a table of expressions and assign each > one a field ID. But as we started thinking about the use cases where we > need expressions it became clear that denormalizing *all* expressions was > adding a lot of complexity for little benefit. For example, CHECK > constraints won’t have reusable expressions and it makes little sense to > create them in two parts (expression and constraint). Similarly, it is > awkward to model a collation sequence as an expression, and we don’t need > to rebuild partition specs just to assign field IDs. However, we also don’t > want to just embed table field IDs in every one of these structures. > > My proposal is to directly model what we want: one table of fields that > are not materialized in the table, but are assigned IDs for stats or other > purposes. This would take a few forms: > > - Partition output value: {"field-id": 102, "type": "partition-value", > "partition-field-id": 1000} > - Collation sequence: {"field-id": 103, "type": "collation-bounds", > "collation-seq-id": 1} > - Value expression results: {"field-id": 104, "type": "expr-value", > "data-type": "long", "expr": <hilbert(col_a, col_b) expr>} > > This representation leaves existing structures alone and is a single place > outside of schema to allocate table field IDs. This can be expanded with > new types later when we want to add new structures, like a cluster-by spec. > > I think this is a fairly clean way to move forward and solve two > challenges that we’re currently hitting. We'll discuss this in the next v4 > metadata sync, but in the meantime please reply with feedback if you have > an opinion. > > Thanks, > > Ryan >
