Glad to see that everyone (up to this point) agrees! Should we ratify the decision in the doc itself?
On Wed, Aug 26, 2026 at 6:15 PM Anurag Mantripragada <[email protected]> wrote: > > The addition of `_pos` column was not a remnant of previous design, but a > conscious choice we made during one of the column updates sync. I'm in favor > of removing it if it makes implementation easier. > > ~ Anurag > > On Wed, Aug 26, 2026 at 7:31 AM Gábor Kaszab <[email protected]> wrote: >> >> Yes, the question is whether we want to persist _pos into the column files. >> We earlier concluded that even though we went with the dense representation >> we persist _pos. However, I think re-opening the question is reasonable, >> because that's just extra noise ATM, and we shouldn't write that field into >> column files. Reading the _pos column would still work regardless if we >> include the base file or not to the query. >> >> Best Regards, >> Gabor >> >> >> Russell Spitzer <[email protected]> ezt írta (időpont: 2026. aug. >> 26., Sze, 16:21): >>> >>> What is the actual argument here? I think having a persisted field doesn't >>> make sense since we expect perfect alignment. We would expect reading the >>> file in isolation with the metadata _pos column should still work right? >>> >>> If we are just discussing removing a persisted value, I'm in favor of that. >>> >>> On Wed, Aug 26, 2026 at 8:37 AM Gábor Kaszab <[email protected]> wrote: >>>> >>>> I hear you, and I share the same opinion. If we don't need such a field >>>> then it's just extra unnecessary complexity to write it. I'm not entirely >>>> convinced on the debugging use of the _pos field. Would be beneficial to >>>> reduce unnecessary noise and confusion by not adding the _pos field. >>>> >>>> Let's discuss this on the next sync! In the meantime, opinions are welcome >>>> here too. >>>> >>>> Thanks, >>>> Gabor >>>> >>>> Andrei Tserakhau via dev <[email protected]> ezt írta (időpont: 2026. >>>> aug. 26., Sze, 15:32): >>>>> >>>>> +1 on this question. >>>>> >>>>> Right now `_pos` column feels more like debug leftovers, it bring some >>>>> confusion for read-side weather it's expected to be readed or not. >>>>> >>>>> I think removing it would make implementation easier. >>>>> >>>>> Best, >>>>> Andrei >>>>> >>>>> On Wed, Aug 26, 2026 at 2:42 PM Leonid Lygin via dev >>>>> <[email protected]> wrote: >>>>>> >>>>>> Thanks for the quick response! >>>>>> >>>>>> My biggest concern with `_pos` is not performance but rather clarity >>>>>> and implementation divergence: >>>>>> >>>>>> 1. including `_pos` is redundant, and (at least for me) provokes a >>>>>> re-read of the row alignment section — "why include `_pos` if files >>>>>> are fully aligned?"; >>>>>> 2. having `_pos` fully duplicate the row position, there are two >>>>>> different legal ways to implement reads: either positionally, or using >>>>>> `_pos`. >>>>>> >>>>>> On Wed, Aug 26, 2026 at 2:35 PM Gábor Kaszab <[email protected]> >>>>>> wrote: >>>>>> > >>>>>> > Hey All, >>>>>> > >>>>>> > Thanks for bringing this up! (for me the initial mail went to spam, >>>>>> > though...) >>>>>> > >>>>>> > Technically, with the dense representation we don't really need the >>>>>> > _pos column in the column files, unless for troubleshooting. While >>>>>> > checking the row counts is good, if they don't match we might get a >>>>>> > better understanding on what the writer missed writing if we had the >>>>>> > _pos col, also the order could be verified. >>>>>> > >>>>>> > Apart from debugging, I think either way is just fine. An additional >>>>>> > detail to consider is that according to my experiments, there isn't >>>>>> > really any storage cost for writing the _pos with delta encoding (e.g. >>>>>> > with Parquet V2). So the conclusion was that since it comes for free, >>>>>> > and might help for debugging, why not write it. >>>>>> > >>>>>> > Should we reopen this question? Any further feedback is welcome. >>>>>> > >>>>>> > Best Regards, >>>>>> > Gabor >>>>>> > >>>>>> > Leonid Lygin via dev <[email protected]> ezt írta (időpont: 2026. >>>>>> > aug. 26., Sze, 14:10): >>>>>> >> >>>>>> >> Definitely agree that including `_pos` raises questions. >>>>>> >> >>>>>> >> If "debugging" is to be understood as figuring out if the column files >>>>>> >> have gaps -- just checking the row counts is good enough for that. Is >>>>>> >> there a lot to be gained from figuring out where exactly the gap is >>>>>> >> occurring? >>>>>> >> >>>>>> >> On Mon, Aug 24, 2026 at 1:57 PM Marco Kroll >>>>>> >> <[email protected]> wrote: >>>>>> >> > >>>>>> >> > Hi all, >>>>>> >> > >>>>>> >> > I just saw the agenda [1] for tomorrow's (2026-08-25) sync and want >>>>>> >> > to +1 the `_pos` column topic. >>>>>> >> > My understanding is that this column exists for two reasons: >>>>>> >> > 1. debugging >>>>>> >> > 2. detect if writers skipped deleted rows >>>>>> >> > >>>>>> >> > My take is that using the dense Null filled representation >>>>>> >> > addresses both of these issues. >>>>>> >> > It implicitly encodes the position, very much like for deletion >>>>>> >> > vectors and since all rows need to be present, comparing the row >>>>>> >> > count of the base file with the column file can be used to verify >>>>>> >> > that all rows were written. >>>>>> >> > >>>>>> >> > The main thing to add to the doc would be that the row order must >>>>>> >> > be identical to the base file. >>>>>> >> > >>>>>> >> > Best >>>>>> >> > Marco >>>>>> >> > >>>>>> >> > [1]: >>>>>> >> > https://docs.google.com/document/d/1Bd7JVzgajA8-DozzeEE24mID_GLuz6iwj0g4TlcVJcs/edit?tab=t.jvm7iiiulf8q#heading=h.rbisiun18esp
