+1 on this question.

Right now `_pos` column feels more like debug leftovers, it bring some
confusion for read-side weather it's expected to be readed or not.

I think removing it would make implementation easier.

Best,
Andrei

On Wed, Aug 26, 2026 at 2:42 PM Leonid Lygin via dev <[email protected]>
wrote:

> Thanks for the quick response!
>
> My biggest concern with `_pos` is not performance but rather clarity
> and implementation divergence:
>
> 1. including `_pos` is redundant, and (at least for me) provokes a
> re-read of the row alignment section — "why include `_pos` if files
> are fully aligned?";
> 2. having `_pos` fully duplicate the row position, there are two
> different legal ways to implement reads: either positionally, or using
> `_pos`.
>
> On Wed, Aug 26, 2026 at 2:35 PM Gábor Kaszab <[email protected]>
> wrote:
> >
> > Hey All,
> >
> > Thanks for bringing this up! (for me the initial mail went to spam,
> though...)
> >
> > Technically, with the dense representation we don't really need the _pos
> column in the column files, unless for troubleshooting. While checking the
> row counts is good, if they don't match we might get a better understanding
> on what the writer missed writing if we had the _pos col, also the order
> could be verified.
> >
> > Apart from debugging, I think either way is just fine. An additional
> detail to consider is that according to my experiments, there isn't really
> any storage cost for writing the _pos with delta encoding (e.g. with
> Parquet V2). So the conclusion was that since it comes for free, and might
> help for debugging, why not write it.
> >
> > Should we reopen this question? Any further feedback is welcome.
> >
> > Best Regards,
> > Gabor
> >
> > Leonid Lygin via dev <[email protected]> ezt írta (időpont: 2026.
> aug. 26., Sze, 14:10):
> >>
> >> Definitely agree that including `_pos` raises questions.
> >>
> >> If "debugging" is to be understood as figuring out if the column files
> >> have gaps -- just checking the row counts is good enough for that. Is
> >> there a lot to be gained from figuring out where exactly the gap is
> >> occurring?
> >>
> >> On Mon, Aug 24, 2026 at 1:57 PM Marco Kroll
> >> <[email protected]> wrote:
> >> >
> >> > Hi all,
> >> >
> >> > I just saw the agenda [1] for tomorrow's (2026-08-25) sync and want
> to +1 the `_pos` column topic.
> >> > My understanding is that this column exists for two reasons:
> >> > 1. debugging
> >> > 2. detect if writers skipped deleted rows
> >> >
> >> > My take is that using the dense Null filled representation addresses
> both of these issues.
> >> > It implicitly encodes the position, very much like for deletion
> vectors and since all rows need to be present, comparing the row count of
> the base file with the column file can be used to verify that all rows were
> written.
> >> >
> >> > The main thing to add to the doc would be that the row order must be
> identical to the base file.
> >> >
> >> > Best
> >> > Marco
> >> >
> >> > [1]:
> https://docs.google.com/document/d/1Bd7JVzgajA8-DozzeEE24mID_GLuz6iwj0g4TlcVJcs/edit?tab=t.jvm7iiiulf8q#heading=h.rbisiun18esp
>

Reply via email to