Got it.  Great direction…!  But, note that the current patch is infinitely
faster for columnar in the sense that now it is possible!

On Fri, Jul 24, 2026 at 2:59 PM Suryaa Charan Shivakumar <
[email protected]> wrote:

> Yes, exactly. The current patchset is only the first step: enabling
> GEOMETRY
>  in the columnar store.
>
> The other optimizations are longer-term ideas that came out of exploring
> the implementation. There are really two related but separable ideas.
>
> The first is improving evaluation over the serialized geometry
> representation, for example avoiding JTS when a function can be answered
> directly from the serialized bytes. *That would benefit both row and column
> storage.*
>
> The second is storing derived spatial metadata such as MBRs, point counts,
> type, or emptiness.* Row storage could benefit from this as well*. The
> columnar-specific advantage is that these values could be read
> independently, allowing some queries to avoid reading the full geometry
> payload entirely and enabling cheaper filtering before exact spatial
> evaluation.
>
> One way Ahmed phrased it while we were discussing this is that *the more
> specific our reads of the binary representation (WKB) become, the more
> efficient execution can be.* I think that's a useful way to think about it.
> Rather than treating WKB as an opaque blob that must always be fully
> materialized into a JTS object, we can progressively exploit more of its
> structure directly. Initially that might mean reading just enough of the
> serialized bytes to answer simple functions, and later exposing derived
> metadata (e.g., MBRs, geometry type, point count) for cheaper execution.
>
> I brought up these longer-term ideas mainly to motivate the current design
> question. By itself the patchset does not provide a meaningful performance
> benefit for geometry workloads.  Geometry has no natural ordering in its
> WKB representation, currently uses delta byte-array encoding, and cannot
> directly benefit from the usual min/max filtering available to ordered
> scalar types. That raises the broader question: beyond projection and/with
> reduced I/O, how should we harness the columnar engine’s capabilities for a
> complex binary type like geometry?
>
> On Fri, Jul 24, 2026 at 2:41 PM Mike Carey <[email protected]> wrote:
>
> > Just to clarify (for me):  The proposed add'l info and streamlining
> > would be orthogonal to column versus row storage as the format, yes?
> >
> > On 7/24/26 10:04 AM, Suryaa Charan Shivakumar wrote:
> > > Hi everyone,
> > >
> > > I wanted to start a discussion around columnar support for GEOMETRY,
> > > motivated by the work in my current PR (
> > > https://asterix-gerrit.ics.uci.edu/c/asterixdb/+/21333). The PR itself
> > is
> > > fairly small, it enables GEOMETRY as a columnar type but while working
> on
> > > it I realized the broader design deserves discussion first. Today,
> > geometry
> > > in columnar storage behaves differently from most primitive types. This
> > > patchset focuses on enabling the columnar store for the GEOMETRY type
> as
> > a
> > > first step. While implementing it, I found that the current serialized
> > > representation limits the performance benefits we can get from columnar
> > > execution. It would be great if someone familiar with the columnar
> layer
> > > could take a look at the patchset and help confirm whether this is the
> > > right direction before I build further geometry-specific optimizations
> on
> > > top of it.
> > > Current flow
> > >
> > > The current implementation stores the serialized geometry (WKB) as an
> > > opaque byte sequence [int32 length][WKB bytes]. During query execution:
> > >
> > > column scan -> read geometry bytes -> copy into tagged Asterix
> > > GEOMETRY value -> deserialize into JTS Geometry -> execute spatial
> > > function
> > >
> > > For eg. SELECT SUM(st_n_points(g)) FROM all_nodes_col; still performs
> > > WKB parsing, JTS object allocation, and JTS traversal for every
> > > geometry.
> > >
> > > Columnar storage avoids reading unrelated fields, but once g is
> > projected,
> > > execution follows nearly the same geometry-materialization path as row
> > > storage. Geometry also currently uses delta byte-array encoding and
> does
> > > not have a meaningful ordering, so it cannot benefit from the normal
> > column
> > > min/max filters.
> > >
> > > This was visible in a PBSM benchmark. Columnar storage read roughly 72%
> > > fewer pages than row storage, but execution was still slightly slower
> > > because most of the runtime was spent after the scan in geometry
> > > materialization, tiling, sorting, and exact predicate evaluation.
> > >
> > > The goal would be to (a) add serialized-geometry fast paths for
> functions
> > > that can be evaluated directly from WKB, such as st_n_points,
> > > st_geometrytype, st_dimension, and st_isempty. The evaluator would fall
> > > back to JTS only when direct serialized evaluation is unsupported. (b)
> > > Store derived geometry metadata alongside the WKB payload, such as:
> type,
> > > point count, empty flag, xmin, ymin, xmax, ymax. This could allow
> scalar
> > > functions to avoid reading the full geometry and allow spatial joins
> and
> > > predicates to use MBR metadata before JTS refinement. Note this might
> > also
> > > be doubling the storage for simple geometry types like point where we
> > would
> > > be storing more metadata than the data itself. The geometry-specific
> > > metadata would be new to the columnar primary store, but the broader
> > > pattern of storing derived metadata to reduce query-time work already
> > > exists.
> > > Existing patterns
> > >
> > > This follows patterns used elsewhere:
> > >
> > >     -
> > >
> > >     PostGIS exposes metadata and bounding boxes from its serialized
> > geometry
> > >     representation.
> > >     -
> > >
> > >     GeoParquet supports geometry metadata and bounding-box covering
> > columns.
> > >     -
> > >
> > >     Sedona uses bounding-box metadata for pruning before JTS
> evaluation.
> > >     -
> > >
> > >     AsterixDB columnar storage already maintains auxiliary filter
> > metadata
> > >     for other types, and spatial indexes already store derived MBR
> > values.
> > >
>

Reply via email to