[
https://issues.apache.org/jira/browse/HIVE-30042?page=com.atlassian.jira.plugin.system.issuetabpanels:all-tabpanel
]
Vikram Ahuja updated HIVE-30042:
--------------------------------
Summary: Fix LLAP encoded/vectorized read path for TIMESTAMP WITH LOCAL
TIME ZONE on ORC tables (was: Select query with ORDER BY clause fails with
TIMESTAMP WITH LOCAL TIME ZONE columns on ORC tables)
> Fix LLAP encoded/vectorized read path for TIMESTAMP WITH LOCAL TIME ZONE on
> ORC tables
> --------------------------------------------------------------------------------------
>
> Key: HIVE-30042
> URL: https://issues.apache.org/jira/browse/HIVE-30042
> Project: Hive
> Issue Type: Bug
> Reporter: Vikram Ahuja
> Assignee: Vikram Ahuja
> Priority: Major
>
> HIVE-30020 fixed the plain (non-LLAP) ORC read/write path for TIMESTAMP WITH
> LOCAL TIME ZONE columns. However, any query that forces the read of such a
> column through LLAP's encoded/cached columnar IO path — e.g. adding an ORDER
> BY on a table with a TIMESTAMPLOCALTZ column, which introduces a Map+Reduce
> shuffle — still fails or silently produces incorrect results, because LLAP's
> specialized decode path duplicates (rather than reuses) ORC's type dispatch,
> and that duplication was never updated for TIMESTAMP_INSTANT/TIMESTAMPLOCALTZ.
> Steps to reproduce:
> {code:java}
> CREATE TABLE timestampltz_orc_format (formatid string, tsval timestamp with
> local time zone) STORED AS ORC;
> INSERT INTO timestampltz_orc_format SELECT * FROM timestampltz_formats;
> SELECT * FROM timestampltz_orc_format ORDER BY formatid;{code}
> This fails during Tez task execution with, progressively:
> {code:java}
> java.lang.RuntimeException: Vectorizaton is not supported for
> datatype:TIMESTAMPLOCALTZ
> java.lang.IllegalArgumentException: LLAP does not support TIMESTAMP_INSTANT
> java.lang.IllegalArgumentException: Unknown type timestamp with local time
> zone('...'){code}
> Root cause: Multiple independent type-dispatch switches in the
> LLAP/ORC-encoded stack had no case for TIMESTAMPLOCALTZ/TIMESTAMP_INSTANT,
> each falling into a default that throws (or, in one case, silently rejects a
> valid schema string):
> 1. VectorizedBatchUtil.createColumnVector — general vectorized-batch column
> allocator used once a shuffle/Reducer stage is introduced. No case for
> TIMESTAMPLOCALTZ.
> 2. EncodedTreeReaderFactory — LLAP's ORC encoded/cached tree-reader
> construction (two switches: category acceptance, and reader-builder). The
> isInstant flag plumbing already existed on TimestampStreamReader but was
> never wired to TIMESTAMP_INSTANT.
> 3. OrcSerde.convertPrimitiveType — schema inference from raw ORC files
> explicitly rejected TIMESTAMPLOCALTZ with a hardcoded throw, even though the
> generic TypeDescription.toString() fallback already produces the correct type
> name.
> 4. OrcEncodedDataConsumer.createColumn — LLAP's per-column vector allocation
> during batch decode. Missing TIMESTAMP_INSTANT case.
> 5. BatchToRowReader.nextValue — the columnar→row bridge used to materialize
> LLAP-decoded batch values back into row objects for downstream consumption.
> No case to build a TimestampLocalTZWritable from the vector.
> 6. LlapRecordReader batch-size/memory weighting heuristic — missing
> TIMESTAMPLOCALTZ entry (cosmetic/accuracy only, not correctness-critical).
--
This message was sent by Atlassian Jira
(v8.20.10#820010)