dwsmith1983 commented on PR #5864:
URL: 
https://github.com/apache/datafusion-comet/pull/5864#issuecomment-5965621333

   > A small fixture with a filter against a same-typed interval literal ..., a 
native function over a dispatched timestamp on the DST rows, and a 
`spark_answer_only` group-by would do it.
   
   Added to `subtract_timestamps.sql`: `WHERE ts1 - ts2 > INTERVAL '1 00:00:00' 
DAY TO SECOND` (same type on both sides, and the fall-back rows drop out since 
they equal exactly one day), `hour(ts2 + INTERVAL '23' HOUR)`, `hour(ts2 + 
INTERVAL '1' DAY)` and `CAST(ts1 + INTERVAL '12' HOUR AS DATE)` over the DST 
rows as `expect_native`, and `spark_answer_only` group-bys on `ts1 - ts2` and 
`d - date'2024-01-01'`. The date group-by sits in the timestamps file because 
`subtract_dates.sql` runs with legacy intervals on, where Spark 3.x refuses a 
`CalendarIntervalType` grouping key (4.0 allows it). Both group-bys keep 
`spark_answer_only`: the partial aggregate runs native, but native shuffle has 
no hash partitioning on an interval key yet, so the exchange and final 
aggregate run in Spark.
   
   > Could we reword them to give the reason that still holds?
   
   Both notes now say `NullPropagation` folds a NULL operand before Comet sees 
the plan and that the column rows cover NULL; the timestamps one adds that the 
folded literal would be a `DayTimeIntervalType` with legacy mode off. The line 
about #5133 is gone from the description too.
   
   > Could you add cases for them to `CometDatetimeExpressionBenchmark` ... and 
post the numbers?
   
   Added a projection and a filter for each operator. Release build, Spark 3.5, 
1M rows, `America/Los_Angeles`, best/avg ms. The fallback run is the same 
benchmark with `spark.comet.exec.scalaUDF.codegen.enabled=false`, which falls 
back at the `Project`. The two runs are separate, so each is shown against its 
own Spark baseline:
   
   | Case | Dispatch: Spark | Dispatch: Comet | Fallback: Spark | Fallback: 
Comet |
   |---|---|---|---|---|
   | `date - date` | 25 / 27 | 21 / 23 (1.2X) | 22 / 25 | 20 / 23 (1.1X) |
   | `date - date`, filter | 24 / 27 | 14 / 16 (1.8X) | 24 / 27 | 20 / 22 
(1.2X) |
   | `ts - ts` | 102 / 108 | 93 / 95 (1.1X) | 102 / 104 | 94 / 98 (1.1X) |
   | `ts - ts`, filter | 109 / 111 | 102 / 103 (1.1X) | 99 / 101 | 91 / 92 
(1.1X) |
   | `ts + interval` | 103 / 105 | 97 / 97 (1.1X) | 102 / 104 | 100 / 101 
(1.0X) |
   | `ts + interval`, filter | 111 / 113 | 101 / 106 (1.1X) | 111 / 112 | 102 / 
103 (1.1X) |
   
   Against Spark, dispatch is at or above the fallback in every case. The date 
filter gains the most, since the comparison stays native; the timestamp cases 
are within run-to-run noise.
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]


---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to