dwsmith1983 commented on PR #5864: URL: https://github.com/apache/datafusion-comet/pull/5864#issuecomment-5965621333
> A small fixture with a filter against a same-typed interval literal ..., a native function over a dispatched timestamp on the DST rows, and a `spark_answer_only` group-by would do it. Added to `subtract_timestamps.sql`: `WHERE ts1 - ts2 > INTERVAL '1 00:00:00' DAY TO SECOND` (same type on both sides, and the fall-back rows drop out since they equal exactly one day), `hour(ts2 + INTERVAL '23' HOUR)`, `hour(ts2 + INTERVAL '1' DAY)` and `CAST(ts1 + INTERVAL '12' HOUR AS DATE)` over the DST rows as `expect_native`, and `spark_answer_only` group-bys on `ts1 - ts2` and `d - date'2024-01-01'`. The date group-by sits in the timestamps file because `subtract_dates.sql` runs with legacy intervals on, where Spark 3.x refuses a `CalendarIntervalType` grouping key (4.0 allows it). Both group-bys keep `spark_answer_only`: the partial aggregate runs native, but native shuffle has no hash partitioning on an interval key yet, so the exchange and final aggregate run in Spark. > Could we reword them to give the reason that still holds? Both notes now say `NullPropagation` folds a NULL operand before Comet sees the plan and that the column rows cover NULL; the timestamps one adds that the folded literal would be a `DayTimeIntervalType` with legacy mode off. The line about #5133 is gone from the description too. > Could you add cases for them to `CometDatetimeExpressionBenchmark` ... and post the numbers? Added a projection and a filter for each operator. Release build, Spark 3.5, 1M rows, `America/Los_Angeles`, best/avg ms. The fallback run is the same benchmark with `spark.comet.exec.scalaUDF.codegen.enabled=false`, which falls back at the `Project`. The two runs are separate, so each is shown against its own Spark baseline: | Case | Dispatch: Spark | Dispatch: Comet | Fallback: Spark | Fallback: Comet | |---|---|---|---|---| | `date - date` | 25 / 27 | 21 / 23 (1.2X) | 22 / 25 | 20 / 23 (1.1X) | | `date - date`, filter | 24 / 27 | 14 / 16 (1.8X) | 24 / 27 | 20 / 22 (1.2X) | | `ts - ts` | 102 / 108 | 93 / 95 (1.1X) | 102 / 104 | 94 / 98 (1.1X) | | `ts - ts`, filter | 109 / 111 | 102 / 103 (1.1X) | 99 / 101 | 91 / 92 (1.1X) | | `ts + interval` | 103 / 105 | 97 / 97 (1.1X) | 102 / 104 | 100 / 101 (1.0X) | | `ts + interval`, filter | 111 / 113 | 101 / 106 (1.1X) | 111 / 112 | 102 / 103 (1.1X) | Against Spark, dispatch is at or above the fallback in every case. The date filter gains the most, since the comparison stays native; the timestamp cases are within run-to-run noise. -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected] --------------------------------------------------------------------- To unsubscribe, e-mail: [email protected] For additional commands, e-mail: [email protected]
