andygrove commented on issue #1819: URL: https://github.com/apache/datafusion-comet/issues/1819#issuecomment-5872360620
Closing this. The conversions that made sense landed in 2025 (#1825, #1826, #1874, #1880, #1887, #1954, #2041), and bit_not, bit_count, space and substring now come straight from datafusion-spark, which was the point of this epic. Comet uses datafusion-spark directly now, so converting our own expressions as a first step toward upstreaming isn't how this work gets done. Most of what's left shouldn't be converted. IfExpr wraps CaseExpr so each branch is evaluated only for the rows that take it. A ScalarUDF receives every argument already evaluated, so IF(b = 0, NULL, a / b) would fail under ANSI, and datafusion-spark's `if` only works through `simplify`, which our planner never calls. GetStructField and GetArrayStructFields were ruled out above, and the nondeterministic expressions need per-partition state. Subquery, UnboundColumn, CheckOverflow and NormalizeNaNAndZero are planner and analyzer plumbing rather than SQL functions. Cast would be a project of its own, and DataFusion already has named_struct. Timestamp truncation is covered by #5103. The user-facing ones that remain (to_json, element_at, array_insert, rlike) are better as separate issues for upstreaming each one, if someone wants to take them on. -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected] --------------------------------------------------------------------- To unsubscribe, e-mail: [email protected] For additional commands, e-mail: [email protected]
