andygrove commented on code in PR #5496: URL: https://github.com/apache/datafusion-comet/pull/5496#discussion_r4155998507
########## docs/source/user-guide/latest/tuning/transitions.md: ########## @@ -38,3 +38,25 @@ columns and nested fields. If profiling shows these conversions dominating a que exceeds the threshold to Spark row-based execution — Comet removes the stage's native operators rather than running a mix of native and fallback operators joined by repeated conversions — which can be cheaper than paying the expensive conversions again and again. + +## Experimental: Direct Columnar-to-Row Conversion + +When the JVM columnar-to-row operator is in use (`spark.comet.exec.columnarToRow.native.enabled=false`), +setting `spark.comet.exec.columnarToRow.direct.enabled=true` enables an experimental converter that writes +values straight from Arrow buffers into Spark's row format without allocating an object per value. This is +most beneficial for decimal-heavy schemas, where the default conversion allocates a `Decimal` object per Review Comment: The PR description reports that the direct converter is about 29% slower at 8192 rows per batch for a mixed `long, int, double, string` schema (12.1 vs 9.4 ns/row), because `UnsafeProjection` is already specialized for that shape. This section only describes the upside. Could we say here that schemas with strings and no expensive decimals can get slower with the flag on, and that it is worth measuring a workload before enabling it? Otherwise someone turning it on for a whole cluster could regress the most common schema shape without knowing why. -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected] --------------------------------------------------------------------- To unsubscribe, e-mail: [email protected] For additional commands, e-mail: [email protected]
