rich7420 commented on PR #6180: URL: https://github.com/apache/datafusion-comet/pull/6180#issuecomment-5876693173
Pushed 5929c0bd4 with bitmap counters for batches of at least 64 rows; smaller batches retain row counting. This preserves per-row short circuit and ANSI errors. Local checks passed: 7 Rust expression tests, 4 Spark 4.1.3 expression tests, all-target workspace Clippy, formatting and RAT. Against 6ad5c0d2, paired component measurements improved about 4.8–5.0x for Double (0% NULL, 5% NaN) and 8.4–8.5x for String (50% NULL) at 8192 rows, 32 columns, threshold 16. The full-query gains are smaller: repeated best times improved roughly 3–9%, with mean improvements around 3–4%. The low-NULL Double case (0% NULL, 5% NaN, aggregating all 32 columns) improved from 396 to 361 ms but still trails Comet fallback at 314 ms. These are local M3 Pro/Spark 4.1.3 measurements, not results from the original issue workload. Further isolation points to the aggregate's `if(isnan(c), 0D, c)` work: retaining the native filter with Spark aggregation recovers the fallback performance, and the 0%-NaN control favors native execution. The IF optimization is still experimental and is not included in this commit. Keeping this draft while reviewing that path; [fork CI for this update](https://github.com/rich7420/datafusion-comet/actions/runs/36470291204) is running. -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected] --------------------------------------------------------------------- To unsubscribe, e-mail: [email protected] For additional commands, e-mail: [email protected]
