dwsmith1983 commented on issue #2545:
URL: 
https://github.com/apache/datafusion-comet/issues/2545#issuecomment-5873367734

   I'll take the Comet side of this. The operator work is upstream now: 
jayzhan211's epic (apache/datafusion#24768), with a sort-merge fallback under 
memory pressure as its first step (apache/datafusion#25217), and #6309 points 
the roadmap there, so nothing here should duplicate it.
   
   Two Comet pieces:
   
   1. A size guard in `RewriteJoin`. Today 
`spark.comet.exec.forceShuffledHashJoin` rewrites every eligible sort-merge 
join, and `getOptimalBuildSide` reads `sizeInBytes` only to pick the build 
side. Rewriting only when the build side's estimate is under a threshold, the 
way Spark's `JoinSelection` gates shuffled hash join with 
`canBuildLocalHashMapBySize`, keeps the hash join for the joins that fit, which 
is where the TPC-H gain is, and leaves the rest on sort-merge join until 
spilling lands. Small change, no native code.
   
   2. Once apache/datafusion#25217 lands and Comet takes that DataFusion 
release: the Comet side of the fallback (its metrics, the `fair_unified` 
interaction, and a memory-limited join matrix on TPC-H and TPC-DS like 
apache/datafusion#24779), so the hash join can go on by default.
   
   @andygrove is that the split you want, or should Comet wait for the fallback?
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]


---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to