mohitgurav20 commented on issue #25932:
URL: https://github.com/apache/datafusion/issues/25932#issuecomment-5929898813

   Hey @mkleen, thanks for the detailed reproduction!
   
   I looked into this and the inconsistent behavior is coming from 
is_single_distinct_agg in single_distinct_to_groupby.rs. Currently, the 
heuristic that checks if the rewrite is actually worth it 
(rewrite_pays_for_count) is gated behind has_count_rollup.
   
   If a query only has sum(DISTINCT x), the optimizer never evaluates if the 
rewrite pays off and just blindly applies it, causing the memory explosion 
you're seeing by building a massive list of (group, distinct value) pairs. When 
you add count(*), has_count_rollup becomes true, the optimizer evaluates the 
heuristic, realizes the rewrite is too expensive, and aborts it — which 
ironically saves memory.
   
   I can pick this up and work on making the optimizer apply these heuristics 
consistently so that sum(DISTINCT x) doesn't blow up memory by default. Let me 
know if you have any thoughts on the best approach for the heuristic!


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]


---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to