mohitgurav20 commented on issue #25932: URL: https://github.com/apache/datafusion/issues/25932#issuecomment-5929898813
Hey @mkleen, thanks for the detailed reproduction! I looked into this and the inconsistent behavior is coming from is_single_distinct_agg in single_distinct_to_groupby.rs. Currently, the heuristic that checks if the rewrite is actually worth it (rewrite_pays_for_count) is gated behind has_count_rollup. If a query only has sum(DISTINCT x), the optimizer never evaluates if the rewrite pays off and just blindly applies it, causing the memory explosion you're seeing by building a massive list of (group, distinct value) pairs. When you add count(*), has_count_rollup becomes true, the optimizer evaluates the heuristic, realizes the rewrite is too expensive, and aborts it — which ironically saves memory. I can pick this up and work on making the optimizer apply these heuristics consistently so that sum(DISTINCT x) doesn't blow up memory by default. Let me know if you have any thoughts on the best approach for the heuristic! -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected] --------------------------------------------------------------------- To unsubscribe, e-mail: [email protected] For additional commands, e-mail: [email protected]
