github-actions[bot] commented on code in PR #68651:
URL: https://github.com/apache/doris/pull/68651#discussion_r4140421957


##########
fe/fe-core/src/main/java/org/apache/doris/nereids/properties/ChildrenPropertiesRegulator.java:
##########
@@ -210,6 +219,55 @@ private boolean 
onePhaseAggWithDistribute(PhysicalHashAggregate<? extends Plan>
                 && children.get(0).getPlan() instanceof PhysicalDistribute;
     }
 
+    /**
+     * Check data-volume gates for bucketed hash aggregation using group-level
+     * statistics available during property regulation. Returns true if the
+     * pattern should be allowed (stats pass or unavailable), false if it 
should
+     * be banned due to unfavorable data characteristics.
+     * Mirrors the checks from the old implementBucketedPhase.
+     */
+    private boolean bucketedDataVolumeGatesPass(PhysicalHashAggregate<? 
extends Plan> aggregate) {
+        Statistics inputStats = 
aggregate.getGroupExpression().get().childStatistics(0);
+        if (inputStats == null) {
+            return true; // no stats → allow (other gates handle eligibility)
+        }
+        Statistics outputStats = aggregate.getGroupExpression().get()
+                .getOwnerGroup().getStatistics();
+        SessionVariable sv = ConnectContext.get().getSessionVariable();
+        double rows = inputStats.getRowCount();
+
+        // Gate 1: minimum input rows
+        if (sv.bucketedAggMinInputRows > 0 && rows < 
sv.bucketedAggMinInputRows) {
+            return false;
+        }
+
+        // Gate 2: high-cardinality GROUP BY columns
+        double highCardThreshold = sv.bucketedAggHighCardThreshold;
+        if (highCardThreshold > 0) {
+            for (Expression groupByKey : aggregate.getGroupByExpressions()) {
+                ColumnStatistic colStat = 
inputStats.findColumnStatistics(groupByKey);
+                if (colStat != null && !colStat.isUnKnown
+                        && colStat.ndv > rows * highCardThreshold) {
+                    return false;
+                }
+            }
+        }
+
+        // Gate 3: max group keys (merge phase cost dominates)
+        if (sv.bucketedAggMaxGroupKeys > 0 && outputStats != null
+                && outputStats.getRowCount() > sv.bucketedAggMaxGroupKeys) {
+            return false;
+        }
+
+        // Gate 4: aggregation output cardinality ratio

Review Comment:
   [P2] Allow bucketed aggregation when group cardinality is only a fallback 
estimate. For a single-BE grouped scan with at least 100000 estimated input 
rows but no NDV for its group key, StatsCalculator assigns output rows = input 
rows / 3. Gate 2 skips the unknown NDV, but this gate compares that synthetic 
estimate with the new default threshold of 0.3 * input rows; 1/3 > 0.3 always 
bans the bucketed candidate. The new regression suite raises the threshold to 
1.0 to exercise fusion. Skip the output-ratio gate when its cardinality came 
from unknown NDV, or choose a default that admits the fallback, and cover the 
default setting in a plan test.



-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]


---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to