mrhhsg opened a new pull request, #68565:
URL: https://github.com/apache/doris/pull/68565

   ### What problem does this PR solve?
   
   Issue Number: None
   
   Related PR: #61495, #65024
   
   Problem Summary: On a single-BE cluster, bucketed hash aggregation is on by
   default and the translator fuses a one-phase GLOBAL aggregate with its
   distribute child into a BucketedAggregationNode. The source side of that
   operator merges the live aggregate states built by different sink instances
   directly, instead of serializing them and deserializing them with the merging
   evaluator as the two-phase plan does. Java and Python UDAFs rely on the 
latter:
   
   - Java UDAF: the extra evaluator clone used by the bucketed source never 
calls
     create(), so its _exec_place stays null and merge()/insert_result_into()
     dereference a null state. Reproduced locally with a Java UDAF
     (`SELECT k, my_udaf(v) FROM t GROUP BY k`): UBSan reports "reference 
binding
     to null pointer of type AggregateJavaUdafData" in AggregateJavaUdaf::merge
     and the query fails / the BE goes down.
   - Python UDAF: merge() builds the rhs state from serialize_data, which is
     only filled on the deserialize path, so the rhs contribution is dropped or
     the Python server RPC fails.
   
   None of the FE gates excluded UDAFs. Add the check to the shared gate
   AggregateUtils.isBucketedHashAggEnabled, which now takes the aggregate and
   returns false when any aggregate function is a Udf (JavaUdaf / PythonUdaf).
   The translator, ChildrenPropertiesRegulator, ChildOutputPropertyDeriver and
   CostModel all go through this gate, so the optimizer also stops preferring
   the one-phase plan for these aggregates and they keep the regular
   aggregation path.
   
   ### Release note
   
   Fix BE crash / wrong result when a Java or Python UDAF is used with GROUP BY
   on a single-BE cluster with bucketed hash aggregation enabled.
   
   ### Check List (For Author)
   
   - Test:
       - Unit Test: BucketedAggregateTranslatorTest (new Python UDAF case under
         agg_phase=0 and agg_phase=1, fails
         without the fix), BucketedAggregateTest, 
ChildOutputPropertyDeriverTest,
         ChildrenPropertiesRegulatorTest, CostModelV1Test
       - Regression test: query_p0/javaudf/test_javaudaf_bucketed_agg (default
         and agg_phase=1 plans; fails on
         the old FE with BUCKETED AGGREGATE in the plan and a BE null deref when
         executed), plus bucketed_hash_agg and percentile_bucketed_agg_merge
   - Behavior changed: Yes (aggregates containing Java/Python UDAFs no longer
     use bucketed hash aggregation)
   - Does this need documentation: No
   
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]


---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to