mrhhsg opened a new pull request, #68565:
URL: https://github.com/apache/doris/pull/68565
### What problem does this PR solve?
Issue Number: None
Related PR: #61495, #65024
Problem Summary: On a single-BE cluster, bucketed hash aggregation is on by
default and the translator fuses a one-phase GLOBAL aggregate with its
distribute child into a BucketedAggregationNode. The source side of that
operator merges the live aggregate states built by different sink instances
directly, instead of serializing them and deserializing them with the merging
evaluator as the two-phase plan does. Java and Python UDAFs rely on the
latter:
- Java UDAF: the extra evaluator clone used by the bucketed source never
calls
create(), so its _exec_place stays null and merge()/insert_result_into()
dereference a null state. Reproduced locally with a Java UDAF
(`SELECT k, my_udaf(v) FROM t GROUP BY k`): UBSan reports "reference
binding
to null pointer of type AggregateJavaUdafData" in AggregateJavaUdaf::merge
and the query fails / the BE goes down.
- Python UDAF: merge() builds the rhs state from serialize_data, which is
only filled on the deserialize path, so the rhs contribution is dropped or
the Python server RPC fails.
None of the FE gates excluded UDAFs. Add the check to the shared gate
AggregateUtils.isBucketedHashAggEnabled, which now takes the aggregate and
returns false when any aggregate function is a Udf (JavaUdaf / PythonUdaf).
The translator, ChildrenPropertiesRegulator, ChildOutputPropertyDeriver and
CostModel all go through this gate, so the optimizer also stops preferring
the one-phase plan for these aggregates and they keep the regular
aggregation path.
### Release note
Fix BE crash / wrong result when a Java or Python UDAF is used with GROUP BY
on a single-BE cluster with bucketed hash aggregation enabled.
### Check List (For Author)
- Test:
- Unit Test: BucketedAggregateTranslatorTest (new Python UDAF case under
agg_phase=0 and agg_phase=1, fails
without the fix), BucketedAggregateTest,
ChildOutputPropertyDeriverTest,
ChildrenPropertiesRegulatorTest, CostModelV1Test
- Regression test: query_p0/javaudf/test_javaudaf_bucketed_agg (default
and agg_phase=1 plans; fails on
the old FE with BUCKETED AGGREGATE in the plan and a BE null deref when
executed), plus bucketed_hash_agg and percentile_bucketed_agg_merge
- Behavior changed: Yes (aggregates containing Java/Python UDAFs no longer
use bucketed hash aggregation)
- Does this need documentation: No
--
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.
To unsubscribe, e-mail: [email protected]
For queries about this service, please contact Infrastructure at:
[email protected]
---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]