David Jacot created KAFKA-21163:
-----------------------------------
Summary: Uniform v2: spread every topic over its subscribers
Key: KAFKA-21163
URL: https://issues.apache.org/jira/browse/KAFKA-21163
Project: Kafka
Issue Type: Improvement
Reporter: David Jacot
Assignee: David Jacot
Fix For: 4.5.0
The uniform assignor, the default server-side assignor of consumer groups
(KIP-848), balances the total number of partitions per member but not the
partitions of each topic. With homogeneous subscriptions, it cuts the
partitions of all the topics, in order, into consecutive slices, so a member
can get all the partitions of a topic while the other members get none of them.
For example, take three members A, B and C, all subscribed to three topics T1,
T2 and T3 of three partitions each. The uniform assignor gives each member all
the partitions of one topic:
||Member||Uniform assignor||Every topic spread over its subscribers||
|A|T1-0, T1-1, T1-2|T1-0, T2-0, T3-0|
|B|T2-0, T2-1, T2-2|T1-1, T2-1, T3-1|
|C|T3-0, T3-1, T3-2|T1-2, T2-2, T3-2|
Every member has three partitions, so the group looks balanced. But the number
of partitions is only a proxy for the load. It is fair to assume that the
partitions of a topic carry about the same throughput, since producers spread
their records over them, but not that different topics do: some topics are much
busier than others. If T1 is busier than T2 and T3, A lags behind while B and C
have little to do. Adding more consumers does not really help either: the
partitions of a topic stay together in consecutive slices, so the load of a
busy topic stays on one member, or a few, instead of being spread over all of
them. Balancing the partitions of every topic over its subscribers balances the
load whatever the throughput of each topic.
The uniform assignor also ignores racks, so members may read from partitions
whose replicas are all in other racks even when a replica is in their own rack,
and it is slow on large groups with heterogeneous subscriptions.
The goal is a new version of the uniform assignor, meant to replace it, with
these properties in priority order:
# *Validity*: every partition of every subscribed topic goes to exactly one of
its subscribers.
# *Spread*: every subscriber of a topic gets the same number of its partitions,
or one more.
# *Balance*: no partition could move to another subscriber of its topic that
has at least two partitions fewer in total. With homogeneous subscriptions, all
members are within one partition of each other.
# *Rack alignment*, when enabled: as many partitions as possible go to a member
in a rack holding one of their replicas, without breaking the properties above.
# *Stickiness*: partitions move only when the properties above require it, and
an assignment that already has them stays unchanged.
# *Determinism*: the assignment only depends on the group and its topics, not
on the order in which they are given.
--
This message was sent by Atlassian Jira
(v8.20.10#820010)