[ 
https://issues.apache.org/jira/browse/IGNITE-28955?page=com.atlassian.jira.plugin.system.issuetabpanels:all-tabpanel
 ]

Anton Vinogradov reassigned IGNITE-28955:
-----------------------------------------

    Assignee:     (was: Anton Vinogradov)

> Continuous query start hangs under load with ZooKeeper discovery
> ----------------------------------------------------------------
>
>                 Key: IGNITE-28955
>                 URL: https://issues.apache.org/jira/browse/IGNITE-28955
>             Project: Ignite
>          Issue Type: Bug
>            Reporter: Anton Vinogradov
>            Priority: Major
>              Labels: ise
>             Fix For: 2.19
>
>
> Starting a continuous query can hang for minutes on a cluster that uses 
> ZooKeeper discovery while the same caches are under a heavy update load.
> h3. Reproduction
> Add a subclass of {{CacheContinuousQueryConcurrentPartitionUpdateTest}} that 
> runs under ZooKeeper discovery (the patch attached to IGNITE-28673 does 
> exactly that) and run {{testConcurrentUpdatesAndQueryStartAtomicCacheGroup}}.
> On current master it does not finish: 318 s against the 300 s test timeout. 
> TCP discovery is not affected - the same test class passes in 45 s.
> h3. Root cause
> The message that completes the routine start and the continuous query 
> notifications share a topic and a pool:
> {code:java}
> // GridContinuousProcessor, sending the start result
> ctx.io().sendToGridTopic(node, TOPIC_CONTINUOUS, new 
> ContinuousRoutineStartResultMessage(...), SYSTEM_POOL);
> // GridContinuousProcessor, sending CQ notifications
> ctx.io().sendToGridTopic(node, TOPIC_CONTINUOUS, msg, SYSTEM_POOL, ackC);
> {code}
> Under a heavy update load the notifications fill the queue, and the single 
> {{ContinuousRoutineStartResultMessage}} waits behind them - head-of-line 
> blocking. {{cache.query(qry)}} waits for that message, so the caller sees a 
> hang. The analysis in IGNITE-28673 shows the same picture from a heap dump: 
> millions of retained {{GridContinuousMessage}} and 
> {{CacheContinuousQueryEntry}} objects and one waiting 
> {{ContinuousRoutineStartResultMessage}}.
> h3. Why only ZooKeeper today
> {{GridContinuousProcessor}} picks the path by {{immutableDiscoCustomMsg = 
> !ctx.discovery().mutableCustomMessages()}}, which reads the 
> {{@DiscoverySpiMutableCustomMessageSupport}} annotation of the configured 
> SPI. {{TcpDiscoverySpi}} has it, so the routine start is acknowledged over 
> discovery and never touches the communication queue. 
> {{ZookeeperDiscoverySpi}} does not, so it already uses the communication path 
> - and hits this.
> h3. Why it matters beyond ZooKeeper
> IGNITE-28673 removes the mutable path and makes the communication one the 
> only path. Merging it without fixing this would extend the hang to TCP 
> discovery, that is, to everyone. So this blocks IGNITE-28673.



--
This message was sent by Atlassian Jira
(v8.20.10#820010)

Reply via email to