WojciechWitos created NIFI-16358:
------------------------------------

             Summary: Unusual RAM consumption after the upgrade
                 Key: NIFI-16358
                 URL: https://issues.apache.org/jira/browse/NIFI-16358
             Project: Apache NiFi
          Issue Type: Bug
          Components: NiFi Stateless
    Affects Versions: 2.12.0
         Environment: 4x nodes,
6x CPU, 
25 GB RAM on each node, 20 configured for NiFi in Boostrap
OS Rhel 9
            Reporter: WojciechWitos


After migration from 2.11 to 2.12 coordinator starts to consume huge amount of 
RAM during startup and crashes down and restarts the application. 

 

On version 2.11 NiFi starts normally and operates on the setup and on normal 
load:
| |20 GB|20 GB|8.23 GB|41.0%|405.88 MB|390.99 MB| |00:05:37.279|
| |20 GB|20 GB|6.14 GB|31.0%|378 MB|361.22 MB| |00:05:36.980|
| |20 GB|20 GB|5.98 GB|30.0%|378 MB|361.36 MB| |00:05:37.553|
| |20 GB|20 GB|6.96 GB|35.0%|380.94 MB|363.52 MB| |00:05:37.503|

 

Coordinator works on the Core load avarage: 2.0-2.7, rest of the nodes operates 
on the 1.0-1.8 core load avarage.

 

After loading flows on coordinator it crashes down with an error:
{code:java}
2026-09-17 10:34:34,601 ERROR [main] org.apache.nifi.web.server.JettyServer 
Failed to start Serverjava.lang.OutOfMemoryError: unable to create native 
thread: possibly out of memory or process/resource limits reached        at 
java.base/java.lang.Thread.start0(Native Method)        at 
java.base/java.lang.Thread.start(Thread.java:1553)        at 
java.base/java.lang.System$2.start(System.java:2578)        at 
java.base/jdk.internal.vm.SharedThreadContainer.start(SharedThreadContainer.java:152)
        at 
java.base/java.util.concurrent.ThreadPoolExecutor.addWorker(ThreadPoolExecutor.java:953)
        at 
java.base/java.util.concurrent.ThreadPoolExecutor.execute(ThreadPoolExecutor.java:1364)
        at 
org.apache.nifi.controller.queue.clustered.partition.StandardRebalancingPartition.submitRebalanceTask(StandardRebalancingPartition.java:225)
        at 
org.apache.nifi.controller.queue.clustered.partition.StandardRebalancingPartition.start(StandardRebalancingPartition.java:189)
        at 
org.apache.nifi.controller.queue.clustered.SocketLoadBalancedFlowFileQueue.startLoadBalancing(SocketLoadBalancedFlowFileQueue.java:320)
        at 
org.apache.nifi.controller.FlowController.onFlowInitialized(FlowController.java:1615)
        at 
org.apache.nifi.controller.StandardFlowService.loadFromConnectionResponse(StandardFlowService.java:967)
        at 
org.apache.nifi.controller.StandardFlowService.load(StandardFlowService.java:491)
        at org.apache.nifi.web.server.JettyServer.start(JettyServer.java:946)   
     at org.apache.nifi.runtime.Application.startServer(Application.java:131)   
     at org.apache.nifi.runtime.Application.run(Application.java:78)        at 
org.apache.nifi.runtime.Application.run(Application.java:60)        at 
org.apache.nifi.NiFi.main(NiFi.java:42)
 {code}
Starts the application again, but other nodes do not change the coordinator - 
and it changes coordinator after some time, but situation is being repeat on 
other nodes also.

On the cluster there are 6500 processors, where 2250 are disabled, 3200 are 
stopped. 

On the other cluster there are 3 nodes, 1050 processors active and 350 
processors in disabled state and the migration went fine and upgrade was 
successful.

 

If something more is needed, let me know I will put requested details as fast 
as it's possible.



--
This message was sent by Atlassian Jira
(v8.20.10#820010)

Reply via email to