Hi,
I agree with Richard for option (1).

Karthick feel free to open a PR.
In case you want proceed, instead of removing the line what about
annotating it with the following pattern?

@Deprecated(since = "3.0.1", forRemoval = true)

Cheers,

Gianluca

On Wed, 5 Aug 2026, 19:39 Richard Zowalla, <[email protected]> wrote:

> I would prefer (1) right now ;-)
>
> > Am 05.08.2026 um 00:18 schrieb Rui Abreu <[email protected]>:
> >
> > Hi! Thanks for spotting it. It seems that config was never used from
> > the beginning.
> > Storm 3.0.0 dropped Clojure in favour of Java, but the port kept the
> > logic as much as possible.
> > There used to be a comment in a Clojure source fle:
> >
> > -        ;; TODO: this is broken. need to maintain a map since last time
> > -        ;; supervisor hearbeats like is done for tasks
> > -        ;; maybe it's ok to trust ephemeral nodes here?
> > -        ;;[[id info]]
> > -        ;; (when (< (time-delta (:time-secs info))
> > -        ;;          (conf NIMBUS-SUPERVISOR-TIMEOUT-SECS))
> > -        ;;   [[id info]]
> > -        ;;   )
> >
> >
> https://github.com/apache/storm/commit/40e092dda0e7affe89c00552bc138a251f095dc0
> >
> > The actual mechanism still relies on the ephemeral nodes (and always
> > had, apparently)
> >
> > - ZooKeeper session timeout (storm.zookeeper.session.timeout, default
> > 20000 ms). This is negotiated with ZK and drives when ZK deletes the
> > ephemeral node after the supervisor stops responding.
> > - Nimbus scheduler tick (nimbus.monitor.freq.secs) — how often Nimbus
> > re-reads ZK and reassigns.
> >
> > supervisor.heartbeat.frequency.secs is misleading because even if it
> > would not fire, Zookeeper client uses pings to keep the session alive,
> > according to what I can gather (this property keeps the Supervisor
> > data read by Nimbus fresh)
> >
> > @Richard Zowalla @Gianluca Graziadei @Julien Nioche would like to hear
> > on thoughts on this.
> >
> > 1- We remove the dead configs/code/documentation and just properly
> > document the actual mechanism
> > 2- We implement a different Supervisor liveness mechanism based on
> > that property that has been never used
> >
> > Either way, this is a bug. @Karthick do you want to open an issue for
> this?
> >
> > Thank you
> >
> > On Tue, 4 Aug 2026 at 05:29, Karthick <[email protected]>
> wrote:
> >>
> >> Hi,
> >> Im checking on the heartbeat flow, The below configuration is not in
> use, as per comment it seems. Am I missing anything? Please guide me.
> >>
> >> /**
> >> * How long before a supervisor can go without heartbeating before
> nimbus considers it dead and stops assigning new work to it.
> >> */
> >> @isInteger
> >> @isPositiveNumber
> >> public static final String NIMBUS_SUPERVISOR_TIMEOUT_SECS =
> "nimbus.supervisor.timeout.secs";
> >>
> >>
> >> On Sat, Jul 25, 2026 at 1:45 AM Gianluca Graziadei <
> [email protected]> wrote:
> >>>
> >>> Hi,
> >>>
> >>> I dug through the JIRA history and your reconciliation holds, though
> with two key refinements: the bottleneck in STORM-2693 was actually
> per-round read-and-recompute overhead on Nimbus rather than Zookeeper write
> pressure, which is why the fix prioritized caching and supervisor reporting
> over a faster store. SupervisorInfo stayed in ZK not for liveness
> monitoring, but because it is shared cluster-state metadata that any
> elected Nimbus leader needs to access for scheduling (content of the
> serialized
> https://github.com/apache/storm/blob/master/storm-client/src/jvm/org/apache/storm/generated/SupervisorInfo.java
> )
> >>>
> >>> Best,
> >>>
> >>> Gianluca
> >>>
> >>>
> >>> Il giorno ven 24 lug 2026 alle ore 12:25 Karthick <
> [email protected]> ha scritto:
> >>>>
> >>>> Thanks, that's a clear and helpful breakdown. The push-for-liveness /
> pull-for-stats split makes sense: keeping a continuous in-memory pulse per
> worker in Nimbus (HeartbeatCache) while serving heavy stats on demand from
> ZK keeps Nimbus lightweight, and the Supervisor Thrift relay keeps those
> frequent liveness writes off ZK.
> >>>>
> >>>> One thing I left out of my original question: the supervisor liveness
> heartbeat — which is also liveness, but it still goes to ZooKeeper
> (ephemeral znode at /supervisors/<id>, via SupervisorHeartbeat), consumed
> by Nimbus under nimbus.supervisor.timeout.secs. So "liveness" isn't
> uniformly on the Thrift path.
> >>>>
> >>>> My working reconciliation is that the deciding factor is volume and
> semantics, not liveness-vs-stats:
> >>>>
> >>>> Worker/executor liveness is high-volume (potentially thousands of
> heartbeats per cluster, every ~1s) → moved off ZK onto the Supervisor
> Thrift relay to avoid write pressure.
> >>>> Supervisor liveness is low-volume (one per node, every ~5s) → cheap
> on ZK, and the ephemeral znode gives automatic crash detection when the
> session dies, which the Thrift path wouldn't provide for free.
> >>>>
> >>>> Does that match the design intent — i.e. supervisor heartbeats stayed
> on ZK deliberately because per-node volume is low and ephemeral-node
> semantics are valuable, whereas per-worker heartbeats were the actual ZK
> scaling problem?
> >>>>
> >>>> Also noted on 3.0.0's ZK read/serialization improvements — thanks for
> the pointer, will look into it.
> >>>>
> >>>>
> >>>>
> >>>> On Thu, Jul 23, 2026 at 4:53 PM Karthick <[email protected]>
> wrote:
> >>>>>
> >>>>> Hi all,
> >>>>>
> >>>>> I'm studying the Storm 2.0 heartbeat/liveness paths and want to
> confirm my understanding of a design decision.
> >>>>>
> >>>>> As I read the 2.0 code, there are two distinct worker-originated
> heartbeats:
> >>>>>
> >>>>> Liveness — the worker writes an LSWorkerHeartbeat to local disk
> (Worker.doHeartBeat); the supervisor reads those files and relays a batch
> to the leader Nimbus over Thrift (ReportWorkerHeartbeats →
> Nimbus.sendSupervisorWorkerHeartbeats → HeartbeatCache), governed by
> nimbus.task.timeout.secs.
> >>>>> Stats — Worker.doExecutorHeartbeats writes a heartbeat object
> (time-secs + uptime + executor stats) to ZooKeeper, which the UI/metrics
> consume (and which HeartbeatCache.updateFromZkHeartbeat can still use for
> liveness on the ZK strategy).
> >>>>>
> >>>>> My understanding is that liveness was moved off ZooKeeper (the 1.x
> model, and later Pacemaker) because high-volume per-worker heartbeat writes
> made ZK a scaling bottleneck, and since heartbeats are ephemeral they don't
> need ZK's persistence/consistency — so the supervisor-relay-over-Thrift
> model removes those writes from ZK entirely.
> >>>>>
> >>>>> A few questions:
> >>>>>
> >>>>> Is that the correct/primary motivation for the Thrift
> supervisor-relay path, or were there other drivers (connection count, watch
> load, Nimbus HA, recovery on leader change)?
> >>>>> Why do executor stats still go through ZooKeeper rather than riding
> the same Thrift path — is it purely that stats are lower-frequency and
> UI-oriented, or is there a stronger reason?
> >>>>> Is there a JIRA / design doc that captures this transition (beyond
> docs/Pacemaker.md) that I could read?
> >>>>>
> >>>>> Thanks for any pointers — trying to make sure I document this
> accurately.
>
>

Reply via email to