Hi folks,

Documentation was retro-fitted to the website.
Thank you

On Thu, 20 Aug 2026 at 00:46, Rui Abreu <[email protected]> wrote:
>
> Created a PR enhancing the documentation and using @Gianluca Graziadei
> suggestion to set it as Deprecated.
> https://github.com/apache/storm/pull/8984
>
> I can retrofit these documentation changes into the live website 
> documentation.
>
> On Wed, 5 Aug 2026 at 21:18, Gianluca Graziadei
> <[email protected]> wrote:
> >
> > Hi,
> > I agree with Richard for option (1).
> >
> > Karthick feel free to open a PR.
> > In case you want proceed, instead of removing the line what about 
> > annotating it with the following pattern?
> >
> > @Deprecated(since = "3.0.1", forRemoval = true)
> >
> > Cheers,
> >
> > Gianluca
> >
> >
> > On Wed, 5 Aug 2026, 19:39 Richard Zowalla, <[email protected]> wrote:
> >>
> >> I would prefer (1) right now ;-)
> >>
> >> > Am 05.08.2026 um 00:18 schrieb Rui Abreu <[email protected]>:
> >> >
> >> > Hi! Thanks for spotting it. It seems that config was never used from
> >> > the beginning.
> >> > Storm 3.0.0 dropped Clojure in favour of Java, but the port kept the
> >> > logic as much as possible.
> >> > There used to be a comment in a Clojure source fle:
> >> >
> >> > -        ;; TODO: this is broken. need to maintain a map since last time
> >> > -        ;; supervisor hearbeats like is done for tasks
> >> > -        ;; maybe it's ok to trust ephemeral nodes here?
> >> > -        ;;[[id info]]
> >> > -        ;; (when (< (time-delta (:time-secs info))
> >> > -        ;;          (conf NIMBUS-SUPERVISOR-TIMEOUT-SECS))
> >> > -        ;;   [[id info]]
> >> > -        ;;   )
> >> >
> >> > https://github.com/apache/storm/commit/40e092dda0e7affe89c00552bc138a251f095dc0
> >> >
> >> > The actual mechanism still relies on the ephemeral nodes (and always
> >> > had, apparently)
> >> >
> >> > - ZooKeeper session timeout (storm.zookeeper.session.timeout, default
> >> > 20000 ms). This is negotiated with ZK and drives when ZK deletes the
> >> > ephemeral node after the supervisor stops responding.
> >> > - Nimbus scheduler tick (nimbus.monitor.freq.secs) — how often Nimbus
> >> > re-reads ZK and reassigns.
> >> >
> >> > supervisor.heartbeat.frequency.secs is misleading because even if it
> >> > would not fire, Zookeeper client uses pings to keep the session alive,
> >> > according to what I can gather (this property keeps the Supervisor
> >> > data read by Nimbus fresh)
> >> >
> >> > @Richard Zowalla @Gianluca Graziadei @Julien Nioche would like to hear
> >> > on thoughts on this.
> >> >
> >> > 1- We remove the dead configs/code/documentation and just properly
> >> > document the actual mechanism
> >> > 2- We implement a different Supervisor liveness mechanism based on
> >> > that property that has been never used
> >> >
> >> > Either way, this is a bug. @Karthick do you want to open an issue for 
> >> > this?
> >> >
> >> > Thank you
> >> >
> >> > On Tue, 4 Aug 2026 at 05:29, Karthick <[email protected]> wrote:
> >> >>
> >> >> Hi,
> >> >> Im checking on the heartbeat flow, The below configuration is not in 
> >> >> use, as per comment it seems. Am I missing anything? Please guide me.
> >> >>
> >> >> /**
> >> >> * How long before a supervisor can go without heartbeating before 
> >> >> nimbus considers it dead and stops assigning new work to it.
> >> >> */
> >> >> @isInteger
> >> >> @isPositiveNumber
> >> >> public static final String NIMBUS_SUPERVISOR_TIMEOUT_SECS = 
> >> >> "nimbus.supervisor.timeout.secs";
> >> >>
> >> >>
> >> >> On Sat, Jul 25, 2026 at 1:45 AM Gianluca Graziadei 
> >> >> <[email protected]> wrote:
> >> >>>
> >> >>> Hi,
> >> >>>
> >> >>> I dug through the JIRA history and your reconciliation holds, though 
> >> >>> with two key refinements: the bottleneck in STORM-2693 was actually 
> >> >>> per-round read-and-recompute overhead on Nimbus rather than Zookeeper 
> >> >>> write pressure, which is why the fix prioritized caching and 
> >> >>> supervisor reporting over a faster store. SupervisorInfo stayed in ZK 
> >> >>> not for liveness monitoring, but because it is shared cluster-state 
> >> >>> metadata that any elected Nimbus leader needs to access for scheduling 
> >> >>> (content of the serialized 
> >> >>> https://github.com/apache/storm/blob/master/storm-client/src/jvm/org/apache/storm/generated/SupervisorInfo.java)
> >> >>>
> >> >>> Best,
> >> >>>
> >> >>> Gianluca
> >> >>>
> >> >>>
> >> >>> Il giorno ven 24 lug 2026 alle ore 12:25 Karthick 
> >> >>> <[email protected]> ha scritto:
> >> >>>>
> >> >>>> Thanks, that's a clear and helpful breakdown. The push-for-liveness / 
> >> >>>> pull-for-stats split makes sense: keeping a continuous in-memory 
> >> >>>> pulse per worker in Nimbus (HeartbeatCache) while serving heavy stats 
> >> >>>> on demand from ZK keeps Nimbus lightweight, and the Supervisor Thrift 
> >> >>>> relay keeps those frequent liveness writes off ZK.
> >> >>>>
> >> >>>> One thing I left out of my original question: the supervisor liveness 
> >> >>>> heartbeat — which is also liveness, but it still goes to ZooKeeper 
> >> >>>> (ephemeral znode at /supervisors/<id>, via SupervisorHeartbeat), 
> >> >>>> consumed by Nimbus under nimbus.supervisor.timeout.secs. So 
> >> >>>> "liveness" isn't uniformly on the Thrift path.
> >> >>>>
> >> >>>> My working reconciliation is that the deciding factor is volume and 
> >> >>>> semantics, not liveness-vs-stats:
> >> >>>>
> >> >>>> Worker/executor liveness is high-volume (potentially thousands of 
> >> >>>> heartbeats per cluster, every ~1s) → moved off ZK onto the Supervisor 
> >> >>>> Thrift relay to avoid write pressure.
> >> >>>> Supervisor liveness is low-volume (one per node, every ~5s) → cheap 
> >> >>>> on ZK, and the ephemeral znode gives automatic crash detection when 
> >> >>>> the session dies, which the Thrift path wouldn't provide for free.
> >> >>>>
> >> >>>> Does that match the design intent — i.e. supervisor heartbeats stayed 
> >> >>>> on ZK deliberately because per-node volume is low and ephemeral-node 
> >> >>>> semantics are valuable, whereas per-worker heartbeats were the actual 
> >> >>>> ZK scaling problem?
> >> >>>>
> >> >>>> Also noted on 3.0.0's ZK read/serialization improvements — thanks for 
> >> >>>> the pointer, will look into it.
> >> >>>>
> >> >>>>
> >> >>>>
> >> >>>> On Thu, Jul 23, 2026 at 4:53 PM Karthick <[email protected]> 
> >> >>>> wrote:
> >> >>>>>
> >> >>>>> Hi all,
> >> >>>>>
> >> >>>>> I'm studying the Storm 2.0 heartbeat/liveness paths and want to 
> >> >>>>> confirm my understanding of a design decision.
> >> >>>>>
> >> >>>>> As I read the 2.0 code, there are two distinct worker-originated 
> >> >>>>> heartbeats:
> >> >>>>>
> >> >>>>> Liveness — the worker writes an LSWorkerHeartbeat to local disk 
> >> >>>>> (Worker.doHeartBeat); the supervisor reads those files and relays a 
> >> >>>>> batch to the leader Nimbus over Thrift (ReportWorkerHeartbeats → 
> >> >>>>> Nimbus.sendSupervisorWorkerHeartbeats → HeartbeatCache), governed by 
> >> >>>>> nimbus.task.timeout.secs.
> >> >>>>> Stats — Worker.doExecutorHeartbeats writes a heartbeat object 
> >> >>>>> (time-secs + uptime + executor stats) to ZooKeeper, which the 
> >> >>>>> UI/metrics consume (and which HeartbeatCache.updateFromZkHeartbeat 
> >> >>>>> can still use for liveness on the ZK strategy).
> >> >>>>>
> >> >>>>> My understanding is that liveness was moved off ZooKeeper (the 1.x 
> >> >>>>> model, and later Pacemaker) because high-volume per-worker heartbeat 
> >> >>>>> writes made ZK a scaling bottleneck, and since heartbeats are 
> >> >>>>> ephemeral they don't need ZK's persistence/consistency — so the 
> >> >>>>> supervisor-relay-over-Thrift model removes those writes from ZK 
> >> >>>>> entirely.
> >> >>>>>
> >> >>>>> A few questions:
> >> >>>>>
> >> >>>>> Is that the correct/primary motivation for the Thrift 
> >> >>>>> supervisor-relay path, or were there other drivers (connection 
> >> >>>>> count, watch load, Nimbus HA, recovery on leader change)?
> >> >>>>> Why do executor stats still go through ZooKeeper rather than riding 
> >> >>>>> the same Thrift path — is it purely that stats are lower-frequency 
> >> >>>>> and UI-oriented, or is there a stronger reason?
> >> >>>>> Is there a JIRA / design doc that captures this transition (beyond 
> >> >>>>> docs/Pacemaker.md) that I could read?
> >> >>>>>
> >> >>>>> Thanks for any pointers — trying to make sure I document this 
> >> >>>>> accurately.
> >>

Reply via email to