This is an automated email from the ASF dual-hosted git repository. reiabreu pushed a commit to branch docs/heartbeat-clarify-deprecate-supervisor-timeout in repository https://gitbox.apache.org/repos/asf/storm.git
commit 4f1b4ecd5e85cd3ed3c103e2ca7107885e4acb71 Author: Rui Abreu <[email protected]> AuthorDate: Wed Aug 19 23:44:05 2026 +0100 Clarify worker/supervisor heartbeat docs and deprecate unused nimbus.supervisor.timeout.secs Documentation still described the pre-2.0 model in which workers/tasks heartbeat directly into ZooKeeper. Since 2.0 (STORM-2693), worker liveness heartbeats are written to local disk and relayed to Nimbus over Thrift, held in an in-memory heartbeat cache; supervisor liveness is an ephemeral ZooKeeper node detected via session expiry. Docs: - Daemon-Fault-Tolerance.md: describe the actual worker heartbeat relay path and the ephemeral-znode supervisor liveness mechanism. - Lifecycle-of-a-topology.md: note (inline in the 0.7.1 walkthrough) that the ZK-directory heartbeat model was replaced in 2.0. - Cluster-State-Serialization.md: clarify worker heartbeats are not persisted in ZooKeeper by default (only under Pacemaker/legacy configuration). Config: - Because supervisor crash detection relies on ephemeral znodes, no Nimbus-side supervisor timeout check exists, so nimbus.supervisor.timeout.secs is never read. Mark the DaemonConfig constant @Deprecated, remove the misleading defaults.yaml entry, and drop two inert references in NimbusClojurePortTest. Co-Authored-By: Claude Opus 4.8 <[email protected]> --- conf/defaults.yaml | 1 - docs/Cluster-State-Serialization.md | 9 +++++++++ docs/Daemon-Fault-Tolerance.md | 4 +++- docs/Lifecycle-of-a-topology.md | 1 + storm-server/src/main/java/org/apache/storm/DaemonConfig.java | 6 ++++++ .../org/apache/storm/daemon/nimbus/NimbusClojurePortTest.java | 2 -- 6 files changed, 19 insertions(+), 4 deletions(-) diff --git a/conf/defaults.yaml b/conf/defaults.yaml index 9682cf8bc..96118e7e5 100644 --- a/conf/defaults.yaml +++ b/conf/defaults.yaml @@ -81,7 +81,6 @@ nimbus.thrift.tls.client.auth.required: true topology.worker.nimbus.thrift.client.use.tls: false nimbus.childopts: "-Xmx1024m" nimbus.task.timeout.secs: 30 -nimbus.supervisor.timeout.secs: 60 nimbus.monitor.freq.secs: 10 nimbus.cleanup.inbox.freq.secs: 600 nimbus.inbox.jar.expiration.secs: 3600 diff --git a/docs/Cluster-State-Serialization.md b/docs/Cluster-State-Serialization.md index b59a0b2f4..8cc404ed0 100644 --- a/docs/Cluster-State-Serialization.md +++ b/docs/Cluster-State-Serialization.md @@ -9,6 +9,15 @@ ZooKeeper (and other configured state stores) such as topology assignments, Nimb summaries, `StormBase` records, log configs, credentials, worker heartbeats, profile requests, errors, etc. +> **Note on worker heartbeats.** Since 2.0 ([STORM-2693](https://issues.apache.org/jira/browse/STORM-2693)), +> worker liveness heartbeats are, by default, *not* persisted in ZooKeeper: workers +> write them to local disk, supervisors relay them to Nimbus over Thrift, and Nimbus +> keeps them in an in-memory heartbeat cache. Worker heartbeats are only written to a +> state store (the `WORKERBEATS_SUBTREE` path) when a heartbeat store such as Pacemaker +> is configured. The serialization described below still applies to those stored +> heartbeats, and to supervisor liveness (`SupervisorInfo`), which is always kept as an +> ephemeral ZooKeeper node. + It is distinct from [tuple serialization](Serialization.html), which covers payloads exchanged between spouts and bolts at runtime via Kryo. diff --git a/docs/Daemon-Fault-Tolerance.md b/docs/Daemon-Fault-Tolerance.md index 8dce601a8..b419e16cc 100644 --- a/docs/Daemon-Fault-Tolerance.md +++ b/docs/Daemon-Fault-Tolerance.md @@ -7,7 +7,7 @@ Storm has several different daemon processes. Nimbus that schedules workers, su ## What happens when a worker dies? -When a worker dies, the supervisor will restart it. If it continuously fails on startup and is unable to heartbeat to Nimbus, Nimbus will reschedule the worker. +When a worker dies, the supervisor will restart it. Worker liveness reaches Nimbus indirectly: each worker writes heartbeats to local disk, and its supervisor relays them to Nimbus over Thrift (this replaced the pre-2.0 model in which workers heartbeat directly into ZooKeeper; see [STORM-2693](https://issues.apache.org/jira/browse/STORM-2693)). If a worker stops heartbeating for longer than `nimbus.task.timeout.secs`, Nimbus reschedules it. A freshly launched worker is given a longer gra [...] ## What happens when a node dies? @@ -19,6 +19,8 @@ The Nimbus and Supervisor daemons are designed to be fail-fast (process self-des Most notably, no worker processes are affected by the death of Nimbus or the Supervisors. This is in contrast to Hadoop, where if the JobTracker dies, all the running jobs are lost. +Supervisor liveness is tracked differently from worker liveness. Each supervisor registers itself as an ephemeral ZooKeeper node (its `SupervisorInfo`, which also carries scheduling metadata such as ports and resources). When a supervisor dies, its ZooKeeper session expires and the ephemeral node disappears, so Nimbus detects the loss directly from ZooKeeper rather than by timing out heartbeats. (This is why there is no active Nimbus-side supervisor heartbeat-timeout setting.) + ## Is Nimbus a single point of failure? If you lose the Nimbus node, the workers will still continue to function. Additionally, supervisors will continue to restart workers if they die. However, without Nimbus, workers won't be reassigned to other machines when necessary (like if you lose a worker machine). diff --git a/docs/Lifecycle-of-a-topology.md b/docs/Lifecycle-of-a-topology.md index fe785f1e4..83e8991c0 100644 --- a/docs/Lifecycle-of-a-topology.md +++ b/docs/Lifecycle-of-a-topology.md @@ -34,6 +34,7 @@ First a couple of important notes about topologies: - Jars and configs are kept on local filesystem because they're too big for Zookeeper. The jar and configs are copied into the path {nimbus local dir}/stormdist/{topology id} - `setup-storm-static` writes task -> component mapping into ZK - `setup-heartbeats` creates a ZK "directory" in which tasks can heartbeat + - (**Since 2.0, STORM-2693**: workers no longer heartbeat directly into ZooKeeper. A worker now writes liveness heartbeats to local disk, and its supervisor relays them to Nimbus over Thrift. See [Daemon Fault Tolerance](Daemon-Fault-Tolerance.html) for the current mechanism.) - Nimbus calls `mk-assignment` to assign tasks to machines [code](https://github.com/apache/storm/blob/0.7.1/src/clj/org/apache/storm/daemon/nimbus.clj#L458) - Assignment record definition is here: [code](https://github.com/apache/storm/blob/0.7.1/src/clj/org/apache/storm/daemon/common.clj#L25) - Assignment contains: diff --git a/storm-server/src/main/java/org/apache/storm/DaemonConfig.java b/storm-server/src/main/java/org/apache/storm/DaemonConfig.java index 6dcc2eb37..ca4c47501 100644 --- a/storm-server/src/main/java/org/apache/storm/DaemonConfig.java +++ b/storm-server/src/main/java/org/apache/storm/DaemonConfig.java @@ -277,7 +277,13 @@ public class DaemonConfig implements Validated { /** * How long before a supervisor can go without heartbeating before nimbus considers it dead and stops assigning new work to it. + * + * @deprecated Unused. Supervisor liveness is tracked via an ephemeral ZooKeeper node (see + * {@code StormClusterState#supervisorHeartbeat}); when a supervisor dies its ZooKeeper session + * expires and the node disappears, so Nimbus detects the loss directly rather than by timing out + * heartbeats. No code reads this value. Retained only for backward compatibility. */ + @Deprecated @IsInteger @IsPositiveNumber public static final String NIMBUS_SUPERVISOR_TIMEOUT_SECS = "nimbus.supervisor.timeout.secs"; diff --git a/storm-server/src/test/java/org/apache/storm/daemon/nimbus/NimbusClojurePortTest.java b/storm-server/src/test/java/org/apache/storm/daemon/nimbus/NimbusClojurePortTest.java index 3acff939a..526cb0924 100644 --- a/storm-server/src/test/java/org/apache/storm/daemon/nimbus/NimbusClojurePortTest.java +++ b/storm-server/src/test/java/org/apache/storm/daemon/nimbus/NimbusClojurePortTest.java @@ -1717,7 +1717,6 @@ public class NimbusClojurePortTest { DaemonConfig.NIMBUS_TASK_LAUNCH_SECS, 60, DaemonConfig.NIMBUS_TASK_TIMEOUT_SECS, 20, DaemonConfig.NIMBUS_MONITOR_FREQ_SECS, 10, - DaemonConfig.NIMBUS_SUPERVISOR_TIMEOUT_SECS, 100, Config.TOPOLOGY_ACKER_EXECUTORS, 0, Config.TOPOLOGY_EVENTLOGGER_EXECUTORS, 0)) .build()) { @@ -1821,7 +1820,6 @@ public class NimbusClojurePortTest { DaemonConfig.NIMBUS_TASK_LAUNCH_SECS, 60, DaemonConfig.NIMBUS_TASK_TIMEOUT_SECS, 20, DaemonConfig.NIMBUS_MONITOR_FREQ_SECS, 10, - DaemonConfig.NIMBUS_SUPERVISOR_TIMEOUT_SECS, 100, Config.TOPOLOGY_ACKER_EXECUTORS, 0, Config.TOPOLOGY_EVENTLOGGER_EXECUTORS, 0)) .build()) {
