How Snaplex node crash detection works

How a Snaplex detects and reports a node crash, why the alert timestamp can lag the actual outage, and how to avoid misattributed crash reports on Groundplex nodes.

Detection architecture

Each Snaplex node maintains a small state marker on local disk while it's running. When a node shuts down cleanly, it removes its own marker before exiting. When a node starts up, it checks whether a marker is already present from a previous process on that same disk path:

  • If no marker is present, the node assumes the path was previously used by a process that shut down cleanly, never existed, or crashed before it had a chance to write its own marker — a node doesn't write its marker the instant it starts, so a crash during that early window leaves no trace either. Either way, it writes its own marker.
  • If a marker is present, the node treats this as evidence that the previous process on this path did not shut down cleanly. It reports a node crash event, attributing the crash to the node identity recorded in that leftover marker, then removes it and writes its own.

This means node crash detection is retrospective: a crash is only detected and reported when a subsequent node starts up and finds the leftover marker. If no new node ever starts on that same disk path, no crash is ever reported for that death.

Why alert timing can look delayed or unrelated

Because detection happens on the next startup, the crash alert's timestamp reflects when the replacement node initialized, not the moment the previous node actually went down. In most cases a replacement node starts quickly, so this distinction isn't noticeable. But if the same underlying disk path is reused later — for example, a persistent volume that is detached from one node and reattached to a different node, possibly hours or days afterward — the new node finds the old marker and reports a crash at that later startup time, attributed to the old node's identity. This can look like a "phantom" crash notification with no obvious connection to the node that actually failed.

Important: A crash alert that is not manually resolved closes automatically after 60 minutes.

Groundplex mitigation

Since Cloudplex infrastructure is managed by SnapLogic, this section applies to Groundplex deployments, where you control node provisioning, shutdown, and storage.

Always shut down JCC gracefully

The marker is only removed by a clean shutdown. If the underlying VM or container is terminated before the JCC process finishes shutting down — for example, a SIGKILL or an abrupt instance termination with no grace period — the marker is left behind, and whatever node starts up next on that disk path reports a crash. Placing a node in maintenance mode drains its pipelines, but does not by itself stop the JCC process; a graceful stop still needs to happen before the node is terminated.

  • On Linux/VM Groundplex nodes, make sure your stop procedure runs jcc.sh stop (as the systemd or init.d service definitions do) rather than killing the process directly. See Start or stop a Groundplex on Linux.
  • On Kubernetes Groundplex nodes, make sure terminationGracePeriodSeconds gives the jcc_prestop.sh PreStop hook enough time to drain pipelines and stop JCC before Kubernetes force-terminates the pod. See Deploy a Groundplex on Kubernetes.

Don't carry the marker across node identities

The marker is a single file, nodeState.json, in the node's log directory ($SL_ROOT/run/log, typically /opt/snaplogic/run/log) alongside JCC and runtime logs you may still need. If you reuse a persistent disk, volume snapshot, or PVC across unrelated node identities — rather than provisioning a clean volume for each new node — remove only that marker file first, or provision a clean volume instead. Clearing the entire log directory is unnecessary and discards diagnostics you may still need. Otherwise, a brand-new node can inherit a leftover marker from a completely different, unrelated node and immediately report a crash under that old identity.

Related settings

  • Two feature flags tune this behavior: Feature flags reference (Metrics and telemetry section) documents CrashEventPusher.WAIT_DURATION (how long an idle node waits before writing its own marker) and CrashEventPusher.AUTO_UPLOAD (whether a support package is uploaded automatically when a crash is detected).
  • For a live, real-time check of node health instead of a retrospective crash report, see Monitor Snaplex health.