How Snaplex node crash detection works
How a Snaplex detects and reports a node crash, why the alert timestamp can lag the actual outage, and how to avoid misattributed crash reports on Groundplex nodes.
Detection architecture
Each Snaplex node maintains a small state marker on local disk while it's running. When a node shuts down cleanly, it removes its own marker before exiting. When a node starts up, it checks whether a marker is already present from a previous process on that same disk path:
- If no marker is present, the node assumes the path was previously used by a process that shut down cleanly, never existed, or crashed before it had a chance to write its own marker — a node doesn't write its marker the instant it starts, so a crash during that early window leaves no trace either. Either way, it writes its own marker.
- If a marker is present, the node treats this as evidence that the previous process on this path did not shut down cleanly. It reports a node crash event, attributing the crash to the node identity recorded in that leftover marker, then removes it and writes its own.
This means node crash detection is retrospective: a crash is only detected and reported when a subsequent node starts up and finds the leftover marker. If no new node ever starts on that same disk path, no crash is ever reported for that death.
Why alert timing can look delayed or unrelated
Because detection happens on the next startup, the crash alert's timestamp reflects when the replacement node initialized, not the moment the previous node actually went down. In most cases a replacement node starts quickly, so this distinction isn't noticeable. But if the same underlying disk path is reused later — for example, a persistent volume that is detached from one node and reattached to a different node, possibly hours or days afterward — the new node finds the old marker and reports a crash at that later startup time, attributed to the old node's identity. This can look like a "phantom" crash notification with no obvious connection to the node that actually failed.
Groundplex mitigation
Since Cloudplex infrastructure is managed by SnapLogic, this section applies to Groundplex deployments, where you control node provisioning, shutdown, and storage.
Always shut down JCC gracefully
The marker is only removed by a clean shutdown. If the underlying VM or container is
terminated before the JCC process finishes shutting down — for example, a
SIGKILL or an abrupt instance termination with no grace period — the
marker is left behind, and whatever node starts up next on that disk path reports a
crash. Placing a node in maintenance mode drains its pipelines, but does not by itself
stop the JCC process; a graceful stop still needs to happen before the node is
terminated.
- On Linux/VM Groundplex nodes, make sure your stop procedure runs
jcc.sh stop(as the systemd or init.d service definitions do) rather than killing the process directly. See Start or stop a Groundplex on Linux. - On Kubernetes Groundplex nodes, make sure
terminationGracePeriodSecondsgives thejcc_prestop.shPreStop hook enough time to drain pipelines and stop JCC before Kubernetes force-terminates the pod. See Deploy a Groundplex on Kubernetes.
Don't carry the marker across node identities
The marker is a single file, nodeState.json, in the node's log directory ($SL_ROOT/run/log, typically /opt/snaplogic/run/log) alongside JCC and runtime logs you may still need. If you reuse a persistent disk, volume snapshot, or PVC across unrelated node identities — rather than provisioning a clean volume for each new node — remove only that marker file first, or provision a clean volume instead. Clearing the entire log directory is unnecessary and discards diagnostics you may still need. Otherwise, a brand-new node can inherit a leftover marker from a completely different, unrelated node and immediately report a crash under that old identity.
Related settings
- Two feature flags tune this behavior: Feature flags reference
(Metrics and telemetry section) documents
CrashEventPusher.WAIT_DURATION(how long an idle node waits before writing its own marker) andCrashEventPusher.AUTO_UPLOAD(whether a support package is uploaded automatically when a crash is detected). - For a live, real-time check of node health instead of a retrospective crash report, see Monitor Snaplex health.