Troubleshoot Snaplex issues
Common Snaplex node issues and where to find the detailed explanation or procedure for each.
Common issues you might encounter with a Snaplex include:
- Why did I receive a node crash notification long after a node was removed?
- A pipeline instance stays visible in Monitor after its node crashes.
- A node fails to start.
- A node is slow, or running out of memory or disk space.
Why did I receive a node crash notification long after a node was removed?
A node crash alert's timestamp reflects when a new node started up and detected the crash, not when the original node actually went down. Node crash detection happens on the next node startup on the same disk path, so if that disk path is reused later — for example, a persistent volume reattached to a different, unrelated node — the new node can report a crash attributed to the old node's identity, at whatever later time it happens to start. For the full explanation and how to avoid this on Groundplex nodes, see How Snaplex node crash detection works.
A pipeline instance stays visible in Monitor after its node crashes
When a JCC node goes down unexpectedly, Monitor can continue to show its pipeline instances as running for up to 8 hours, because the node never reported its final state. See Pipeline visibility after a node failure for why this happens and how to interpret it.
A node fails to start
Startup failures — such as clock skew, a port already in use, or insufficient file descriptors — raise their own alerts distinct from a node crash alert. See Node initialization alerts for the full list of startup alert types and their resolutions.
A node is slow, or running out of memory or disk space
Use the Node diagnostics table in Monitor (Infrastructure > node details panel > Additional details tab) to compare a node's current CPU, memory, and disk values against their recommended ranges. See Troubleshooting for details, or run the built-in diagnostic tool directly on the node: Run Snaplex diagnostics.