Check Snaplex health

Use the SnapGPT Assistant to analyze the health of your Snaplexes over a selected time range and receive AI-generated insights about issues and suggested remediation.

Use the SnapGPT Check Snaplex health skill on the Infrastructure page in Monitor to analyze Snaplex health including high resource utilization, crash events, node alerts and others for the time range you select. SnapGPT groups detected issues by Snaplex so you can quickly identify which nodes require attention and act directly from the panel.


Check Snaplex health skill in the SnapGPT Assistant panel on the Infrastructure page

Note: SnapGPT must be enabled for your environment. Contact your Environment admin if SnapGPT is not available.
  1. In Monitor, click Infrastructure in the left navigation pane.
  2. Click the SnapGPT button from the header menu to open the SnapGPT Assistant panel.
  3. Click Check Snaplex health.
  4. The wizard guides you through four steps: Action, Time range, Snaplex, and Analysis. Select any of the following time ranges to analyze:
    • Last 24 hours: Hourly resolution, fastest.
    • Last 48 hours: Hourly resolution, 2-day trend.
    • Last 72 hours: Hourly resolution, 3-day trend.

    Time range selection step in the Check Snaplex health wizard

    The wizard advances to step 3, Snaplex, and displays a list of Snaplexes in your environment.

  5. In the Snaplex step, view all Snaplexes in a card format. Each Snaplex card displays:
    • The Snaplex name and the number of active nodes out of the total configured nodes.
    • A health indicator: Critical, Warning, or Healthy on the Snaplex card.
    • Snaplexes are sorted by severity: critical first, then warning, then healthy. Use the Search Snaplexes field to filter by name.
    • The following issue types are shown as a chip on each of the Snaplex cards:
      • Node Down: The control plane has not received a heartbeat from the node in the last 15 minutes.
      • Node Restart: The node was restarted during the selected period.
      • Node Alerts: One or more alerts were raised on the node.
      • High Resource Utilization: Peak resource utilization exceeded the threshold.
      • Node Crash: The node crashed at least once during the selected period.
      • Node Interruption: One or more pipelines running on the node were interrupted during the selected period.
      • Node Metrics Unstable: Some metrics are not available for the selected period, or the metric values are invalid.

    Snaplex selection step showing Snaplex cards with health indicators and issue chips

  6. Click any of the Snaplex cards to analyze.
  7. Review the analysis view for the selected Snaplex.

    Analysis view showing issue groups and affected nodes for the selected Snaplex

    The top of the view shows a header card with the Snaplex name and active node count. Below it, issues are grouped by type, such as High Resource Utilization and Node Down. All issue groups are expanded by default. Click a group header to collapse or expand it. Each issue group displays:

    • An icon identifying the issue type, the name of the issue type, and the number of affected nodes.
    • A severity badge: CRITICAL or WARNING.
    • A brief description of the condition detected. For example, Node restarts detected on 2 node(s) or 2 inactive node(s) detected.
    • An Affected Nodes section listing each affected node by name, with a status line specific to the issue type, such as:
      • Node Restart: the timestamp of the last restart (for example, restarted at Jul 24, 08:12:07 UTC).
      • Node Down: the timestamp of the last heartbeat received (for example, last heartbeat at Jul 18, 00:42:51 UTC).
      • High Resource Utilization: the peak utilization percentage (for example, utilization: max disk /home at 72.64%).

    Affected Nodes section showing per-node status lines for each issue type

  8. Hover over the node rows to view any of the following buttons for further action:
    • Open in Metrics: Available for Node Down, Node Restart, High Resource Utilization, and Node Metrics Unstable. Opens the Metrics page in Monitor with the node and time range pre-selected.
    • Open in Designer: Available for Node Interruption. Opens the affected pipeline execution in Designer.
    • Click to Analyze: Available on each alert under Node Alert. SnapGPT explains the alert and suggests possible fixes.
  9. To return to the full Snaplex list, click All Snaplexes at the top of the analysis view.