Check Snaplex health

Use the SnapGPT Assistant to analyze the health of your Snaplexes over a selected time range and receive AI-generated insights about issues and suggested remediation.

SnapGPT must be enabled for your environment. Contact your Customer Success Manager (CSM) if SnapGPT is not available.

The Check Snaplex health skill is available on the Infrastructure page in Monitor. It analyzes Snaplex telemetry — including CPU and memory utilization, node availability, crash events, pipeline interruptions, and node alerts — for the time range you select. The SnapGPT Assistant groups detected issues by Snaplex so you can quickly identify which nodes require attention and act directly from the panel.

The wizard guides you through four steps: Action, Time range, Snaplex, and Analysis.

  1. In Monitor, click Infrastructure in the left navigation pane.
  2. Click the SnapGPT button in the Monitor header to open the SnapGPT Assistant panel.

    If this is the first time you are opening the SnapGPT Assistant in Monitor, the Welcome to SnapGPT for Observability! screen appears. Review the guidelines and click Understood, Let's Begin to continue.

    The panel then displays the How SnapGPT can help home screen. Under Actions available now, the Check Snaplex health skill is shown with the description Identify potential issues and suggested fixes on your Snaplexes.

  3. Click Check Snaplex health.
    The wizard opens at step 2, Time range. The subtitle Pick a time range, a Snaplex, then start the analysis is displayed below the panel title. The step indicator shows step 1 (Action) as completed, step 2 (Time range) as active, and steps 3 (Snaplex) and 4 (Analysis) as inactive.
  4. Select a time range to analyze.

    Three options are available:

    • Last 24 hours — Hourly resolution, fastest.
    • Last 48 hours — Hourly resolution, 2-day trend.
    • Last 72 hours — Hourly resolution, 3-day trend.
    The wizard advances to step 3, Snaplex, and displays a list of Snaplexes in your environment.
  5. In the Snaplex step, review the Snaplex list and click the Snaplex you want to investigate.

    The SnapGPT Assistant analyzes all Snaplexes in your environment immediately after you select a time range. The results are ready when step 3 opens. Each Snaplex card shows:

    • The Snaplex name and the number of active nodes out of the total configured nodes (for example, 2/34 Active Node).
    • A health indicator: Critical, Warning, or Healthy on the Snaplex card. Issue groups within the analysis view show the same status in all caps: CRITICAL or WARNING.
    • Issue type chips for each category of problem detected, such as Node Down, Node Restart, High Resource Utilization, or Node Metrics Unstable.

    Snaplexes are sorted by severity: critical first, then warning, then healthy. Use the Search Snaplexes field to filter by name. The step indicator shows steps 1 and 2 as completed.

    The wizard advances to step 4, Analysis, showing the detailed issue breakdown for the selected Snaplex.
  6. Review the analysis view for the selected Snaplex.

    The top of the view shows a header card with the Snaplex name and active node count. Below it, issues are grouped by type. All issue groups are expanded by default; click a group header to collapse or expand it. Each issue group shows:

    • An icon identifying the issue type, the issue type name, and the number of affected nodes (for example, 2 affected). Issue types include Node Restart, Node Down, High Resource Utilization, Node Metrics Unstable, and Node Alerts.
    • A severity badge: CRITICAL or WARNING.
    • A brief description of the condition detected (for example, Node restarts detected on 2 node(s) or 2 inactive node(s) detected).
    • An Affected Nodes section listing each affected node by name, with a status line specific to the issue type:
      • Node Restart: the timestamp of the last restart (for example, restarted at Jul 24, 08:12:07 UTC).
      • Node Down: the timestamp of the last heartbeat received (for example, last heartbeat at Jul 18, 00:42:51 UTC).
      • High Resource Utilization: the peak utilization percentage (for example, utilization: max disk /home at 72.64%).

    Hover over any node row to reveal the Open in Metrics button for that node.

  7. Optional: For Node Alerts issues, hover over a node row in the Affected Nodes list and click Analyze to generate an on-demand AI analysis for that node.

    A loading indicator is displayed while the analysis is generated. Results are cached so that repeated requests for the same node return immediately. If the SnapGPT service is temporarily unavailable, an error message is displayed in the node row.

    The AI analysis appears inline within the node row, including a root cause explanation and one or more suggested fixes. The primary fix is marked with a Recommended badge and appears first.
  8. Optional: To view a node's metrics, hover over the node row in the Affected Nodes list and click Open in Metrics.
    The main page navigates to the Metrics page within the same tab. The metrics view is pre-filtered to the specific node and time range you selected in the wizard. The SnapGPT Assistant panel remains open alongside the metrics charts, so you can refer to the analysis while reviewing the node data.
  9. Optional: To view an interrupted pipeline in Designer, click Open in Designer on a pipeline interruption issue card.
    Designer opens in a new browser tab showing the affected pipeline at the runtime that was interrupted. The SnapGPT Assistant panel remains open on the Monitor page.
  10. Optional: To return to the full Snaplex list, click All Snaplexes at the top of the analysis view.

The SnapGPT Assistant provides a prioritized view of Snaplex health issues with AI-generated explanations and recommended remediation steps, enabling you to identify and address infrastructure problems quickly.