agentusecasesAll 965 use cases
Software & DevOps

Diagnose Kubernetes alerts with MCP servers

An n8n agent triggered by Alertmanager queries Kubernetes, Grafana and other MCP servers and posts a diagnostic report with root-cause hypotheses in a Mattermost thread.

Done withn8n

What they did
An Alertmanager webhook starts an n8n workflow. It deduplicates alerts over 48 hours and fetches the triggering PromQL expression from the Prometheus Rules API. An AI agent then investigates through Kubernetes, Grafana, DigitalOcean and GitHub MCP servers, plus a Qdrant store of internal docs. It is read-only. The workflow finds the original Mattermost alert post and replies in its thread.
How it went
The source gives no measured results. It describes a report with a summary, an event timeline, up to two root-cause hypotheses and troubleshooting steps. The author is still weighing extra data sources, such as database stats and pipeline status.
Worth knowing
Limit tool retries to one per failed call to avoid loops, and strip noisy labels like job and instance from the prompt. You must configure your own MCP endpoints and credentials.

Try it yourself with n8n

Build an n8n workflow that receives Alertmanager webhooks and starts an AI agent connected to my MCP servers for [Kubernetes, Grafana, other tools]. The agent should investigate the alert read-only, then post a diagnostic report with evidence and ranked root-cause hypotheses as a thread reply in [Mattermost channel]. You're done when a sample alert produces a report; it must not change anything in the cluster, and show me the report before the first post.

Read the original ↗

Source: community.n8n.io · Undated

Five of these in your inbox every morning

The best things people got an AI agent to do, each with the prompt to try it.

More like this