agentusecasesAll 965 use cases
Software & DevOps

Diagnose and roll back bad config during outages

Google SREs use Gemini CLI to pull logs, find a bad config push, generate a reverting changelist, and file postmortem action items via MCP.

Done withGemini

What they did
In a simulated outage, an SRE opened Gemini CLI, which pulled incident details, correlated time series and analyzed logs through Google's internal ProdAgent tools. It proposed a restart playbook, and the SRE approved it. The restart failed, so the agent checked the service's source and recent changes, found a bad config push, and drafted a reverting changelist. A custom command then wrote the postmortem and filed action items via MCP.
How it went
The agent found the faulty config push in under two minutes, and the SRE-approved revert rolled out and the service recovered. The article gives no other measurements, and the scenario was simulated.
Worth knowing
Humans approved each production change, and the agent could only use typed, policy-checked tools. Several tools here are Google-internal, so you'd need your own MCP servers.

Try it yourself with Gemini

An outage is happening in [service name]. Pull the recent logs from [log source or command], find the most recent config change that lines up with the errors, and explain your evidence. Then draft a revert of that change as a patch or pull request, and write a list of postmortem action items. Stop once the revert and action items are ready, and show me both before you push, deploy, or file anything.

Read the original ↗

Source: cloud.google.com · Undated

Five of these in your inbox every morning

The best things people got an AI agent to do, each with the prompt to try it.

More like this