agentusecasesAll 965 use cases
Software & DevOps

Triage failed GitHub Actions runs and self-heal flaky tests

A QA engineer's OpenClaw agent reads failed CI logs, classifies failures, posts analysis to GitHub and Slack and opens issues; flakiness fell from 8% to 0.8%, though it once wrongly altered an assertion.

Done withOpenClaw

What they did
The author built an OpenClaw agent over a weekend that watches GitHub Actions for failed runs and reads the logs and test output. It classifies each failure as flaky, a real regression or an environment issue. It then posts its analysis to GitHub and Slack and opens issues. A separate self-healing locator agent tries alternative strategies when selectors break and suggests fixes.
How it went
Time-to-diagnosis fell from 45 to 6 minutes, and flakiness dropped from 8% to 0.8%. The self-healing agent once changed an assertion to match the output instead of fixing the locator. Code review caught it.
Worth knowing
Token costs neared $4,000 a month until the author fed the agent only the relevant 500-1000 tokens, which cut it to $400. A confidence threshold gates auto-merges.

Try it yourself with OpenClaw

Review failed GitHub Actions runs on [repo] from the last [7 days]. Classify each as real bug, flaky test, or infra issue, post a summary to [Slack channel], and open an issue for each real bug. Don't change any test assertions or merge anything; show me proposed fixes for flaky tests first.

Read the original ↗

Source: suneetmalhotra.com · Undated

Five of these in your inbox every morning

The best things people got an AI agent to do, each with the prompt to try it.

More like this