agentusecasesAll 965 use cases
Security, Science & Hardware

Reverse engineer a patch into a full exploit chain

A researcher used Claude Opus in Claude Code with Playwright MCP to decompile code, find vulnerabilities, and build a working exploit.

Done withClaude Code

What they did
Kev Breen gave Claude Opus in Claude Code (with Ghidra/ilspy, semgrep, Playwright and Proxmox tools, root on an isolated code-server harness) a single prompt to investigate a newly reported PaperCut NG zero-day, build a lab VM, and patch-diff the vendor fix. The agent autonomously decompiled jars, traced IoCs, found an auth-bypass and SQLi, used Playwright to drive the admin UI, and iteratively built and self-tested PoC scripts, with the human mainly steering direction and challenging weak conclusions.
How it went
Within 90 minutes it produced a working unauthenticated RCE chain (auth bypass, SQLi, arbitrary file write, code execution), later finding a third bypass that defeated the emergency patch too; total run used 10 prompts, 293 tool calls, 224M tokens over about two hours, though it initially hit a dead-end chain and made a false 'validated' claim the researcher had to catch.
Worth knowing
Hit an Opus 5 safety guardrail when checking exploitation logs, forcing a fallback to Opus 4.8 mid-session.

Try it yourself with Claude Code

Given this [software patch or diff], analyze what changed, decompile or inspect the relevant binaries, and identify the security vulnerability it fixes. Build a proof-of-concept exploit demonstrating the issue in a safe, isolated test environment, and stop before using it against anything outside that environment.

Read the original ↗

Source: techanarchy.net · Aug 27, 2026

Five of these in your inbox every morning

The best things people got an AI agent to do, each with the prompt to try it.

More like this