agentusecasesAll 965 use cases
Security, Science & Hardware

Hunt business-logic flaws and open fix PRs

Parallel Devin agents each read part of a codebase, find business-logic and auth bypass flaws, confirm them in a sandbox, then open remediation PRs (vendor description).

Done withDevin

What they did
Parallel agents each take a segment of the codebase and reason across files for business-logic flaws, chained auth bypasses and cross-service exploit paths. Devin then combines the findings into full attack paths and reproduces each one in an isolated sandbox. Once a vulnerability is confirmed, it writes the patch and opens a PR for review. Scan profiles can be generated from existing threat model documents.
How it went
On Cognition's own 50-vulnerability benchmark, Devin Security scored 72% recall at $90.23 per run, against 68% for Claude Security at $131.87. This is a vendor benchmark, not an independent one.
Worth knowing
The first full scan sets the baseline. Later scans cover only changed code, so cost falls over time, and batch size per profile controls depth and cost.

Try it yourself with Devin

Review my codebase [repo name] for business-logic flaws and authentication or authorization bypasses. Split it into parts and examine each, then try to confirm every suspected issue in a sandbox, not on production. Finish with a ranked list of confirmed findings, and open a draft pull request with a fix for each one. Do not merge anything or touch live systems.

Read the original ↗

Source: cognition.com · Undated

Five of these in your inbox every morning

The best things people got an AI agent to do, each with the prompt to try it.

More like this