Debug low-level cryptography code
Cryptography engineer Filippo Valsorda's write-up on Claude Code debugging low-level crypto.
Done withClaude Code
- What they did
- Filippo Valsorda ran Claude Code (Opus 4.1, no system prompt) on his new Go ML-DSA implementation. His prompt named the test command and code directory, described the failing symptom, and asked it to think hard. It read the code and traced the bug. It also tested a hypothesis with a small test of its own. He then repeated this on two earlier signing bugs, using a fresh session for each.
- How it went
- It found all three bugs in one shot without help. The Verify bug took minutes: high bits were taken twice. Its proposed fixes were mediocre, so Valsorda threw them away and rewrote them himself.
- Worth knowing
- Use well-scoped failing tests, start a fresh session for each independent bug, and treat the agent's output as a pointer to the bug, not a fix to merge.
Try it yourself with Claude Code
I have a failing test in my cryptography code at [path], where [describe the symptom, e.g. wrong output for certain inputs]. Investigate why it fails, compare against [reference implementation or test vectors], and propose a fix. Finished means the tests pass and you've explained the root cause, and you should show me the diff before changing any files.
Discussion on HN · Nov 1, 2025
Five of these in your inbox every morning
The best things people got an AI agent to do, each with the prompt to try it.
More like this
Turn tickets into merged pull requests
Pulls tasks from GitHub Issues or Linear, then runs Claude Code or Codex in isolated Kubernetes pods to carry each one through to a PR.
Manage an inbox and build courses and budgets with Claude Code
A non-technical user filters their inbox to emails needing replies, turned 15 homeschool PDFs into interactive narrated courses, and built a budget dashboard.
Build a 2D platformer game without coding
A content creator with no coding experience used Claude Code to build a black-and-white wave-combat platformer in HTML canvas and played it live.
- Claude
Find exploitable bugs in smart contracts
Anthropic's red team reports agents finding $4.6M worth of blockchain smart contract exploits.
Control and calibrate a robotic hand with OpenClaw
An OpenClaw agent running Claude drove a 16-joint printed hand, checked its actions with a USB camera, calibrated firmware and narrated in Telegram.
Find security flaws in open-source software
Google's Big Sleep agent reported about 20 security flaws in open-source software, following its SQLite find.