Debug overnight with a long-running agent
The author leaves Codex CLI running for hours to add debug prints and test hypotheses against a reproducing unit test, often only to find the bug.
Done withCodex
- What they did
- The author starts Codex running for a few hours overnight on a bug they can't explain. The key requirement is a deterministic unit test that reproduces the bug through the app's front door. The agent then grinds through the tedious work on its own: adding debug prints throughout the code and testing hypotheses about the cause.
- How it went
- The source gives no numbers or specific results. The author calls it an ideal use case and says they do it routinely, though it doesn't say how often the agent actually finds the bug.
- Worth knowing
- Write a deterministic reproducing test first, entering through the app's front door; the approach depends on the agent having a repeatable pass/fail signal to test hypotheses against.
Try it yourself with Codex
There is a bug in [project] that reproduces with the test [test name or path]. Form hypotheses about the cause, add temporary debug prints, run the test, and keep testing hypotheses until you find the root cause. Stop when you can explain the cause with evidence from the output, then list the debug prints you added and remove them. Propose a fix but do not apply it without asking me.
Source: HN · Undated
Five of these in your inbox every morning
The best things people got an AI agent to do, each with the prompt to try it.
More like this
Turn tickets into merged pull requests
Pulls tasks from GitHub Issues or Linear, then runs Claude Code or Codex in isolated Kubernetes pods to carry each one through to a PR.
Recommend a hotel from client notes using Codex
A travel advisor had Codex use client notes to recommend a Belmond hotel in Italy, checking live booking engines and generating a detailed Word doc.
Build a racing game from a single prompt
Codex made a racing game with multiple racers, eight maps and items from one prompt, acting as designer, developer and QA tester by playing it.
- Claude
Build a C compiler with a team of parallel agents
Anthropic's engineering write-up on tasking Opus 4.6 agent teams with building a C compiler.
Configure a local AI lab and build a knowledge graph by chat
Users had Hermes Agent set up a local AI stack (liteLLM, Postgres, Prometheus, Grafana) and install Gbrain as an MCP server to build a personal knowledge graph.
Click through apps to verify its own work
With computer use, Devin built and played a desktop game, ordered on Amazon, and tested its own work by clicking through the app.