agentusecasesAll 965 use cases
Software & DevOps

Crawl documentation into clean Markdown for AI agents

The author built a documentation crawler for their AI agents that pulled 60 pages of clean Markdown in 49 seconds.

Done with a custom-built or unnamed agent

What they did
The author built pulpie-mcp, an MCP server with four tools: fetch, save, crawl and list. It runs a local extraction model behind a backend that starts automatically. Claude was asked to crawl a 60-page docs site. It found pages through robots.txt and sitemaps, used four concurrent workers, and wrote Markdown files with frontmatter to disk instead of returning them to the context window.
How it went
Total time was about 49 seconds: roughly 10 for backend cold start and 37 for crawling and saving 60 pages. Untested on client-rendered pages, which may come back empty. PDFs aren't supported.
Worth knowing
The model is licensed CC BY-NC 4.0 (non-commercial), and it has only been tested on Linux with an NVIDIA GPU, using about 420 MB of VRAM.

Try it yourself with Claude Code

Build a crawler that takes the documentation site [docs URL], follows its internal links, and saves each page as clean Markdown (no nav or footer) in a folder called docs-md, one file per page. Respect robots.txt and rate limit politely. Done when you have crawled the whole site and given me a count of pages saved and a list of any that failed.

Read the original ↗

Source: DEV · Oct 5, 2026

Five of these in your inbox every morning

The best things people got an AI agent to do, each with the prompt to try it.

More like this