agentusecasesAll 965 use cases
Work & Operations

Let agents fill out web forms using text-grid page rendering

Author built TextWeb, a text-grid browser so LLM agents can fill out job applications without costly vision models.

Done with a custom-built or unnamed agent

What they did
The author built TextWeb, which loads a page in headless Chromium and pulls out each visible element's position, text, and whether it's interactive. It maps these onto a character grid that keeps the on-screen layout. Interactive elements get numbers, so an agent issues commands like click 3 or type 7 followed by text. Any LLM can read it.
How it went
Each page comes out as roughly 2-5KB of text instead of a roughly 1MB screenshot, so no vision model is needed. The post gives no results from actual job applications and asks for feedback on the grid format.
Worth knowing
It installs with npm install -g textweb and plugs in through an MCP server, function-calling definitions, LangChain, CrewAI, an HTTP API, a CLI, or a Node.js library.

Try it yourself with Claude Code

Write a small tool that renders a web page as a text grid, listing each form field, button, and link with a reference number, so an LLM agent can fill out forms without screenshots. Test it on [job application or form URL] using placeholder data and show me the filled fields. Never submit the form.

Read the original ↗

Source: HN · Feb 19, 2026

Five of these in your inbox every morning

The best things people got an AI agent to do, each with the prompt to try it.

More like this