agentusecasesAll 965 use cases
Software & DevOps

Build a phone voice agent that answers in under half a second

A from-scratch voice agent averaging about 400ms end to end, with streaming speech-to-text, LLM, and text-to-speech and clean barge-ins.

Done with a custom-built or unnamed agent

What they did
A FastAPI server takes Twilio's audio over a WebSocket and sends it to Deepgram Flux, which handles transcription and turn detection. When a turn ends, the transcript goes to an LLM, tokens stream into ElevenLabs TTS, and the audio goes straight back to Twilio. A barge-in cancels the LLM and TTS and flushes Twilio's buffer. Pre-connected TTS sockets are kept warm.
How it went
Run locally, latency was about 1.7s. Deploying in the EU cut it to about 790ms, slightly better than Vapi's estimate of about 840ms. Swapping gpt-4o-mini for Groq's llama-3.3-70b brought it to about 400ms. The first turn was slower.
Worth knowing
The build took about a day and roughly $100 in API credits. Keep the server in the same region as Twilio, Deepgram and ElevenLabs, because location halved latency.

Try it yourself with Claude Code

Help me build a phone voice agent in [language, e.g. Python] for [use case, e.g. answering calls for a dental office] using streaming speech-to-text, an LLM, and streaming text-to-speech, with the goal of responding in under 500ms and handling interruptions cleanly. Scaffold the project, wire up [providers, e.g. Deepgram, an LLM API, ElevenLabs, Twilio], and add latency logging. Finish when a test call works and the log shows average latency; ask before making any real phone calls or purchasing numbers.

Read the original ↗

Discussion on HN · Mar 2, 2026

Five of these in your inbox every morning

The best things people got an AI agent to do, each with the prompt to try it.

More like this