Voice agents
A voice agent is an AI app you talk to out loud; it listens, works out what you want, can use tools, and answers in speech.
A voice agent is an AI app you talk to instead of type to. You ask a question or give it a job out loud, and it answers out loud. OpenAI’s guide says to pick how sound goes in and out first. After that, building the agent works much like building a text chatbot. The new part is sound: it has to listen and speak.
OpenAI’s guide says the biggest choice is how the talking part connects to the thinking part and the tools.
One option is a chained pipeline: three steps in a row. Speech-to-text (a model that writes down what it hears) turns your words into a transcript. The agent workflow, the same kind you would build for text, works on that transcript and produces a reply. Text-to-speech (a model that reads text aloud) turns the reply into sound. Because every stage passes plain text to the next, you can inspect or change it in between. You can also replace any one stage without changing the others.
Another option is speech-to-speech. Here one model works directly with audio, keeps track of the conversation and can call tools. OpenAI’s Realtime API works this way. OpenAI suggests it for agents that need three things: barge-in (letting you talk over the agent), a short wait before the first sound, and natural turn taking. In a browser, the app connects over WebRTC. On a server, it uses a WebSocket connection. OpenAI also offers GPT-Live, which can hear you while it is still talking. It hands the heavy thinking and tool calls to a separate system behind the scenes.
Google’s version is the Gemini Live API. It handles live voice and vision conversations with little delay. It takes in nonstop audio, pictures and text over one WebSocket. Audio goes in as raw 16-bit PCM at 16kHz. It comes back at 24kHz.
The hardest part is turn-taking: deciding when you have finished speaking. Both vendors use voice activity detection, or VAD, which spots when speech starts and stops. In OpenAI’s speech-to-speech sessions it is on by default and can be turned off. The basic mode watches for silence. A shorter silence setting ends turns faster. A “semantic” mode looks at meaning too: it uses your words to judge if you are done. If your sentence trails off, it waits longer than after a clear statement.
VAD also makes interruptions work. In the Gemini Live API, when VAD hears you cut in, the reply that was being generated is cancelled and thrown away. Only the part that already reached your app stays in the conversation history.
Imagine you phone a booking line. You say, “Move my appointment to next week.” VAD notices you have stopped. The agent understands the request, checks the calendar through a tool, and replies, “I can do Tuesday at ten. Does that work?” You interrupt with “Thursday instead”, and it stops mid-sentence and tries again.
OpenAI recommends measuring how long callers sit waiting before they hear a helpful reply. They also suggest comparing similar calls: the median (the typical wait) and the 95th percentile (a wait that only about one call in twenty is slower than). Slowdowns can come from connecting, the model thinking, tool calls, holding audio in a waiting area (buffering) or playing sound back, so builders time each part separately.
Voice agents have clear limits. Speech recognition can mishear, so OpenAI advises testing across accents, background noise, names and numbers. An agent may say “Booked” when nothing was saved, so spoken confirmations should be checked against what the system actually did. A faster reply is not always a better one: it can come with a wrong tool call or a failed task. Risky actions may still need safety checks or a human’s approval. And OpenAI’s usage policies require telling listeners that a text-to-speech voice is AI-generated, not a human.
Talking to software is harder than typing to it.
Follow one spoken request from microphone to reply.
- 1 · listenThe app sends your voice to the model as a live audio stream.
- 2 · detectVoice activity detection marks when your speech starts and stops.
- 3 · thinkA model works out what you asked and can call tools.
- 4 · speakThe answer is turned into audio and played back.
Good turn detection is what makes a voice agent feel like a conversation instead of a walkie-talkie.
| Who | What they ask | What it works with |
|---|---|---|
| Booking line | “Can callers book an appointment by speaking?” | Whether the correct appointment was saved |
- It lets people finish tasks by speaking instead of typing.
- It can detect when a speaker starts and stops talking.
- It lets users interrupt the model while it is talking.
- It can call tools during a spoken conversation.
- It does not guarantee it heard names, numbers or accents correctly.
- Spoken confirmations still need checking against what the system actually saved.
- A faster reply can still come with a wrong tool call or a failed task.
- It does not remove the need for safety checks and human approval on risky actions.
Sources used
This explainer is written in original language. The links below support its factual claims.
- docsVoice agents, OpenAI · read 28 Sept 2026
- docsGetting started with the Realtime API, OpenAI · read 28 Sept 2026
- docsVoice activity detection (VAD), OpenAI · read 28 Sept 2026
- docsText to speech, OpenAI · read 28 Sept 2026
- docsGemini Live API overview, Google AI for Developers · read 28 Sept 2026
- docsLive API capabilities guide, Google AI for Developers · read 28 Sept 2026