GPT-Live
GPT-Live
GPT-Live is an OpenAI voice model that listens to the caller’s audio and generates its own speech. It keeps listening while it speaks, so it can respond to an interruption or a new detail while it’s talking.
In Vapi, a GPT-Live assistant pairs that voice model with a second model that does the task work, so the conversation can carry on while your tools run.
Private beta. GPT-Live must be enabled for your Vapi organization. See Access. If you already have a Vapi assistant or squad, see Migrate to GPT-Live.
How it works
A GPT-Live assistant has two parts:
When the speaker needs something done, such as checking a calendar, it delegates the task to the reasoner. The reasoner works from the conversation so far and your instructions. Meanwhile the speaker can continue listening and responding to the caller. When the result comes back, the speaker fits it into the current conversation.
Because GPT-Live hears and speaks directly, you don’t configure a separate transcriber or text-to-speech provider.
What changes with GPT-Live
For the caller, the conversation can feel less like taking turns. They can speak over the assistant, add a detail while it’s talking, or ask something else while it’s checking.
For you, more of the design lives in the two prompts:
- Decide what happens while work runs. Which work to start early, what the assistant can usefully ask or answer while it waits, and how it brings the result back into the conversation.
- Make delegation explicit. The speaker only starts work when it delegates, so tell it which requests need the reasoner.
- Plan for generated speech. All speech, including the greeting, is generated, so exact wording isn’t guaranteed.
- Keep saying and doing separate. The assistant saying something happened is different from your service confirming it.
Design conversations works through these decisions with an appointment-booking example. As an assistant takes on more kinds of work, reasoner skills let you organize each procedure and its tools behind the same conversation.
In this guide
What Vapi provides
Vapi connects GPT-Live to your phone numbers and tools, and provides call records and observability according to your recording and data settings:
- Browser, Twilio, Vapi SIP, and WebSocket calls.
- Function, API request, and MCP tools, plus built-in end-call, transfer, and DTMF tools.
- Speaker and reasoner prompts, reasoner skills, personality packs, and 22 voices.
- Idle-message hooks and HTTP live call control.
- Transcripts, recordings, structured outputs, scorecards, Boards, Monitoring, and Voice Simulations.
Before you commit, check whether your assistant depends on something GPT-Live doesn’t support:
The full list is in Settings and compatibility.
These pages describe Vapi’s GPT-Live integration. OpenAI’s API and ChatGPT offer some capabilities, such as web search, custom voices, and image input, that aren’t available through Vapi. For background on the model itself, see OpenAI’s GPT-Live documentation.
Access
GPT-Live is in private beta and must be enabled for your organization. If you don’t have access yet, join the waitlist. Joining doesn’t guarantee early access.
Cost
A GPT-Live call is billed for voice time by the second, including silence and waiting, plus Vapi’s platform fee, the reasoner’s token usage, and telephony. See Cost for a worked example and how to check a real call.