GPT-Live settings and compatibility
GPT-Live settings and compatibility
This page is the reference for GPT-Live’s settings and supported features. For how to use them in a conversation, start with Design conversations.
Jump to: Compatibility · Model settings · Skills · Personality · Voices · Greetings · Idle messages · Calls · Live control
Compatibility
GPT-Live supports a different set of features from Vapi’s other voice architectures. These limits describe Vapi’s GPT-Live integration. A capability in OpenAI’s API or in ChatGPT isn’t necessarily available through Vapi.
Calls and voice
Transfers and keypad input
Tools
Give tools unique names. Classic tool messages, such as request-start messages, aren’t spoken. Guide what the assistant says while work runs in the speaker prompt.
Assistant features
Recordings and post-call data
GPT-Live supports post-call analysis, structured outputs, scorecards, Boards, and Monitoring. Availability depends on your configuration and retained call data:
Settings that no longer apply
Classic transcriber selection, text-to-speech settings, and start and stop speaking plans don’t configure a GPT-Live conversation. General model fields such as temperature and maxTokens don’t tune it either.
Model settings
These fields belong to the assistant object. See Create Assistant for the full schema.
Prompt defaults and overrides
- Explicit speaker instructions take precedence over
model.systemPromptand system messages inmodel.messages. If you omit the speaker instructions, Vapi uses that classic configuration. - Omitting reasoner instructions uses Vapi’s default reasoner prompt. Custom instructions replace that prompt entirely. When skills are configured, Vapi adds skill-loading guidance and the content of loaded skills.
- An explicit empty string stays empty. Updating one prompt doesn’t update the other.
- Personality packs append guidance to the speaker prompt.
- Vapi adds a summary of skill names and descriptions to the speaker prompt. The reasoner manages the catalog, loads skills, and uses their instructions and tools.
Reasoner skills
Set model.reasoner.skills to an array of skills. See Reasoner skills for how to design them.
An assistant can have up to 20 skills.
Skills are stored with the assistant. Tools in model.tools and model.toolIds stay available whether or not a skill is loaded. Each delegation starts from the catalog and loads the skills it needs, and the reasoner can unload a skill when it finishes that work or changes tasks. A saved tool ID that doesn’t exist can stop calls from starting.
Personality and language
Set model.speaker.personalityPacks to an array of pack IDs. Each pack appends speaking-style guidance to the speaker prompt:
Packs are prompt guidance, not fixed speed or pause controls. Test one at a time. See Speaking style for how to choose.
Write the speaker prompt and firstMessage in the language you want the assistant to speak. For Bossa or Tempo, for example:
Keep delegation triggers in the prompt in the same language. Try names, dates, and product terms on real calls, and add pronunciation guidance where needed.
Voices
Choose any of these 22 voices with voice.provider: "openai" and voice.voiceId. Save the assistant and start a new call to change voices. In-call voice changes aren’t supported.
Voices with regional tags
Use the tags to narrow your choices, then listen to the samples. Bossa and Tempo previews are in Brazilian Portuguese. The other previews are in English. Test your preferred voice with your own prompts to hear how it handles the accent and vocabulary you need.
More voices
Greetings and duration
The spoken greeting may differ from the text you supply.
Idle messages
Idle messages use customer.speech.timeout assistant hooks. See When the caller goes quiet for a complete example.
How timing works:
- A quiet period starts when the conversation becomes active, and again after the assistant speaks or the reasoner finishes work.
- Activity is measured from speech that appears in the transcript. Background noise doesn’t count.
- All hooks count from the start of the same quiet period. A check-in doesn’t postpone later hooks, and each hook fires at most once per quiet period. For another check-in in the same quiet period, add a separate hook with a different
timeoutSeconds. - When the caller speaks, check-ins stop until the next quiet period starts. A check-in that was already being spoken may still be heard.
- If the caller asks for a moment and the assistant replies, that reply starts a new quiet period.
- Check-ins wait while the reasoner is actively working. Once an async tool has been dispatched, check-ins can resume even though its result is still pending.
Timing follows conversation activity, so test the delays with your prompts and tools.
Supported idle-message hook actions
Other actions, including saved toolId references, function calls, and transfers, aren’t run by GPT-Live hooks. The endCall tool’s own messages aren’t used, so configure the goodbye as a separate say action. Hook messages affect the speaker, not the reasoner. Hook actions are separate from HTTP live call control.
Connect a call
Start with the dashboard’s Talk button, then connect your call channel. GPT-Live supports browser WebRTC, native Twilio, Vapi SIP, and raw WebSocket audio.
Phone calls
For inbound Twilio calls, import a Twilio phone number and assign your assistant to it. For SIP, follow the SIP guide.
For an outbound Twilio test, send this body to Create Call, using a customer number you control:
phoneNumberId is the ID Vapi gave your imported Twilio number, not Twilio’s own phone number SID.
WebSocket configuration
Send this body to POST https://api.vapi.ai/call with your Vapi private API key:
Connect to transport.websocketCallUrl in the response. Send and receive binary frames of mono, signed 16-bit little-endian audio at 24 kHz. The default format is rejected, so set sampleRate to 24000 explicitly. Browser and phone connections negotiate their own formats.
To end the call, close the WebSocket, have the assistant call endCall, or send an HTTP end-call request to the live control URL. Send controls over HTTP, not as JSON frames on the audio WebSocket. See WebSocket transport for the general transport.
Live call control
Your application can end an active GPT-Live call, add context for the speaker, or start a cold transfer. Send an HTTP POST to the call’s monitor.controlUrl, returned in the call object. These commands aren’t accepted as audio WebSocket text frames or SDK data-channel messages.
Set CONTROL_URL to the returned URL. Call control must be enabled in the assistant’s monitorPlan: setting controlEnabled to false disables it. If controlAuthenticationEnabled is true, include Authorization: Bearer YOUR_VAPI_PUBLIC_API_KEY from the same organization on each request. The examples below assume control authentication is disabled. Keep the control URL private.
End the call
Append speaker context
Use append-context with a kind and nonblank content of up to 16,000 characters:
This context goes to the speaker only. It doesn’t reach the reasoner or start reasoning or tool work, and it isn’t added to the transcript as speech. thinking doesn’t request speech, but the speaker can use it and repeat it later, so don’t send secrets that way.
A successful response means the context was submitted. It doesn’t guarantee exact wording, immediate speech, or completed playback.
Cold transfer
On native Twilio and Vapi SIP calls, send a phone or SIP destination:
For SIP, use a destination such as {"type":"sip","sipUri":"sip:agent@example.com"}. Your application chooses the destination in this request. A reasoner using a transferCall tool chooses only from that tool’s configured destinations.
Cold transfer isn’t supported on browser or raw WebSocket calls. Don’t include pre-transfer content, warm-transfer plans, generated summaries, or fallback plans. A successful response means the carrier accepted the transfer, not that the destination answered.
Responses and unsupported controls
Successful commands return HTTP 200 with {"status":"ok"}. Rejected requests return an error, including when the call isn’t active or is already transferring. If a 502 response reports an unconfirmed outcome, don’t retry automatically: the action may have reached the provider. A 503 response means no action was taken and is safe to retry.
GPT-Live doesn’t support the HTTP say, add-message, or mute and unmute controls. Use append-context for speaker guidance. The say and message.add actions in idle-message hooks are assistant configuration, not HTTP controls.
Moved sections
These sections used to be on this page.
Before you start
Prerequisites are now in the Quickstart and Build with tools.
Create an assistant
See the Quickstart for a first assistant, and Build with tools for connecting tools.
Make a test call
See Talk to it and Check what happened.
Connect your tools
See Build with tools and the tool response contract.
MCP tools
Slow and asynchronous requests
Review and monitor calls
Cost
See Cost.