GPT-Live settings and compatibility

Beta
Supported features, fields, voices, hooks, call connections, and live controls

This page is the reference for GPT-Live’s settings and supported features. For how to use them in a conversation, start with Design conversations.

Jump to: Compatibility · Model settings · Skills · Personality · Voices · Greetings · Idle messages · Calls · Live control

Compatibility

GPT-Live supports a different set of features from Vapi’s other voice architectures. These limits describe Vapi’s GPT-Live integration. A capability in OpenAI’s API or in ChatGPT isn’t necessarily available through Vapi.

Calls and voice

FeatureSupport
Browser WebRTC, native Twilio, and Vapi SIPSupported
Raw WebSocketSupported with mono, signed 16-bit little-endian PCM at 24 kHz. Set the rate explicitly
Native Telnyx or VonageNot supported
22 preset OpenAI voicesSupported. See voices
Custom or cloned voices, or other voice providersNot supported
Changing voice during a callNot supported. Save the assistant and start a new call
Model and voice fallbacksNot supported. GPT-Live also can’t be a fallback model
Reconnecting to an existing GPT-Live sessionNot supported. Start a new call and check the status of pending work

Transfers and keypad input

FeatureSupport
Blind transferSupported on native Twilio and Vapi SIP, to phone or SIP destinations. Use a transferCall tool with configured destinations, or an HTTP transfer request
Transfer on browser or raw WebSocket callsNot supported
Warm transfer or generated transfer summaryNot supported
Transfer fallback plans or SIP bye verbNot supported
SIP transfer with extension dialingNot supported. Use a direct destination
Outgoing DTMFSupported on SIP using RTP DTMF. Pass digits as a string, for example "0011*#"
Outgoing DTMF on Twilio, browser, or raw WebSocket callsNot supported
SIP INFO DTMFNot supported
Incoming keypad collectionNot supported. Leave keypadInputPlan.enabled disabled

Tools

Tool typeSupport
function, apiRequest, mcpSupported. Called by the reasoner
endCallSupported
transferCall, dtmfSupported on the connections listed above
Other tool types, including the Query tool, handoff, SMS, voicemail, and integration toolsNot supported. Use a function, API request, or MCP tool instead

Give tools unique names. Classic tool messages, such as request-start messages, aren’t spoken. Guide what the assistant says while work runs in the speaker prompt.

Assistant features

FeatureSupport
Squads and assistant handoffsNot supported. See Migrate to GPT-Live
Reasoner skillsSupported. See Reasoner skills
Knowledge basesNot available to the reasoner. Use a retrieval tool
Exact or prerecorded speechNot supported. All speech, including greetings and idle check-ins, is generated
Audio URL greetings, generated first-message modeNot supported
Idle-message hooksSupported with say.prompt, message.add, and inline endCall. See Idle messages
Other assistant hook eventsNot covered by idle-message support. Check before relying on them
Live call controlEnd call, append speaker context, and cold transfer over HTTP. See Live call control
Voice SimulationsSupported in Voice mode with a single assistant. See Test and improve

Recordings and post-call data

GPT-Live supports post-call analysis, structured outputs, scorecards, Boards, and Monitoring. Availability depends on your configuration and retained call data:

Setting or featureBehavior
Recording disabled in the assistantNo recording is created
Recording-consent plan configuredRecording is disabled, because GPT-Live doesn’t collect that consent
Transcript-based analysisRequires transcripts
Audio-based analysisRequires a recording
Monitoring with Zero Data RetentionNot supported
Live listen or monitor socketsNot supported. Post-call Monitoring doesn’t provide live audio

Settings that no longer apply

Classic transcriber selection, text-to-speech settings, and start and stop speaking plans don’t configure a GPT-Live conversation. General model fields such as temperature and maxTokens don’t tune it either.

Model settings

These fields belong to the assistant object. See Create Assistant for the full schema.

FieldValues or behavior
model.provideropenai
model.modelgpt-live-1
model.speaker.instructionsSpeaker prompt. See prompt defaults
model.speaker.personalityPacksArray of eager-listener, idle-hummer, bouncy, and unhurried
model.reasoner.provideropenai. Defaults to OpenAI when omitted
model.reasoner.modelgpt-5.6-sol, gpt-5.6-terra, or gpt-5.6-luna. Defaults to gpt-5.6-terra
model.reasoner.reasoningEffortnone, low, medium, high, xhigh, or max. Defaults to low. Higher effort can increase response time
model.reasoner.instructionsReasoner prompt. Omit to use Vapi’s default
model.reasoner.skillsArray of skills. See Reasoner skills
model.toolsInline base tool definitions, available whether or not a skill is loaded
model.toolIdsIDs of saved base tools
voice.provideropenai
voice.voiceIdOne of the 22 supported voices

Prompt defaults and overrides

  • Explicit speaker instructions take precedence over model.systemPrompt and system messages in model.messages. If you omit the speaker instructions, Vapi uses that classic configuration.
  • Omitting reasoner instructions uses Vapi’s default reasoner prompt. Custom instructions replace that prompt entirely. When skills are configured, Vapi adds skill-loading guidance and the content of loaded skills.
  • An explicit empty string stays empty. Updating one prompt doesn’t update the other.
  • Personality packs append guidance to the speaker prompt.
  • Vapi adds a summary of skill names and descriptions to the speaker prompt. The reasoner manages the catalog, loads skills, and uses their instructions and tools.

Reasoner skills

Set model.reasoner.skills to an array of skills. See Reasoner skills for how to design them.

FieldBehaviorLimit
nameUnique within the assistant. Start with a lowercase letter; use lowercase letters, digits, hyphens, or underscores. The speaker and reasoner both see it64 characters
descriptionTells the reasoner when to load the skill; also included in the speaker’s skill summary1,024 characters
contentThe full procedure. Only the reasoner sees it, once the skill is loaded32,000 characters
toolsOptional inline tool definitions, available only while the skill is loaded20 tools
toolIdsOptional IDs of saved tools, available only while the skill is loaded20 tools

An assistant can have up to 20 skills.

Skills are stored with the assistant. Tools in model.tools and model.toolIds stay available whether or not a skill is loaded. Each delegation starts from the catalog and loads the skills it needs, and the reasoner can unload a skill when it finishes that work or changes tasks. A saved tool ID that doesn’t exist can stop calls from starting.

Personality and language

Set model.speaker.personalityPacks to an array of pack IDs. Each pack appends speaking-style guidance to the speaker prompt:

PackWhat its guidance encourages
eager-listenerBrief acknowledgments at natural openings, with room for the caller to continue
idle-hummerSoft humming or wordless sounds during quiet moments while waiting for tools
bouncyMore expressive pitch, emphasis, and rhythm, adapting to the caller’s mood
unhurriedA slower pace, clear articulation, and space between ideas
{
"speaker": {
"instructions": "Speak warmly and briefly. Slow down for dates and times.",
"personalityPacks": ["unhurried"]
}
}

Packs are prompt guidance, not fixed speed or pause controls. Test one at a time. See Speaking style for how to choose.

Write the speaker prompt and firstMessage in the language you want the assistant to speak. For Bossa or Tempo, for example:

Fale em português do Brasil, com respostas curtas e claras.
Faça uma pergunta por vez. Só confirme um agendamento
quando a ferramenta confirmar que ele foi realizado.

Keep delegation triggers in the prompt in the same language. Try names, dates, and product terms on real calls, and add pronunciation guidance where needed.

Voices

Choose any of these 22 voices with voice.provider: "openai" and voice.voiceId. Save the assistant and start a new call to change voices. In-call voice changes aren’t supported.

Voices with regional tags

Use the tags to narrow your choices, then listen to the samples. Bossa and Tempo previews are in Brazilian Portuguese. The other previews are in English. Test your preferred voice with your own prompts to hear how it handles the accent and vocabulary you need.

VoiceAPI IDLanguageRegional influencePresentationSourcePreview
QuartzquartzEnglishAustralianFeminineGeneratedListen to Quartz
RipplerippleEnglishAustralianMasculineNaturalListen to Ripple
VespervesperEnglishBritishMasculineNaturalListen to Vesper
WillowwillowEnglishIrishFeminineNaturalListen to Willow
StonestoneEnglishIrishMasculineNaturalListen to Stone
GleamgleamEnglishNorth AmericanFeminineNaturalListen to Gleam
MeridianmeridianEnglishNorth AmericanMasculineNaturalListen to Meridian
BossabossaPortugueseBrazilianFeminineNaturalListen to Bossa
TempotempoPortugueseBrazilianMasculineNaturalListen to Tempo
BeaconbeaconEnglishFilipinoMasculineGeneratedListen to Beacon
DeltadeltaEnglishSouthern U.S.FeminineGeneratedListen to Delta
CindercinderEnglishSouthern U.S.MasculineGeneratedListen to Cinder

More voices

VoiceAPI IDPreview
AlloyalloyListen to Alloy
AshashListen to Ash
BalladballadListen to Ballad
CedarcedarListen to Cedar
CoralcoralListen to Coral
EchoechoListen to Echo
MarinmarinListen to Marin
SagesageListen to Sage
ShimmershimmerListen to Shimmer
VerseverseListen to Verse

Greetings and duration

Field or featureBehavior
firstMessageMode: "assistant-speaks-first"Speaks a greeting generated from the text firstMessage
firstMessageMode: "assistant-waits-for-user"Waits for the caller
firstMessageText greeting. Without usable text, the assistant waits
maxDurationSecondsMaximum call duration in seconds
Audio URL greetingNot supported. Not played as a greeting
Generated first-message modeNot supported. The assistant waits for the caller

The spoken greeting may differ from the text you supply.

Idle messages

Idle messages use customer.speech.timeout assistant hooks. See When the caller goes quiet for a complete example.

Hook optionBehavior
options.timeoutSecondsSet explicitly for each hook. Accepts 2–1000 seconds
options.triggerMaxCountLimits how many times the hook fires across quiet periods. Accepts 1–10. Defaults to 3
options.triggerResetModeDefaults to never, which keeps the count for the whole call. Set onUserSpeech to reset the count when the caller speaks

How timing works:

  • A quiet period starts when the conversation becomes active, and again after the assistant speaks or the reasoner finishes work.
  • Activity is measured from speech that appears in the transcript. Background noise doesn’t count.
  • All hooks count from the start of the same quiet period. A check-in doesn’t postpone later hooks, and each hook fires at most once per quiet period. For another check-in in the same quiet period, add a separate hook with a different timeoutSeconds.
  • When the caller speaks, check-ins stop until the next quiet period starts. A check-in that was already being spoken may still be heard.
  • If the caller asks for a moment and the assistant replies, that reply starts a new quiet period.
  • Check-ins wait while the reasoner is actively working. Once an async tool has been dispatched, check-ins can resume even though its result is still pending.

Timing follows conversation activity, so test the delays with your prompts and tools.

Supported idle-message hook actions

ActionGPT-Live behavior
say with promptGenerates one spoken response guided by the prompt. The model chooses the wording. If the caller interrupts, it answers the caller without resuming the check-in
message.addAdds context to the speaker and requests a response by default. Set triggerResponseEnabled: false to add context without requesting speech. System and developer messages become speaker instructions. Other roles become speaker context
tool with inline tool: { "type": "endCall" }Ends the call, after an accompanying say action if there is one. Once the hook fires, caller speech doesn’t cancel the hangup

Other actions, including saved toolId references, function calls, and transfers, aren’t run by GPT-Live hooks. The endCall tool’s own messages aren’t used, so configure the goodbye as a separate say action. Hook messages affect the speaker, not the reasoner. Hook actions are separate from HTTP live call control.

Connect a call

Start with the dashboard’s Talk button, then connect your call channel. GPT-Live supports browser WebRTC, native Twilio, Vapi SIP, and raw WebSocket audio.

Phone calls

For inbound Twilio calls, import a Twilio phone number and assign your assistant to it. For SIP, follow the SIP guide.

For an outbound Twilio test, send this body to Create Call, using a customer number you control:

{
"assistantId": "YOUR_ASSISTANT_ID",
"phoneNumberId": "YOUR_TWILIO_PHONE_NUMBER_ID",
"customer": { "number": "+14155550100" }
}

phoneNumberId is the ID Vapi gave your imported Twilio number, not Twilio’s own phone number SID.

WebSocket configuration

Send this body to POST https://api.vapi.ai/call with your Vapi private API key:

{
"assistantId": "YOUR_ASSISTANT_ID",
"transport": {
"provider": "vapi.websocket",
"audioFormat": {
"format": "pcm_s16le",
"container": "raw",
"sampleRate": 24000
}
}
}

Connect to transport.websocketCallUrl in the response. Send and receive binary frames of mono, signed 16-bit little-endian audio at 24 kHz. The default format is rejected, so set sampleRate to 24000 explicitly. Browser and phone connections negotiate their own formats.

To end the call, close the WebSocket, have the assistant call endCall, or send an HTTP end-call request to the live control URL. Send controls over HTTP, not as JSON frames on the audio WebSocket. See WebSocket transport for the general transport.

Live call control

Your application can end an active GPT-Live call, add context for the speaker, or start a cold transfer. Send an HTTP POST to the call’s monitor.controlUrl, returned in the call object. These commands aren’t accepted as audio WebSocket text frames or SDK data-channel messages.

Set CONTROL_URL to the returned URL. Call control must be enabled in the assistant’s monitorPlan: setting controlEnabled to false disables it. If controlAuthenticationEnabled is true, include Authorization: Bearer YOUR_VAPI_PUBLIC_API_KEY from the same organization on each request. The examples below assume control authentication is disabled. Keep the control URL private.

End the call

curl --fail-with-body -X POST "$CONTROL_URL" \
-H "Content-Type: application/json" \
-d '{"type":"end-call"}'

Append speaker context

Use append-context with a kind and nonblank content of up to 16,000 characters:

KindPurposeExample content
commentaryInformation for the speaker to convey in its own wordsYour appointment is confirmed for Tuesday at 2 PM.
thinkingContext for the speaker to use, without prompting speechThe caller has already verified their account.
instructionsGuidance for the speaker’s behavior, including asking it to try saying somethingTell the caller their order is ready, then ask whether they need anything else.
curl --fail-with-body -X POST "$CONTROL_URL" \
-H "Content-Type: application/json" \
-d '{"type":"append-context","kind":"instructions","content":"Tell the caller their order is ready, then ask whether they need anything else."}'

This context goes to the speaker only. It doesn’t reach the reasoner or start reasoning or tool work, and it isn’t added to the transcript as speech. thinking doesn’t request speech, but the speaker can use it and repeat it later, so don’t send secrets that way.

A successful response means the context was submitted. It doesn’t guarantee exact wording, immediate speech, or completed playback.

Cold transfer

On native Twilio and Vapi SIP calls, send a phone or SIP destination:

curl --fail-with-body -X POST "$CONTROL_URL" \
-H "Content-Type: application/json" \
-d '{"type":"transfer","destination":{"type":"number","number":"+14155550100","transferPlan":{"mode":"blind-transfer"}}}'

For SIP, use a destination such as {"type":"sip","sipUri":"sip:agent@example.com"}. Your application chooses the destination in this request. A reasoner using a transferCall tool chooses only from that tool’s configured destinations.

Cold transfer isn’t supported on browser or raw WebSocket calls. Don’t include pre-transfer content, warm-transfer plans, generated summaries, or fallback plans. A successful response means the carrier accepted the transfer, not that the destination answered.

Responses and unsupported controls

Successful commands return HTTP 200 with {"status":"ok"}. Rejected requests return an error, including when the call isn’t active or is already transferring. If a 502 response reports an unconfirmed outcome, don’t retry automatically: the action may have reached the provider. A 503 response means no action was taken and is safe to retry.

GPT-Live doesn’t support the HTTP say, add-message, or mute and unmute controls. Use append-context for speaker guidance. The say and message.add actions in idle-message hooks are assistant configuration, not HTTP controls.

Moved sections

These sections used to be on this page.

Before you start

Prerequisites are now in the Quickstart and Build with tools.

Create an assistant

See the Quickstart for a first assistant, and Build with tools for connecting tools.

Make a test call

See Talk to it and Check what happened.

Connect your tools

See Build with tools and the tool response contract.

MCP tools

See Use existing MCP tools.

Slow and asynchronous requests

See Slow and external work.

Review and monitor calls

See Watch production calls.

Cost

See Cost.