Create Assistant
Authentication
Authenticate server-side requests with a private Vapi API key. Create or copy a key from the Vapi Dashboard and send it in the Authorization header as Bearer <token>. Keep private API keys out of client-side code and public repositories.
Request
This is the first message that the assistant will say. This can also be a URL to a containerized audio file (mp3, wav, etc.).
If unspecified, assistant will wait for user to speak and use the model to respond once they speak.
Set to true to allow the user to interrupt the assistant while it speaks the first message. Default is false.
This is the mode for the first message. Default is ‘assistant-speaks-first’.
Use:
- ‘assistant-speaks-first’ to have the assistant speak first.
- ‘assistant-waits-for-user’ to have the assistant wait for the user to speak first.
- ‘assistant-speaks-first-with-model-generated-message’ to have the assistant speak first with a message generated by the model based on the conversation state. (
assistant.model.messagesat call start,call.messagesat squad transfer points).
@default ‘assistant-speaks-first’
These are the settings to configure or disable voicemail detection. Alternatively, voicemail detection can be configured using the model.tools=[VoicemailTool]. By default, voicemail detection is disabled.
These are the messages that will be sent to your Client SDKs. Default is conversation-update,function-call,hang,model-output,speech-update,status-update,transfer-update,transcript,tool-calls,user-interrupted,voice-input,workflow.node.started,assistant.started. You can check the shape of the messages in ClientMessage schema.
These are the messages that will be sent to your Server URL. Default is conversation-update,end-of-call-report,function-call,hang,speech-update,status-update,tool-calls,transfer-destination-request,handoff-destination-request,user-interrupted,assistant.started. You can check the shape of the messages in ServerMessage schema.
This is the maximum number of seconds that the call will last. When the call reaches this duration, it will be ended.
@default 600 (10 minutes)
This determines whether the model’s output is used in conversation history rather than the transcription of assistant’s speech.
@default false
This enables filtering of noise and background speech while the user is talking.
Features:
- Smart denoising using Krisp
- Fourier denoising
Smart denoising can be combined with or used independently of Fourier denoising.
Order of precedence:
- Smart denoising
- Fourier denoising
This is the plan for artifacts generated during assistant’s calls. Stored in call.artifact.
This is the plan for when the assistant should start talking.
You should configure this if you’re running into these issues:
- The assistant is too slow to start talking after the customer is done speaking.
- The assistant is too fast to start talking after the customer is done speaking.
- The assistant is so fast that it’s actually interrupting the customer.
This is the plan for when assistant should stop talking on customer interruption.
You should configure this if you’re running into these issues:
- The assistant is too slow to recognize customer’s interruption.
- The assistant is too fast to recognize customer’s interruption.
- The assistant is getting interrupted by phrases that are just acknowledgments.
- The assistant is getting interrupted by background noises.
- The assistant is not properly stopping — it starts talking right after getting interrupted.
This is the plan for real-time monitoring of the assistant’s calls.
Usage:
- To enable live listening of the assistant’s calls, set
monitorPlan.listenEnabledtotrue. - To enable live control of the assistant’s calls, set
monitorPlan.controlEnabledtotrue. - To attach monitors to the assistant, set
monitorPlan.monitorIdsto the set of monitor ids.
This is where Vapi will send webhooks. You can find all webhooks available along with their shape in ServerMessage schema.
The order of precedence is:
- assistant.server.url
- phoneNumber.serverUrl
- org.serverUrl
This is the plan for analysis of assistant’s calls. Stored in call.analysis.
Response
This is the ISO 8601 date-time string of when the assistant was created.
This is the ISO 8601 date-time string of when the assistant was last updated.
This is the first message that the assistant will say. This can also be a URL to a containerized audio file (mp3, wav, etc.).
If unspecified, assistant will wait for user to speak and use the model to respond once they speak.
Set to true to allow the user to interrupt the assistant while it speaks the first message. Default is false.
This is the mode for the first message. Default is ‘assistant-speaks-first’.
Use:
- ‘assistant-speaks-first’ to have the assistant speak first.
- ‘assistant-waits-for-user’ to have the assistant wait for the user to speak first.
- ‘assistant-speaks-first-with-model-generated-message’ to have the assistant speak first with a message generated by the model based on the conversation state. (
assistant.model.messagesat call start,call.messagesat squad transfer points).
@default ‘assistant-speaks-first’
These are the settings to configure or disable voicemail detection. Alternatively, voicemail detection can be configured using the model.tools=[VoicemailTool]. By default, voicemail detection is disabled.
These are the messages that will be sent to your Client SDKs. Default is conversation-update,function-call,hang,model-output,speech-update,status-update,transfer-update,transcript,tool-calls,user-interrupted,voice-input,workflow.node.started,assistant.started. You can check the shape of the messages in ClientMessage schema.
These are the messages that will be sent to your Server URL. Default is conversation-update,end-of-call-report,function-call,hang,speech-update,status-update,tool-calls,transfer-destination-request,handoff-destination-request,user-interrupted,assistant.started. You can check the shape of the messages in ServerMessage schema.
This is the maximum number of seconds that the call will last. When the call reaches this duration, it will be ended.
@default 600 (10 minutes)
This determines whether the model’s output is used in conversation history rather than the transcription of assistant’s speech.
@default false
This is the latest version label (e.g. v3) of the assistant in the
version history. null while the org is not yet
onboarded to versioning, or for assistants that have not yet been
published under it.
Read-only. Present only when a model this configuration uses is deprecated or retired in Vapi’s model deprecation registry, judged on the day of the response. Each entry names the slot (for example model or model.fallbackModels[1]), provider, stored model, and deprecation and retirement dates as YYYY-MM-DD in UTC. replacementStatus is available with a replacementModel when a replacement can be recommended, or manual-action-required with no replacement model when eligibility is unknown or no eligible replacement exists. manual-action-required can be transient when compliance context is unavailable; re-fetch before acting. HIPAA-required configurations, sparse drafts, and squads with unresolved assistant references currently require manual action. HIPAA requirements include the organization and assistant settings, including HIPAA with data retention. Recommendations reflect the response-time decision; they do not confirm a swap or authorize future execution. Examples: an available recommendation includes {"replacementStatus":"available","replacementModel":"gpt-5"}; a blocked recommendation includes {"replacementStatus":"manual-action-required"}. Ignored if sent back in a create or update request.
This enables filtering of noise and background speech while the user is talking.
Features:
- Smart denoising using Krisp
- Fourier denoising
Smart denoising can be combined with or used independently of Fourier denoising.
Order of precedence:
- Smart denoising
- Fourier denoising
This is the plan for artifacts generated during assistant’s calls. Stored in call.artifact.
This is the plan for when the assistant should start talking.
You should configure this if you’re running into these issues:
- The assistant is too slow to start talking after the customer is done speaking.
- The assistant is too fast to start talking after the customer is done speaking.
- The assistant is so fast that it’s actually interrupting the customer.
This is the plan for when assistant should stop talking on customer interruption.
You should configure this if you’re running into these issues:
- The assistant is too slow to recognize customer’s interruption.
- The assistant is too fast to recognize customer’s interruption.
- The assistant is getting interrupted by phrases that are just acknowledgments.
- The assistant is getting interrupted by background noises.
- The assistant is not properly stopping — it starts talking right after getting interrupted.
This is the plan for real-time monitoring of the assistant’s calls.
Usage:
- To enable live listening of the assistant’s calls, set
monitorPlan.listenEnabledtotrue. - To enable live control of the assistant’s calls, set
monitorPlan.controlEnabledtotrue. - To attach monitors to the assistant, set
monitorPlan.monitorIdsto the set of monitor ids.
This is where Vapi will send webhooks. You can find all webhooks available along with their shape in ServerMessage schema.
The order of precedence is:
- assistant.server.url
- phoneNumber.serverUrl
- org.serverUrl
This is the plan for analysis of assistant’s calls. Stored in call.analysis.