Path to production

Build, publish, test, and monitor an assistant

Overview

A working first call is the start. Before real customers call your assistant, publish a version, test it, and set up a way to catch problems. This page walks through that path and links to the detailed guide for each step.

The path to production:

  1. Build an assistant: Create an assistant, choose a model preset, and try it in Chat.
  2. Publish and test: Define what success looks like, publish a version, then test your setup with Simulations.
  3. Go live: Get a phone number, assign it to your assistant, and place a controlled call.
  4. Monitor: Review calls and get alerts when quality drops.
  5. Iterate: Change a draft, publish, and retest. As your team grows, manage your configuration as code.

Prerequisites

1. Build an assistant

In the Dashboard, open Assistants. To start from a template, select the down arrow next to Create Assistant, then choose a template. To start from scratch, select Create Assistant. Then choose the transcriber, model, and voice.

Model Presets bundle a transcriber, model, and voice for a specific goal. New assistants use Balanced by default. Assistants created from a template use the preset that best fits the template.

PresetChoose it when
BalancedYou’re not sure which preset fits. Start here.
High IntelligenceConversations are complex or high-stakes.
Ultra FastResponse speed matters most.
Cost SaverCalls are simple and high-volume.

To compare individual components, use the Performance Metrics shown for latency, cost, and quality.

Vapi saves your edits as a draft. To try draft changes, select the down arrow next to Talk in the assistant editor, then select Chat. Chat uses the configuration shown in the editor. See Test draft changes in Chat.

2. Publish and test

One good test call doesn’t prove an assistant is ready. Define what success looks like, publish a version, then use Simulations to test your setup.

1

Define success with structured outputs

Structured outputs extract specific data from each conversation, such as whether the caller’s issue was resolved or an appointment was booked. Simulations use structured outputs as their success criteria. Attach the same structured outputs to your assistant in the Artifact Plan section so they also run on every call. Start with the Structured outputs quickstart.

2

Publish a version

Select Publish, review the changes, then select Publish or Quick Publish. Publishing creates a new version and makes it the current version. See Versioning assistants.

3

Test your setup with Simulations

A simulation runs an AI tester through a complete conversation with your assistant and returns a pass or fail result. Save the conversations that matter most as a simulation suite, then run the suite again after every change. Start with the Simulations quickstart.

Simulations test the assistant’s latest published version, not your draft. Publish your changes before you run a simulation suite.

Simulations run unmocked tools for real. Use sandbox integrations or mock tools that could create bookings, send messages, or affect real customers.

To check a single decision, such as which tool the assistant calls and with which arguments, add Evals.

To plan which conversations to cover, see Testing voice agents and Plan test coverage.

3. Go live

When the assistant passes your simulations, connect it to real callers.

1

Get a phone number

Create a free Vapi number, or import a number from Twilio, Telnyx, or DIDWW. With free Vapi numbers, only inbound calls are supported. Import a number to make outbound or international calls.

2

Assign the number to your assistant

In the Dashboard, open Phone Numbers and select the number. Under Inbound Settings, choose your assistant, then select Save. Inbound calls use the assistant’s current published version.

3

Place a controlled call

Call the number yourself to check the full phone path, including audio and live integrations, before customers do.

When the call ends, open it in Logs → Calls and check the Structured Outputs section to see whether the call met your success criteria. Results appear a few seconds after the call ends.

Outbound and web calls also use the current version by default. To use a specific version, pass assistantVersion with assistantId when you create a call.

4. Monitor

After launch, watch real calls so you find problems before customers report them.

  • Review calls: Open Logs → Calls to see each call’s transcript, recording, and how it ended. To understand an ended call, see Call ended reasons.
  • Track outcomes: Your structured outputs extract the same data from every live call, so you can see which calls met your success criteria.
  • Set up alerts: Monitors check your call data against thresholds you set and alert your team by email, Slack, or webhook. Effectiveness and compliance monitors use structured outputs.
  • Track trends: Build charts of call metrics with Boards.

5. Iterate

Use what you learn from production to improve the assistant.

  1. Turn a confirmed production failure into a regression test.
  2. Edit the draft and try the change in Chat.
  3. Publish a new version and rerun your simulation suite.

After the assistant has a phone number, a published version takes live calls right away, before you rerun your simulations. If a new version causes problems, restore an earlier version. Restoring creates a new current version immediately. To test changes before they reach live calls, use config as code.

Manage your configuration as code

As your team grows, use config as code (GitOps) to keep assistants, squads, tools, structured outputs, and simulations as files in git. Every change goes through a pull request, simulation suites can run against the pull request before it merges, and you promote changes from a development org to production instead of editing production directly. See Development, staging and production.

Next steps