Migrate to GPT-Live
Migrate to GPT-Live
Move a copy of your assistant first, compare it with the original, and switch traffic after it meets the same goals. This guide covers a single assistant. If you use a squad, see From a squad for the current limits.
Before you start
Check that your assistant can move
Some features don’t carry over. If your assistant depends on one of these, plan how to replace it before you begin:
- Voices must be one of the OpenAI voices GPT-Live supports. Other voice providers and custom or cloned voices aren’t available.
- Squads and assistant handoffs aren’t supported. There is no direct squad conversion; see the current limits.
- Model and voice fallbacks must be removed. GPT-Live can’t be a fallback model either.
- Exact or prerecorded speech isn’t available. GPT-Live generates all of its speech, including the greeting.
- Keypad input from callers isn’t supported.
- Knowledge bases and the Query tool aren’t available. Use a function, API request, or MCP tool for retrieval.
- Transfers are blind only, on native Twilio and Vapi SIP calls. Native Telnyx and Vonage numbers aren’t supported.
The full list is in Settings and compatibility.
Save representative calls
Pick a handful of calls that represent what the assistant has to get right: a straightforward success, a caller who changes their mind, a failure your tools return, and anything specific to your business. For each, note the starting state, the expected tool actions, the expected outcome, and a recording from the current assistant. These are what you’ll compare against.
Keep the original running
Leave your phone numbers on the original assistant until the GPT-Live version has passed the same calls. Work on a copy, test it on its own, and switch traffic when you’re ready.
From a single assistant
1. Take inventory
Write down what the assistant does and what it depends on:
2. Convert a copy
Convert a copy, so the original keeps serving callers while you test:
Duplicate the assistant
In Assistants, open the assistant’s menu and select Duplicate. Don’t attach a phone number to the copy.
Because conversion publishes immediately, working on a duplicate is what keeps callers on the original until you’re ready.
To go back, select Revert to Classic. It restores the assistant’s Classic configuration as a draft, and the published version stays GPT-Live until you select Publish.
You can also migrate by hand through the API: create a new assistant with the GPT-Live configuration from the steps below.
3. Split the prompt
A classic assistant has one prompt that covers both how to talk and how to do the work. GPT-Live has two. The speaker prompt shapes the conversation and says when to delegate. The reasoner prompt holds the procedures and tool rules. Neither sees the other, so rules that affect both belong in both, worded for each role.
For example, a classic prompt for an appointment assistant at Example Service Studio, a fictional business, might read:
Split it by who does what:
Three things changed besides the split:
- Explicit delegation triggers. The speaker only starts work when it delegates, so every kind of task, including ending the call, is listed.
- Waiting is designed. In the classic flow, the assistant read times out after a lookup and then asked for the name. The speaker now asks for the name while the lookup runs.
- Results are for speaking. The reasoner returns short facts that the speaker can use directly.
4. Map the settings
5. Redesign the waits
Look at the places where your current flow waits for a tool before it moves on. With GPT-Live, the conversation can keep going during that time, so decide what it should cover. Start work as soon as its inputs are known, then ask for something the next step will need:
Caller: A consultation downtown on Friday.
Assistant: I’ll check Friday downtown. Who should I put the appointment under?
Caller: Sam Lee.
Assistant: Thanks, Sam. Friday I have 10:30, 2:30, or 4.
When the next step is blocked, the conversation can still help. Here the booking is running, and the caller asks about something the assistant already knows:
Assistant: That’s Friday at 2:30, downtown, under Sam Lee. Shall I book it?
Caller: Yes. Do I need to bring anything?
Assistant: Nothing’s required for a consultation. Bring any questions you’d like to cover, and it helps to arrive five minutes early.
Assistant: You’re booked: Friday at 2:30 downtown, under Sam Lee.
The preparation answer doesn’t depend on the booking, so it can come while the booking runs. “You’re booked” waits for the result. For the speaker to answer like this, give it a few stable facts in its prompt, such as arrival and preparation guidance. When there’s nothing useful to say, one acknowledgment and a pause is better than filler. Design conversations covers these choices in depth.
6. Compare with the original
Run your saved calls against the copy. For each one, check:
- The outcome. Your service’s state shows the same result as before, with no duplicate or missing actions.
- The tool calls. The call’s messages show the expected tool calls and arguments.
- The conversation. The recording shows the assistant reached the goal without making the caller repeat themselves, and that waits sounded reasonable.
Record differences you intended, such as a shorter call, separately from regressions. Test and improve describes how to repeat these calls with Voice Simulations.
7. Roll out and roll back
When the copy passes, assign your phone number to it. Keep the original assistant, unchanged, until you’re confident in the new one.
If you need to take a converted assistant back to Classic, select Revert to Classic, then Publish. Until you publish, callers still reach the GPT-Live version.
From a squad
GPT-Live doesn’t support squads or assistant handoffs. Switch to GPT Live converts one assistant; it doesn’t convert a squad or preserve its routing and handoff behavior.
Keep your existing squad if your call flow depends on those capabilities. The single-assistant steps above aren’t a replacement for a squad migration.