Model Presets for transcribers, models, and voices
Model Presets are curated configurations that bundle a transcriber, model, and voice for a specific goal. Use a preset for a dependable setup or as a starting point for further customization.
How it works
New assistants use Balanced as the default. When you start from one of Vapi’s use-case templates, the assistant uses the preset that best fits the template.
Presets are a starting point rather than a lock - you can fine-tune or swap out any component to better fit your use case.
Presets at a glance
Choose a preset
Choose the preset that best matches your use case. Start with Balanced if you are unsure.
Balanced
Balanced is the default for new assistants and the best starting point for most use cases. It provides a strong mix of quality, responsiveness, and cost.
Choose Balanced when any of the following apply.
- You’re building a new assistant and aren’t sure which preset fits.
- Your use case covers support, scheduling, FAQs, or qualification.
- You want an all-around dependable setup out of the box.
If you later need more reasoning, faster responses, or lower cost, switch to the preset built for that.
High Intelligence
High Intelligence prioritizes reasoning and conversation quality. Use it when the assistant needs to handle nuance, follow multi-step logic, or reason reliably across tools and context.
Choose High Intelligence when any of the following apply.
- Conversations are complex, open-ended, or high-stakes.
- The assistant needs to reason through multi-step problems or use tools reliably.
- Accuracy and quality matter more than speed or cost.
More capable models respond a little slower and cost more per minute than Balanced. See how latency works for why more capable models take longer to respond.
Ultra Fast
Ultra Fast prioritizes low latency and responsive conversations. Use it when response speed matters more than complex reasoning.
Choose Ultra Fast when any of the following apply.
- Responsiveness is the priority and replies should feel fast.
- Calls are high-volume and relatively straightforward.
- The flow is scripted or transactional rather than open-ended.
The fastest models are smaller and less capable, so Ultra Fast may not suit complex reasoning or heavy tool use. Choose Balanced or High Intelligence if difficult tasks produce quality issues.
Cost Saver
Cost Saver prioritizes the lowest cost per minute. Use it for high-volume calls when cost is the main constraint.
Choose Cost Saver when any of the following apply.
- Call volume is high and you’re optimizing spend.
- Interactions are simple or scripted.
- A small quality trade-off is acceptable in exchange for lower cost.
The lowest-cost models are less capable, so this preset fits simpler interactions best. See how cost works for what drives cost per minute.
The Customized state
Swapping a model out of any component moves your assistant to Customized and the Performance Metrics displayed will update to match the new configuration.
Apply a Model Preset
Choose a preset
Choose an option under Model Presets. The transcriber, model, voice, and displayed totals update to match the preset.
Verify the configuration
The option you chose is highlighted under Model Presets. Confirm the component panels and displayed totals match the preset or your custom configuration.