> For clean Markdown of any page, append .md to the page URL.
> For a complete documentation index, see https://docs.vapi.ai/llms.txt.
> For AI client integration (Claude Code, Cursor, etc.), connect to the MCP server at https://docs.vapi.ai/_mcp/server.

# Model Presets for transcribers, models, and voices

> Model presets optimize for balanced performance, reasoning, low latency, or low cost. Learn when to use each preset and how to apply or customize its configuration.

**Model Presets** are curated configurations that bundle a transcriber, model, and voice for a specific goal. Use a preset for a dependable setup or as a starting point for further customization.

## How it works

New assistants use **Balanced** as the default. When you start from one of Vapi’s use-case templates, the assistant uses the preset that best fits the template.

Presets are a starting point rather than a lock - you can fine-tune or swap out any component to better fit your use case.

## Presets at a glance

| Preset                | Optimizes for                                      | Best for                                          |
| --------------------- | -------------------------------------------------- | ------------------------------------------------- |
| **Balanced**          | A strong all-round mix of quality, speed, and cost | Most assistants; start here if you're unsure      |
| **High Intelligence** | Reasoning and conversation quality                 | Complex, nuanced, or high-stakes conversations    |
| **Ultra Fast**        | Lowest latency                                     | Fast, responsive, high-volume conversations       |
| **Cost Saver**        | Lowest cost per minute                             | Simple, high-volume calls where cost matters most |

## Choose a preset

Choose the preset that best matches your use case. Start with **Balanced** if you are unsure.

### Balanced

**Balanced** is the default for new assistants and the best starting point for most use cases. It provides a strong mix of quality, responsiveness, and cost.

Choose **Balanced** when any of the following apply.

* You're building a new assistant and aren't sure which preset fits.
* Your use case covers support, scheduling, FAQs, or qualification.
* You want an all-around dependable setup out of the box.

> **Note**
>
> If you later need more reasoning, faster responses, or lower cost, switch to the preset built for that.

### High Intelligence

**High Intelligence** prioritizes reasoning and conversation quality. Use it when the assistant needs to handle nuance, follow multi-step logic, or reason reliably across tools and context.

Choose **High Intelligence** when any of the following apply.

* Conversations are complex, open-ended, or high-stakes.
* The assistant needs to reason through multi-step problems or use tools reliably.
* Accuracy and quality matter more than speed or cost.

> **Note**
>
> More capable models respond a little slower and cost more per minute than **Balanced**. See [how latency works](/assistants/model-intelligence/understanding-latency) for why more capable models take longer to respond.

### Ultra Fast

**Ultra Fast** prioritizes low latency and responsive conversations. Use it when response speed matters more than complex reasoning.

Choose **Ultra Fast** when any of the following apply.

* Responsiveness is the priority and replies should feel fast.
* Calls are high-volume and relatively straightforward.
* The flow is scripted or transactional rather than open-ended.

> **Note**
>
> The fastest models are smaller and less capable, so **Ultra Fast** may not suit complex reasoning or heavy tool use. Choose **Balanced** or **High Intelligence** if difficult tasks produce quality issues.

### Cost Saver

**Cost Saver** prioritizes the lowest cost per minute. Use it for high-volume calls when cost is the main constraint.

Choose **Cost Saver** when any of the following apply.

* Call volume is high and you're optimizing spend.
* Interactions are simple or scripted.
* A small quality trade-off is acceptable in exchange for lower cost.

> **Note**
>
> The lowest-cost models are less capable, so this preset fits simpler interactions best. See [how cost works](/assistants/model-intelligence/understanding-cost) for what drives cost per minute.

## The Customized state

Swapping a model out of any component moves your assistant to **Customized** and the **Performance Metrics** displayed will update to match the new configuration.

## Apply a Model Preset

#### Open the Assistant tab

Open your assistant in the Vapi Dashboard, then open **Assistant**.

#### Choose a preset

Choose an option under **Model Presets**. The transcriber, model, voice, and displayed totals update to match the preset.

#### Customize components if needed

Click the pencil icon on the transcriber, model, or voice panel. Use the dropdown menu in the settings panel to switch out the model or provider. Your assistant moves to **Customized**. **Performance Metrics** update to reflect your choices.

#### Publish your changes

Click **Publish** to apply your changes.

## Verify the configuration

The option you chose is highlighted under **Model Presets**. Confirm the component panels and displayed totals match the preset or your custom configuration.

## Related

#### [Model Intelligence for transcribers, models, and voices](/assistants/model-intelligence/overview)

Learn how presets and performance metrics help you choose components.

#### [Performance metrics and methodology reference](/assistants/model-intelligence/metrics-methodology)

How every latency, cost, and quality metric is sourced.