> For clean Markdown of any page, append .md to the page URL.
> For a complete documentation index, see https://docs.vapi.ai/llms.txt.
> For AI client integration (Claude Code, Cursor, etc.), connect to the MCP server at https://docs.vapi.ai/_mcp/server.

# How voice agent component costs and estimates work

> Understand how Vapi estimates transcriber, model, and voice costs per minute, including prompt caching assumptions and the effects of conversation history.

The cost to run a voice agent mainly depends on the costs of the transcriber, LLM, and voice. Each component uses a different billing unit, so we've made some assumptions in order to provide a way for you to compare estimated cost per minute.

> **Note**
>
> The **Cost** value under **Performance Metrics** estimates costs for model comparison; it is not an exact quote. See [Vapi pricing](https://vapi.ai/pricing) for current rates, or contact your account team for contract-specific rates.

To review costs for a completed call, open its [call log](/observability/logs/call-logs) and check the **Call Cost** tab when available.

## Key concepts

| Term                     | Definition                                                                                                        |
| ------------------------ | ----------------------------------------------------------------------------------------------------------------- |
| Input tokens per minute  | Tokens sent to the model each minute based on your prompt size, tool-definition size, and a typical request rate. |
| Effective input rate     | The per-token input rate blended between standard and cached rates when the model supports prompt caching.        |
| Output tokens per minute | Tokens the model generates each minute based on typical Vapi call data.                                           |

## Component billing units

Each component uses a different billing unit.

| Component   | Billing unit                      |
| ----------- | --------------------------------- |
| Transcriber | Minutes of audio                  |
| Model       | 1 million input and output tokens |
| Voice       | Characters of spoken text         |

### Transcriber cost

Transcription is billed per minute of audio. Vapi transcribes audio from both the caller and the assistant. The estimate therefore includes two audio channels. This cost scales with call duration rather than your prompt or configuration.

### Voice cost

The Voice model is billed per character of spoken text. The estimate uses the number of characters spoken in a typical minute. Cost increases when the assistant speaks more.

### LLM cost

Model cost varies more than transcriber or voice cost because it heavily depends on the assistant's configuration. Vapi calculates the estimate from the assistant's actual prompt and tool definitions, so two assistants using the same LLM model can show different cost estimates.

The estimate uses each provider's input and output token rates.

| Calculation                         | Formula                                        |
| ----------------------------------- | ---------------------------------------------- |
| **Input cost**                      | Input tokens per minute × effective input rate |
| **Output cost**                     | Output tokens per minute × output rate         |
| **Estimated model cost per minute** | Input cost + output cost                       |

#### Input tokens per minute

**Input tokens per minute** is the largest cost factor you can control. Vapi calculates it from your [system prompt](/prompting-guide), [tool definitions](/tools), and a typical number of model requests per minute. Longer prompts and more tools increase the estimate, while leaner configurations reduce it.

Vapi estimates 5 requests per turn. We use number of requests rather than conversational turns, since a single turn can trigger multiple requests, and providers bill each request.

#### Effective input rate

**Effective input rate** accounts for prompt caching. Cached input is stable prompt content that providers reuse across requests. Many providers offer a discounted rate for this input.

When a model supports caching, Vapi blends its standard and cached rates based on an assumed 50% cache hit rate. Otherwise, Vapi uses the standard rate. The calculation uses the actual rates for each model and provider.

#### Output tokens per minute

Vapi estimates 150 **output tokens per minute** based on typical generated speech in production call data. It prices those tokens at the model’s output rate.

## Estimate limitations

Two assumptions can cause actual model cost to differ from the estimate.

| Scenario                             | Typical effect              | Reason                                                       |
| ------------------------------------ | --------------------------- | ------------------------------------------------------------ |
| Large prompts with caching supported | Actual cost is often lower  | Actual cache use is usually higher than the estimate assumes |
| Long calls                           | Actual cost is often higher | Growing conversation history is sent again with each request |

These deliberate simplifications keep estimates comparable across models. Vapi may refine them over time.

## Reduce cost

You can reduce cost in four ways.

* Shorten your system prompt and tool definitions. This often has the largest effect on model cost.
* Use models that support prompt caching for large, stable prompts. Providers charge less for prompt content reused from cache, which lowers model cost.
* Choose a lower-cost model when the use case allows. The [Cost Saver preset](/assistants/model-intelligence/presets#cost-saver) optimizes components for the lowest cost per minute.
* Keep calls focused. Shorter, more focused conversations cost less and tend to provide a better experience.

## Related

#### [Model presets for transcribers, models, and voices](/assistants/model-intelligence/presets)

Compare presets or apply Cost Saver.

#### [Performance metrics and methodology reference](/assistants/model-intelligence/metrics-methodology)

See how Vapi calculates model cost estimates.

#### [Call cost breakdown](/observability/logs/call-logs)

Open a completed call and check its **Call Cost** tab for the per-component breakdown when available.