Performance metrics and methodology reference
Vapi measures latency from production calls. Cost metrics are calculated estimates. Quality metrics come from Vapi or third-party benchmarks.
Source summary
Latency
Vapi reports the median response time for each component based on production calls.
The displayed total is simply the sum of the three component medians, so treat it as a rough estimate rather than an exact figure.
The total excludes endpointing and transport time.
See how latency works for help understanding latency and why more capable models tend to be slower.
Cost
In order to estimate cost, we came up with a number of assumptions based on actual Vapi data. However, costs vary greatly depending on actual call performance and agent behavior. The displayed cost is an estimated rate per minute, use it to compare models rather than predict your exact bill.
For contracted customers, certain pricing may vary depending on your agreement. Please contact your account team for the most accurate estimate.
LLM cost depends on the assistant’s configuration. Vapi uses the assistant’s prompt and tool definitions, so two assistants on the same model can show different estimates. The calculation also uses each provider’s cached-input rate when prompt caching is supported.
The estimate assumes a moderate cache-hit rate, so actual costs for large prompts are often lower. It excludes growing conversation history, so actual costs for long calls are often higher.
See how cost works for the formulas and assumptions behind each estimate.
Quality
Each component has a quality metric that helps you compare performance.
Transcriber word error rate (WER)
Word error rate (WER) measures the percentage of words a transcriber gets wrong. Lower values are better. A WER of 5% means about 1 in 20 words is transcribed incorrectly.
Vapi sources WER from Daily, a third-party speech-to-text benchmark. This provides a consistent comparison across providers instead of relying on self-reported accuracy.
Intelligence metric for models
The Intelligence metric measures a model’s reasoning and task capability. Vapi sources the score from the Artificial Analysis Intelligence Index. This third-party benchmark scores models from 0 to 100 across reasoning and knowledge tasks.
Voice agents run LLMs with reasoning turned off to keep latency low. The displayed scores reflect reasoning-off performance. The smartest voice agents run at around 20-30 intelligence score, be sure to compare models to each other for a relative benchmark.
Humanness metric for voices
The Humanness metric measures how natural a voice sounds. Vapi measures it with the first-party Vapi Humanness Index. The index uses a blind listening test to score voices from 1 to 100. A higher score means the voice is harder to distinguish from a human.
Data refresh schedule
Vapi refreshes performance metrics weekly as models change and more call data becomes available. The metrics reflect each model at its last update rather than a continuous live feed.