Performance metrics and methodology reference

See how Vapi sources and calculates latency, cost, and quality metrics.

Vapi measures latency from production calls. Cost metrics are calculated estimates. Quality metrics come from Vapi or third-party benchmarks.

Source summary

MetricComponentWhat it measuresSource
LatencyTranscriber, model, voiceTypical response timeMedian from Vapi production calls
CostTranscriber, model, voiceEstimated cost per minuteCalculated estimate based on typical usage
Word error rate (WER)TranscriberThe percentage of words the transcriber gets wrongDaily benchmark
IntelligenceModelReasoning and task capabilityArtificial Analysis benchmark
HumannessVoiceHow human the voice soundsVapi Humanness Index

Latency

Vapi reports the median response time for each component based on production calls.

ComponentLatency measurement
TranscriberTime to convert speech to text
ModelTime to first token
VoiceTime to first audio from the text-to-speech provider

The displayed total is simply the sum of the three component medians, so treat it as a rough estimate rather than an exact figure.

The total excludes endpointing and transport time.

See how latency works for help understanding latency and why more capable models tend to be slower.

Cost

In order to estimate cost, we came up with a number of assumptions based on actual Vapi data. However, costs vary greatly depending on actual call performance and agent behavior. The displayed cost is an estimated rate per minute, use it to compare models rather than predict your exact bill.

For contracted customers, certain pricing may vary depending on your agreement. Please contact your account team for the most accurate estimate.

ComponentEstimation basis
TranscriberAudio minutes for the caller and assistant
ModelPrompt size, tool definitions, prompt caching, and provider rates
VoiceCharacters spoken based on typical speaking volume

LLM cost depends on the assistant’s configuration. Vapi uses the assistant’s prompt and tool definitions, so two assistants on the same model can show different estimates. The calculation also uses each provider’s cached-input rate when prompt caching is supported.

The estimate assumes a moderate cache-hit rate, so actual costs for large prompts are often lower. It excludes growing conversation history, so actual costs for long calls are often higher.

See how cost works for the formulas and assumptions behind each estimate.

Quality

Each component has a quality metric that helps you compare performance.

Transcriber word error rate (WER)

Word error rate (WER) measures the percentage of words a transcriber gets wrong. Lower values are better. A WER of 5% means about 1 in 20 words is transcribed incorrectly.

Vapi sources WER from Daily, a third-party speech-to-text benchmark. This provides a consistent comparison across providers instead of relying on self-reported accuracy.

Intelligence metric for models

The Intelligence metric measures a model’s reasoning and task capability. Vapi sources the score from the Artificial Analysis Intelligence Index. This third-party benchmark scores models from 0 to 100 across reasoning and knowledge tasks.

Voice agents run LLMs with reasoning turned off to keep latency low. The displayed scores reflect reasoning-off performance. The smartest voice agents run at around 20-30 intelligence score, be sure to compare models to each other for a relative benchmark.

Humanness metric for voices

The Humanness metric measures how natural a voice sounds. Vapi measures it with the first-party Vapi Humanness Index. The index uses a blind listening test to score voices from 1 to 100. A higher score means the voice is harder to distinguish from a human.

Data refresh schedule

Vapi refreshes performance metrics weekly as models change and more call data becomes available. The metrics reflect each model at its last update rather than a continuous live feed.