What's New: Week of August 10, 2026
-
Simulations: You can now build simulation suites and run them against your assistants or squads to test and evaluate their behavior, with detailed results for each run.
-
Unified Provider and Model Picker: The assistant editor now combines provider and model into a single picker for your LLM, voice, and transcriber.
-
Chat and Session Log Export: You can now export selected rows from chat and session logs and use a bulk-actions bar, matching the call logs experience.
-
SIP Transfer Improvements: Blind call transfers now support a configurable
fallbackPlan, SIP verb, and dial timeout, and cold-transfer outcomes now appear in call logs. -
Composer Improvements: Composer is now more capable, reliable, and easier to use when building voice agents.
- Upload files for it to reference, read current Vapi documentation directly, and follow new Vapi-built skills for building assistants, choosing models, managing tools, configuring phone numbers, and debugging calls.
- Longer, multi-step tasks are far more reliable, with mid-task error recovery and progress that survives page refreshes and thread switches.
- Connect Google Calendar from the conversation, clickable resource mentions with copyable IDs, and cleaner, collapsible activity timelines.
What's New: Week of August 3, 2026
-
Conditional Structured Outputs: You can now add conditions to a structured output so it only generates when the condition is met, and any skipped outputs are surfaced in the assistant preview, call logs, and sessions.
-
Deepgram Aura-2 German Voices: New German voices are now available for Deepgram’s Aura-2 voice.
What's New: Week of July 27, 2026
New Models
-
Anthropic Claude Sonnet 5: Anthropic’s Claude Sonnet 5 is now available as a model for assistants.
-
OpenAI GPT-5.5 “Instant”: OpenAI’s GPT-5.5 (
gpt-5.5andchat-latest) is now available as a model for assistants. -
OpenAI GPT-5.6 Models: OpenAI’s GPT-5.6 models (
sol,terra, andluna) are now available for assistants. -
Inworld TTS-2: Inworld’s TTS-2 voice is now available for assistants.
What's New: Week of July 20, 2026
-
Model Intelligence: You can now set your assistant’s transcriber, llm, and voice models in one click with a Model Preset (choose between Balanced, High Intelligence, Ultra Fast, or Cost Saver), and see the latency, cost, and quality metrics for your chosen models so you can compare options and optimize with data.
-
Recording Download URLs in the End of Call Report: The end of call report now includes short-lived presigned download URLs for your call recordings and logs, so you can download them directly.
-
End of Call Reports with Zero Data Retention: End of call reports are now reliably delivered for transient assistants running under Zero Data Retention.
What's New: Week of July 13, 2026
-
Playback Speed for Recordings: Call recording players now include a playback-speed control.
-
Download Every Recording Type: You can now download every recording a call produced: mono, stereo, separate assistant and customer tracks, video, and packet capture.
-
VAD Transitions in Call Logs: The call log now surfaces voice-activity-detection transitions with a per-phase latency breakdown.
-
HIPAA Compliance: xAI is now HIPAA-compliant across its model, voice, and transcriber.
What's New: Week of June 29, 2026
-
OpenAI Realtime v2: OpenAI’s latest Realtime v2 model is now available for assistants.
-
AI-Generated Tool Failure and Completion Messages: You can now set
role: 'system'on request-failed tool messages, and the dashboard has a new UI for configuring AI-generated messages when tools fail or complete. This gives assistants more natural responses when tools hit errors. -
Call Logs Improvements:
- Significant latency improvements across both the dashboard and the
/callsAPI endpoints. - The call detail flyout now shows which assistant or squad handled the call, includes a link to the assistant or squad, and indicates which phone number was used.
- In squad or handoff calls, transcript messages now show which assistant said what, making it easier to trace conversation flow.
- Significant latency improvements across both the dashboard and the
-
MCP Child Tools in Dashboard: When you connect an MCP server, the dashboard tool form now lists all child tools it discovers, so you can see exactly what capabilities your MCP server exposes.
What's New: Week of June 22, 2026
-
Gemini 3.5 Flash and 3.1 Flash-Lite: Google’s Gemini 3.5 Flash and 3.1 Flash-Lite models are now available for assistants.
-
Soniox stt-rt-v5: A new
stt-rt-v5real-time speech-to-text model is available for the Soniox transcriber. -
Discord Login Removed: Discord is no longer offered as a login or signup option.
-
Variable Values in Handoff Webhooks: After an assistant handoff, webhook payloads now include the assistant’s configured
variableValues.
What's New: Week of June 15, 2026
-
Claude Haiku (Global): Now available as a model through Amazon Bedrock for assistants.
-
Pronunciation Dictionary Management in Voice Config: Pronunciation dictionaries configured via the API can now be viewed and managed directly in an assistant’s voice settings in the dashboard.
- Stale dictionary references are flagged when a voice changes to a model that cannot apply them.
-
Rotating Tool Messages: You can now configure multiple message variants for a tool, and the assistant picks one at random so longer calls feel less repetitive.
-
Dynamic Variables in Test Calls: A dialog lets you set dynamic variable values before starting a test call from the dashboard.
-
Concurrency and Rate Limits in Organization Settings: Your call concurrency cap and API request rate limit now appear as read-only fields in Organization Settings.
-
Fixes and Improvements:
- Cartesia voice overrides in squads now apply correctly instead of falling back to a hard-coded default.
- Duplicate tools sharing the same
function.nameare de-duplicated during model streaming, preventing duplicate tool calls. - The call concurrency chart in analytics now renders correctly.
What's New: Week of June 1, 2026
-
xAI Speech-to-Text and Text-to-Speech: xAI is now available as a transcriber (STT) and voice (TTS) provider for assistants.
-
Upgraded Vapi Voices: A new text-to-speech model powering Vapi Voices makes them sound more authentic, human, and consistent — at ~50% lower cost.
- Existing deployments don’t change automatically — opt in by setting
version: 2on the voice configuration via the API or Dashboard. See Vapi Voices for supported voices and audio samples.
- Existing deployments don’t change automatically — opt in by setting
-
Pronunciation Dictionaries in the Dashboard: The Assistants view now supports creating new pronunciation dictionaries directly from the dashboard.
-
Phone Number Fixes: Improvements to phone number creation and listing.
- The Phone Numbers list no longer breaks on rows with no number or SIP URI.
- Creating a phone number now validates its Vapi identifier up front.
What's New: Week of May 25, 2026
- Dashboard Performance: Front-end infrastructure improvements for faster page loads and a snappier feel across the dashboard.
What's New: Week of May 18, 2026
-
Responsive UI Polish: A round of UI adjustments to make the app work better across viewport sizes.
-
New Composer-Based Onboarding Flow: A new onboarding experience built on top of the Assistant Builder and powered by Composer is rolling out to select users as part of a phased release.
What's New: Week of May 11, 2026
- New Assistant Builder Experience: An updated, streamlined assistant configuration experience, now available to all users.
- UI optimizations for the Phone Numbers page were also made to align its look and interactions with the new experience.
What's New: Week of May 4, 2026
- Soniox — General Availability: The Soniox transcriber is now rolled out to all customers. Configure it on any assistant via
assistant.transcriber(provider:soniox) for low-latency, multilingual real-time speech-to-text.
What's New: Week of April 27, 2026
- Deepgram Flux — Multilingual Support: Full support for Deepgram’s multilingual Flux model. Multilingual agents can now leverage the same smart turn-taking that powers the English Flux transcriber, making cross-lingual conversations feel more fluid and natural.
What's New: Week of April 20, 2026
-
Logs UX Refresh: New filter layout plus a round of UX improvements — improved date picker, active row is clearly highlighted across all log views when the flyout is opened, log tables are fully keyboard-accessible, sortable
costanddurationcolumns, pagination, and more. -
Squads
contextEngineeringPlanHandoff Type —previousAssistantMessages: Forwards only the conversation history from before the current assistant’s session. The current assistant’s own messages and tool calls are excluded entirely from the handoff payload. See the updated handoff context configuration docs. -
assistant.speechStartedEvent — Live Captions & Word-Level Timing (GA): A new opt-in message fires as the assistant begins speaking each segment, carrying the full turn text,turn,source(model/force-say/custom-voice), and optional timing:- Per-word alignment on ElevenLabs
- Cursor-based word-progress on Minimax (set
voice.subtitleType: "word", with correct CJK handling) - Text-only fallback on all other providers
Subscribe by adding
"assistant.speechStarted"to your assistant’sclientMessagesand/orserverMessages— now GA with no feature flag. Use it for live captions, karaoke-style highlighting, or any UI that needs to stay in sync with assistant audio. Fully backward-compatible; no existing messages changed. -
Autofallbacks on Transcribers: Let Vapi pick the best transcriber to fall back to if your primary one fails — even mid-call. Opt in by setting
assistant.transcriber.fallbackPlan.autoFallback.enabledtotrue. See the updated transcriber fallback plan docs.
What's New: Week of April 13, 2026
-
Monitoring — GA: Automated call quality monitoring is now generally available. Detect issues with trigger-based rules, get alerts when something goes wrong, and surface resolution suggestions — all from the dashboard.
What's New: October 2025 – March 2026
Here’s a summary of major items shipped from October 2025 through March 2026.
Breaking Changes & API Cleanup
-
Legacy Endpoint Removal: The following deprecated endpoints have been removed as part of our API modernization effort:
/logs- Use call artifacts and monitoring instead/workflow/{id}- Access workflows through the main workflow endpoints/test-suiteand related paths - Replaced by the new evaluation system/knowledge-baseand related paths - Integrated into model configurations
-
Knowledge Base Architecture Change: The
knowledgeBaseIdproperty has been removed from all model configurations. This affects:XaiModel,GroqModel,GoogleModelOpenAIModel,AnthropicModel,CustomLLMModel- All other model provider configurations
-
Transcriber Property Deprecation:
AssemblyAITranscriber.wordFinalizationMaxWaitTimeandFallbackAssemblyAITranscriber.wordFinalizationMaxWaitTimeare now deprecated:- Use smart endpointing plans for better speech timing control
- More precise conversation flow management
- Enhanced end-of-turn detection capabilities
-
Schema Path Cleanup: Removed numerous unused schema paths from model configurations to simplify the API structure and improve performance. This cleanup affects internal schema references but doesn’t impact your existing integrations.
-
New v2 API: We are introducing a new API version v2. These changes are part of our ongoing effort to:
- Simplify the API structure for better developer experience
- Remove redundant and deprecated functionality
- Complete the transition to new evaluation and compliance systems
- Improve API performance and maintainability
Evaluation Execution & Results Processing
-
Evaluation Execution Engine: Run comprehensive assistant evaluations with
EvalRunandCreateEvalRunDTO. Execute your mock conversations against live assistants and squads to validate performance and behavior in controlled environments. -
Multiple Evaluation Models: Choose from various AI models for LLM-as-a-judge evaluation:
EvalOpenAIModel: GPT models including GPT-4.1, o1-mini, o3, and regional variantsEvalAnthropicModel: Claude models with optional thinking features for complex evaluationsEvalGoogleModel: Gemini models from 1.0 Pro to 2.5 Pro for diverse evaluation needsEvalGroqModel: High-speed inference models including Llama and custom optionsEvalCustomModel: Your own evaluation models with custom endpoints
-
Evaluation Results: Comprehensive result tracking with
EvalRunResult:status: Pass/fail evaluation outcomesmessages: Complete conversation transcript from the evaluationstartedAtandendedAt: Precise timing information for performance analysis
-
Target Flexibility: Run evaluations against different targets:
EvalRunTargetAssistant: Test individual assistants with optional overridesEvalRunTargetSquad: Evaluate entire squad performance and coordination
-
Evaluation Status Tracking: Monitor evaluation progress with detailed status information:
running: Evaluation in progressended: Evaluation completedqueued: Evaluation waiting to start- Detailed
endedReasonincluding success, error, timeout, and cancellation states
-
Judge Configuration: Optimize evaluation accuracy with model-specific settings:
maxTokens: Recommended 50-10000 tokens (1 token for simple pass/fail responses)temperature: 0-0.3 recommended for LLM-as-a-judge to reduce hallucinations
Voicemail Detection & Handling Improvements
-
Enhanced Beep Detection: Improve voicemail detection accuracy with
CreateVoicemailToolDTO.beepDetectionEnabledspecifically for Twilio-based calls. This feature detects the characteristic beep sound that indicates voicemail recording has started. -
Workflow Voicemail Integration: Configure comprehensive voicemail handling in workflows with enhanced message and detection capabilities:
Workflow.voicemailMessage: Custom messages for voicemail scenarios (up to 1000 characters)Workflow.voicemailDetection: Configurable detection methods for different providers
-
Assistant Voicemail Enhancement: Improved voicemail handling in assistant configurations with
Assistant.voicemailMessageandAssistant.voicemailDetectionfor consistent behavior across all conversation types. -
Multiple Detection Methods: Choose from various voicemail detection providers:
- Google:
GoogleVoicemailDetectionPlanfor AI-powered detection - OpenAI:
OpenAIVoicemailDetectionPlanfor intelligent voicemail recognition - Twilio:
TwilioVoicemailDetectionPlanfor carrier-level detection - Vapi:
VapiVoicemailDetectionPlanfor integrated detection
- Google:
-
Beep Detection for Call Flows: The new beep detection capability works specifically with Twilio transport, providing reliable voicemail identification when traditional detection methods may not be sufficient.
-
Voicemail Tool Configuration: Enhanced tool rejection and messaging capabilities ensure appropriate handling when voicemail is detected, with configurable responses based on your business requirements.