AssemblyAI
Use AssemblyAI Universal-Streaming for real-time speech-to-text in Vapi.
Universal-Streaming
Universal-Streaming is AssemblyAI’s purpose-built speech-to-text model that delivers ultra-fast, immutable transcripts in ~300ms with intelligent endpointing and superior accuracy for voice agents. It eliminates common pain points like misheard account numbers, awkward pauses, and premature cutoffs, enabling more natural and successful voice interactions.
Speech-to-text
Set transcriber.provider to assembly-ai. AssemblyAI’s Universal-Streaming model is selected automatically, so no model field is required.
For additional API configuration options, review the AssemblyAITranscriber fields in the Create Assistant API reference.
Configure in the dashboard
This guide details how to setup AssemblyAI as a transcriber for your assistant.
Supported languages
Vapi supports AssemblyAI’s English and multilingual Universal-Streaming modes.
Keyterms prompting is not supported with the multilingual model.


