Gladia

Use Gladia for real-time speech-to-text in your Vapi voice agent.

Speech-to-text

Set transcriber.provider to gladia. Vapi supports Gladia’s real-time transcription models, automatic or manual language selection, partial transcripts, audio enhancement, and custom vocabulary controls.

$curl -X PATCH "https://api.vapi.ai/assistant/ASSISTANT_ID" \
> -H "Authorization: Bearer $VAPI_API_KEY" \
> -H "Content-Type: application/json" \
> -d '{
> "transcriber": {
> "provider": "gladia",
> "model": "solaria-1",
> "languageBehaviour": "manual",
> "language": "en",
> "region": "us-west",
> "receivePartialTranscripts": true
> }
> }'

For additional API configuration options, review the GladiaTranscriber fields in the Create Assistant API reference.

Configuration

FieldRequiredDescription
providerYesSet to gladia.
modelNoGladia model to use. Defaults to fast.
languageBehaviourNoHow Gladia selects languages. Defaults to automatic single language.
languageFor manual single-language transcriptionOne supported language code.
languagesFor manual multilingual transcriptionAn array of supported language codes.
regionNoProcessing region: us-west or eu-west.
receivePartialTranscriptsNoSet to true to receive partial transcripts.

Supported models

These are the Gladia models currently supported by Vapi.

ModelModel ID
Solariasolaria-1
Accurateaccurate
Fastfast

Language selection

Language behaviorValueHow to configure languages
Automatic, one languageautomatic single languageGladia detects one language for the stream. This is the Vapi default.
Automatic, multiple languagesautomatic multiple languagesGladia can detect multiple languages in the stream.
ManualmanualSet language for one language or languages for multiple languages.

The following language codes are currently accepted by Vapi for Gladia.

LanguageLanguage code
Afrikaansaf
Albaniansq
Amharicam
Arabicar
Armenianhy
Assameseas
Azerbaijaniaz
Bashkirba
Basqueeu
Belarusianbe
Bengalibn
Bosnianbs
Bretonbr
Bulgarianbg
Catalanca
Chinesezh
Croatianhr
Czechcs
Danishda
Dutchnl
Englishen
Estonianet
Faroesefo
Finnishfi
Frenchfr
Galiciangl
Georgianka
Germande
Greekel
Gujaratigu
Haitian Creoleht
Hausaha
Hawaiianhaw
Hebrewhe
Hindihi
Hungarianhu
Icelandicis
Indonesianid
Italianit
Japaneseja
Javanesejv
Kannadakn
Kazakhkk
Khmerkm
Koreanko
Laolo
Latinla
Latvianlv
Lingalaln
Lithuanianlt
Luxembourgishlb
Macedonianmk
Malagasymg
Malayms
Malayalamml
Maltesemt
Maorimi
Marathimr
Mongolianmn
Myanmar (Burmese)my
Nepaline
Norwegianno
Norwegian Nynorsknn
Occitanoc
Pashtops
Persianfa
Polishpl
Portuguesept
Punjabipa
Romanianro
Russianru
Sanskritsa
Serbiansr
Shonasn
Sindhisd
Sinhalasi
Slovaksk
Sloveniansl
Somaliso
Spanishes
Sundanesesu
Swahilisw
Swedishsv
Tagalogtl
Tajiktg
Tamilta
Tatartt
Telugute
Thaith
Tibetanbo
Turkishtr
Turkmentk
Ukrainianuk
Urduur
Uzbekuz
Vietnamesevi
Welshcy
Yiddishyi
Yorubayo

Transcription options

FieldType or rangeDescription
transcriptionHintString, up to 600 charactersSupplies context-specific words, names, or technical terms.
prosodyBooleanIncludes detected non-speech events such as laughter or music. Defaults to false.
audioEnhancerBooleanPreprocesses audio to improve accuracy at the cost of additional latency. Defaults to false.
confidenceThreshold0 to 1Discards transcripts below the threshold. Defaults to 0.4.
endpointing0.01 to 10 secondsTime to wait before treating speech as ended.
speechThreshold0 to 1Adjusts speech-detection sensitivity.
customVocabularyEnabledBooleanEnables Gladia custom vocabulary.
customVocabularyConfigObjectSupplies vocabulary items and an optional default intensity.