OpenAI Realtime

Build voice assistants with OpenAI’s native speech-to-speech models for ultra-low latency conversations

Overview

OpenAI’s Realtime API enables developers to use a native speech-to-speech model. Unlike other Vapi configurations which orchestrate a transcriber, model and voice API to simulate speech-to-speech, OpenAI’s Realtime API natively processes audio in and audio out.

In this guide, you’ll learn to:

  • Choose the right realtime model for your use case
  • Configure voice assistants with realtime capabilities
  • Implement best practices for production deployments
  • Optimize prompts specifically for realtime models

Available models

OpenAI offers three realtime models, each with different capabilities and cost/performance trade-offs:

ModelStatusBest ForKey Features
gpt-realtime-2025-08-28ProductionProduction workloadsProduction Ready
gpt-4o-realtime-preview-2024-12-17PreviewDevelopment & testingBalanced performance/cost
gpt-4o-mini-realtime-preview-2024-12-17PreviewCost-sensitive appsLower latency, reduced cost

Voice options

Realtime models support a specific set of OpenAI voices optimized for speech-to-speech:

Standard Voices

Available across all realtime models:

  • alloy - Neutral and balanced
  • echo - Warm and engaging
  • shimmer - Energetic and expressive
Realtime-Exclusive Voices

Only available with realtime models:

  • marin - Professional and clear
  • cedar - Natural and conversational

The following voices are NOT supported by realtime models: ash, ballad, coral, fable, onyx, and nova.

Configuration

Configure an API Request tool

Use an API Request tool to fetch current weather directly from WeatherAPI. You don’t need a custom server or a tool-calls webhook handler.

Create a WeatherAPI account, get an API key, and replace YOUR_WEATHERAPI_KEY in the tool’s url field below. For the TypeScript and Python examples, set VAPI_API_KEY to your Vapi private API key and run the code on your server.

WeatherAPI authenticates through a URL query parameter. Do not expose the configured URL or your WeatherAPI key in public repositories or browser code. Redact the WeatherAPI key from request logs.

The tool’s body schema defines the location argument the model supplies. For this GET request, Vapi inserts that argument into the URL and sends no HTTP request body. The url_encode filter encodes spaces and other special characters in the city name.

{
"model": {
"provider": "openai",
"model": "gpt-realtime-2025-08-28",
"messages": [
{
"role": "system",
"content": "You are a concise, friendly weather assistant. If the caller has not provided a location, ask for one. If the city is ambiguous, ask for the missing region or country before using getWeather. Call getWeather for each new current-weather request, including a request for another city. Pass the complete location, preserving any region/state and country the caller supplied. Use only the latest successful result for the requested location and report the returned location with the weather. If the returned city, region, or country conflicts with the request, clarify before reporting weather. Differences in spelling or formatting alone are not a location mismatch. If the lookup fails or current-weather data is missing, explain that current weather is unavailable. Do not invent weather or reuse an earlier result after a failed lookup."
}
],
"temperature": 0.7,
"maxTokens": 250,
"tools": [
{
"type": "apiRequest",
"name": "getWeather",
"description": "Get current weather for the caller's complete requested location. Call getWeather for every request for current weather, and use the returned location to confirm that the result matches the request.",
"method": "GET",
"url": "https://api.weatherapi.com/v1/current.json?key=YOUR_WEATHERAPI_KEY&q={{location|url_encode}}",
"timeoutSeconds": 10,
"body": {
"type": "object",
"properties": {
"location": {
"type": "string",
"description": "The complete location requested by the caller, preserving the city and any supplied region/state and country; for example, Springfield, Illinois, USA."
}
},
"required": ["location"]
}
}
]
},
"voice": {
"provider": "openai",
"voiceId": "alloy"
}
}

Test the weather tool

After you create the assistant, test these cases:

  • Ask for the current weather in a city and include its region or country. Confirm that getWeather receives the complete location and the spoken answer matches the returned location and weather.
  • Ask about another city during the same call. Confirm that getWeather runs again and the answer uses the new result.
  • Ask about an ambiguous city. Confirm that the assistant requests the missing region or country before calling getWeather.
  • Trigger a failed lookup after a successful one. Confirm that the assistant explains the failure without guessing or reusing the earlier weather.

Using realtime-exclusive voices

To use the enhanced voices only available with realtime models:

{
"voice": {
"provider": "openai",
"voiceId": "marin" // or "cedar"
}
}

Handling instructions

Unlike traditional OpenAI models, realtime models receive instructions through the session configuration. Vapi automatically converts your system messages to session instructions during WebSocket initialization.

The system message in your model configuration is automatically optimized for realtime processing:

  1. System messages are converted to session instructions
  2. Instructions are sent during WebSocket session initialization
  3. The instructions field supports the same prompting strategies as system messages

Prompting best practices

Realtime models benefit from different prompting techniques than text-based models. These guidelines are based on OpenAI’s official prompting guide.

General tips

  • Iterate relentlessly: Small wording changes can significantly impact behavior
  • Use bullet points over paragraphs: Clear, short bullets outperform long text blocks
  • Guide with examples: The model closely follows sample phrases you provide
  • Be precise: Ambiguity or conflicting instructions degrade performance
  • Control language: Pin output to a target language to prevent unwanted switching
  • Reduce repetition: Add variety rules to avoid robotic phrasing
  • Capitalize for emphasis: Use CAPS for key rules to make them stand out

Prompt structure

Organize your prompts with clear sections for better model comprehension:

# Role & Objective
You are a customer service agent for Acme Corp. Your goal is to resolve issues quickly.
# Personality & Tone
- Friendly, professional, and empathetic
- Speak naturally at a moderate pace
- Keep responses to 2-3 sentences
# Instructions
- Greet callers warmly
- Ask clarifying questions before offering solutions
- Always confirm understanding before proceeding
# Tools
Use the available tools to look up account information and process requests.
# Safety
If a caller becomes aggressive or requests something outside your scope,
politely offer to transfer them to a specialist.

Realtime-specific techniques

Control the model’s speaking pace with explicit instructions:

## Pacing
- Deliver responses at a natural, conversational speed
- Do not rush through information
- Pause briefly between key points

Migration guide

Transitioning from standard STT/TTS to realtime models:

1

Update your model configuration

Change your model to one of the realtime options:

{
"model": {
"provider": "openai",
"model": "gpt-realtime-2025-08-28" // Changed from gpt-4
}
}
2

Verify voice compatibility

Ensure your selected voice is supported (alloy, echo, shimmer, marin, or cedar)

3

Remove transcriber configuration

Realtime models handle speech-to-speech natively, so transcriber settings are not needed

4

Test function calling

Your existing function configurations work unchanged with realtime models

5

Optimize your prompts

Apply realtime-specific prompting techniques for best results

Best practices

Model selection strategy

Best for production workloads requiring:

  • Structured outputs for form filling or data collection
  • Complex function orchestration
  • Highest quality voice interactions
  • Responses API integration

Best for development and testing:

  • Prototyping voice applications
  • Balanced cost/performance during development
  • Testing conversation flows before production

Best for cost-sensitive applications:

  • High-volume voice interactions
  • Simple Q&A or routing scenarios
  • Applications where latency is critical

Performance optimization

  • Temperature settings: Use 0.5-0.7 for consistent yet natural responses
  • Max tokens: Set appropriate limits (200-300) for conversational responses
  • Voice selection: Test different voices to match your brand personality
  • Function design: Keep function schemas simple for faster execution

Error handling

Handle edge cases gracefully:

{
"messages": [{
"role": "system",
"content": "If you don't understand the user, politely ask them to repeat. Never make assumptions about unclear requests."
}]
}

Current limitations

Be aware of these limitations when implementing realtime models:

  • Knowledge Bases are not currently supported with the Realtime API
  • Endpointing and Interruption models are managed by Vapi’s orchestration layer
  • Custom voice cloning is not available for realtime models
  • Some OpenAI voices (ash, ballad, coral, fable, onyx, nova) are incompatible
  • Transcripts may have slight differences from traditional STT output

Additional resources

Next steps

Now that you understand OpenAI Realtime models: