Implement the custom endpoint

Handle Vapi's retrieval request and return documents or a direct response from your own server

Vapi sends your server a request whenever the assistant needs information from the knowledge base. Your endpoint decides what comes back: a set of documents for the model to read, or a finished answer to speak.

Implementing the Custom Endpoint

Your custom knowledge base server must handle POST requests at the configured URL and return structured responses.

Request Structure

Vapi will send requests to your endpoint with the following structure:

Request Format
{
"message": {
"type": "knowledge-base-request",
"messages": [
{
"role": "user",
"content": "What is your return policy?"
},
{
"role": "assistant",
"content": "I'll help you with information about our return policy."
},
{
"role": "user",
"content": "How long do I have to return items?"
}
]
// Additional metadata fields about the call or chat will be included here
}
}

Response Options

Your endpoint can respond in two ways:

Option 1: Return Documents for AI Processing

Return an array of relevant documents that the AI will use to formulate a response:

Document Response
{
"documents": [
{
"content": "Our return policy allows customers to return items within 30 days of purchase for a full refund. Items must be in original condition with tags attached.",
"similarity": 0.92,
"uuid": "doc-return-policy-1" // optional
},
{
"content": "Extended return periods apply during holiday seasons - customers have up to 60 days to return items purchased between November 1st and December 31st.",
"similarity": 0.78,
"uuid": "doc-return-policy-holiday" // optional
}
]
}

Option 2: Return Direct Response

Return a complete response that the assistant will speak directly:

Direct Response
{
"message": {
"role": "assistant",
"content": "You have 30 days to return items for a full refund. Items must be in original condition with tags attached. During the holiday season (November 1st to December 31st), you get an extended 60-day return period."
}
}

Implementation Examples

Here are complete server implementations in different languages:

import express from 'express';
import crypto from 'crypto';
const app = express();
app.use(express.json());
// Your knowledge base data (replace with actual database/vector store)
const documents = [
{
id: "return-policy-1",
content: "Our return policy allows customers to return items within 30 days of purchase for a full refund. Items must be in original condition with tags attached.",
category: "returns"
},
{
id: "shipping-info-1",
content: "We offer free shipping on orders over $50. Standard shipping takes 3-5 business days.",
category: "shipping"
}
];
app.post('/kb/search', (req, res) => {
try {
// Verify webhook secret (recommended)
const signature = req.headers['x-vapi-signature'];
const secret = process.env.VAPI_WEBHOOK_SECRET;
if (signature && secret) {
const expectedSignature = crypto
.createHmac('sha256', secret)
.update(JSON.stringify(req.body))
.digest('hex');
if (signature !== `sha256=${expectedSignature}`) {
return res.status(401).json({ error: 'Invalid signature' });
}
}
const { message } = req.body;
if (message.type !== 'knowledge-base-request') {
return res.status(400).json({ error: 'Invalid request type' });
}
// Get the latest user message
const userMessages = message.messages.filter(msg => msg.role === 'user');
const latestQuery = userMessages[userMessages.length - 1]?.content || '';
// Simple keyword-based search (replace with vector search)
const relevantDocs = documents
.map(doc => ({
...doc,
similarity: calculateSimilarity(latestQuery, doc.content)
}))
.filter(doc => doc.similarity > 0.1)
.sort((a, b) => b.similarity - a.similarity)
.slice(0, 3);
// Return documents for AI processing
res.json({
documents: relevantDocs.map(doc => ({
content: doc.content,
similarity: doc.similarity,
uuid: doc.id
}))
});
} catch (error) {
console.error('Knowledge base search error:', error);
res.status(500).json({ error: 'Internal server error' });
}
});
function calculateSimilarity(query: string, content: string): number {
// Simple similarity calculation (replace with proper vector similarity)
const queryWords = query.toLowerCase().split(' ');
const contentWords = content.toLowerCase().split(' ');
const matches = queryWords.filter(word =>
contentWords.some(cWord => cWord.includes(word))
).length;
return matches / queryWords.length;
}
app.listen(3000, () => {
console.log('Custom Knowledge Base server running on port 3000');
});

Advanced Implementation Patterns

Vector Database Integration

For production use, integrate with a proper vector database:

import { PineconeClient } from '@pinecone-database/pinecone';
import OpenAI from 'openai';
const pinecone = new PineconeClient();
const openai = new OpenAI();
app.post('/kb/search', async (req, res) => {
try {
const { message } = req.body;
const latestQuery = getLatestUserMessage(message);
// Generate embedding for the query
const embedding = await openai.embeddings.create({
model: 'text-embedding-ada-002',
input: latestQuery
});
// Search vector database
const index = pinecone.Index('knowledge-base');
const searchResults = await index.query({
vector: embedding.data[0].embedding,
topK: 5,
includeMetadata: true
});
// Format response
const documents = searchResults.matches.map(match => ({
content: match.metadata.content,
similarity: match.score,
uuid: match.id
}));
res.json({ documents });
} catch (error) {
console.error('Vector search error:', error);
res.status(500).json({ error: 'Search failed' });
}
});

Security and Best Practices

Performance Optimization

Response time is critical: Your endpoint should respond in milliseconds (ideally under ~50ms) for optimal user experience. While Vapi allows up to 10 seconds timeout, slower responses will significantly affect your assistant’s conversational flow and response quality.

Cache frequently requested documents and implement request timeouts to ensure fast response times. Consider using in-memory caches, CDNs, or pre-computed embeddings for faster retrieval.

Error Handling

Always handle errors gracefully and return appropriate HTTP status codes:

app.post('/kb/search', async (req, res) => {
try {
// Your search logic here
} catch (error) {
console.error('Search error:', error);
// Return empty documents rather than failing
res.json({
documents: [],
error: "Search temporarily unavailable"
});
}
});