Virbe Documentation

Use a Custom LLM Endpoint

Route your virtual being's conversations through your own LLM backend using a Custom Endpoint conversational engine.

Sometimes the built-in Virbe AI models aren't enough – you may have your own fine-tuned model, a proprietary RAG system, or a backend that orchestrates multiple AI services. This tutorial shows how to connect a Virbe virtual being to a custom HTTP endpoint that handles conversation logic.

What you'll build: A virtual being that sends user messages to your own backend API, receives a response, and speaks it back to the user.

Time: approximately 30 minutes (plus time to build your backend)

Prerequisites:

  • A running HTTP endpoint that accepts a conversation payload and returns a response text
  • A Virbe account with a profile and conversation pipeline

How Custom Endpoint Works

The Custom Endpoint conversational engine sends an HTTP POST request to your API for each user turn. Your API receives the user's message, processes it (using whatever logic you choose – an LLM, a rules engine, a RAG pipeline), and returns a response text. Virbe then delivers that text as the virtual being's response.

This is the most flexible integration option – your endpoint can:

  • Call any LLM (including models not natively supported in Virbe)
  • Apply custom RAG retrieval from your own vector database
  • Look up data in internal systems before generating a response
  • Apply business-specific rules, guardrails, or filters
  • Maintain custom state or context per user

Step 1 – Build Your Endpoint

Your endpoint must accept an HTTP POST with the Virbe conversation payload and return response actions – either as a Server-Sent Events stream (/stream) or as a JSON array (/batch). The full request and response contract, including all action types, is documented on the Custom Endpoint engine page – treat that page as the source of truth for the payload shapes.

Start from the official sample

Rather than building from scratch, start from the official sample repository: github.com/VirbeHQ/virbe-api-integration. It is a TypeScript monorepo (Fastify) that implements a working Custom Endpoint backend and includes the Virbe API types as Zod schemas (packages/virbe-dtos), so your request and response payloads are validated against the real contract.

git clone https://github.com/VirbeHQ/virbe-api-integration.git
cd virbe-api-integration
pnpm install
pnpm dev

The sample API runs at http://localhost:3000; replace the response logic with your own LLM call, RAG retrieval, or business rules.

Deploy your endpoint to a publicly accessible HTTPS URL. Virbe cannot reach localhost – your endpoint must be accessible from the internet or your Virbe deployment's network.


Step 2 – Configure a Custom Endpoint Engine in Virbe

  1. Go to Configure → Conversational Engines
  2. Click + Add engine
  3. Name it (e.g., "My LLM Backend")
  4. Select Custom Endpoint as the provider
  5. Click Add
  6. Enter your Endpoint URL – the base URL of your server (e.g., https://api.yourcompany.com/conversation/message); Virbe automatically appends /stream or /batch depending on the response mode
  7. If your endpoint requires authentication, set the optional Bearer Token – it is sent in the Authorization header of every request
  8. Click Save

Step 3 – Add a Call Conv AI Node to Your Pipeline

  1. Open the Conversation Editor
  2. In the On User Input handler (or the relevant flow), add a Call Conv AI node
  3. In the node configuration, select the custom engine you just created
  4. Connect the Call Conv AI node to a Text node to deliver the response, or to a Route by Signal/If-Else Router if you want to perform additional routing based on the response

A typical simple flow:

Flow Entrypoint
    │
    └── On User Input
              │
              └── Call Conv AI (My LLM Backend)
                        │
                        └── Text ({{var.convAiResponse}})

The Call Conv AI node stores the engine's response text in {{var.convAiResponse}} by default, which you can then pass to a Text node.


Step 4 – Test the Integration

  1. Open the Test conversation panel
  2. Enable Show system messages
  3. Send a test message and verify it reaches your backend (check your server logs)
  4. Verify the response is returned correctly and spoken by the virtual being
  5. Test edge cases: empty responses, slow responses, error responses

Step 5 – Secure Your Endpoint

For production, ensure your endpoint is secure:

  • Use HTTPS – Virbe will not connect to plain HTTP endpoints
  • Add authentication – require an API key or shared secret in request headers. Virbe's Custom Endpoint allows you to configure custom headers for each request.
  • Validate the payload – check that incoming requests are actually from Virbe (validate the profile ID or a shared secret)
  • Rate limit your endpoint to prevent abuse

Extending This Pattern

Streaming responses: The Custom Endpoint supports streaming natively – implement the /stream variant and emit each action as a Server-Sent Event, so the virtual being can start speaking before your LLM finishes generating. See the stream mode documentation for the event format.

Session memory: Use the conversationId field from the request payload to maintain per-conversation context in your backend (conversation history, user preferences, gathered information).

Passing Virbe variables: If you need to pass data gathered earlier in the pipeline (e.g., a product name from a Find Records node), use the Call Webhook node instead of Call Conv AI – it gives you full control over the request body and allows variable interpolation.

On this page