Speech & AI
Configure the speech-to-text, text-to-speech, and AI model engines that power your virtual being's voice and intelligence.
The Three Engine Types
Together, these three engines are what let a virtual being hear, think, and speak. Configure at least one of each type and mark it default before testing a voice conversation end to end.
What's in this section
Speech-to-Text
Convert user voice input into text, using providers like Azure Cognitive, Google Speech, NVIDIA Riva, or ElevenLabs.
AI Models
Configure LLMs for response generation (Tool Agent, LLM Response nodes) and embedding models for Knowledge Base retrieval.
Text-to-Speech
Convert text responses into spoken audio, using Azure Text-to-Speech or ElevenLabs.
How the pieces fit together
A voice interaction flows through all three engine types in sequence: Speech-to-Text transcribes what the user says, an AI Model processes that text and generates a response (optionally drawing on the Knowledge Base via RAG), and Text-to-Speech converts the response back into audio using the voice configured in the Persona.
Each engine type supports multiple providers, and you can configure more than one engine per type – useful for testing alternatives or supporting different languages. The engine marked default is used whenever a node or Persona setting doesn't specify one explicitly.
Configuring engines here makes them available for selection elsewhere – TTS voices are assigned per language in Personas, and AI Models are selected per node in the Conversation Editor. Set up your engines first, then wire them into Personas and Flows.