Customer Experience
Tune how end users experience the Virbe Kiosk – avatar appearance and animation, microphone setup, and speech detection mode.
Field Configuration
These settings are best tuned on-site, not in the office – acoustics, lighting, and viewing distance vary enough between installations that defaults rarely hold up once the kiosk is in its final location.
In This Section
Avatars
Select and configure the Metahuman avatar – appearance, camera framing, animations, and focus/defocus behavior.
Microphone Settings
Configure Windows audio settings and STT engine options for reliable speech recognition in your kiosk environment.
Speech Capture: VAD, Hold to Talk, Tap to Speak
Choose between always-on Voice Activity Detection, hold-to-talk, and tap-to-speak speech capture.
Getting the Experience Right
These three areas interact closely in a physical installation. The avatar's camera framing and scene should match your display and enclosure. Microphone settings determine whether speech is transcribed accurately at all – a poorly tuned microphone will undermine even the best-configured avatar and conversation logic. The speech capture mode is largely a function of ambient noise: quiet spaces (offices, museums) suit hands-free VAD, moderately noisy public spaces work well with Tap to Speak, and very loud environments (trade shows, food courts) often need Hold to Talk to avoid false triggers.
Test the full experience on-site before go-live – acoustics, lighting, and foot traffic patterns vary enough between locations that settings tuned in one environment may need adjustment in another.