Speech Capture: VAD, Hold to Talk, Tap to Speak
Configure how the Virbe Kiosk detects user speech – always-on voice activity detection, hold-to-talk, or tap-to-speak activation.
The Virbe Kiosk can detect user speech in three different modes:
| Mode | Activation | Speech end detection |
|---|---|---|
| Voice Activity Detection (VAD) | None – always listening | Automatic (VAD) |
| Hold to Talk | User presses and holds a button while speaking | Button release |
| Tap to Speak | User taps once to start listening | Automatic (VAD) – waits for speech start and end |
Each approach has trade-offs depending on the environment.
Voice Activity Detection (VAD)
VAD is the default mode. The kiosk continuously monitors the audio input and starts processing speech as soon as it detects voice activity above a threshold.
Advantages
- Hands-free – the user simply speaks; no button is needed
- Natural interaction – feels more like talking to a person
- No additional hardware required
Disadvantages
- Sensitive to background noise – in loud environments, VAD can trigger on background speech, music, or other sounds
- False starts – the kiosk may begin processing before the user has finished their thought if there is a brief pause
When to use VAD
VAD works well in:
- Quiet environments (offices, hotel lobbies, museum halls)
- Environments with consistent, predictable background noise levels
- Scenarios where the user's proximity can be detected (face detection or proximity sensor) to activate listening only when someone is in front of the kiosk
Tap to Speak
In Tap to Speak mode, the user taps once (on-screen microphone button or touch area) to activate listening. From that point the kiosk behaves like VAD: it waits for the user's speech to start, detects when it ends automatically, and then processes the utterance. The user does not need to hold anything while speaking.
Advantages
- Noise-resistant activation – the kiosk is not listening until the user explicitly taps, eliminating false triggers from passers-by, music, or announcements
- Hands-free while speaking – unlike Hold to Talk, the user speaks naturally after a single tap; no need to keep a button pressed
- Clear interaction model – the tap gives users an unambiguous "now it's listening" moment
- More accessible than Hold to Talk – a single tap is easier than press-and-hold for users with limited motor control
Disadvantages
- Speech end detection is still VAD-based – in very loud environments, background noise can delay or confuse the end-of-speech detection even though activation was explicit
- Requires one touch interaction – slightly less seamless than pure VAD in quiet spaces
When to use Tap to Speak
Tap to Speak is the practical middle ground for moderately noisy public spaces – retail floors, showrooms, event booths – where always-on VAD produces too many false triggers, but Hold to Talk feels too mechanical. It is often the best default for public kiosk deployments.
Hold to Talk
In Hold to Talk mode (also known as press-to-talk / PTT), the microphone only captures speech while the user presses and holds a button – a physical button or a capacitive touch area. The kiosk listens for as long as the button is held and processes the speech when it is released.
Advantages
- Most noise-resistant – both the start and the end of capture are explicit, so background noise cannot trigger or cut off the interaction
- Clear interaction model – users know exactly when the kiosk is listening
- Suitable for very loud environments – trade shows, airport terminals, food courts
Disadvantages
- May require additional hardware – a physical button, capacitive touch sensor, or footswitch
- Least natural – users must learn to press and hold while speaking
- Accessibility consideration – press-and-hold may be harder for users with limited motor control (consider Tap to Speak instead)
How to configure a physical Hold to Talk button
- Connect a physical button to the kiosk via USB (most USB buttons emulate a HID keyboard key)
- In the Virbe profile → Settings → Capture HID input, enable HID input capture
- Configure the signal name for the button press event
- In the Conversation Editor, handle the button-press signal to start and stop recording:
- On button press → start recording (using a Custom Action or Virbe SDK call)
- On button release → stop recording and submit the audio for STT processing
- The STT result is then processed through the normal pipeline
Physical-button Hold to Talk requires custom signal handling in the Conversation Editor and may require a Custom Endpoint or JavaScript SDK call depending on the kiosk app version. Contact Virbe support for guidance on your specific kiosk version's capabilities.
Choosing a Mode
| Environment | Recommended mode |
|---|---|
| Quiet office, hotel lobby, museum | VAD |
| Retail floor, showroom, event booth | Tap to Speak |
| Trade show hall, airport terminal, food court | Hold to Talk (or Tap to Speak with careful microphone selection) |
Hybrid Approach: Face Detection + Explicit Activation
A practical pattern for noisy environments that still want a smooth first contact:
- Use face detection (
On face-detectedhandler) to activate the kiosk when a user approaches - On detection, display a prompt ("Tap to speak" or "Speak now") and use Tap to Speak, Hold to Talk, or VAD depending on the noise level
- Use the profile's Defocus timeout to return to idle after the user leaves
This pattern reduces false VAD triggers (because the kiosk is only listening when a face is detected) while keeping the initial interaction inviting.
VAD Sensitivity Tuning
Both VAD mode and Tap to Speak rely on voice activity detection (for speech onset and end). If you experience false triggers or missed speech:
- Too sensitive (many false triggers): The kiosk activates on background sounds. Increase the VAD threshold or switch to Tap to Speak / Hold to Talk.
- Too insensitive (missing speech onset): The kiosk does not detect the user starting to speak. Lower the VAD threshold, check microphone gain, or move the microphone closer to the expected speaker position.
VAD threshold settings are part of the STT engine configuration. Refer to your STT provider's documentation (Azure Cognitive Services or Google Speech-to-Text) for specifics on sensitivity adjustment.