Virbe Documentation

Microphone Settings

Configure microphone sensitivity, gain, and audio settings in Windows and the Virbe profile for optimal speech recognition in your kiosk environment.

Getting microphone settings right is the difference between a kiosk that works reliably and one that mishears users constantly. This page covers how to configure both the Windows audio settings and the Virbe profile settings to get the best possible speech recognition in your environment.

Windows Audio Settings

Set the Correct Default Microphone

Windows must be configured to use the correct microphone as the default recording device:

  1. Right-click the speaker icon in the taskbar → Sounds (or Open Sound settings)
  2. Select the Recording tab
  3. Right-click your kiosk microphone → Set as Default Device and Set as Default Communication Device
  4. Speak into the microphone and verify the level meter shows input activity

If multiple audio input devices appear (e.g., the display's built-in microphone and a dedicated kiosk microphone), make sure only the intended microphone is set as default. Disable unused microphones to prevent accidental selection.

Microphone Level

Set the microphone level in Windows to capture speech clearly without clipping:

  1. In the Recording devices panel, double-click your microphone → Levels tab
  2. Start with a level of 80–90 and test
  3. Speak at normal conversational volume from the expected user position
  4. The level meter in the Recording devices panel should peak around -12 to -6 dBFS during speech (the meter should show a green to amber range, not constantly red)
  5. If speech is too quiet (meter barely moves), increase the level; if clipping (meter stays red), decrease it

Windows Enhancements

Windows applies audio processing effects to microphone input by default. These can either help or hurt speech recognition depending on the environment:

  1. Double-click your microphone in the Recording devices panel → Enhancements tab
  2. Settings to consider:
    • Noise suppression – generally helpful in noisy environments; can sometimes over-suppress soft speech in quieter environments
    • Echo cancellation – should be enabled to prevent TTS output from being picked up as user speech
    • Acoustic echo cancellation (AEC) – similar to echo cancellation; enable if available
  3. If Windows enhancements are causing problems, try disabling all enhancements and using a microphone with built-in DSP instead

STT Engine Microphone Configuration

The Speech-to-Text (STT) engine used in the Virbe profile determines how audio is processed and transcribed. Configure the STT engine in Configure → Speech and Language → Speech-to-Text.

Key STT settings that affect microphone performance:

Language

Ensure the language configured in the STT engine matches the language you expect users to speak. If you support multiple languages, ensure the languages are added to the profile under Configure → Profiles → [Your Profile] → Language and the STT engine supports them.

Profanity filter

The profanity filter (available on Azure STT) replaces profanities with asterisks in the transcript. Enable or disable based on your deployment context.

Auto punctuation

Auto punctuation (available on Google STT) adds punctuation to transcribed speech. This can improve LLM response quality by giving the model better-formatted text to work with.


Testing Microphone Performance

Before finalising the kiosk installation:

  1. Open the Test conversation panel in the Virbe Conversation Editor
  2. Enable Show system messages to see transcription results in the transcript
  3. Stand at the expected user position and speak several test phrases
  4. Check whether the STT engine correctly transcribes your speech
  5. Test with typical background noise for the environment (e.g., play music at the volume level of the store)

If transcription accuracy is poor:

  • Check the Windows microphone level (too low = missed words; too high = distortion)
  • Reposition the microphone closer to the user's expected position
  • Consider a higher-quality or beamforming microphone for the noise level of the environment
  • Try a different STT engine if available (Azure vs Google can have different performance characteristics for specific accents, languages, or acoustic conditions)

Push-to-Talk vs Voice Activity Detection

By default, the Virbe Kiosk uses Voice Activity Detection (VAD) to automatically detect when a user starts speaking. In noisy environments, VAD can trigger false positives (background noise interpreted as speech onset).

As an alternative, switch to an explicit activation mode – Tap to Speak (single tap, then automatic speech detection) or Hold to Talk (button held while speaking). See Speech Capture: VAD, Hold to Talk, Tap to Speak for details.

On this page