Virbe Documentation

Data and Privacy

What Virbe processes and retains, how it flows through third-party AI providers, and your obligations as the operator.

This page provides an overview of how data flows through the Virbe platform and what obligations apply to organisations operating a virtual being. Unless regulated by the separate Enterprise agreement, for contractual and legally binding details, refer to Virbe's Data Processing Agreement (DPA) and Privacy Policy, which govern the relationship between Virbe and its customers.


Data the Virbe platform processes

When a virtual being handles a conversation, the following categories of data pass through or are retained by the platform:

Conversation content
User messages and virtual being responses are processed in real time to execute conversation logic. Conversation transcripts are stored and accessible in Conversations for 12 months by default (DPA, Annex I), unless you configure a shorter or longer retention period or an Enterprise agreement defines another one.

Analytics and usage data
Aggregate usage metrics (conversation counts, duration, node activity) are retained for the analytics period defined in your agreement and displayed in Analytics.

Configuration data Your Pipeline logic, system instructions, knowledge base content, and profile settings are stored in Virbe and form part of your account data.

Audio data (kiosk and voice) For voice-enabled deployments, audio is processed by the configured STT (Speech-to-Text) provider. Depending on your STT provider and configuration, audio may or may not be retained by the provider – consult your STT provider's data processing terms.


How data flows: touchpoints, hosting, and third-party services

Every conversation follows the same basic path, regardless of use case:

1. Touchpoint
The end user interacts through a touchpoint – the Web Widget in a browser, a kiosk application, or another app built on the Virbe APIs. The touchpoint captures user input (text messages, and voice audio in voice-enabled deployments) and sends it to the Virbe Hub.

2. Virbe Hub – Virbe servers or your Azure tenant
The Hub is where conversation logic executes and where data is stored: conversation transcripts, Pipeline configuration, and Knowledge Base content. Where the Hub physically runs depends on your hosting option:

  • Virbe-hosted – the Hub runs on Virbe's cloud infrastructure and is operated by Virbe.
  • Azure-hosted (self-hosted) – the Hub runs inside your own Azure tenant. Conversation data, configuration, and Knowledge Base content reside in infrastructure you control. Virbe has no standing access to this data – only a Virbe system administrator can log into the instance, and only at the Customer's request (for example, for support or maintenance).

3. Third-party services – depending on your configuration
The Hub calls external services only where your configuration requires them, and each service receives only the data needed for its function:

ServiceWhat is sentWhen it is called
STT providerThe user's voice audioVoice-enabled deployments, on each spoken user turn
TTS providerThe virtual being's response textVoice-enabled deployments, on each spoken response
LLM / AI model providerUser message, conversation history, System Instruction, retrieved Knowledge Base contentWhen a Tool Agent, LLM Response, or Call Conv AI node executes (detailed below)
Webhooks and custom endpointsThe payload and variables you configure in the nodeWhen a Call Webhook node or a Custom Endpoint engine you have set up executes

No third-party service receives conversation data unless you have configured it in your deployment. Which specific providers are involved – and under which terms – is therefore determined by your configuration choices.


How data flows to AI model providers

When a conversation turn reaches a Tool Agent, LLM Response, or Call Conv AI node, Virbe sends a request to the AI model provider you have configured. That request typically contains:

  • The user's message (or the processed equivalent)
  • Relevant conversation history (based on the History message count setting)
  • Your System Instruction and any injected context variables
  • Retrieved Knowledge Base content (if RAG is enabled)

This data is processed by the provider under their data processing terms. For regulated industries, you must ensure your AI model provider offers appropriate data processing guarantees. Azure OpenAI, for example, provides EU data residency options and enterprise DPA terms that may be required for financial services or healthcare deployments.

Who holds the provider relationship
By default, you configure AI, STT, and TTS providers using your own API credentials – the provider relationship, including its data processing terms, is then directly between you and the provider. Under a separate Enterprise agreement, Virbe can instead provision the provider credentials and act as an intermediary (proxy) between your deployment and the provider. In that model, the provider acts as Virbe's subprocessor, and responsibility for the provider's services is regulated in the Enterprise agreement on a back-to-back basis: Virbe passes the underlying provider's service commitments (such as availability and support) through to you, and does not assume service obligations beyond what the provider itself offers. For data protection, the provider is Virbe's sub-processor under the DPA, and Virbe remains responsible for it as required by the GDPR.

Review and accept the data processing terms of every AI model provider, STT provider, and TTS provider you configure before going live with a production deployment.


Virbe's role: data processor or data controller

Virbe's role under GDPR (and equivalent regulations) depends on the category of data and on your hosting model:

End-user conversation data – you are the controller.
As the operator, you decide why and how your virtual being processes end-user data, which makes you the data controller for that data. On a Virbe-hosted deployment, Virbe stores and processes conversation data on your behalf and on your instructions, acting as a data processor under the Data Processing Agreement. On an Azure-hosted deployment, conversation data remains within your own Azure tenant, and Virbe does not process it in the ordinary course – access to the instance is limited to a Virbe system administrator logging in at your request (for example, for support or maintenance).

Your account data – Virbe is the controller.
For data about you as a customer – Dashboard user accounts, login credentials, billing details, and support communication – Virbe determines the purposes of processing and acts as the data controller, as described in Virbe's Privacy Policy.

In neither role does Virbe:

  • use your conversation data, Knowledge Base content, or configuration to train AI models – neither Virbe's own nor any third party's;
  • claim ownership of your data – conversation content, Knowledge Base content, and configuration remain yours;
  • sell your data, share it for advertising purposes, or use it for any purpose other than providing, securing, and supporting the service.

Third-party providers you configure (STT, TTS, AI models, webhook targets) process data under their own terms, not Virbe's. Many enterprise providers – Azure OpenAI, for example – contractually commit not to train models on customer data, but you must verify this for every provider you enable.


End-user personal data

As the operator of a virtual being, you determine what personal data your virtual being collects from users. This makes you the data controller for end-user personal data processed through your deployment, with the associated obligations under GDPR and equivalent regulations.

Key considerations:

Virtual beings built on the Virbe platform are not designed or intended to collect, store, or process personal data of end users beyond what is strictly necessary for conversation execution. End users should be explicitly warned not to share personal data — including names, identity numbers, account numbers, financial data, health information, or other identifying information — in conversations with the virtual being.

Inform users appropriately Your privacy policy and/or consent flow should cover data collected through the virtual being. If the virtual being logs transcripts of conversations, this should be disclosed.

Minimise data collection Only collect personal data that is necessary for the virtual being's function. Avoid asking for data – names, email addresses, account numbers – unless the conversation flow genuinely requires it.

Avoid sensitive data categories Do not configure flows that solicit special category data (health, financial, biometric) unless you have the appropriate legal basis and safeguards in place, particularly for financial services deployments.

PII in conversation logs
Conversation logs accessible in the Monitor section may contain personal data entered by users. Ensure access to these logs is restricted to authorised personnel and that your retention and deletion practices are consistent with your privacy policy.

This warning should be displayed to the end user before or at the start of the interaction. In web widget deployments, it is recommended to use the AI Disclaimer setting in Profile settings to display a pre-conversation message. For kiosk deployments, it's good practice to include this information in the first message to the user, and additionally have it accessible on-site, where the physical device is situated.

An example of a warning text (adapt as required for your regulatory context):

"This is an AI assistant. Do not share personal data, account numbers, passwords, or sensitive information in this conversation. Conversations may be logged by [company] for quality purposes."

These recommendations do not constitute legal advice – check and conform to the regional laws.

Even with this warning in place, users may inadvertently include personal data in their messages. Conversation transcripts accessible in the Conversations monitor should be treated as potentially containing personal data. Access must be restricted to authorised personnel, and retention and deletion practices must be consistent with your privacy policy and applicable data protection law.


Automated PII detection and removal (available on request)

Beyond the preventive measures above, the platform can be extended with an automated mechanism that detects personal data in conversation content and removes or substitutes it. This capability is not enabled by default – it is activated on the Customer's request as part of the deployment setup.

Where it operates
The mechanism processes incoming user text – messages the user typed, or the transcript produced by the STT engine for spoken input – before the text is written to the database and before it enters the Conversation Pipeline. Because redaction happens at this earliest text-processing stage, the detected personal data is absent from everything downstream: stored conversation transcripts, the conversation context sent to AI model providers, and payloads passed to webhooks or external engines.

Two implementation options are available:

Azure PII detection
For deployments running on Azure infrastructure, incoming text is processed through Azure AI Language PII detection, which identifies personal data entities – names, phone numbers, addresses, identity and account numbers, and other recognised categories – and removes or masks them.

Bespoke PII mechanism
Where the regulatory or contractual context requires a different approach – specific entity types, specific languages, or substitution (pseudonymisation) instead of removal – a custom PII detection and removal/substitution mechanism can be designed and implemented in agreement with the Customer, typically as part of an Enterprise engagement.

What it does not cover
The mechanism operates on text only. In voice deployments, the user's raw audio necessarily reaches the configured STT provider first in order to be transcribed – redaction applies to the resulting transcript, not the audio stream. It also does not process the virtual being's own generated responses.

Automated PII detection is probabilistic – no mechanism identifies every instance of personal data in free-form conversation, particularly across languages and unusual formats. Enabling PII removal reduces the amount of personal data retained; it does not by itself guarantee that no personal data is processed or stored. The end-user warning, access restrictions, and retention practices described above remain necessary regardless of whether this mechanism is enabled.

Contact Virbe to discuss enabling PII detection and removal for your deployment, including the scope of detected entities and whether detected data should be removed or substituted.


Knowledge base content

Content you add to the Knowledge Base – website pages, text documents, table records – is stored wherever your Virbe instance is deployed, and processed there to generate vector embeddings. On the Virbe-hosted option, that's Virbe's cloud infrastructure; on a self-hosted deployment (e.g. Azure-hosted), it's your own infrastructure, under your control. This content is owned and managed by you regardless of hosting model. It's generally not advised to add content that contains personal data of individuals unless you have a lawful basis for processing that data in this context.


Deletion and data subject requests

For questions about data deletion, subject access requests, or exercising rights under GDPR or equivalent law, refer to your Virbe Data Processing Agreement or contact Virbe support. The DPA governs data retention periods and the mechanics of data deletion on request.

Beyond the contractual mechanics, the platform provides two operational cleanup capabilities:

Manual deletion
The content of individual conversations can be deleted manually from the Conversations log – either for specific conversations or for all of them.

Scheduled automatic cleanup
Conversation content can also be cleaned up automatically on a recurring schedule, according to criteria defined for your deployment. The typical configuration is age-based: clear or obfuscate the text of messages older than a defined threshold – for example 15 minutes, 1 hour, or 1 day – so that message content is retained only for the window your retention policy allows, while it remains available for live monitoring and handover during the conversation itself.

Combined with the automated PII detection and removal mechanism, scheduled cleanup allows implementing short retention windows for deployments where minimising stored personal data is a priority – for example public kiosk installations or regulated environments.

On this page