Virbe Documentation

Embedding Controls

Fine-tune how the Knowledge Base chunks and embeds your documents to improve retrieval accuracy. Added in v2.4.16.

Embedding controls let you tune how the Knowledge Base processes your documents into vector embeddings. These settings are available from the Knowledge Base Settings page under the Storage usage widget.


How Embeddings Work

When you add content to the Knowledge Base:

  1. The document is split into chunks (fixed-size segments of text, with optional overlap)
  2. Each chunk is sent to the embedding model, which converts it into a vector (a numeric representation of meaning)
  3. The vectors are stored in the embedding index

At runtime, when a user asks a question, the query is also embedded using the same model, and the most semantically similar chunks are retrieved and passed to the LLM as context.

Chunking and embedding settings directly affect the quality of this retrieval step.


Settings Reference

Access embedding controls at Knowledge Base → expand Storage usage widget → Settings.

AI Model

Selects which AI model configuration provides the embedding model. This must be an AI model configured in Configurations → AI Models.

Only AI model configurations that include an embedding model are available in the dropdown.

Embedding Model

Specifies the exact embedding model to use.

Available models depend on the selected AI model provider. Common options:

ModelProviderNotes
text-embedding-3-smallOpenAIDefault. Cost-efficient, excellent for most use cases.
text-embedding-3-largeOpenAIHigher dimensional – improved accuracy for complex or technical content at higher cost.
text-embedding-ada-002OpenAIPrevious generation. Use text-embedding-3-small for new projects.

Changing the embedding model after documents have already been embedded will cause a mismatch between existing embeddings and new query embeddings. You must re-embed all documents after changing this setting for retrieval to work correctly.

Chunk Size

Defines the maximum number of tokens in each chunk. Default: 1,000 tokens.

SettingWhen to use
500–800 tokensDocuments with many short, distinct facts (product specs, FAQs)
1,000 tokensGeneral-purpose content – a good starting point
1,500–2,000 tokensLonger explanations where context must be preserved across paragraphs

Smaller chunks improve precision when facts are dense; larger chunks improve context when topics span multiple paragraphs.

Chunk Overlap

Defines how many tokens are shared between adjacent chunks. Default: 200 tokens.

Overlap prevents information from being lost at chunk boundaries – the final 200 tokens of one chunk are also the opening 200 tokens of the next. This ensures that a sentence split across a boundary still appears complete in at least one chunk.

Increase overlap (300–400 tokens) if you observe incomplete answers that appear to be missing the continuation of a concept.

Processing Limit

The maximum number of document chunks processed per embedding batch. This controls how many chunks are sent to the embedding model in a single API request.

The default is suitable for most plans. Lowering this value can help if you experience rate limiting errors from your embedding model provider during large document imports.


Applying Changes

After changing any embedding setting:

  1. Click Save on the Configuration page
  2. Return to the Knowledge Base document list
  3. Re-embed affected documents by clicking Edit and Save on them – the system will regenerate all vector embeddings using the new settings

Re-embedding is required for the new settings to take effect. Documents embedded with the old settings will not be automatically updated – retrieval will use the old vectors until you re-embed.

For guidance on choosing the right settings for your content, see Best Practices.

On this page