Embedding Controls
Fine-tune how the Knowledge Base chunks and embeds your documents to improve retrieval accuracy. Added in v2.4.16.
Embedding controls let you tune how the Knowledge Base processes your documents into vector embeddings. These settings are available from the Knowledge Base Settings page under the Storage usage widget.
How Embeddings Work
When you add content to the Knowledge Base:
- The document is split into chunks (fixed-size segments of text, with optional overlap)
- Each chunk is sent to the embedding model, which converts it into a vector (a numeric representation of meaning)
- The vectors are stored in the embedding index
At runtime, when a user asks a question, the query is also embedded using the same model, and the most semantically similar chunks are retrieved and passed to the LLM as context.
Chunking and embedding settings directly affect the quality of this retrieval step.
Settings Reference
Access embedding controls at Knowledge Base → expand Storage usage widget → Settings.
AI Model
Selects which AI model configuration provides the embedding model. This must be an AI model configured in Configurations → AI Models.
Only AI model configurations that include an embedding model are available in the dropdown.
Embedding Model
Specifies the exact embedding model to use.
Available models depend on the selected AI model provider. Common options:
| Model | Provider | Notes |
|---|---|---|
text-embedding-3-small | OpenAI | Default. Cost-efficient, excellent for most use cases. |
text-embedding-3-large | OpenAI | Higher dimensional – improved accuracy for complex or technical content at higher cost. |
text-embedding-ada-002 | OpenAI | Previous generation. Use text-embedding-3-small for new projects. |
Changing the embedding model after documents have already been embedded will cause a mismatch between existing embeddings and new query embeddings. You must re-embed all documents after changing this setting for retrieval to work correctly.
Chunk Size
Defines the maximum number of tokens in each chunk. Default: 1,000 tokens.
| Setting | When to use |
|---|---|
| 500–800 tokens | Documents with many short, distinct facts (product specs, FAQs) |
| 1,000 tokens | General-purpose content – a good starting point |
| 1,500–2,000 tokens | Longer explanations where context must be preserved across paragraphs |
Smaller chunks improve precision when facts are dense; larger chunks improve context when topics span multiple paragraphs.
Chunk Overlap
Defines how many tokens are shared between adjacent chunks. Default: 200 tokens.
Overlap prevents information from being lost at chunk boundaries – the final 200 tokens of one chunk are also the opening 200 tokens of the next. This ensures that a sentence split across a boundary still appears complete in at least one chunk.
Increase overlap (300–400 tokens) if you observe incomplete answers that appear to be missing the continuation of a concept.
Processing Limit
The maximum number of document chunks processed per embedding batch. This controls how many chunks are sent to the embedding model in a single API request.
The default is suitable for most plans. Lowering this value can help if you experience rate limiting errors from your embedding model provider during large document imports.
Applying Changes
After changing any embedding setting:
- Click Save on the Configuration page
- Return to the Knowledge Base document list
- Re-embed affected documents by clicking Edit and Save on them – the system will regenerate all vector embeddings using the new settings
Re-embedding is required for the new settings to take effect. Documents embedded with the old settings will not be automatically updated – retrieval will use the old vectors until you re-embed.
For guidance on choosing the right settings for your content, see Best Practices.