Knowledge Base Best Practices
Practical guidance for building a Knowledge Base that retrieves accurate, relevant answers and minimises hallucination.
RAG quality depends on the quality of the content you add, how that content is chunked, and how well your queries are matched to it.
Content quality
Write for the question, not the source
Source documents like marketing brochures or legal contracts are written for human readers, not AI retrieval. Check whether each document contains the kind of information a user would actually ask about before importing it.
Restructure or summarise content where possible so that each text document covers one topic clearly. A 50-word paragraph that directly answers "What are your opening hours?" will outperform a 5,000-word PDF that mentions opening hours in passing.
Avoid redundant or contradictory content
Duplicate information (the same fact in multiple documents) can cause the LLM to receive conflicting context. This leads to hedged, uncertain-sounding answers or, worse, incorrect ones if two sources disagree. Before adding a new document, check whether the information already exists in the Knowledge Base.
Keep content current
Outdated content is a common source of incorrect answers. Establish a review schedule for your Knowledge Base documents – quarterly at minimum for fast-changing areas like pricing, product specs, or policies. The website crawler can be re-run on demand via Crawl manually to pick up changes, but Text and Table documents require manual editing.
Use Table documents for structured lookups
If users ask questions like "What is the price of product X?" or "Where is the branch in city Y?", store this data as a Table document rather than embedding it in prose. Table documents can be queried precisely with the Find Records node, returning exact values rather than relying on RAG to extract them from text.
Chunking strategy
Understand what chunking does
When you add a document, its content is split into chunks and each chunk is embedded as a vector. At runtime, the user's query is also embedded, and the most semantically similar chunks are retrieved and sent to the LLM as context.
If chunks are too small, they lack the context needed to form a complete answer. If chunks are too large, they include irrelevant information that confuses the LLM and uses up the context window.
Default settings are a reasonable starting point
The default chunk size of 1,000 tokens and overlap of 200 tokens work well for most general-purpose knowledge. Start here and only adjust if you observe specific retrieval problems.
When to reduce chunk size
Reduce chunk size (500–750 tokens) when your documents contain many short, distinct facts – for example, a product catalogue where each entry is a few sentences. Smaller chunks reduce the chance of retrieving a chunk that contains the right product but the wrong facts.
When to increase chunk size
Increase chunk size (1,500–2,000 tokens) when your documents contain information that only makes sense in context – for example, troubleshooting guides where symptoms and solutions are described together over multiple paragraphs. Larger chunks keep related information together.
Use overlap to preserve context across chunk boundaries
Chunk overlap ensures that sentences at the boundary of one chunk are also included at the start of the next, preventing important context from being split across two chunks that might not both be retrieved. The default of 200 tokens handles most cases; increase to 300–400 if you notice incomplete answers that seem like they're missing the "continuation" of a thought.
After changing chunk size or overlap, re-embed your documents from the Knowledge Base Configuration page. Existing embeddings are not automatically updated.
Website crawling
Limit crawl scope to relevant pages
The website crawler follows links from the URLs you provide. If you add a root domain (e.g. https://yourcompany.com), it will crawl every page it can reach – including blog posts, legal pages, careers pages, and other content that may not be relevant to your virtual being's purpose.
Use a sitemap URL or add only the specific paths relevant to your use case (e.g. https://yourcompany.com/products, https://yourcompany.com/support).
Avoid crawling dynamic or login-gated content
The crawler cannot log in, fill forms, or execute JavaScript. Pages that require authentication or load content via JavaScript (single-page apps, dashboards) will not be crawled correctly. Use Text or Table documents for content that cannot be publicly crawled.
Set parallel requests conservatively
The default of 1 parallel request is safe for most websites. Increasing this speeds up large crawls but can trigger rate limiting on the target server. For internal or own-hosted documentation sites, 2–3 parallel requests is usually fine. For third-party websites, keep it at 1.
Re-crawl after content changes
Website changes are not automatically detected. After updating your product pages, documentation, or other crawled content, open the website source in the Knowledge Base and click Crawl manually to pick up the changes and generate updated embeddings.
Retrieval quality
Test retrieval with representative questions
After loading content, use the Test conversation panel in the Conversation Editor to ask questions a real user would ask. Enable Show system messages to see which knowledge chunks were retrieved. If the wrong chunks are retrieved, or none at all, consider:
- Rewriting the document to use language closer to how users phrase questions
- Reducing chunk size to make individual facts more targetable
- Adding a Table document for structured data the LLM is trying to extract from prose
Use Knowledge Base filtering to reduce noise
If your virtual being has different Knowledge Base collections for different topics (e.g. products, policies, FAQs), use filter overrides in the Tool Agent or LLM Response node to restrict retrieval to the relevant collection at runtime. Retrieving from all collections simultaneously can surface irrelevant chunks that distract the LLM.
Combine RAG with deterministic logic
For critical, high-stakes information (pricing, legal commitments, product specification), do not rely on RAG alone. Use a Find Records node to query a Table document directly, and pass the result to a Text node for exact, verbatim responses. Reserve LLM Response with RAG for explanatory, conversational, or exploratory content.
Storage and limits
Monitor the Storage usage panel in the Knowledge Base sidebar. It shows:
- Storage % – total document storage used vs your plan limit
- Embeddings % – total percentage of vector embeddings generation - how much information is already indexed
- Web crawl % – storage used by crawled web content
If you are approaching your plan limits, archive or delete documents that are no longer needed. Unused crawled pages from old website versions are a common source of wasted storage.
Operator responsibility for Knowledge Base content
The organisation that configures and operates the virtual being is responsible for the accuracy, completeness, and currency of the content added to the Knowledge Base. This includes:
- Factual accuracy — the virtual being's responses are grounded in the content you provide. Inaccurate, outdated, or misleading content in the Knowledge Base will produce inaccurate, outdated, or misleading responses. Virbe does not verify or warrant the content you add.
- Completeness — if the Knowledge Base does not contain the answer to a common user question, the LLM may hallucinate an answer rather than saying "I don't know." Ensure coverage is adequate for the virtual being's intended scope, and instruct the model to acknowledge gaps rather than guess.
- Compliance with applicable law — do not add content that infringes third-party intellectual property, contains unlawful material, or that you do not have the right to use for this purpose.
- Change control — when the underlying source material changes (policies, products, prices, procedures), update the Knowledge Base promptly. Stale content is one of the most common causes of incorrect virtual being responses in production. It is recommended to establish a review and update schedule as part of your operating procedure.
- Post-deployment changes — changes made to the Knowledge Base after the virtual being has been deployed and accepted are made at the operator's risk. Virbe recommends testing all Knowledge Base changes in a staging profile before publishing to production, and maintaining a record of what was changed, by whom, and when.