Skip to main content

What are you looking for?

Explore our services and discover how we can help you achieve your goals

RAG architecture for internal knowledge assistants: the design decisions behind grounded answers

A reliable internal knowledge assistant is a retrieval system first and a language model second. Index approved sources with their access rights and dates, retrieve with hybrid keyword and vector search, filter by the asker's permissions before generation, rerank, answer only from cited passages, refuse when evidence is missing, and score every change on real questions.

Book a solution review See the related service

Reviewed by David (CEO) · Updated 29 Sep 2026 · 10 min read

star

This guide is for architects and engineering leads designing the retrieval pipeline behind an employee-facing assistant. It explains the decisions and their defaults. If you want the engineering delivered, see our enterprise knowledge and RAG service; for the packaged product view, see the AI knowledge assistant. Knowledge assistants are one of the AI steps covered in our AI automation guide for business operations.

In this guide

Why retrieval decides the quality

Retrieval-augmented generation, introduced by Lewis and colleagues in 2020, pairs a language model with an external index: the system first finds the passages relevant to a question, then asks the model to answer from them. For company knowledge this matters twice. The model was never trained on your contracts or procedures, and those documents change every week.

The consequence is easy to miss: when an assistant gives a wrong answer, the cause is usually upstream of the model. The right passage was never indexed, was split in the wrong place, ranked below noise, or filtered out. Swapping the model rarely fixes that. The decisions below are where quality is won or lost.

The pipeline decisions at a glance

Stage Decision Sensible default Revisit when
Ingestion What goes in, with which metadata Approved sources only, each with owner, access rights and last-updated date A source has no owner or no reliable dates
Chunking Where documents are split Split on the document's own structure (headings, sections, table rows) and keep the heading path Answers need whole procedures or long tables
Retrieval How candidates are found Hybrid keyword and vector search, merged by rank fusion Queries are almost all codes and names, or almost all paraphrases
Reranking Which candidates reach the model A reranker scores the candidates and keeps a short, diverse set Latency or cost limits are tight
Permissions Who may see which passage Filter by the asker's rights inside the retrieval query Sources have no machine-readable access rights
Generation How the answer is formed Answer only from the passages, cite each claim, refuse when unsupported Users need calculations over live data
Evaluation How quality is proven A set of real questions with expected sources, run on every change The question mix shifts

Ingestion and chunking

Start with a source inventory rather than a crawl. For each system record the owner, the access model, the update frequency and whether it holds the current version of the truth. Duplicates and outdated copies are the most common cause of confident wrong answers, so decide which source wins before indexing anything.

Chunking is the decision with the largest effect on retrieval. Fixed-size windows are simple but cut procedures and tables in half. Structure-aware chunking splits on headings, list items or table rows and stores the heading path with each chunk, so a passage titled "Refunds > Damaged goods > Outside the EU" can be found and cited on its own. For long procedures, index small chunks for matching but hand the model the parent section, so the answer sees the whole step list. Keep document metadata on every chunk: source, owner, access rights, version and date.

Retrieval: keyword, vector or hybrid

Vector search finds passages that mean the same thing in different words. Keyword search finds exact matches: product codes, contract numbers, names and internal jargon that embeddings blur. Company questions mix both, which is why hybrid retrieval is the sensible default. Microsoft's documentation for its search service describes the common pattern: run full-text and vector queries in parallel and merge the two ranked lists with reciprocal rank fusion.

A reranker then reads each candidate together with the question and reorders them. It is slower than first-stage retrieval, so it runs only on the candidates, and it is often the cheapest single improvement to answer quality. Keep the final set short and diverse. Liu and colleagues showed in "Lost in the Middle" that models use information at the start and end of a long context better than information in the middle, so more context is not automatically better.

Permissions: filter before the model sees anything

An assistant must never quote a document the asker could not open. Enforce that in retrieval, not in the prompt: store access rights with each chunk, pass the asker's identity and groups into the search query, and drop anything they may not see before generation. Filtering after generation, or asking the model to withhold restricted content, leaks information through summaries and hints.

OWASP lists vector and embedding weaknesses as LLM08 in its 2025 Top 10 for LLM applications, including unauthorised access through the index, cross-tenant leaks in shared vector stores and poisoned content. Its mitigations match this design: fine-grained, permission-aware access to the vector store, validated ingestion from trusted sources, content classified by access level and immutable logs of retrieval. When an assistant can also act, not only answer, the controls in our AI agent security guide apply as well.

Freshness and deletion

Company knowledge changes, and the index must follow. Re-index on change events where the source system provides them and on a schedule where it does not. Store the last-updated date with each chunk and show it next to the citation, so a user can see that a policy passage is two years old. Deletion matters as much as updates: when a document is withdrawn or a person's access is revoked, the index must reflect it within the time your data policy promises, and retention of questions and answers needs its own rule.

Grounded answers and refusals

The generation step is a contract: answer from the passages, cite each claim to a passage, and say "I don't know" when the passages do not support an answer, pointing to the document owner. NIST's Generative AI Profile names confabulation, the confident statement of false content, among the core risks of generative AI; citations and refusals are how users check for it. For regulated topics such as HR, legal or finance, route answers to a draft-for-review queue instead of replying directly.

Evaluating a RAG system

  1. Collect real questions

    Take them from tickets, chat logs and interviews, not from the documents, and record the source that should answer each one.

  2. Score retrieval separately

    For each question, check whether the expected source is in the retrieved set and how high it ranks. Most failures show up here.

  3. Score the answer

    Check faithfulness to the passages, relevance to the question and whether citations point to the right source. The Ragas framework proposed reference-free measures along these dimensions; human review of a sample stays necessary.

  4. Test refusals and permissions

    Include questions with no answer in the sources and questions a test user must not see answered.

  5. Run the set on every change

    Chunking, ranking, prompt and model changes all move results; release only when the set holds.

  6. Learn from production

    Unanswered and low-rated questions show document owners what is missing and feed the next test set.

RAG, long context, fine-tuning or search: alternatives and selection criteria

Approach Good at Weak at Choose it when
RAG Large, changing, permissioned document sets with citations Needs ingestion, permissions and evaluation work Answers must come from many sources that change
Long-context prompting A few documents read in full Cost per question, attention to the middle, permissions The whole corpus fits one prompt and one audience
Fine-tuning Tone, format and domain vocabulary Fresh facts, citations, access control Style matters more than current facts
Enterprise search Finding documents fast Synthesising an answer People need the document, not a summary

Four criteria decide it: how often the knowledge changes, whether different people may see different documents, whether answers must be traceable to a source, and how many documents are involved. In practice the approaches combine: RAG for facts, a light fine-tune or instructions for format, and search as the fallback. Choosing which knowledge problem to solve first is part of an enterprise AI readiness assessment.

AI in the RAG pipeline

AI appears at several points of the pipeline, not only in the final answer: embedding models for vector search, rerankers, optional query rewriting, and models that grade answers during evaluation. Each is a component to version and test. Netbase works with the major commercial and open-source AI models, chosen per project, and builds pipelines so the model can be replaced without re-indexing everything. The chat surface can start from the AI chatbot and WorkChat integrator in the Netbase productized module library, and data and model choices are covered in our data and AI stack. Each item below states how mature it is at Netbase.

What delivery record exists, and what does not

  • What exists. Netbase has delivered retrieval-augmented knowledge assistants, document AI and MLOps pipelines for clients that are not named; the anonymised RAG record describes the kind of system and Netbase's role. Delivery follows Netbase's security practices, including role-based access control, MFA for admin dashboards, TLS in transit and AES at rest, secure code review and vulnerability scanning.
  • What does not. The record publishes no client name, corpus size, accuracy, adoption or time-saved figure, so none is offered here. The defaults on this page are design guidance, not measured results from that project.

Limits of this guide

  • The defaults suit internal assistants over documents and records in one organisation; public chatbots, multi-tenant products and regulated advice need further controls.
  • Research papers and vendor documentation are cited for methods, not as endorsements; retrieval quality depends on your content and must be measured on your questions.
  • Personal data in the index brings data-protection duties that this guide does not replace; take legal advice where the GDPR or other privacy law applies.

Plan the next step with a Netbase consultant

Frequently asked questions

There is no single size. Split on the document's own structure, keep the heading path with each chunk, and test sizes on your evaluation set; the right answer differs between policies, contracts and tickets.

Usually not. Company questions mix meaning and exact terms such as codes and names, so hybrid keyword and vector retrieval with a reranker is the safer default.

Store access rights with every chunk and filter by the asker's rights inside the retrieval query, before the model sees anything, then test with users who must not see certain answers.

Fine-tuning shapes style and vocabulary but does not keep facts current, cite sources or respect access rights. For changing company knowledge, RAG is the base and fine-tuning an optional addition.

Next step

Share the sources your people search today, who may see what and ten real questions, and we will book a solution review to sketch the retrieval design and its evaluation set. You can also see AI automation and agents or more Netbase insights.

AI automation and agents that keep people in charge AI automation and agents that keep people in charge

Netbase provides AI automation and agent development for operations teams that want repetitive, multi-step work done by software while people keep approval over the decisions that matter. We combine rule-based workflow automation with AI steps where they add value, design the human approval points in, and measure return against a baseline taken before the build.

Learn More
line
Enterprise AI knowledge assistants: RAG engineering with access control Enterprise AI knowledge assistants: RAG engineering with access control

Netbase engineers retrieval-augmented generation (RAG) over company knowledge: the ingestion, search, permission checks, cited answers and evaluation that let a language model answer from your own documents and systems. It is a growth capability with access control and evaluation designed in from the first prototype, and Netbase has delivered a RAG knowledge assistant for a client that is not named.

Learn More
line
AI knowledge assistant: answers from company knowledge, with sources and access control AI knowledge assistant: answers from company knowledge, with sources and access control

An AI knowledge assistant is an internal tool that answers employees' questions from your company's own documents and systems, cites the source behind every answer and shows each person only what they may see. Netbase offers it as a growth capability, designed with retrieval, evaluation and human review from the first pilot, and has delivered one for an unnamed client.

Learn More
line
Contact Netbase

Discuss a project

Netbase JSC helps organizations design, build, modernize, and operate digital products and AI-enabled business systems.
Project enquiries

[email protected]

WhatsApp

+84 937 869 689

Office address

91 Nguyen Chi Thanh, Dong Da, Hanoi, Vietnam

Get in touch

Tell us what you want to build, modernize, or operate.

Tell us what you want to build, modernize, or operate.

Contact Netbase