RAG in the enterprise: what makes a good document assistant
A document assistant promises to answer questions from company documents, with sources to back it up. The technique behind it, RAG, is quick to set up. What takes time is getting answers that are correct, up to date, and that respect everyone's access rights. This guide is about that part.
Published September 22, 2026
The principle, in one minute
A language model does not know your documents. RAG (retrieval-augmented generation) works around that limit: before answering, the system searches your documents for relevant passages and gives them to the model along with the question. The principle was formalised in 2020 in a research paper by Lewis et al..
- IngestDocuments are collected from their sources (file shares, document management, wiki, email) and their text is extracted.
- Chunk and indexEach document is split into passages, indexed for keyword search and search by meaning.
- RetrieveFor each question, the system finds the most relevant passages the user is allowed to see.
- Answer with citationsThe model writes the answer from those passages and shows where each piece of information comes from.
The model only steps in at the last stage. The quality of a document assistant is therefore mostly decided upstream, in what it is given to read. In the architecture of an internal platform, described in the guide On-premise generative AI in the enterprise, it is also the layer that demands the most operational work over time.
Why so many document assistants disappoint
When an assistant answers badly, the reflex is to blame the model. Most of the time, the model has faithfully summarised the wrong passage. The causes lie in the document chain.
- Outdated or duplicate sources: three versions of the same procedure, two of them obsolete, and the assistant confidently quotes the wrong one.
- Failed extraction: scanned PDFs without text recognition, flattened tables, headers and footers mixed into the content.
- Chunking that cuts the meaning: a condition on one side, its exception in the next passage, never retrieved together.
- Meaning-only search: it finds related ideas but misses exact references such as a contract number, a product code or the name of a regulatory article. Combining keyword search with search by meaning fixes this.
- No refusals: when nothing relevant is found, the assistant makes something up instead of saying it does not know.
Changing the model only fixes these defects at the margin. They are fixed through work on the sources, tuning of the chain and the instructions given to the model.
Access rights: the non-negotiable condition
A document assistant that ignores access rights turns an internal search engine into an information leak: asking the right question is enough to get an excerpt from an HR file or a board memo. The risk is all the more serious because nobody sees it: no alert, just a well-written answer.
Three requirements follow:
- Filtering happens at retrieval, not at display. A passage the user has no right to must never reach the model, even if it does not appear in the answer.
- Rights follow those of the source. A document removed from a share, or an employee who changes department, must be reflected in the index without undue delay.
- The scope is explicit. The user, or the agent working on their behalf, knows which collections the search covers.
The same logic applies when an agent fetches information directly from business tools: the guide Connecting an AI agent to your systems covers the rights it acts with.
The criteria of a good document assistant
Before choosing a tool or approving a project, use this table as a review grid.
| Criterion | Question to ask | Warning sign |
|---|---|---|
| Cited sources | Does every answer point to the exact passage and the original document? | A list of documents with no passage, or no source at all |
| Refusals | What does the tool say when the information does not exist? | A plausible answer with no source |
| Freshness | How long before a modified document is taken into account? | Manual or monthly reindexing |
| Search | Does it find a contract number or an exact reference? | Only search by meaning is available |
| Formats | Does it handle scanned PDFs, tables, slide decks? | Indexed documents that are empty or unreadable |
| Rights | Do the source's rights apply to the search? | A single index, visible to everyone |
A tool that fails the first two rows is not a document assistant, it is a generator of plausible text.
Measure quality before rolling out
A document assistant is judged on a set of real questions, not on a demo. Gather a few dozen questions asked by business teams, each with the document that contains the answer.
Then measure two things separately. First, retrieval: is the right passage among those the system found? Then the answer: is it correct, faithful to the passage and properly sourced? This split tells you where to act. If retrieval fails, changing the model will not help; if retrieval succeeds and the answer is wrong, the problem lies in the instructions or in the model.
Add questions that have no answer in the corpus, to check that the assistant can say so. Replay the whole set on every change: new source, new chunking, new model. And start small: one well-kept collection for one business team beats indexing the whole network share at once.
RAG reduces made-up answers, it does not eliminate them. Citations are there to be checked: a user who never clicks on the sources trusts the tool without any control.
Frequently asked questions
Do we have to choose between RAG and fine-tuning?
To make your documents known to the model, RAG is the right tool: it updates as documents change and cites its sources. Fine-tuning changes the model's behaviour, not its searchable knowledge; it is better suited to enforcing a style or a format. Choosing the model itself is covered in the guide Open source or proprietary LLMs.
Do we need a dedicated vector database?
Not necessarily. A specialised vector database is one option, but a relational database you already run, such as PostgreSQL with the pgvector extension, may be enough. The criterion is mostly operational: backups, monitoring and the skills available in your teams.
Can a document assistant work with no outbound traffic?
Yes. Indexing, retrieval and the model can all run inside your perimeter. For sensitive documents it is even the easiest choice to justify, since the excerpts sent to the model never leave your information system.
Where SmartAGT fits
SmartAGT is deployed inside your perimeter, with a 100% local model if you wish, and no outbound traffic at all. Supported vector databases are PostgreSQL with pgvector, and Pinecone.
Connectors such as SharePoint, OneDrive, Google Drive and Confluence act with the rights of the connected account. Examples are on the Use cases page.