Making an AI answer from your own documents: how RAG works

Most business AI projects rest on the same technique: find the right passages in the organisation's documents, then ask the model to answer from them. Here is how it works, where it goes wrong, and why sensitive documents are better kept on site.

A good source, a wrong answer

A passenger asks the chat assistant on the Air Canada website about fares for bereavement travel. The assistant tells them they can pay the ticket at the normal price and claim the reduction afterwards, within 90 days. In the same message, it provides a link to the airline's official page. That page says the opposite: the bereavement fare does not apply to requests submitted after the journey.

The British Columbia Civil Resolution Tribunal ruled in February 2024. Air Canada argued that its assistant was a separate entity, responsible for its own statements; the tribunal found the argument "remarkable" and dismissed it. The airline had to pay the difference between the price paid and the bereavement fare. For the tribunal, it made no difference whether the information came from a static page or from an assistant: the company answers for what appears on its website.

The episode sums up the problem that retrieval-augmented generation, or RAG, sets out to solve. The exact document existed. What was missing was a system that answers from it, and shows where each statement comes from.

The principle: an open-book exam

The term appeared in 2020 in a paper by Patrick Lewis and eleven co-authors, presented at the NeurIPS conference. The authors distinguish two kinds of memory. Parametric memory is what the model absorbed during training, written into its parameters. Non-parametric memory is an external index, in their experiment a version of Wikipedia, which the system consults at the moment of answering. Their conclusion: combining the two produces more precise and more factual text than the model alone.

The most telling comparison is that of an exam. A student who has learned everything by heart answers quickly, but gets wrong what they retained badly and ignores everything that has changed since they revised. A student allowed to open their binder first looks for the right page, then writes from what they read. RAG puts the language model in the position of the second student.

The CNIL sums up the interest of the method: it produces answers enriched by external data, more specific and easier to update than the model itself. This is the decisive point for a company. Retraining a model on its documents is expensive and has to be redone at every update. ANSSI adds a security argument: once data has been built into training, it becomes impossible to apply access rights to it, whereas documents stored alongside the model can remain subject to the usual rules of who sees what.

In practice, a RAG chains four operations: preparing the documents, converting them into coordinates, retrieving the useful passages, then writing an answer that cites them.

Step 1: prepare and split the documents

Everything starts with text extraction. A word-processing file or a web page is easy to read; a scanned PDF requires character recognition, and a table or diagram needs special handling. This thankless work determines a good part of the final quality: badly extracted text will give wrong answers, however powerful the model.

The text is then split into pieces of a few paragraphs, which practitioners call chunks. Splitting is necessary because the search must point to a precise passage, not a two-hundred-page report. It also creates a trap, well described by Anthropic in a September 2024 publication: an isolated chunk loses its context. The example used is a financial report whose extract says that revenue grew by 3% over the quarter, without saying which company or which quarter it refers to.

The remedies are well known. The chunks are made to overlap slightly, headings and articles are respected rather than cutting in the middle of a sentence, and each chunk is given a short description of its context: source document, section, date. This metadata will later be used to filter, to cite and to control access.

Step 2: translate meaning into coordinates

Each chunk then passes through an embedding model, distinct from the model that will write the answer. Its role is to turn a text into a list of numbers, a vector, that represents its meaning. Two passages that talk about the same thing, even in different words, get vectors that are close together.

The most accurate image is that of a map. Each passage receives coordinates, not on two axes like a latitude and a longitude, but on hundreds or thousands of dimensions. On this map, "notice of resignation" ends up close to "time limit to observe before leaving the company", although the two phrases have almost no words in common. This is what distinguishes semantic search from the keyword search of a classic internal search engine.

These vectors look abstract. They are less so than one might think. In a paper presented at the EMNLP 2023 conference, researchers showed that 92% of short texts, of 32 tokens, could be reconstructed exactly from their embeddings alone, thereby recovering full names in clinical notes. A vector index must therefore be protected with the same care as the documents it is drawn from.

Step 3: store, then retrieve

The vectors are stored in a vector database, together with the original text and its metadata. When a user asks a question, it is converted into a vector by the same embedding model, then the database looks for the chunks whose coordinates are closest. The principle is like that of a librarian who, instead of looking for the exact word in the titles, goes straight to the shelf where the works on the subject are found.

Semantic search alone has its weaknesses. It can miss an exact reference, a contract number or an article code, which a keyword search finds easily. Serious systems therefore combine the two, then often add a reranking step: a second model re-reads the shortlisted candidates and sorts them by their real relevance to the question asked.

Anthropic put figures on the effect of these improvements on its own test sets. With a simple vector search, the right information was missing from the top twenty results in 5.7% of cases. Adding context to each chunk brought this rate down to 3.7%. Combined with a BM25-type keyword search, it fell to 2.9%, then to 1.9% with reranking. These are one provider's measurements on its own data, but the order of magnitude shows that retrieval, more than writing, determines the quality of a RAG.

Step 4: write the answer and cite sources

The few selected passages are inserted into the prompt sent to the language model, along with the question and explicit instructions: answer only from the extracts provided, state the document and page for each claim, and say clearly when the information is not there. The model then writes an answer in plain language, and the interface displays links to the cited passages.

Citation changes the nature of the tool. Without it, the user must take the machine's word. With it, they can check in one click, and the organisation can trace afterwards where a disputed answer came from. In the Air Canada case, the link was present but contradicted the text; a well-tuned system must produce an answer faithful to the source it cites, and the user must have the reflex of opening it.

Nor should the model be drowned in extracts. A study published in 2023 in the Transactions of the Association for Computational Linguistics, titled "Lost in the Middle", showed that models make better use of information placed at the beginning or end of the text they are given, and markedly worse when it sits in the middle. A few relevant passages are better than a pile of approximate ones.

RAG is not always necessary, moreover. Anthropic notes that below about 200,000 tokens, or roughly 500 pages, it may be simpler to insert the whole knowledge base into the prompt, provided the model accepts a context window that long. For a set of internal rules and a few procedures, the question arises. For the archives of a firm, it does not.

Why almost every project starts here

A language model, however good, knows neither a company's internal procedures, nor its contracts, nor its product sheets, nor the history of its files. RAG is the most direct way of giving it access to this knowledge without modifying it. Adding a document means indexing it; removing an obsolete version means deleting it from the index. The operation is a matter of document management, not of new training.

The uses are similar from one sector to another. An HR department answers employees' recurring questions about the collective agreement or leave. A customer support team finds the right technical sheet. A legal department queries a contract base to spot termination clauses. A local authority helps its staff find their way through resolutions and regulations accumulated over the years. In each of these cases, the value lies less in the writing than in the ability to find the right passage.

Hallucinations decrease, they do not disappear

RAG is often presented as the remedy for invented answers. A study by the RegLab and the HAI institute at Stanford University, published in May 2024 and covering more than 200 legal questions, strongly qualifies this promise. The legal research tools from LexisNexis and Thomson Reuters, built on RAG and some of which claimed to be free of hallucinations, produced incorrect information in more than 17% of cases for Lexis+ AI and Ask Practical Law AI, and in more than 34% of cases for Westlaw AI-Assisted Research. GPT-4 used alone did markedly worse, with 58 to 82% errors on this type of question.

The authors identify several causes: a retrieval step that brings back texts that do not apply, a difficulty specific to legal reasoning, and the model's tendency to go along with the question, even when it rests on a false premise. RAG therefore moves the risk more than it removes it. If retrieval brings back the wrong passage, the model writes a fluent answer from the wrong passage.

The safeguards are mainly organisational. One builds a set of real questions whose right answers are known, and measures the system against it before deploying it, then at every change. One allows it to answer that it does not know. One reserves human validation for answers that commit the organisation. In its support for a France Travail project, the CNIL insisted on this last point: train staff to spot errors, and give them the time to depart from the tool's suggestion.

Access rights: the assistant sees what the index sees

The most underestimated risk is not error, it is internal leakage. If all of an organisation's documents are indexed in a single database, an intern can ask the assistant for their director's salary and obtain an extract from a payslip. The model forced nothing; it read what the database passed to it.

ANSSI makes this an explicit requirement: an AI system must respect, in its answers, the access restrictions specific to each user. It specifies that this is feasible for documents stored alongside the model, depending on the capabilities of the storage tool, and recommends regularly reviewing the configured rights, for example every month. Microsoft, in its Copilot readiness documentation, puts it in its own way: the assistant respects existing permissions, which amounts to exposing all the overly broad shares accumulated over the years. The vendor also recalls that SharePoint's sharing settings are, by default, the most permissive.

OWASP, the reference foundation in application security, devotes a category to vector and embedding weaknesses in the 2025 edition of its ranking of the ten main risks linked to language models. It covers leaks between users or customers who share a single vector database, poisoning of the database by booby-trapped documents, and the embedding inversion mentioned above. The example cited speaks for itself: a CV containing, in white text on a white background, an instruction meant to make the screening system recommend the candidate. Every indexed document will end up in front of the model; one must know where it comes from.

The practical consequence is simple to state and long to implement. Each indexed chunk must carry the rights of its source document, the search must filter according to the identity of the person asking, and every query must be logged.

Outdated documents, outdated answers

A RAG answers from what it has been given. If the index contains the 2019 memo and the one that replaced it, it may cite either with the same assurance. The case is common in organisations where documents pile up without ever being withdrawn: old price lists, abandoned procedures, outdated contract templates.

The remedy is documentary before it is technical. A person must be made responsible for each corpus, documents must be dated, what is replaced must be archived, and the index must automatically follow additions, modifications and deletions. Microsoft itself recommends archiving inactive sites so that its assistant relies on up-to-date content. On the interface side, displaying the date of each cited source makes it possible to spot at a glance an answer based on an old text.

Why run it locally for sensitive documents

By construction, a RAG handles the organisation's most valuable documents. With an online service, these documents, their vectors, the users' questions and the answers pass through or stay with the provider. In its July 2024 questions and answers on generative AI, the CNIL recommends favouring on-site deployment when personal or sensitive data is processed, to limit the risks of extraction by a third party. ANSSI, in its recommendation R34, advises against online generative AI tools for any professional use involving sensitive data.

The public service offers a documented example. Within its sandbox on AI in public services, the CNIL supported in 2024 the "Conseils Personnalisés" project of France Travail, a tool that suggests training courses to advisers from a RAG fed by the training catalogue and the jobseeker's profile. The model chosen, Mixtral from Mistral AI, was installed on site. The CNIL stresses in its recommendations that a model used in the provider's environment presents risks of loss of confidentiality of personal data.

For some professions, the question touches on professional secrecy. Article 226-13 of the French Criminal Code punishes with one year's imprisonment and a €15,000 fine the disclosure of confidential information by a person entrusted with it by reason of their profession. Law firms and accounting firms are concerned first and foremost. An HR department that indexes individual files, or a local authority that handles social assistance requests, for their part handle personal data whose disclosure would cause real harm to the people concerned.

Local means the whole chain stays on site: text extraction, the embedding model, the vector database, the language model and the logs. Keeping the model in-house while entrusting the vectors to an outside service protects little, since these vectors make it possible to reconstruct part of the texts. In return, the organisation takes on operations: backups, updates, access monitoring.

Where to start

A successful first project comes down to few things: a limited and well-kept corpus, for example one department's procedures, clear access rights, a set of real questions to measure quality, and users who know they must open the cited sources. Difficulties rarely come from the model. They come from the documents, their state, their versions and the question of who has the right to read them.

This is the approach HOMN takes: the documents, the index and the model stay within the organisation's walls, and recourse to an outside service remains an explicit choice. Whatever tool is chosen, the quality of the answers will depend first on that of the corpus and the rigour of the rights.

What is RAG?

RAG, short for retrieval-augmented generation, is a technique that first searches for the relevant passages in a document base, then asks a language model to answer from those passages. It makes it possible to query your own documents without retraining the model and to cite the source of each answer.

Does RAG eliminate hallucinations?

No, it reduces them without eliminating them. A Stanford study published in 2024 measured more than 17% incorrect answers on commercial legal tools built on RAG. If retrieval brings back a wrong passage, the model writes a fluent answer from that passage, hence the importance of citations and human checking.

Do you need to train an AI on your documents for it to know them?

Generally, no. RAG leaves the model as it is and supplies it with the useful extracts at the moment of each question, which makes it possible to update the base by adding or removing documents. It also makes it possible to keep access rights per document, which ANSSI considers impossible once the data has been built into training.

Can you build a chatbot on your documents without sending them to the cloud?

Yes. Text extraction, the embedding model, the vector database and the language model can all run on a server installed on the organisation's premises. This is the option the CNIL recommends favouring when personal or sensitive data is processed.