On this page8 sections
Retrieval-augmented generation (RAG) is a way of getting a large language model to answer questions from your own documents. When someone asks a question, the system first searches your documents for the most relevant passages, then gives those passages to the model along with the question and instructs it to answer using only that material, citing where each part came from. The model supplies the language; your documents supply the facts. That makes answers current, specific to your business and checkable, without training a model on your data.
This post explains how RAG works step by step, what it is good for, what it costs to run, and the mistakes that make RAG systems give poor answers.
Why not just ask a language model directly?
A general language model has learned from a huge amount of public text, but it knows nothing about your price list, your contracts, your internal policies or last month's product changes. Asked about them, it either says it does not know or, worse, produces a plausible answer that is wrong.
You could try to train a model on your documents, but that is slow, expensive and has to be repeated whenever the documents change. It also makes it hard to see where an answer came from. RAG keeps your documents separate and looks things up at the moment of the question, so updating knowledge is as simple as updating the documents.
How does RAG work, step by step?
A RAG system has two phases: preparing the documents once, and answering each question.
Preparing the documents
- Collect the sources. Policies, manuals, product sheets, contracts, help articles, past support tickets: whatever the answers should come from.
- Extract clean text. Turn PDFs, web pages and office files into plain text, keeping headings and the document name. Our post on extracting data from PDFs covers why this step is harder than it looks.
- Split into chunks. Break each document into passages of a few paragraphs, ideally along headings, so each chunk is about one thing.
- Create embeddings. An embedding model turns each chunk into a list of numbers that represents its meaning. Passages about similar topics get similar numbers.
- Store them in an index, usually a vector database, together with the original text and a link back to the source.
Answering a question
- Embed the question with the same model.
- Retrieve the chunks whose meaning is closest to the question. Many systems combine this with ordinary keyword search, which is better at exact terms such as product codes.
- Rank the results and keep the best few.
- Generate. Send the question and the selected chunks to the language model, with instructions to answer only from them, cite each source, and say so when the answer is not there.
- Show the answer with its citations, so the reader can open the original passage.
What is RAG good for in a business?
- Internal knowledge search. Staff ask questions in plain language and get answers from policies, procedures and manuals, with links to the source.
- Customer support assistance. Agents get suggested answers drawn from help articles and past resolved tickets, which they check and send.
- Sales and proposals. Pulling accurate product details, security answers and past proposal text into new documents.
- Contract and document review. Finding the relevant clauses across many agreements.
- Onboarding. New staff can ask questions without interrupting colleagues for every small thing.
The common thread is a body of written knowledge that people currently search by hand or by asking someone.
What does RAG cost to run?
The main costs are:
- Building the pipeline: connecting sources, cleaning text, chunking, and building the interface.
- Embedding the documents, which is cheap and mostly a one-off cost, repeated only for changed documents.
- Model calls for each answer, which scale with the number of questions and the amount of text sent each time.
- Hosting the index and the application.
- Ongoing care: adding new sources, removing outdated ones, and checking answer quality.
For most small and mid-sized businesses, the build and the upkeep cost more than the model usage.
Why do RAG systems give bad answers?
Most failures happen before the language model writes a word, in retrieval and in the data itself:
- Poor text extraction. Tables and scanned pages turned into garbled text cannot be retrieved correctly.
- Bad chunking. Chunks that cut a topic in half, or merge unrelated topics, confuse retrieval.
- Outdated documents sitting in the index next to current ones, so the system cites last year's policy.
- Missing keyword search, so questions about exact codes or names find nothing.
- Weak instructions, letting the model fill gaps from general knowledge instead of saying the answer is not in the documents.
- No permissions. If everyone can query everything, confidential documents can surface in answers to people who should not see them.
How do you know if a RAG system is working?
Build a test set before launch: fifty or more real questions with known correct answers and the documents they should come from. Measure two things separately:
- Retrieval. Did the right passage appear in the results?
- Answer quality. Was the answer correct, complete, and supported by the cited passage?
Separating them tells you where to fix things. If retrieval fails, work on extraction, chunking and search. If retrieval succeeds but answers are wrong, work on instructions and the model.
Rerun the test set after every change, and keep collecting questions the system answered badly in real use.
How should a business start with RAG?
- Pick one body of documents and one group of users, such as support agents and the help centre.
- Clean up the documents first, removing duplicates and outdated versions.
- Build a small version with citations and an "I don't know" answer.
- Test it against real questions and measure.
- Widen the sources and users once the results hold up.
A narrow system that is right and shows its sources earns trust. A broad system that is sometimes wrong in ways nobody can check loses it quickly.
Working with Syntora Ai
Syntora Ai builds retrieval systems over company documents with citations, permissions and measured answer quality, and we tell clients plainly when a simpler search would do the job better. If you want staff or customers to get answers from your own documents, write to hello@syntorahq.ai or see our AI systems practice.