RAG Best Practices: Lessons From Production Retrieval Systems
RAG best practices from real production systems: chunking, metadata, hybrid search, reranking, evaluation, access control and citations users can trust.

The most important RAG best practices are unglamorous: clean and well-structured source documents, sensible chunking with rich metadata, hybrid search plus a reranker, access control enforced at retrieval time, answers that cite their sources, and an evaluation set that tells you whether a change helped. The language model is rarely the bottleneck. Retrieval is.
These are the lessons we apply when building retrieval-augmented generation systems for teams that need answers from manuals, contracts, tickets and wikis.
Start with the questions, not the documents
Before indexing anything, collect 50–100 real questions from the people who will use the system, and for each one note the document (and ideally the passage) that answers it. This list does three jobs:
- it shows which sources actually matter,
- it reveals question types you did not expect (comparisons, “latest version of…”, yes/no policy checks),
- it becomes your evaluation set.
Teams that skip this step end up tuning by feel, and every change is an argument.
Document preparation: garbage in, garbage out
Parse properly
PDFs, slides and scanned documents are where most quality is lost. Use a parser that keeps headings, lists and tables, and check a sample of the output by eye. Tables flattened into a stream of numbers produce confident wrong answers.
Clean and deduplicate
Remove headers, footers, navigation text and boilerplate that appears on every page. Deduplicate near-identical versions, or tag them so you can prefer the latest.
Keep structure
Store the document title, section heading path (“Handbook > Leave > Parental leave”), page number and URL with every chunk. You will need them for filtering, ranking and citations.
Chunking that respects meaning
There is no universal chunk size. What works:
- Split on structure first — headings, sections, list items, table rows — then by length.
- Keep chunks self-contained. A chunk that says “see the table above” is useless on its own. Prepend the heading path or a one-line document summary so each chunk carries its context.
- Use modest overlap so sentences at boundaries are not lost.
- Treat special content differently. FAQs work well as one question-answer pair per chunk; code works best split by function or class; tables can be stored whole or row by row with the header repeated.
- Test two or three chunking strategies against your evaluation set rather than guessing.
Metadata is a retrieval feature
Metadata is how you answer “what is the current policy for Germany?” rather than returning a 2021 policy for another country. Useful fields:
| Metadata | Used for |
|---|---|
| Source, title, URL, page | Citations and deep links |
| Section path | Context for the model, better ranking |
| Date, version, status | Prefer latest, exclude drafts and superseded docs |
| Product, region, department, language | Filters from user context or the question |
| Access groups / owner | Permission-aware retrieval |
| Document type | Different chunking and prompts per type |
Extract filters from the question when you can (a small LLM call works well) and apply them before or during the vector search.
Hybrid search and reranking
Why hybrid search
Dense vector search is good at meaning and paraphrase. Keyword search (BM25) is good at exact terms: product codes, error messages, names, clause numbers. Real questions need both. Hybrid search runs both and merges the results, commonly with reciprocal rank fusion. Most mature vector stores, and PostgreSQL with pgvector plus full-text search, support this. We compare the options in vector databases compared.
Why reranking
Retrieve broadly (say, the top 30–50 candidates), then use a cross-encoder reranker or an LLM to reorder them and keep the best handful for the prompt. Reranking is often the single biggest quality improvement for the least effort, because it fixes the “right document was retrieved at position 17” problem.
Other retrieval upgrades worth testing
- Query rewriting — turn a vague or conversational follow-up into a standalone search query.
- Multi-query — generate several phrasings and merge results for broad questions.
- Parent-document retrieval — match on small chunks, but give the model the larger surrounding section.
- Agentic retrieval — let the model search again when the first results are weak. Useful for complex questions, but it adds latency and cost. See agentic AI design patterns.
Generation: grounded answers with citations
- Instruct the model to answer only from the provided context and to say “I couldn’t find this in the documents” when the context doesn’t contain the answer. Test that it actually does.
- Cite at the passage level. Each claim should point to a document and section the user can open. Validate that cited IDs exist in the retrieved set before showing them.
- Show the sources in the UI, not just inline markers. Users trust answers they can check.
- Keep the prompt short and put the retrieved context in a clear, labelled block. Include the metadata (title, date) so the model can prefer recent sources.
- Handle conflicts explicitly. If two sources disagree, ask the model to say so rather than choose silently.
Access control: enforce it at retrieval time
This is the part teams most often get wrong. If a user cannot open a document in the source system, the RAG system must never retrieve it for them.
- Store access groups on every chunk and filter in the search query, using the authenticated user’s identity.
- Never rely on the prompt (“don’t reveal HR documents”) for security.
- Sync permission changes from the source systems regularly, and handle deletions quickly.
- Log who asked what and which documents were used, with retention rules that match your privacy policy. Our GDPR and AI Act checklist covers the legal side, and private RAG for business covers keeping the whole stack on infrastructure you control.
Evaluate retrieval and answers separately
Measure the two halves of the system on their own, so you know which one to fix.
| Stage | What to measure | How |
|---|---|---|
| Retrieval | Is the right passage in the top k? (recall@k, MRR) | Compare retrieved IDs to your labelled answers |
| Answer faithfulness | Is every claim supported by the retrieved context? | LLM-as-judge with spot checks by a human |
| Answer relevance | Does it actually answer the question? | LLM-as-judge plus user feedback |
| Citations | Do cited passages support the claim? | Automated check plus sampling |
| Refusals | Does it say “not found” when it should? | Include unanswerable questions in the set |
| Latency and cost | p50/p95 response time, cost per answer | Tracing |
Open-source tools such as Ragas, DeepEval and promptfoo can automate much of this. Run the evaluation on every change to chunking, embeddings, prompts or models. The broader method is in our LLM evaluation guide.
Operations: keep the index fresh
- Incremental indexing triggered by changes in the source system, not a nightly full rebuild.
- Version your pipeline. Record which parser, chunker and embedding model produced each chunk, so you can re-index cleanly when you change one.
- Re-embed everything when you switch embedding models. Vectors from different models are not comparable.
- Monitor “no answer” and thumbs-down rates by source and topic; they point at missing or badly parsed documents.
- Cache frequent questions and embeddings to cut cost and latency.
RAG best practices checklist
- 50–100 real questions with labelled source passages
- Parser output checked by eye, tables preserved
- Structure-aware chunking with heading context
- Rich metadata: dates, versions, access groups, section paths
- Hybrid search (vector + keyword) with a reranker
- Permission filters applied in the retrieval query
- Answers grounded in context with passage-level citations
- “Not found” behaviour tested with unanswerable questions
- Retrieval and answer quality measured separately, in CI
- Incremental indexing and monitoring in production
When RAG is not the answer
RAG is for knowledge that lives in documents and changes over time. If you need the model to adopt a style or format, or to perform a narrow task consistently, fine-tuning may fit better; if the data is structured, a text-to-SQL or API tool may beat document retrieval. We cover the trade-offs in fine-tuning vs RAG.
Getting a RAG system into production
A working demo takes days; a system people trust takes careful retrieval engineering, permissions and evaluation. That is the focus of our private RAG service: we start from your real questions and documents, build an evaluation set, and improve retrieval step by step with measurable results. If you want to discuss your documents and use case, contact us.