AI Solutions

RAG Best Practices: Lessons From Production Retrieval Systems

RAG best practices from real production systems: chunking, metadata, hybrid search, reranking, evaluation, access control and citations users can trust.

GPTLabAI team 7 min read

The most important RAG best practices are unglamorous: clean and well-structured source documents, sensible chunking with rich metadata, hybrid search plus a reranker, access control enforced at retrieval time, answers that cite their sources, and an evaluation set that tells you whether a change helped. The language model is rarely the bottleneck. Retrieval is.

These are the lessons we apply when building retrieval-augmented generation systems for teams that need answers from manuals, contracts, tickets and wikis.

Start with the questions, not the documents

Before indexing anything, collect 50–100 real questions from the people who will use the system, and for each one note the document (and ideally the passage) that answers it. This list does three jobs:

  • it shows which sources actually matter,
  • it reveals question types you did not expect (comparisons, “latest version of…”, yes/no policy checks),
  • it becomes your evaluation set.

Teams that skip this step end up tuning by feel, and every change is an argument.

Document preparation: garbage in, garbage out

Parse properly

PDFs, slides and scanned documents are where most quality is lost. Use a parser that keeps headings, lists and tables, and check a sample of the output by eye. Tables flattened into a stream of numbers produce confident wrong answers.

Clean and deduplicate

Remove headers, footers, navigation text and boilerplate that appears on every page. Deduplicate near-identical versions, or tag them so you can prefer the latest.

Keep structure

Store the document title, section heading path (“Handbook > Leave > Parental leave”), page number and URL with every chunk. You will need them for filtering, ranking and citations.

Chunking that respects meaning

There is no universal chunk size. What works:

  • Split on structure first — headings, sections, list items, table rows — then by length.
  • Keep chunks self-contained. A chunk that says “see the table above” is useless on its own. Prepend the heading path or a one-line document summary so each chunk carries its context.
  • Use modest overlap so sentences at boundaries are not lost.
  • Treat special content differently. FAQs work well as one question-answer pair per chunk; code works best split by function or class; tables can be stored whole or row by row with the header repeated.
  • Test two or three chunking strategies against your evaluation set rather than guessing.

Metadata is a retrieval feature

Metadata is how you answer “what is the current policy for Germany?” rather than returning a 2021 policy for another country. Useful fields:

Metadata Used for
Source, title, URL, page Citations and deep links
Section path Context for the model, better ranking
Date, version, status Prefer latest, exclude drafts and superseded docs
Product, region, department, language Filters from user context or the question
Access groups / owner Permission-aware retrieval
Document type Different chunking and prompts per type

Extract filters from the question when you can (a small LLM call works well) and apply them before or during the vector search.

Hybrid search and reranking

Dense vector search is good at meaning and paraphrase. Keyword search (BM25) is good at exact terms: product codes, error messages, names, clause numbers. Real questions need both. Hybrid search runs both and merges the results, commonly with reciprocal rank fusion. Most mature vector stores, and PostgreSQL with pgvector plus full-text search, support this. We compare the options in vector databases compared.

Why reranking

Retrieve broadly (say, the top 30–50 candidates), then use a cross-encoder reranker or an LLM to reorder them and keep the best handful for the prompt. Reranking is often the single biggest quality improvement for the least effort, because it fixes the “right document was retrieved at position 17” problem.

Other retrieval upgrades worth testing

  • Query rewriting — turn a vague or conversational follow-up into a standalone search query.
  • Multi-query — generate several phrasings and merge results for broad questions.
  • Parent-document retrieval — match on small chunks, but give the model the larger surrounding section.
  • Agentic retrieval — let the model search again when the first results are weak. Useful for complex questions, but it adds latency and cost. See agentic AI design patterns.

Generation: grounded answers with citations

  • Instruct the model to answer only from the provided context and to say “I couldn’t find this in the documents” when the context doesn’t contain the answer. Test that it actually does.
  • Cite at the passage level. Each claim should point to a document and section the user can open. Validate that cited IDs exist in the retrieved set before showing them.
  • Show the sources in the UI, not just inline markers. Users trust answers they can check.
  • Keep the prompt short and put the retrieved context in a clear, labelled block. Include the metadata (title, date) so the model can prefer recent sources.
  • Handle conflicts explicitly. If two sources disagree, ask the model to say so rather than choose silently.

Access control: enforce it at retrieval time

This is the part teams most often get wrong. If a user cannot open a document in the source system, the RAG system must never retrieve it for them.

  • Store access groups on every chunk and filter in the search query, using the authenticated user’s identity.
  • Never rely on the prompt (“don’t reveal HR documents”) for security.
  • Sync permission changes from the source systems regularly, and handle deletions quickly.
  • Log who asked what and which documents were used, with retention rules that match your privacy policy. Our GDPR and AI Act checklist covers the legal side, and private RAG for business covers keeping the whole stack on infrastructure you control.

Evaluate retrieval and answers separately

Measure the two halves of the system on their own, so you know which one to fix.

Stage What to measure How
Retrieval Is the right passage in the top k? (recall@k, MRR) Compare retrieved IDs to your labelled answers
Answer faithfulness Is every claim supported by the retrieved context? LLM-as-judge with spot checks by a human
Answer relevance Does it actually answer the question? LLM-as-judge plus user feedback
Citations Do cited passages support the claim? Automated check plus sampling
Refusals Does it say “not found” when it should? Include unanswerable questions in the set
Latency and cost p50/p95 response time, cost per answer Tracing

Open-source tools such as Ragas, DeepEval and promptfoo can automate much of this. Run the evaluation on every change to chunking, embeddings, prompts or models. The broader method is in our LLM evaluation guide.

Operations: keep the index fresh

  • Incremental indexing triggered by changes in the source system, not a nightly full rebuild.
  • Version your pipeline. Record which parser, chunker and embedding model produced each chunk, so you can re-index cleanly when you change one.
  • Re-embed everything when you switch embedding models. Vectors from different models are not comparable.
  • Monitor “no answer” and thumbs-down rates by source and topic; they point at missing or badly parsed documents.
  • Cache frequent questions and embeddings to cut cost and latency.

RAG best practices checklist

  • 50–100 real questions with labelled source passages
  • Parser output checked by eye, tables preserved
  • Structure-aware chunking with heading context
  • Rich metadata: dates, versions, access groups, section paths
  • Hybrid search (vector + keyword) with a reranker
  • Permission filters applied in the retrieval query
  • Answers grounded in context with passage-level citations
  • “Not found” behaviour tested with unanswerable questions
  • Retrieval and answer quality measured separately, in CI
  • Incremental indexing and monitoring in production

When RAG is not the answer

RAG is for knowledge that lives in documents and changes over time. If you need the model to adopt a style or format, or to perform a narrow task consistently, fine-tuning may fit better; if the data is structured, a text-to-SQL or API tool may beat document retrieval. We cover the trade-offs in fine-tuning vs RAG.

Getting a RAG system into production

A working demo takes days; a system people trust takes careful retrieval engineering, permissions and evaluation. That is the focus of our private RAG service: we start from your real questions and documents, build an evaluation set, and improve retrieval step by step with measurable results. If you want to discuss your documents and use case, contact us.

7 min

How Much Does an AI Chatbot for Business Cost?

What drives AI chatbot cost for a business: build, model API usage, hosting and maintenance, build vs buy, plus a simple formula to estimate your token costs.

Read article

Have a project in mind? Let’s talk.

Whether you run a business or a research group, tell us what you need built, fixed or evaluated. You get a free consultation and a clear written estimate — no obligation.

  • Free consultation
  • Written scope and estimate
  • We reply within one working day
Contact us