Fine-Tuning vs RAG vs Prompting: Which Should You Use?
Fine-tuning vs RAG vs prompt engineering: a practical decision guide with a comparison table, costs, risks and when to combine approaches.

Fine-tuning vs RAG is one of the first questions teams ask when an off-the-shelf model is not good enough. The short answer: start with prompting, add retrieval-augmented generation (RAG) when the model needs knowledge it does not have, and consider fine-tuning only when you need to change behaviour — format, style or a narrow skill — that prompting cannot reliably deliver. Most successful business applications we see use prompting plus RAG, and only a minority need fine-tuning.
The three options in one paragraph each
Prompting (prompt engineering). You change the instructions, examples and structure of what you send to the model. No training, no infrastructure. Modern models follow detailed instructions well, and a good system prompt with a few examples solves more problems than people expect. See our prompt engineering best practices.
RAG. At question time, you search your own documents, pick the most relevant passages and pass them to the model with the question. The model answers from that context and can cite it. Knowledge lives in your database, so updating it is as easy as re-indexing a document.
Fine-tuning. You train a model further on examples of inputs and desired outputs. The model’s weights change, so the behaviour is “baked in”. This can be done on a hosted model through a provider’s fine-tuning API (where offered) or on an open-weight model you run yourself.
Fine-tuning vs RAG vs prompting: comparison table
| Prompting | RAG | Fine-tuning | |
|---|---|---|---|
| What it changes | Instructions per request | Knowledge available per request | Model behaviour (weights) |
| Best for | Tasks the model can already do | Answering from your documents, fresh or private data | Consistent format, tone, narrow classification, smaller cheaper model |
| Upfront effort | Hours to days | Days to weeks | Weeks, plus data preparation |
| Data needed | A few examples | Your documents | Hundreds to thousands of high-quality labelled examples |
| Updating knowledge | Edit the prompt | Re-index documents (minutes) | Retrain |
| Citations / traceability | No | Yes, answers can link to sources | No |
| Hallucination control | Moderate | Strong when retrieval is good | Does not solve it on its own |
| Per-request cost | Grows with prompt length | Adds retrieval + context tokens | Can be lower (shorter prompts, smaller model) |
| Access control per user | N/A | Yes, filter documents per user | No — everything trained is in the model |
| Vendor lock-in | Low | Low | Medium to high |
When prompting is enough
Try a well-structured prompt first if:
- The task is general (summarising, drafting, extracting fields, classifying into a handful of categories).
- You can describe what “good” looks like in plain language and a few examples.
- The knowledge needed is already general knowledge, or fits comfortably in the prompt.
Measure it before moving on. Many teams jump to fine-tuning because the first prompt was vague, not because prompting hit a ceiling.
When you need RAG
Choose RAG when the problem is the model does not know your stuff:
- Internal policies, product manuals, contracts, tickets, wikis.
- Information that changes weekly or daily.
- Answers must cite a source so staff or customers can verify them.
- Different users may see different documents (permissions).
- Data must stay in infrastructure you control — a private RAG setup can keep documents and the vector database on your own servers.
A common misconception is that fine-tuning is the way to “teach the model our documents”. It is a poor fit: the model does not reliably memorise facts from training examples, you cannot easily remove or update a fact, you lose citations, and per-user permissions become impossible. Retrieval handles knowledge better. Our RAG best practices article goes deeper into chunking, hybrid search and evaluation.
When fine-tuning is worth it
Fine-tuning earns its cost when the problem is behaviour, not knowledge:
- Strict, consistent output format that prompting gets right most but not all of the time, at high volume.
- A particular voice or style that is hard to describe but easy to show with many examples.
- Narrow, high-volume tasks (for example, classifying support tickets into your own taxonomy) where a fine-tuned small model can match a large general model at lower cost and latency.
- Shortening prompts: if every request carries a long list of instructions and examples, fine-tuning can move that into the model, reducing input tokens.
- Domain language in specialised fields where the base model misreads terms, although RAG with good glossaries often helps too.
Before you commit, check:
- Do you have the data? You need a clean, representative set of examples, reviewed by someone who knows the domain. Garbage in, confident garbage out.
- Can you evaluate it? Keep a held-out test set and compare the fine-tuned model against a strong prompted baseline. Our LLM evaluation guide explains how.
- Who maintains it? When the base model is upgraded or deprecated, you may need to fine-tune again. Budget for that.
- Where does the training data go? If examples contain personal data, fine-tuning on a third-party platform is a data-processing decision with privacy implications. See our GDPR and AI Act checklist.
Combining approaches
These are not mutually exclusive. Common combinations:
- Prompting + RAG: the default for internal knowledge assistants and support chatbots.
- RAG + fine-tuned model: retrieval supplies the facts; a fine-tuned model reliably produces your format or tone and makes better use of retrieved context.
- Fine-tuned small model as a router or classifier in front of a larger prompted model, to cut costs (see LLM cost optimization).
Common mistakes we see
- Fine-tuning to add facts. The model may repeat some of them, but it will also blend them with what it already “knows”, and you cannot trace where an answer came from.
- Skipping the baseline. Without a well-prompted baseline measured on the same test set, you cannot tell whether fine-tuning helped at all.
- Blaming the model for bad retrieval. When a RAG system gives wrong answers, the cause is usually chunking, missing metadata or weak search, not the language model. Inspect what was retrieved before changing anything else.
- Stuffing everything into the prompt. Long context windows make it tempting to skip retrieval, but you pay for every token on every request and quality can drop when the relevant passage is buried.
- Forgetting maintenance. Prompts, indexes and fine-tuned models all drift as your data and the underlying models change. Plan for regular re-evaluation.
A simple decision flow
- Write a clear prompt with examples. Build a small evaluation set. Measure.
- Are failures caused by missing or outdated knowledge? Add RAG.
- Are failures caused by inconsistent format, tone or a narrow skill, and do you have hundreds of good examples? Consider fine-tuning.
- Is the main problem cost or latency at high volume? Try a smaller model with a better prompt first, then consider fine-tuning that smaller model.
- Re-measure after every change. Keep the simplest approach that meets your quality bar.
Key takeaways
- Prompting first. It is cheap, fast to change and often enough.
- RAG is for knowledge: private, changing or citable information, with per-user permissions.
- Fine-tuning is for behaviour: format, style, narrow tasks, or making a small model do a big model’s job.
- Fine-tuning does not replace retrieval for facts and does not by itself stop hallucinations.
- Always compare against a strong baseline on the same evaluation set.
Choosing the right approach for your project
If you are not sure which approach fits your data and use case, we are happy to look at a sample of your documents and requirements and recommend the simplest option that will work. Our AI solutions cover prompting, RAG pipelines and fine-tuning when it is justified. Contact us to talk it through.