"Should we use RAG or fine-tune a model?" comes up early in most AI projects we're involved in. It's a reasonable question with a slightly misleading framing: the two techniques solve different problems, and many features need neither, or both.
This guide explains what each technique is good at, where each one goes wrong, and a practical order of work that keeps you from spending weeks on the wrong approach.
Start with the problem, not the technique
Before choosing an approach, be specific about what the model is getting wrong today, with a good prompt and a capable model:
- It doesn't know things: your products, your policies, this week's data, a customer's history. That's a knowledge problem.
- It knows enough but behaves wrong: wrong format, wrong tone, inconsistent judgment on a narrow task. That's a behavior problem.
RAG addresses knowledge. Fine-tuning addresses behavior. Mixing them up is the most common source of wasted effort.
Write the problem down as a sentence, with examples. "The assistant gives outdated return-policy answers" is a knowledge problem. "The classifier puts refund requests in the billing queue about one time in ten" is a behavior problem. The sentence usually makes the choice obvious.
What RAG does well
Retrieval-augmented generation fetches relevant information at request time (documents, records, search results) and gives it to the model alongside the question. The model answers using that context.
RAG is usually the right starting point when:
- The information changes. Update the source, and answers update immediately.
- You need citations. Answers can point to the documents they came from.
- Access varies by user. Retrieval can be filtered to what each user is allowed to see.
- You have a lot of material. Far more than fits in a prompt.
The quality of a RAG system depends mostly on retrieval, not on the model. Poor chunking, weak search and missing metadata produce confident answers based on the wrong documents. Most of the engineering effort in a good RAG system goes into getting the right context in front of the model.
// Filter by permission before retrieval, not after generation
const chunks = await search(query, {
filter: { tenantId: user.tenantId, visibility: { in: user.roles } },
limit: 8,
});Where RAG systems go wrong
The failures are predictable, which makes them fixable:
| Symptom | Usual cause | What helps |
|---|---|---|
| Answers cite the wrong document | Weak search, no metadata filters | Hybrid keyword and vector search, filters by product, date or type |
| Answers miss information that exists | Chunks split tables or sections apart | Chunk along document structure; keep headings with their content |
| Answers mix old and new policies | Outdated documents still indexed | Index ownership, expiry dates, removing superseded versions |
| Answers sound right but aren't supported | Model fills gaps when retrieval fails | Instruct it to say when context is missing; test for that case |
| Users see data they shouldn't | Access checked after retrieval, or not at all | Filter by permission inside the search query |
Measure retrieval separately from answers. For a set of test questions, record which documents should be retrieved, and check whether they are. If retrieval is wrong, no amount of prompt work on the answer step will fix it.
What fine-tuning does well
Fine-tuning trains a model further on your examples, changing how it behaves. It's a good fit when:
- You need a consistent output format or style that prompting can't hold reliably.
- The task is narrow and repeated at high volume, such as classification or extraction.
- You want a smaller, cheaper or faster model to match a larger one on one specific task.
Fine-tuning is a poor way to teach facts. The model may absorb some, but it can't cite them, you can't easily update them, and it will still make things up with confidence. It also adds operational work: training data curation, evaluation, and retraining when base models change.
What fine-tuning needs
Teams often underestimate the inputs. Before starting, you need:
- Hundreds or more high-quality examples of the exact task, reviewed by someone who knows what a correct answer looks like. Inconsistent examples teach inconsistent behavior.
- A held-out evaluation set the model never trains on, so you can measure real improvement.
- A plan for base-model changes. When the provider releases a new base model, you'll want to retrain and re-evaluate, or accept that your tuned model falls behind.
If collecting the examples sounds hard, that's useful information. The same examples are what you'd need to evaluate any approach, so the work isn't wasted either way.
A comparison
| RAG | Fine-tuning | |
|---|---|---|
| Solves | Missing or changing knowledge | Inconsistent behavior on a narrow task |
| Keeps up with new information | Yes, update the source | No, requires retraining |
| Citations | Yes | No |
| Per-user access control | Yes, at retrieval | No |
| Upfront effort | Retrieval pipeline and indexing | Curated training examples |
| Ongoing effort | Keeping sources and indexes healthy | Retraining and re-evaluation |
What about long context windows?
Modern models accept very long prompts, and it's tempting to skip retrieval and send everything. For a small, stable set of documents, that can be a perfectly good design: simpler to build and easy to reason about.
It stops working well when the material is large, changes often or differs by user. Long prompts cost more and respond more slowly, quality can drop when the relevant passage is buried in a lot of irrelevant text, and you still need a way to send each user only what they're allowed to see. At that point you're building retrieval anyway.
Try this first
Before either, spend a few days on:
- A strong prompt with examples. Few-shot examples fix many behavior problems outright.
- Structured output. Most major providers support schema-constrained output, which removes a whole class of formatting issues.
- A capable model. Test with a strong model first to learn what's possible, then optimize for cost.
- An evaluation set. Fifty to a few hundred real examples with expected outputs. Without this, you can't tell whether any change helped.
The evaluation set is the most valuable artifact in any AI project. It turns "it seems better" into a number, and it's what lets you change models, prompts or architecture later with confidence.
A simple decision rule
If the model gives wrong answers because it lacks information, add retrieval. If it has the information and still behaves inconsistently after good prompting and structured output, consider fine-tuning. If you can't tell which, build the evaluation set first.
Three examples
- An internal policy assistant. Staff ask about HR, travel and security policies that change a few times a year. This is a knowledge problem with per-department access. RAG, with documents owned by the teams that write them.
- Support ticket routing. Every ticket goes into one of a dozen queues, at high volume. The model knows enough; it needs consistent judgment. Start with a prompt, examples and structured output. If accuracy plateaus below what you need, a fine-tuned smaller model is a reasonable next step.
- A sales assistant that drafts proposals. It needs current product and pricing documents (knowledge) and your house style (behavior). Start with RAG and style examples in the prompt. Fine-tuning for style is rarely worth it until the rest works well.
When to combine them
Some systems benefit from both: RAG for knowledge, plus a fine-tuned model that's better at using retrieved context in your domain's format. This is worth considering once you have a working RAG system, a solid evaluation set and a specific, measured gap that prompting hasn't closed. It's rarely the right place to start.
The short version
If the model lacks knowledge, use retrieval. If it lacks consistency on a narrow task, consider fine-tuning. In both cases, start with a good prompt, a strong model and an evaluation set, and let measured results decide what comes next.