Skip to content
NewSecurity Audit for AI-Built Apps. Fixed scope, clear report
GeekTech
AI Engineering

RAG vs Fine-Tuning: How to Choose

When retrieval-augmented generation is the right approach, when fine-tuning earns its cost, and why most business AI features should start with neither.

Updated 6 min read

GeekTech Engineering Team18 years building and securing software

"Should we use RAG or fine-tune a model?" comes up early in most AI projects we're involved in. It's a reasonable question with a slightly misleading framing: the two techniques solve different problems, and many features need neither, or both.

This guide explains what each technique is good at, where each one goes wrong, and a practical order of work that keeps you from spending weeks on the wrong approach.

Start with the problem, not the technique

Before choosing an approach, be specific about what the model is getting wrong today, with a good prompt and a capable model:

  • It doesn't know things: your products, your policies, this week's data, a customer's history. That's a knowledge problem.
  • It knows enough but behaves wrong: wrong format, wrong tone, inconsistent judgment on a narrow task. That's a behavior problem.

RAG addresses knowledge. Fine-tuning addresses behavior. Mixing them up is the most common source of wasted effort.

Write the problem down as a sentence, with examples. "The assistant gives outdated return-policy answers" is a knowledge problem. "The classifier puts refund requests in the billing queue about one time in ten" is a behavior problem. The sentence usually makes the choice obvious.

What RAG does well

Retrieval-augmented generation fetches relevant information at request time (documents, records, search results) and gives it to the model alongside the question. The model answers using that context.

RAG is usually the right starting point when:

  • The information changes. Update the source, and answers update immediately.
  • You need citations. Answers can point to the documents they came from.
  • Access varies by user. Retrieval can be filtered to what each user is allowed to see.
  • You have a lot of material. Far more than fits in a prompt.

The quality of a RAG system depends mostly on retrieval, not on the model. Poor chunking, weak search and missing metadata produce confident answers based on the wrong documents. Most of the engineering effort in a good RAG system goes into getting the right context in front of the model.

// Filter by permission before retrieval, not after generation
const chunks = await search(query, {
  filter: { tenantId: user.tenantId, visibility: { in: user.roles } },
  limit: 8,
});

Where RAG systems go wrong

The failures are predictable, which makes them fixable:

SymptomUsual causeWhat helps
Answers cite the wrong documentWeak search, no metadata filtersHybrid keyword and vector search, filters by product, date or type
Answers miss information that existsChunks split tables or sections apartChunk along document structure; keep headings with their content
Answers mix old and new policiesOutdated documents still indexedIndex ownership, expiry dates, removing superseded versions
Answers sound right but aren't supportedModel fills gaps when retrieval failsInstruct it to say when context is missing; test for that case
Users see data they shouldn'tAccess checked after retrieval, or not at allFilter by permission inside the search query

Measure retrieval separately from answers. For a set of test questions, record which documents should be retrieved, and check whether they are. If retrieval is wrong, no amount of prompt work on the answer step will fix it.

What fine-tuning does well

Fine-tuning trains a model further on your examples, changing how it behaves. It's a good fit when:

  • You need a consistent output format or style that prompting can't hold reliably.
  • The task is narrow and repeated at high volume, such as classification or extraction.
  • You want a smaller, cheaper or faster model to match a larger one on one specific task.

Fine-tuning is a poor way to teach facts. The model may absorb some, but it can't cite them, you can't easily update them, and it will still make things up with confidence. It also adds operational work: training data curation, evaluation, and retraining when base models change.

What fine-tuning needs

Teams often underestimate the inputs. Before starting, you need:

  • Hundreds or more high-quality examples of the exact task, reviewed by someone who knows what a correct answer looks like. Inconsistent examples teach inconsistent behavior.
  • A held-out evaluation set the model never trains on, so you can measure real improvement.
  • A plan for base-model changes. When the provider releases a new base model, you'll want to retrain and re-evaluate, or accept that your tuned model falls behind.

If collecting the examples sounds hard, that's useful information. The same examples are what you'd need to evaluate any approach, so the work isn't wasted either way.

A comparison

RAGFine-tuning
SolvesMissing or changing knowledgeInconsistent behavior on a narrow task
Keeps up with new informationYes, update the sourceNo, requires retraining
CitationsYesNo
Per-user access controlYes, at retrievalNo
Upfront effortRetrieval pipeline and indexingCurated training examples
Ongoing effortKeeping sources and indexes healthyRetraining and re-evaluation

What about long context windows?

Modern models accept very long prompts, and it's tempting to skip retrieval and send everything. For a small, stable set of documents, that can be a perfectly good design: simpler to build and easy to reason about.

It stops working well when the material is large, changes often or differs by user. Long prompts cost more and respond more slowly, quality can drop when the relevant passage is buried in a lot of irrelevant text, and you still need a way to send each user only what they're allowed to see. At that point you're building retrieval anyway.

Try this first

Before either, spend a few days on:

  1. A strong prompt with examples. Few-shot examples fix many behavior problems outright.
  2. Structured output. Most major providers support schema-constrained output, which removes a whole class of formatting issues.
  3. A capable model. Test with a strong model first to learn what's possible, then optimize for cost.
  4. An evaluation set. Fifty to a few hundred real examples with expected outputs. Without this, you can't tell whether any change helped.

The evaluation set is the most valuable artifact in any AI project. It turns "it seems better" into a number, and it's what lets you change models, prompts or architecture later with confidence.

A simple decision rule

If the model gives wrong answers because it lacks information, add retrieval. If it has the information and still behaves inconsistently after good prompting and structured output, consider fine-tuning. If you can't tell which, build the evaluation set first.

Three examples

  • An internal policy assistant. Staff ask about HR, travel and security policies that change a few times a year. This is a knowledge problem with per-department access. RAG, with documents owned by the teams that write them.
  • Support ticket routing. Every ticket goes into one of a dozen queues, at high volume. The model knows enough; it needs consistent judgment. Start with a prompt, examples and structured output. If accuracy plateaus below what you need, a fine-tuned smaller model is a reasonable next step.
  • A sales assistant that drafts proposals. It needs current product and pricing documents (knowledge) and your house style (behavior). Start with RAG and style examples in the prompt. Fine-tuning for style is rarely worth it until the rest works well.

When to combine them

Some systems benefit from both: RAG for knowledge, plus a fine-tuned model that's better at using retrieved context in your domain's format. This is worth considering once you have a working RAG system, a solid evaluation set and a specific, measured gap that prompting hasn't closed. It's rarely the right place to start.

The short version

If the model lacks knowledge, use retrieval. If it lacks consistency on a narrow task, consider fine-tuning. In both cases, start with a good prompt, a strong model and an evaluation set, and let measured results decide what comes next.

Insights

Related insights

More on AI Engineering →
  • AI Engineering

    When Does a Company Need a Fractional AI CTO?

    What a fractional AI CTO does, the signs you need one, how it compares with hiring or consultants, and how to set up the engagement so it delivers.

    5 min read

  • AI Governance

    How to Write a Company AI Use Policy

    A practical structure for an AI acceptable-use policy that staff will read: approved tools, data rules, review duties, disclosure and new-tool requests.

    5 min read

  • AI Governance

    EU AI Act: What SMBs Actually Need to Know

    A practical guide to the EU AI Act for small and mid-sized companies: which rules apply to you, what to do first, and what you can safely deprioritize.

    6 min read

Get practical AI insights, monthly

AI security, automation and governance, written by engineers. No spam, unsubscribe any time.

Book a free 30-min call