Intermediate 12 minStrategy

Fine-tuning vs. RAG: which should you choose for your use case ?

Fine-tuning or RAG: the question comes up every time you want to specialize a local LLM for a domain, style, or business data. The two approaches solve different problems, and choosing the wrong one is costly in GPU time and maintenance. This guide gives you the criteria to decide quickly and explains why the right answer is often “both.”

By Mohamed Meguedmi·Update 2026-08-27·Tested on Windows, macOS, and Linux

#Why this question keeps coming up

You installed Ollama, chose a 7B or 14B model, and now want it to “know” your domain: your internal documentation, industry terminology, case-law decisions, and support tickets. Two paths are available—fine-tuning or RAG—and the community often discusses them as if they were interchangeable alternatives. They aren’t.

The trap: fine-tuning has an aura of “real AI,” so people imagine a model becoming an expert. RAG looks improvised, like automated “copy and paste.” The industrial reality is the opposite—RAG has become the standard for 80% of enterprise use cases, while fine-tuning is reserved for specific problems where it provides what RAG cannot.

i
The one-sentence summary
RAG = giving the model dynamic access to external knowledge. Fine-tuning = changing the model's intrinsic behavior (style, format, reasoning, specialized language).

#The two approaches in 1 minute

The Local RAG Kit

Your documents, your AI: a reliable local RAG over your PDFs, notes and mail — nothing leaves your machine.

  • Lifetime online access
  • PDF + files
  • Lifetime updates

#RAG (Retrieval-Augmented Generation)

RAG indexes your documents (PDF, Markdown, database, code) in a vector database (Chroma, Qdrant, Weaviate). When you ask a question, the system finds relevant passages by semantic similarity, injects them into the prompt, and the LLM generates its answer based on them. The model remains generic; the context becomes specialized.

What this changes
The model can cite your data, provide sources for its answers, and access knowledge it did not have during training.
What this doesn’t change
Response style, tone, ability to reason in a precise format, and highly specialized technical vocabulary.
Marginal cost of adding a document
A few seconds: we're indexing the new document, that's all.

#Fine-tuning (LoRA / QLoRA / full)

Fine-tuning retrains the model (or some of its weights, through LoRA) on a dataset of input/output pairs representative of what you want it to do. The knowledge and behavior are absorbed into the model's weights.

What this changes
Default behavior: style, output format, conventions, deep domain vocabulary, implicit reasoning.
What this does not change (for the better)
Access to fresh factual information — a model fine-tuned in March on your procedures won't know about procedures written in April.
Marginal cost of adding a document
A new full training run for every significant dataset update.
→
Quick thought experiment
If the question is "how can the model know X?", the answer is almost always RAG. If the question is "how can the model answer like that?", it is probably fine-tuning.

#Decision matrix

Rather than saying “it depends,” here are the criteria that really settle the question, line by line.

Factual knowledge that changes frequently
RAG. Fine-tuning is obsolete as soon as the dataset is updated. All product documentation, ticket databases, FAQs, and case law fall into this category.
Specific style, tone, output format
Fine-tuning. No number of examples in a prompt can replace 500 well-constructed training pairs for anchoring a strict JSON format, a corporate tone, or a report structure.
Extremely specialized business vocabulary
Fine-tuning, especially when pretraining provides poor coverage of the language (French legal, medical, or dialectal language). RAG is not enough if the model does not understand the terms to begin with.
Source traceability and citation
RAG. You can display “according to document X, paragraph Y.” With a fine-tuned model, it is impossible to prove where a claim came from.
Ultra-confidential data, never in shared RAM
Fine-tuning with weights stored locally. RAG requires injecting passages into the context for every query—which can be a problem on shared infrastructure.
Fast response in under 200ms (chatbot, inline agent)
Fine-tuning. RAG adds 100–500 ms of retrieval plus a longer context to process. For critical real-time workloads, that matters.
Huge knowledge volume (> 100k pages)
RAG. You cannot reasonably fine-tune on 100k documents—and even if you did, the model would hallucinate details.
High-quality training dataset available
If you don't have at least 500–1,000 clean input/output pairs, fine-tuning will do more to degrade the model than improve it. Start with RAG.

#5 typical use cases

#1. Support chatbot for product documentation

Verdict
RAG, without hesitation.
Why
The documentation changes constantly (new features, fixes, deprecations). A fine-tune would be obsolete in 3 weeks, and the client wants a sourced answer ("see section X of the manual"), not an opaque assertion.
Typical stack
Ollama (Qwen 3.5 9B to fit in 8 GB, or Mistral Small 24B in 16 GB for more polished French) + Qdrant/Chroma + nomic-embed-text + Open WebUI or AnythingLLM.

#2. Structured information extractor (invoices, resumes, contracts)

Verdict
Fine-tuning, or advanced prompting with JSON mode.
Why
The output format must be exactly the same every time (same fields, same types, same default values). Even a well-written prompt drifts in 5% of cases, which breaks a pipeline. A LoRA trained on 800 annotated examples fixes this permanently.
Typical stack
Unsloth or Axolotl for training, GGUF export, Ollama deployment. If the “knowledge” base covering document types evolves, you can combine it with lightweight RAG.

#3. Legal assistant for French case law

Verdict
Hybrid — RAG first, then fine-tuning if the vocabulary remains limited.
Why
The case-law corpus is enormous and changes every month (RAG is mandatory for up-to-date decisions). But French legal vocabulary is poorly covered by most open-weight models, and a lightweight fine-tune (LoRA on 2–3,000 examples of legal Q&A) significantly improves understanding of the terminology before RAG comes into play.
Typical stack
Légifrance/Doctrine as the source → Qdrant + BGE reranker → LoRA fine-tuned 14B LLM on French legal jargon.

#4. Code generator adapted to an internal codebase

Verdict
RAG (reading repo files), with no fine-tuning except in very specific cases.
Why
A codebase changes every day. A fine-tune would be outdated every sprint. The best coding assistants (Continue.dev, Aider) dynamically read the relevant files using RAG over the AST or code embeddings.
Typical stack
Continue.dev + Qwen3-Coder 30B-A3B (qwen3-coder:30b, MoE 256k ctx, 3B active, so fast) or Devstral 24B via Ollama, with retrieval integrated into the plugin.

#5. In-house editorial style (newsletter, reports, product sheets)

Verdict
Pure fine-tuning.
Why
The content is new every time (nothing to “retrieve”), but the tone, structure, sentence rhythm, and use of “you” and subheadings must be absolutely consistent. That is exactly what fine-tuning handles well.
Typical stack
200–500 well-written articles → Alpaca or ChatML format → QLoRA on Qwen 3.5 9B with Unsloth → GGUF export, modelfile Ollama with an additional system prompt.

#The hidden costs of the two approaches

Public comparisons often boil down to "the cost of a GPU for 4 hours of training." Operational reality is tougher on both sides.

#Hidden costs of RAG

Chunking quality
Document splitting affects everything. If it's done poorly, retrieval returns out-of-context snippets, the LLM hallucinates, and no one understands why. It's rarely "just put the PDFs in Chroma": you often need section-based chunking, and sometimes OCR preprocessing for scans.
Cumulative latency
Query embedding + vector search + (optional) reranker + extended context for the LLM. On a poorly optimized setup, this can go from 400ms (LLM alone) to 2–3s (full RAG). Budget for it from the design stage.
Database maintenance
When a document is deleted or updated, you need to remove it or reindex it. For external sources (web, API), plan a refresh job. The more the stack evolves, the more work this requires.
Embedding model quality
For French, the default embeddings (text-embedding-ada like) are mediocre. nomic-embed-text, BGE-M3, or Solon make a difference—but that means knowing the options.

#Hidden costs of fine-tuning

Dataset preparation
This is 80% of the work. Collecting, cleaning, formatting into instruction/response pairs, deduplicating, and balancing classes. In a 4-week fine-tuning project, plan on 3 weeks of data preparation and 1 week of training.
Risk of regression
A poorly calibrated fine-tune degrades the model's general capabilities ("catastrophic forgetting"). The model becomes good at your task and useless at everything else. You need to test it on a general benchmark before and after.
Retraining with every change
The dataset grows over time. Each release requires rerunning 2–12 hours of GPU work, revalidating, and redeploying. Versioning datasets and checkpoints becomes mandatory.
Training hardware
Inferring with a 7B Q4 requires 5 GB of VRAM, but training it (even with QLoRA) requires at least 12–16 GB. Fine-tuning has a higher hardware entry threshold than inference.
!
Common underestimation
Teams getting started overestimate the cost of RAG (“you have to index everything”) and underestimate the cost of fine-tuning (“we have 200 examples, that will be enough”). Reality is the opposite: a basic RAG setup takes 2 days, while a useful fine-tune requires 2 to 4 weeks full-time.

#The hybrid approach: RAG + fine-tuning

The most effective architectures don’t choose—they combine. Fine-tuning defines how the model talks about your domain; RAG gives it access to what it needs to know at a given time.

Fine-tune for style and format
200–1000 examples that anchor the tone (corporate, technical, legal), response format (JSON, structured Markdown), and stance (always cite the source, never make anything up).
RAG for factual knowledge
Documentation, ticket database, case law, codebase—anything that changes and needs to be findable and citable.
Guardrails in the system prompt
The system prompt reminds the model to refuse to answer if the RAG context is empty or contradictory. Essential for limiting hallucinations.
→
Implementation order
ALWAYS start with RAG alone, using a good system prompt. Measure. If the quality is insufficient (wrong tone, inconsistent format, poor command of the domain vocabulary), then add targeted fine-tuning for the identified shortcomings. Doing the reverse wastes weeks.

#Quick decision in 3 questions

  1. 01
    Question 1 — Do your data change more than once a month?
    If so: RAG is mandatory. Fine-tuning cannot keep pace without becoming an operational nightmare.
  2. 02
    Question 2 — Do you have at least 500 high-quality input/output pairs verified by humans?
    If not, start with RAG. Fine-tuning on 100 hastily assembled examples degrades the model. If you plan to accumulate more, set up RAG first and use its logs to build the dataset.
  3. 03
    Question 3 — Is the problem “knowing something” or “responding in a certain way”?
    Knowing → RAG. Answering in a particular way → fine-tuning. Both → hybrid. It is the simplest framework, and it is right 9 times out of 10.

#Common pitfalls to avoid

"Fine-tune to learn facts"
Most common mistake. A fine-tune is not a knowledge base—it interpolates from the examples it has seen but hallucinates on precise details (numbers, dates, references) far more than RAG.
“RAG without a reranker” on corpora > 10k chunks
Vector search alone retrieves volume but not always relevance. A cross-encoder reranker (BGE, mxbai) on the top 20 transforms quality, at an additional cost of 50-100ms.
Confusing RAG with long context
“I’ll just put the entire document in the prompt.” Beyond 8k useful tokens, quality drops sharply (“lost in the middle”). A well-chunked RAG system beats a naive long context once you reach a certain volume.
Fine-tune an already aligned model with an unaligned dataset
If you fine-tune an “instruct” model with examples that do not use the same prompt structure, you break alignment and the model becomes erratic. Always follow the model’s template (ChatML, Alpaca, Mistral, Llama-3 chat).
Wanting a public benchmark to decide
No generic benchmark will tell you whether YOUR use case is better suited to RAG or fine-tuning. Set up an internal eval with 30-50 representative questions, and measure before moving to production.

#Go further

Once the decision is made, the corresponding guides cover the practical implementation, including pitfalls and optimizations.

Implementing RAG without coding
Open WebUI or AnythingLLM let you set up a document RAG system in a few hours, without writing Python—useful for validating the approach before industrializing it.
Local LoRA / QLoRA fine-tuning
The dedicated guide covers Unsloth, dataset format, choosing between LoRA and QLoRA, and GGUF export for Ollama. One RTX 3090 is enough for a 7B model.
Optimize an existing RAG
Reranking, hybrid BM25 + vector search, and chunking strategies—three levers that take a RAG system from "works" to production.
Did this guide help you?

Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.