Fine-tuning vs. RAG: which should you choose for your use case ?
Fine-tuning or RAG: the question comes up every time you want to specialize a local LLM for a domain, style, or business data. The two approaches solve different problems, and choosing the wrong one is costly in GPU time and maintenance. This guide gives you the criteria to decide quickly and explains why the right answer is often “both.”
#Why this question keeps coming up
You installed Ollama, chose a 7B or 14B model, and now want it to “know” your domain: your internal documentation, industry terminology, case-law decisions, and support tickets. Two paths are available—fine-tuning or RAG—and the community often discusses them as if they were interchangeable alternatives. They aren’t.
The trap: fine-tuning has an aura of “real AI,” so people imagine a model becoming an expert. RAG looks improvised, like automated “copy and paste.” The industrial reality is the opposite—RAG has become the standard for 80% of enterprise use cases, while fine-tuning is reserved for specific problems where it provides what RAG cannot.
#The two approaches in 1 minute
Your documents, your AI: a reliable local RAG over your PDFs, notes and mail — nothing leaves your machine.
- Lifetime online access
- PDF + files
- Lifetime updates
#RAG (Retrieval-Augmented Generation)
RAG indexes your documents (PDF, Markdown, database, code) in a vector database (Chroma, Qdrant, Weaviate). When you ask a question, the system finds relevant passages by semantic similarity, injects them into the prompt, and the LLM generates its answer based on them. The model remains generic; the context becomes specialized.
- What this changes
- The model can cite your data, provide sources for its answers, and access knowledge it did not have during training.
- What this doesn’t change
- Response style, tone, ability to reason in a precise format, and highly specialized technical vocabulary.
- Marginal cost of adding a document
- A few seconds: we're indexing the new document, that's all.
#Fine-tuning (LoRA / QLoRA / full)
Fine-tuning retrains the model (or some of its weights, through LoRA) on a dataset of input/output pairs representative of what you want it to do. The knowledge and behavior are absorbed into the model's weights.
- What this changes
- Default behavior: style, output format, conventions, deep domain vocabulary, implicit reasoning.
- What this does not change (for the better)
- Access to fresh factual information — a model fine-tuned in March on your procedures won't know about procedures written in April.
- Marginal cost of adding a document
- A new full training run for every significant dataset update.
#Decision matrix
Rather than saying “it depends,” here are the criteria that really settle the question, line by line.
- Factual knowledge that changes frequently
- RAG. Fine-tuning is obsolete as soon as the dataset is updated. All product documentation, ticket databases, FAQs, and case law fall into this category.
- Specific style, tone, output format
- Fine-tuning. No number of examples in a prompt can replace 500 well-constructed training pairs for anchoring a strict JSON format, a corporate tone, or a report structure.
- Extremely specialized business vocabulary
- Fine-tuning, especially when pretraining provides poor coverage of the language (French legal, medical, or dialectal language). RAG is not enough if the model does not understand the terms to begin with.
- Source traceability and citation
- RAG. You can display “according to document X, paragraph Y.” With a fine-tuned model, it is impossible to prove where a claim came from.
- Ultra-confidential data, never in shared RAM
- Fine-tuning with weights stored locally. RAG requires injecting passages into the context for every query—which can be a problem on shared infrastructure.
- Fast response in under 200ms (chatbot, inline agent)
- Fine-tuning. RAG adds 100–500 ms of retrieval plus a longer context to process. For critical real-time workloads, that matters.
- Huge knowledge volume (> 100k pages)
- RAG. You cannot reasonably fine-tune on 100k documents—and even if you did, the model would hallucinate details.
- High-quality training dataset available
- If you don't have at least 500–1,000 clean input/output pairs, fine-tuning will do more to degrade the model than improve it. Start with RAG.
#5 typical use cases
#1. Support chatbot for product documentation
- Verdict
- RAG, without hesitation.
- Why
- The documentation changes constantly (new features, fixes, deprecations). A fine-tune would be obsolete in 3 weeks, and the client wants a sourced answer ("see section X of the manual"), not an opaque assertion.
- Typical stack
- Ollama (Qwen 3.5 9B to fit in 8 GB, or Mistral Small 24B in 16 GB for more polished French) + Qdrant/Chroma + nomic-embed-text + Open WebUI or AnythingLLM.
#2. Structured information extractor (invoices, resumes, contracts)
- Verdict
- Fine-tuning, or advanced prompting with JSON mode.
- Why
- The output format must be exactly the same every time (same fields, same types, same default values). Even a well-written prompt drifts in 5% of cases, which breaks a pipeline. A LoRA trained on 800 annotated examples fixes this permanently.
- Typical stack
- Unsloth or Axolotl for training, GGUF export, Ollama deployment. If the “knowledge” base covering document types evolves, you can combine it with lightweight RAG.
#3. Legal assistant for French case law
- Verdict
- Hybrid — RAG first, then fine-tuning if the vocabulary remains limited.
- Why
- The case-law corpus is enormous and changes every month (RAG is mandatory for up-to-date decisions). But French legal vocabulary is poorly covered by most open-weight models, and a lightweight fine-tune (LoRA on 2–3,000 examples of legal Q&A) significantly improves understanding of the terminology before RAG comes into play.
- Typical stack
- Légifrance/Doctrine as the source → Qdrant + BGE reranker → LoRA fine-tuned 14B LLM on French legal jargon.
#4. Code generator adapted to an internal codebase
- Verdict
- RAG (reading repo files), with no fine-tuning except in very specific cases.
- Why
- A codebase changes every day. A fine-tune would be outdated every sprint. The best coding assistants (Continue.dev, Aider) dynamically read the relevant files using RAG over the AST or code embeddings.
- Typical stack
- Continue.dev + Qwen3-Coder 30B-A3B (qwen3-coder:30b, MoE 256k ctx, 3B active, so fast) or Devstral 24B via Ollama, with retrieval integrated into the plugin.
#5. In-house editorial style (newsletter, reports, product sheets)
- Verdict
- Pure fine-tuning.
- Why
- The content is new every time (nothing to “retrieve”), but the tone, structure, sentence rhythm, and use of “you” and subheadings must be absolutely consistent. That is exactly what fine-tuning handles well.
- Typical stack
- 200–500 well-written articles → Alpaca or ChatML format → QLoRA on Qwen 3.5 9B with Unsloth → GGUF export, modelfile Ollama with an additional system prompt.
#The hidden costs of the two approaches
Public comparisons often boil down to "the cost of a GPU for 4 hours of training." Operational reality is tougher on both sides.
#Hidden costs of RAG
- Chunking quality
- Document splitting affects everything. If it's done poorly, retrieval returns out-of-context snippets, the LLM hallucinates, and no one understands why. It's rarely "just put the PDFs in Chroma": you often need section-based chunking, and sometimes OCR preprocessing for scans.
- Cumulative latency
- Query embedding + vector search + (optional) reranker + extended context for the LLM. On a poorly optimized setup, this can go from 400ms (LLM alone) to 2–3s (full RAG). Budget for it from the design stage.
- Database maintenance
- When a document is deleted or updated, you need to remove it or reindex it. For external sources (web, API), plan a refresh job. The more the stack evolves, the more work this requires.
- Embedding model quality
- For French, the default embeddings (text-embedding-ada like) are mediocre. nomic-embed-text, BGE-M3, or Solon make a difference—but that means knowing the options.
#Hidden costs of fine-tuning
- Dataset preparation
- This is 80% of the work. Collecting, cleaning, formatting into instruction/response pairs, deduplicating, and balancing classes. In a 4-week fine-tuning project, plan on 3 weeks of data preparation and 1 week of training.
- Risk of regression
- A poorly calibrated fine-tune degrades the model's general capabilities ("catastrophic forgetting"). The model becomes good at your task and useless at everything else. You need to test it on a general benchmark before and after.
- Retraining with every change
- The dataset grows over time. Each release requires rerunning 2–12 hours of GPU work, revalidating, and redeploying. Versioning datasets and checkpoints becomes mandatory.
- Training hardware
- Inferring with a 7B Q4 requires 5 GB of VRAM, but training it (even with QLoRA) requires at least 12–16 GB. Fine-tuning has a higher hardware entry threshold than inference.
#The hybrid approach: RAG + fine-tuning
The most effective architectures don’t choose—they combine. Fine-tuning defines how the model talks about your domain; RAG gives it access to what it needs to know at a given time.
- Fine-tune for style and format
- 200–1000 examples that anchor the tone (corporate, technical, legal), response format (JSON, structured Markdown), and stance (always cite the source, never make anything up).
- RAG for factual knowledge
- Documentation, ticket database, case law, codebase—anything that changes and needs to be findable and citable.
- Guardrails in the system prompt
- The system prompt reminds the model to refuse to answer if the RAG context is empty or contradictory. Essential for limiting hallucinations.
#Quick decision in 3 questions
- 01Question 1 — Do your data change more than once a month?If so: RAG is mandatory. Fine-tuning cannot keep pace without becoming an operational nightmare.
- 02Question 2 — Do you have at least 500 high-quality input/output pairs verified by humans?If not, start with RAG. Fine-tuning on 100 hastily assembled examples degrades the model. If you plan to accumulate more, set up RAG first and use its logs to build the dataset.
- 03Question 3 — Is the problem “knowing something” or “responding in a certain way”?Knowing → RAG. Answering in a particular way → fine-tuning. Both → hybrid. It is the simplest framework, and it is right 9 times out of 10.
#Common pitfalls to avoid
- "Fine-tune to learn facts"
- Most common mistake. A fine-tune is not a knowledge base—it interpolates from the examples it has seen but hallucinates on precise details (numbers, dates, references) far more than RAG.
- “RAG without a reranker” on corpora > 10k chunks
- Vector search alone retrieves volume but not always relevance. A cross-encoder reranker (BGE, mxbai) on the top 20 transforms quality, at an additional cost of 50-100ms.
- Confusing RAG with long context
- “I’ll just put the entire document in the prompt.” Beyond 8k useful tokens, quality drops sharply (“lost in the middle”). A well-chunked RAG system beats a naive long context once you reach a certain volume.
- Fine-tune an already aligned model with an unaligned dataset
- If you fine-tune an “instruct” model with examples that do not use the same prompt structure, you break alignment and the model becomes erratic. Always follow the model’s template (ChatML, Alpaca, Mistral, Llama-3 chat).
- Wanting a public benchmark to decide
- No generic benchmark will tell you whether YOUR use case is better suited to RAG or fine-tuning. Set up an internal eval with 30-50 representative questions, and measure before moving to production.
#Go further
Once the decision is made, the corresponding guides cover the practical implementation, including pitfalls and optimizations.
- Implementing RAG without coding
- Open WebUI or AnythingLLM let you set up a document RAG system in a few hours, without writing Python—useful for validating the approach before industrializing it.
- Local LoRA / QLoRA fine-tuning
- The dedicated guide covers Unsloth, dataset format, choosing between LoRA and QLoRA, and GGUF export for Ollama. One RTX 3090 is enough for a 7B model.
- Optimize an existing RAG
- Reranking, hybrid BM25 + vector search, and chunking strategies—three levers that take a RAG system from "works" to production.
Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.