Advanced 25 minFine-tuning

Local LLM fine-tuning: step-by-step LoRA / QLoRA

Local LLM fine-tuning with LoRA and QLoRA lets you adapt an 8–9B model (such as Qwen 3.5 9B) to your business vocabulary, writing style, or a specific response format — all on a used RTX 3090. No A100 cluster required. This guide takes you from no experience to a GGUF model loaded in Ollama, covering dataset formatting, the Unsloth notebook step by step, and the final export.

By Mohamed Meguedmi·Update 2026-08-27·Tested on Windows, macOS, and Linux

#Why fine-tune a local LLM?

A base LLM is versatile but imperfect for your specific use case. Prompt engineering and RAG solve 80% of cases—fine-tuning is for the remaining 20%. You should consider fine-tuning when your model must respond in a strict format (JSON, tags, internal structure), follow a corporate style (tone, vocabulary), master a specialized field (medical, legal, or business-specific technical), or reproduce behaviors that no prompt can truly stabilize.

Conversely, don’t use fine-tuning to inject factual knowledge (RAG does it better and stays up to date), fix a single hallucination (a better prompt is enough), or compete with GPT-5 on benchmarks (you won’t win).

i
The rule of thumb
If you can express what you want in fewer than 5 examples in a prompt, use a system prompt. If you have 50+ examples and the model still drifts, fine-tune it.

#LoRA vs. full fine-tuning: why LoRA wins

The Local AI Kit

Your private ChatGPT, free, on your own machine in an hour — LM Studio, Ollama, Open WebUI, your documents, no cloud.

  • Lifetime online access
  • PDF + files
  • Lifetime updates

Full fine-tuning updates every model parameter. For a 7B model in FP16, that's 14 GB of weights plus ~30 GB for the gradients and Adam optimizer. Total: 40 to 60 GB of VRAM minimum. Inaccessible without an A100 80GB or H100.

LoRA (Low-Rank Adaptation) freezes the original weights and trains only small rank-r adaptation matrices inserted into the attention layers. In practice, for a 7B with r=16, you train ~20 million parameters instead of 7 billion—0.3% of the model. The VRAM required for gradients and the optimizer collapses.

Full fine-tune 7B
~60 GB of VRAM, 2–4× A100, several hours, 14 GB output file.
LoRA 7B
~16 GB of VRAM, one RTX 4080 is enough, 30 min to 2h, adapter only 30–200 MB.
QLoRA 7B
~6 GB VRAM, RTX 3060 12 GB is sufficient, with nearly identical quality to classic LoRA.
→
Practically equivalent
On modest-sized datasets (< 10,000 examples), LoRA reaches 95-99% of the quality of a full fine-tune. The difference becomes visible only on tasks that are very far from pretraining. For your first fine-tune, don't even ask yourself the question: LoRA.

#QLoRA: the VRAM revolution

QLoRA takes the idea further: the base model is loaded in 4-bit (NF4, NormalFloat 4-bit) instead of FP16. During training, the weights remain quantized; only the FP16 LoRA adapters receive gradients. Result: a 7B fits in 5-6 GB of VRAM, a 13B in 10 GB, and a 70B in 48 GB.

The quality loss is negligible thanks to two techniques: NF4 quantization (calibrated to the weights' Gaussian distribution) and double quantization (the quantization constants are quantized themselves). In practice, across most benchmarks, QLoRA is within 1% of FP16 LoRA.

i
The right default
For a first local fine-tuning run, QLoRA is almost always the right choice. You save 60-70% of VRAM with an imperceptible loss of quality. Switch to classic LoRA only if you are targeting a critical production use case with RTX 4090 or more.

#VRAM required by model size

The figures below are realistic minimums with Unsloth, batch size 2, a 2048-token sequence, and gradient checkpointing enabled. Add 20% headroom for spikes and the OS.

2–3B model (Qwen 3.5 2B, Granite 4.2 3B)
QLoRA: 4 GB VRAM · LoRA FP16: 8 GB · GTX 1660 6 GB or RTX 3050 is sufficient for QLoRA.
8–9B model (Granite 4.2 8B, Qwen 3.5 9B)
QLoRA: 6 GB VRAM · LoRA FP16: 16 GB · RTX 3060 12 GB comfortable, RTX 3090 ideal.
12B model (Gemma 4 12B)
QLoRA: 10 GB VRAM · LoRA FP16: 28 GB · RTX 3090/4090 24 GB in QLoRA, A100 40GB in LoRA.
27–35B model (Qwen 3.8 27B, Qwen 3.6 35B-A3B)
QLoRA: 22 GB VRAM · LoRA FP16: impossible on consumer hardware · RTX 3090/4090 or RTX 5090 in strict QLoRA.
70B model (high-end dense)
QLoRA: 48 GB VRAM · 2× RTX 3090 or 1× A100 80GB · multi-GPU required with Unsloth Pro.

For this guide, the target is an 8–9B model in QLoRA on RTX 3090 24 GB. It is the most universal combination: RTX 3060 12 GB is also sufficient, but a used 3090 (variable price) supports every size up to 12B comfortably in QLoRA. A RTX 4090 or 5090 is 1.5 to 2× faster but not essential.

#Hardware and software requirements

NVIDIA GPU with CUDA 11.8+
RTX 3060 12 GB minimum, RTX 3090 recommended. AMD ROCm GPUs work partially, but Unsloth is optimized for CUDA — stay NVIDIA for this first guide.
Python 3.10 or 3.11
Not 3.12 (bitsandbytes incompatibilities at the time of writing). Use conda or pyenv.
32 GB of system RAM
16 GB may be enough, but initial model loading and dataset preparation are more comfortable with 32 GB.
50 GB of disk space
Base model + checkpoints + merged adapter + GGUF Q4_K_M export. Allow plenty of room.
Connection successful
The model's initial 8-9B download is 5-8 GB. Just once.

On the software side, we use Unsloth (github.com/unslothai/unsloth), a framework that rewrites Triton kernels to achieve 2× the speed and 60% less memory than standard Hugging Face PEFT. Installation in a dedicated environment:

Unsloth installation
conda create -n unsloth python=3.11 -y
conda activate unsloth
pip install --upgrade pip
pip install "unsloth[cu121-torch240] @ git+https://github.com/unslothai/unsloth.git"
pip install --no-deps trl peft accelerate bitsandbytes
!
CUDA versions
Adapt cu121-torch240 to your setup: cu118 for CUDA 11.8, cu124 for CUDA 12.4. Check with nvidia-smi in the upper-right corner. The wrong version will cause cryptic errors on the first import.

#Prepare your dataset in JSONL format

The standard format for instruction fine-tuning is JSONL (one JSON object per line). Three variants are common:

Alpaca format (the simplest)
{"instruction": "Traduis en français formel.", "input": "Hey, what's up?", "output": "Bonjour, comment allez-vous ?"}
{"instruction": "Résume en une phrase.", "input": "Le chat noir a sauté...", "output": "Un chat noir saute sur la table."}
ShareGPT format (multi-turn)
{"conversations": [
  {"from": "system", "value": "Tu es un assistant juridique."},
  {"from": "human", "value": "Qu'est-ce qu'une clause léonine ?"},
  {"from": "gpt", "value": "Une clause léonine est une disposition contractuelle..."}
]}

For a first fine-tune, choose Alpaca: a single turn, clear format, natively supported by Unsloth. A few crucial rules for quality:

Minimum volume
300 examples to observe an effect, 1000-5000 for a solid fine-tune, 10,000+ for serious work. Fewer than 300, and you're wasting your time.
Quality > quantity
100 excellent examples beat 5000 noisy examples. Review them. Have someone else review them. The model literally learns your dataset, including its flaws.
Input diversity
If all your inputs start with “Translate,” the model won’t know how to do anything else. Vary your phrasing.
Balanced lengths
If all your outputs are 2 sentences long, the model will no longer know how to produce a long answer. Mix them up.
Reserve 10% for validation
Randomly split train.jsonl and val.jsonl to measure overfitting.
→
Generate a synthetic dataset
Don’t have 1000 examples on hand? Use a large model (Claude, GPT-5, or Qwen 3.8 27B locally) to generate synthetic instruction/output pairs from your internal documents. This is the standard technique known as “self-instruct.” Allow 1 to 2 hours of generation for 1000 examples.

#The annotated Unsloth notebook

Here’s a complete, minimal script for fine-tuning Qwen 3.5 9B with QLoRA on your dataset. Run it as a .py file or in a Jupyter notebook.

finetune.py — step 1: loading
from unsloth import FastLanguageModel
import torch

MODEL = "unsloth/Qwen3.5-9B-Instruct-bnb-4bit"
MAX_SEQ = 2048

model, tokenizer = FastLanguageModel.from_pretrained(
    model_name = MODEL,
    max_seq_length = MAX_SEQ,
    dtype = None,           # auto : bf16 sur Ampere+, fp16 sinon
    load_in_4bit = True,    # QLoRA : modèle quantifié 4-bit
)

Unsloth hosts pre-quantized 4-bit versions of most popular models on Hugging Face (unsloth/ prefix). Faster downloads, immediate startup. The first run downloads ~5 GB.

finetune.py — step 2: LoRA config
model = FastLanguageModel.get_peft_model(
    model,
    r = 16,                  # rang de la décomposition LoRA
    target_modules = [
        "q_proj", "k_proj", "v_proj", "o_proj",
        "gate_proj", "up_proj", "down_proj",
    ],
    lora_alpha = 16,
    lora_dropout = 0,
    bias = "none",
    use_gradient_checkpointing = "unsloth",  # économie VRAM
    random_state = 42,
)

Rank r=16 is a good default. Increase it to 32 or 64 if you have a lot of data (>10k) and a domain very far from the pretraining data. A higher rank = more trainable parameters = more capacity but a higher risk of overfitting.

finetune.py — step 3: dataset
from datasets import load_dataset

ALPACA_PROMPT = """### Instruction:
{}

### Input:
{}

### Réponse:
{}"""

EOS = tokenizer.eos_token

def format_prompt(ex):
    texts = [
        ALPACA_PROMPT.format(i, inp or "", out) + EOS
        for i, inp, out in zip(ex["instruction"], ex["input"], ex["output"])
    ]
    return {"text": texts}

ds = load_dataset("json", data_files="train.jsonl", split="train")
ds = ds.map(format_prompt, batched=True)
!
Don't forget the EOS token
Without EOS at the end of each example, the model learns never to stop. Classic symptom: your fine-tuned model generates endlessly, repeats itself, or continues with an invented new question. Always verify that tokenizer.eos_token is properly added.
finetune.py — step 4: training
from trl import SFTTrainer
from transformers import TrainingArguments

trainer = SFTTrainer(
    model = model,
    tokenizer = tokenizer,
    train_dataset = ds,
    dataset_text_field = "text",
    max_seq_length = MAX_SEQ,
    args = TrainingArguments(
        per_device_train_batch_size = 2,
        gradient_accumulation_steps = 4,    # batch effectif = 8
        warmup_steps = 10,
        num_train_epochs = 3,
        learning_rate = 2e-4,
        fp16 = not torch.cuda.is_bf16_supported(),
        bf16 = torch.cuda.is_bf16_supported(),
        logging_steps = 10,
        optim = "adamw_8bit",
        weight_decay = 0.01,
        lr_scheduler_type = "linear",
        seed = 42,
        output_dir = "outputs",
    ),
)

trainer.train()

On RTX 3090 with a dataset of 1,000 examples and 3 epochs, allow 20 to 40 minutes. On RTX 4090, 12 to 25 minutes. On RTX 5090, about 8 to 15 minutes. Monitor the loss: it should decrease steadily and then stabilize. If it rises on the validation set, you are overfitting.

→
When to stop
3 epochs is a good default for a dataset of 1k-5k examples. Beyond 5 epochs, the risk of overfitting becomes serious: the model memorizes instead of generalizing. If you have more than 20k examples, reduce this to 1-2 epochs.

#GGUF export to Ollama

Your LoRA adapter is in outputs/. It is a file of a few dozen MB, separate from the base model. To use it in Ollama, there are two steps: merge the adapter into the model, then convert it to quantized GGUF.

finetune.py — step 5: export GGUF Q4_K_M
model.save_pretrained_gguf(
    "mon-modele-gguf",
    tokenizer,
    quantization_method = "q4_k_m",  # le défaut recommandé
)

Unsloth handles everything: adapter + base merging, conversion via llama.cpp, Q4_K_M quantization. The result: a mon-modele-gguf/unsloth.Q4_K_M.gguf file of about 5.3 GB for a 9B. Q5_K_M and Q8_0 are available if you want more precision (and more weight).

To load it into Ollama, which listens on localhost:11434, create a minimal Modelfile and then register the model:

Modelfile
cat > Modelfile <<EOF
FROM ./mon-modele-gguf/unsloth.Q4_K_M.gguf

TEMPLATE """### Instruction:
{{ .Prompt }}

### Réponse:
"""

PARAMETER temperature 0.7
PARAMETER stop "### Instruction:"
EOF

ollama create mon-modele -f Modelfile
ollama run mon-modele
i
The template must match
The Ollama template must exactly reproduce the format used during training (here, Alpaca with ### Instruction: and ### Response:). Otherwise, the model receives prompts it has never seen and hallucinates. If you used ShareGPT or ChatML, adapt accordingly.

#Tips and troubleshooting

OutOfMemoryError at startup
Reduce per_device_train_batch_size to 1 and increase gradient_accumulation_steps proportionally. Reduce max_seq_length to 1024 if your examples are short.
Loss that won’t decrease
Learning rate too low (try 5e-4) or broken prompt format. Print 2-3 examples of ds[0]["text"] and visually check that they look like what you want.
Exploding loss (NaN)
Learning rate too high. Go back to 1e-4. Or enable bf16 if you were using fp16 on Ampere+ — fp16 tends to overflow.
Model that repeats forever
EOS token missing during training, or the stop parameter missing from the Modelfile. The two issues often compound each other.
Visible overfitting
Training loss goes down, validation loss goes up. Stop at the epoch where they diverge. Reduce epochs or increase lora_dropout to 0.05.
GGUF export crashes
Unsloth downloads llama.cpp on the fly the first time—make sure git, cmake, and build-essential are installed on Linux. On Mac, xcode-select --install.
Model that ignores new instructions
Dataset not varied enough or LoRA rank too low. Set r=32 and lora_alpha=64, then run it again.
→
Keep the adapter separate
To iterate quickly, also keep an unfused version of the adapter (~100 MB) with model.save_pretrained("adapter"). You can reuse it later with another base model or combine it with other adapters without downloading everything again.

#Go further

You have a first fine-tuned model running in Ollama. Three natural directions to explore:

Choose the right quantization
You exported in Q4_K_M by default. The guide to choosing your quantization compares Q4_K_M, Q5_K_M, Q8_0, and FP16—useful for deciding whether your fine-tune’s quality warrants Q5 or Q8.
RAG rather than fine-tuning for knowledge
If you want the model to know your documents (rather than just imitate a style), the Local RAG Introduction guide explains why RAG is almost always preferable to fine-tuning for injecting facts.
Hardware for moving to larger models
To fine-tune a 32B or 70B model with QLoRA, you need 24 GB and then 48 GB of VRAM. The which LLM for 24 GB of VRAM guide maps out what is feasible on RTX 3090, 4090, and RX 7900 XTX.
Did this guide help you?

Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.