Local LLM fine-tuning: step-by-step LoRA / QLoRA
Local LLM fine-tuning with LoRA and QLoRA lets you adapt an 8–9B model (such as Qwen 3.5 9B) to your business vocabulary, writing style, or a specific response format — all on a used RTX 3090. No A100 cluster required. This guide takes you from no experience to a GGUF model loaded in Ollama, covering dataset formatting, the Unsloth notebook step by step, and the final export.
#Why fine-tune a local LLM?
A base LLM is versatile but imperfect for your specific use case. Prompt engineering and RAG solve 80% of cases—fine-tuning is for the remaining 20%. You should consider fine-tuning when your model must respond in a strict format (JSON, tags, internal structure), follow a corporate style (tone, vocabulary), master a specialized field (medical, legal, or business-specific technical), or reproduce behaviors that no prompt can truly stabilize.
Conversely, don’t use fine-tuning to inject factual knowledge (RAG does it better and stays up to date), fix a single hallucination (a better prompt is enough), or compete with GPT-5 on benchmarks (you won’t win).
#LoRA vs. full fine-tuning: why LoRA wins
Your private ChatGPT, free, on your own machine in an hour — LM Studio, Ollama, Open WebUI, your documents, no cloud.
- Lifetime online access
- PDF + files
- Lifetime updates
Full fine-tuning updates every model parameter. For a 7B model in FP16, that's 14 GB of weights plus ~30 GB for the gradients and Adam optimizer. Total: 40 to 60 GB of VRAM minimum. Inaccessible without an A100 80GB or H100.
LoRA (Low-Rank Adaptation) freezes the original weights and trains only small rank-r adaptation matrices inserted into the attention layers. In practice, for a 7B with r=16, you train ~20 million parameters instead of 7 billion—0.3% of the model. The VRAM required for gradients and the optimizer collapses.
- Full fine-tune 7B
- ~60 GB of VRAM, 2–4× A100, several hours, 14 GB output file.
- LoRA 7B
- ~16 GB of VRAM, one RTX 4080 is enough, 30 min to 2h, adapter only 30–200 MB.
- QLoRA 7B
- ~6 GB VRAM, RTX 3060 12 GB is sufficient, with nearly identical quality to classic LoRA.
#QLoRA: the VRAM revolution
QLoRA takes the idea further: the base model is loaded in 4-bit (NF4, NormalFloat 4-bit) instead of FP16. During training, the weights remain quantized; only the FP16 LoRA adapters receive gradients. Result: a 7B fits in 5-6 GB of VRAM, a 13B in 10 GB, and a 70B in 48 GB.
The quality loss is negligible thanks to two techniques: NF4 quantization (calibrated to the weights' Gaussian distribution) and double quantization (the quantization constants are quantized themselves). In practice, across most benchmarks, QLoRA is within 1% of FP16 LoRA.
#VRAM required by model size
The figures below are realistic minimums with Unsloth, batch size 2, a 2048-token sequence, and gradient checkpointing enabled. Add 20% headroom for spikes and the OS.
- 2–3B model (Qwen 3.5 2B, Granite 4.2 3B)
- QLoRA: 4 GB VRAM · LoRA FP16: 8 GB · GTX 1660 6 GB or RTX 3050 is sufficient for QLoRA.
- 8–9B model (Granite 4.2 8B, Qwen 3.5 9B)
- QLoRA: 6 GB VRAM · LoRA FP16: 16 GB · RTX 3060 12 GB comfortable, RTX 3090 ideal.
- 12B model (Gemma 4 12B)
- QLoRA: 10 GB VRAM · LoRA FP16: 28 GB · RTX 3090/4090 24 GB in QLoRA, A100 40GB in LoRA.
- 27–35B model (Qwen 3.8 27B, Qwen 3.6 35B-A3B)
- QLoRA: 22 GB VRAM · LoRA FP16: impossible on consumer hardware · RTX 3090/4090 or RTX 5090 in strict QLoRA.
- 70B model (high-end dense)
- QLoRA: 48 GB VRAM · 2× RTX 3090 or 1× A100 80GB · multi-GPU required with Unsloth Pro.
For this guide, the target is an 8–9B model in QLoRA on RTX 3090 24 GB. It is the most universal combination: RTX 3060 12 GB is also sufficient, but a used 3090 (variable price) supports every size up to 12B comfortably in QLoRA. A RTX 4090 or 5090 is 1.5 to 2× faster but not essential.
#Hardware and software requirements
- NVIDIA GPU with CUDA 11.8+
- RTX 3060 12 GB minimum, RTX 3090 recommended. AMD ROCm GPUs work partially, but Unsloth is optimized for CUDA — stay NVIDIA for this first guide.
- Python 3.10 or 3.11
- Not 3.12 (bitsandbytes incompatibilities at the time of writing). Use conda or pyenv.
- 32 GB of system RAM
- 16 GB may be enough, but initial model loading and dataset preparation are more comfortable with 32 GB.
- 50 GB of disk space
- Base model + checkpoints + merged adapter + GGUF Q4_K_M export. Allow plenty of room.
- Connection successful
- The model's initial 8-9B download is 5-8 GB. Just once.
On the software side, we use Unsloth (github.com/unslothai/unsloth), a framework that rewrites Triton kernels to achieve 2× the speed and 60% less memory than standard Hugging Face PEFT. Installation in a dedicated environment:
#Prepare your dataset in JSONL format
The standard format for instruction fine-tuning is JSONL (one JSON object per line). Three variants are common:
For a first fine-tune, choose Alpaca: a single turn, clear format, natively supported by Unsloth. A few crucial rules for quality:
- Minimum volume
- 300 examples to observe an effect, 1000-5000 for a solid fine-tune, 10,000+ for serious work. Fewer than 300, and you're wasting your time.
- Quality > quantity
- 100 excellent examples beat 5000 noisy examples. Review them. Have someone else review them. The model literally learns your dataset, including its flaws.
- Input diversity
- If all your inputs start with “Translate,” the model won’t know how to do anything else. Vary your phrasing.
- Balanced lengths
- If all your outputs are 2 sentences long, the model will no longer know how to produce a long answer. Mix them up.
- Reserve 10% for validation
- Randomly split train.jsonl and val.jsonl to measure overfitting.
#The annotated Unsloth notebook
Here’s a complete, minimal script for fine-tuning Qwen 3.5 9B with QLoRA on your dataset. Run it as a .py file or in a Jupyter notebook.
Unsloth hosts pre-quantized 4-bit versions of most popular models on Hugging Face (unsloth/ prefix). Faster downloads, immediate startup. The first run downloads ~5 GB.
Rank r=16 is a good default. Increase it to 32 or 64 if you have a lot of data (>10k) and a domain very far from the pretraining data. A higher rank = more trainable parameters = more capacity but a higher risk of overfitting.
On RTX 3090 with a dataset of 1,000 examples and 3 epochs, allow 20 to 40 minutes. On RTX 4090, 12 to 25 minutes. On RTX 5090, about 8 to 15 minutes. Monitor the loss: it should decrease steadily and then stabilize. If it rises on the validation set, you are overfitting.
#GGUF export to Ollama
Your LoRA adapter is in outputs/. It is a file of a few dozen MB, separate from the base model. To use it in Ollama, there are two steps: merge the adapter into the model, then convert it to quantized GGUF.
Unsloth handles everything: adapter + base merging, conversion via llama.cpp, Q4_K_M quantization. The result: a mon-modele-gguf/unsloth.Q4_K_M.gguf file of about 5.3 GB for a 9B. Q5_K_M and Q8_0 are available if you want more precision (and more weight).
To load it into Ollama, which listens on localhost:11434, create a minimal Modelfile and then register the model:
#Tips and troubleshooting
- OutOfMemoryError at startup
- Reduce per_device_train_batch_size to 1 and increase gradient_accumulation_steps proportionally. Reduce max_seq_length to 1024 if your examples are short.
- Loss that won’t decrease
- Learning rate too low (try 5e-4) or broken prompt format. Print 2-3 examples of ds[0]["text"] and visually check that they look like what you want.
- Exploding loss (NaN)
- Learning rate too high. Go back to 1e-4. Or enable bf16 if you were using fp16 on Ampere+ — fp16 tends to overflow.
- Model that repeats forever
- EOS token missing during training, or the stop parameter missing from the Modelfile. The two issues often compound each other.
- Visible overfitting
- Training loss goes down, validation loss goes up. Stop at the epoch where they diverge. Reduce epochs or increase lora_dropout to 0.05.
- GGUF export crashes
- Unsloth downloads llama.cpp on the fly the first time—make sure git, cmake, and build-essential are installed on Linux. On Mac, xcode-select --install.
- Model that ignores new instructions
- Dataset not varied enough or LoRA rank too low. Set r=32 and lora_alpha=64, then run it again.
#Go further
You have a first fine-tuned model running in Ollama. Three natural directions to explore:
- Choose the right quantization
- You exported in Q4_K_M by default. The guide to choosing your quantization compares Q4_K_M, Q5_K_M, Q8_0, and FP16—useful for deciding whether your fine-tune’s quality warrants Q5 or Q8.
- RAG rather than fine-tuning for knowledge
- If you want the model to know your documents (rather than just imitate a style), the Local RAG Introduction guide explains why RAG is almost always preferable to fine-tuning for injecting facts.
- Hardware for moving to larger models
- To fine-tune a 32B or 70B model with QLoRA, you need 24 GB and then 48 GB of VRAM. The which LLM for 24 GB of VRAM guide maps out what is feasible on RTX 3090, 4090, and RX 7900 XTX.
Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.