Advanced 14 minFine-tuning

Unsloth: fine-tuning an LLM on a large public

Direct response

Unsloth is a free Python library (Apache 2.0 for the core, with more than 76,000 stars on GitHub) that rewrites the compute-intensive steps of LoRA and QLoRA fine-tuning: up to twice as fast with 70% less VRAM, according to its publisher. Its documentation gives a minimum of 6 GB for an 8-billion-parameter model in 4-bit QLoRA: a 12 GB card is enough, and GGUF export leads to Ollama.

Unsloth makes fine-tuning an open model practical on a consumer graphics card: the same methods used elsewhere, but with rewritten compute kernels that reduce the time and memory required. A 7- or 8-billion-parameter model can therefore be fine-tuned on a 12 GB card, and the result can be exported to GGUF to run in Ollama. The question that comes before everything else remains: do you really need to fine-tune?

By Mohamed Meguedmi·Update 2026-09-29·Tested on Windows, macOS, and Linux

#The question to settle before installing anything

Yes, you can fine-tune a 7- or 8-billion-parameter model on a consumer GPU: Unsloth’s documentation gives a minimum of 5 to 6 GB of VRAM in 4-bit QLoRA, so a 12 GB card is suitable, with room for context and the batch. Unsloth rewrites the expensive LoRA and QLoRA fine-tuning steps to achieve, according to the vendor, up to twice the speed with 70% less VRAM, a figure that applies only to certain notebooks. The result can be exported to GGUF to run in Ollama. But the real question comes before installation: fine-tuning changes a model’s behavior (tone, format, register), not what it knows up to date. If you want it to know your documents, build document retrieval first. Also expect dataset preparation to take longer than training.

Fine-tuning changes how a model behaves. Document retrieval changes what it knows when answering. Confusing the two can cost you weeks.

What you want, and the corresponding tool
Your needsThe right answer
Answers grounded in your documentsA RAG pipeline: documents can change every day
A consistent output formatFine-tuning or constrained generation
A house style, a professional registerFine-tuning
The vocabulary of a specialized fieldFine-tuning, with enough examples
Information that changes frequentlyRAG, as always: retraining for a price makes no sense
Shorter or longer responsesSystem prompt first; never train for this
i
If the honest answer is “I want it to know our documentation”
Stop here and build a document search system first. It can be updated by adding a file and cites its sources, which fine-tuning does not. The Unsloth guide argues for a different interpretation, however: according to it, fine-tuning can inject knowledge and reproduce everything RAG does, while the reverse is not true. That is true in principle; the practical argument for RAG lies elsewhere—in updates, traceability, and the cost of retraining after every change.

#What Unsloth changes in practice

The Local AI Kit

Your private ChatGPT, free, on your own machine in an hour — LM Studio, Ollama, Open WebUI, your documents, no cloud.

  • Lifetime online access
  • PDF + files
  • Lifetime updates

The method is not new: the model weights are frozen, and only thin matrices added to each weight are trained—about 1% of the parameters according to Unsloth's guide. LoRA keeps the original model in 16-bit, while QLoRA quantizes it to 4-bit, saving 75% of memory. This is the standard approach; it does not belong to Unsloth.

What the library provides is the implementation, with rewritten compute kernels for the heavy stages. The project exists in three forms: Unsloth Desktop, a native application; Unsloth Studio, a web interface; and Unsloth Core, the code-driven version discussed in this article. It is released under a dual license: Apache 2.0 for the core, and AGPL-3.0 for certain optional components such as the Studio interface. It claims to train language, diffusion, speech-synthesis, and embedding models twice as fast with 70% less VRAM, without loss of accuracy. On mixture-of-experts (MoE) architectures, the published gain is even higher: up to 12 times faster with more than 35% VRAM savings on some recent models. As with all vendor figures, these describe the best case: the README itself shows that the “twice as fast, 70%” claim applies only to certain notebooks.

Gains reported by Unsloth according to the notebook (vendor figures)
NotebookAdvertised speedAdvertised VRAM
Llama 3.1 (8B) Alpaca2 times faster70% less
gpt-oss (20B)2 times faster70% less
Qwen3.5 (4B), vision1.5 times faster60% less
Gemma 4 (E2B), vision1.5 times faster50% less
Orpheus-TTS (3B)1.5 times faster50% less
embeddinggemma (300M)2 times faster20% less

#Which card for which model

Minimums published by Unsloth for fine-tuning (GB of VRAM)
Model sizeQLoRA (4-bit)LoRA (16-bit)Reading for a consumer-grade card
3 billion3,58Comfortable on any recent card with QLoRA
7 billion519QLoRA on 8 GB; 16-bit LoRA requires a 24 GB card
8 billion622QLoRA on 8 GB or more; 12 GB leaves some headroom
14 billion8,533QLoRA on 12 GB; 16-bit LoRA is out of reach for a 24 GB card
27 billion2264QLoRA on 24 GB, with no headroom; 16-bit LoRA is impossible on a single card
32 billion2676Beyond 24 GB, even with QLoRA
70 billion41164Beyond the reach of a consumer graphics card

The documentation specifies that these figures are absolute minimums: depending on the model, more may be required. It identifies an excessively large batch size as a frequent cause of exhausted memory, recommending reducing it to 1, 2, or 3, and recommends a context length of 2048 for initial tests.

Supported hardware according to Unsloth’s documentation
PlatformWhat the documentation says
NVIDIA, with Unsloth CoreLinux and Windows; compute capability 7.0 minimum (V100, T4, RTX 20 and later, A100, H100), including Blackwell and DGX Spark. The GTX 1070 and 1080 work, slowly.
AMD and IntelDedicated guides; training works on these GPUs, with both Core and Studio.
MacUnsloth Studio supports training, MLX, and GGUF inference (macOS 12 or later). For Core, Apple Silicon (MLX) support is listed as “in preparation.”
Without a GPUStudio works for chatting with GGUF models and for data preparation, not training.
Multiple GPUsSupported by Accelerate and DeepSpeed (FSDP, DDP), with manual setup; the device_map="balanced" option distributes a model that is too large across multiple cards.
i
Check today's page
This landscape is changing quickly: the README now announces training on AMD GPUs under Windows, WSL, and Linux, as well as a Desktop application. Check the current documentation before buying hardware.
!
Studio exposed on the network: server tools enabled
Unsloth's README warns that Studio's server-side tools are enabled by default. Exposing the interface (--secure, non-local host, LAN access) requires an administrator password, while --disable-tools disables these tools. Keep it on 127.0.0.1 until you need to open it up.

#The dataset is the real work

A fine-tuning run never fails because of the library. It fails because of the data.

Format
A two-column dataset—usually question and answer—in the conversation structure expected by the model. The guide recommends starting with an Instruct model: it directly accepts conversation templates (ChatML, ShareGPT) and requires less data than a base model. Consistency matters more than volume.
Volume
The Unsloth guide recommends at least 100 lines, and more than 1,000 lines for better results. Tone and format transfer with few examples; a true business behavior requires more, and they must be consistent with one another.
Quality
The model imitates what it is shown, including errors. A systematic flaw in the data becomes a systematic flaw in the model.
Control set
Examples the model never sees during training. Without them, it is impossible to distinguish learning from simple memorization.
No code
Unsloth Studio offers Data Recipes that transform PDFs, CSVs, or DOCX files into datasets through a visual workflow, with a preview before launching the full build.
!
Overfitting looks like success
The Unsloth guide puts it in one sentence: if the loss drops to 0, it may be overfitting, so you need to check validation. According to the guide, a loss between 0.5 and 1.0 is a good sign in many cases, but it depends on the dataset and task. A memorized model answers your test questions perfectly and everything else worse than before. Keep 20% of the data for testing, aim for 1 to 3 epochs to limit overfitting, and monitor the validation set rather than the training curve.

#From adapter to usable model

The process takes three steps: load a model in 4-bit, train it with one of the official notebooks, then export it. The first code block loads the model and adds the LoRA adapters; the actual training is handled by one of Unsloth’s notebooks, which the guide recommends copying into your local environment.

Load a 4-bit model with Unsloth Core
from unsloth import FastLanguageModel

model, tokenizer = FastLanguageModel.from_pretrained(
    model_name="unsloth/Qwen3-8B-unsloth-bnb-4bit",  # quantification dynamique 4 bits d'Unsloth
    max_seq_length=2048,
    load_in_4bit=True,
)
model = FastLanguageModel.get_peft_model(model, r=16, lora_alpha=16)

The name suffix matters. According to the guide, a model ending in unsloth-bnb-4bit is Unsloth’s dynamic 4-bit quantization: it uses slightly more VRAM than standard BitsAndBytes quantization, but offers significantly higher precision. A name ending only in bnb-4bit refers to the standard version. The guide also notes that training and serving benefit from using the same precision: serving in 4-bit means training in 4-bit.

Official guide default settings
ParameterGuide valueGood to know
per_device_train_batch_size2Larger: better GPU utilization, but training is slowed by filling; prefer gradient_accumulation_steps
gradient_accumulation_steps4Simulates a larger batch without additional memory
max_steps60Quick trial value; for actual training, replace with num_train_epochs between 1 and 3
learning_rate2e-4Lower for slower, more precise tuning: try 1e-4, 5e-5, or 2e-5
max_seq_length2048Recommended length for testing. Setting it higher “to be safe” reserves memory for sequences absent from your data: measure the actual length of your examples first.
  1. 01
    Train
    The result is a LoRA adapter, approximately 100 MB in Unsloth's example, not a complete model.
  2. 02
    Evaluate
    Against the control set and against the original model. “Better” must be demonstrated, not assumed. The guide notes that automatic evaluation tools may poorly reflect your criteria.
  3. 03
    Merge or keep separate
    The adapter can remain separate and be swapped, or it can be merged into the weights with: model.save_pretrained_merged using save_method="merged_16bit".
  4. 04
    Export to GGUF and quantize
    model.save_pretrained_gguf produces the file, with q4_k_m, q8_0, or f16 as options. That's what makes the result runnable by Ollama or llama.cpp on ordinary hardware.
Export a trained model to GGUF (in the same session)
model.save_pretrained_gguf(
    "mon-modele-q4", tokenizer, quantization_method="q4_k_m"
)
# puis, avec le Modelfile fourni : ollama create mon-modele -f Modelfile
!
The exported model responds with nonsense in Ollama
A common case described in the Unsloth documentation: the model works well in Unsloth, then produces gibberish, endless generations, or repetitions after export. The most common cause is a chat template different from the one used during training. You must use the same template during training and inference, use the correct end-of-sequence token, and check that the engine does not add an extra beginning-of-sequence token. The documentation specifies that Unsloth automatically creates the Ollama Modelfile with the template used during fine-tuning.

#Common pitfalls

Starting from a base model without realizing it
The guide recommends Instruct models: they support conversation templates and require less data. A base model requires a different format (Alpaca, Vicuna) and more examples.
Forget the conversation template
Examples formatted differently from what the model expects produce training that is technically successful but practically useless.
Measure on training examples
This is the mistake that leads to spectacular results being announced and disappointing models being delivered.
Believing it will replace document research
What we teach the model becomes outdated when your information changes, and the model does not cite its sources. For changing facts, document search remains safer.
Underestimating example cleanup
Near-identical duplicates in the dataset bias training toward those specific cases without anything indicating it in the tracked metrics.
Change several settings at once
Changing the learning rate, adapter rank, and sequence length in the same run makes it impossible to know which one produced the observed change.

#Concrete use cases

Consistent-tone customer support
A model fine-tuned on real interactions responds in the brand's voice, without repeated style instructions in every prompt.
Extraction in a strict format
A fine-tuner trained on input-output examples fixes the output format more reliably than a prompt alone, which is useful before sending the result to a downstream system.

These cases have one thing in common: each can be measured. Before starting a training run, state the success criterion in one verifiable sentence—“the JSON format is respected in at least 95% of the outputs in the control set,” for example—rather than relying on a general impression of improvement. That criterion, not intuition, will tell you whether the result justifies the time invested.

#The real cost beyond the price of the card

The graphics card is only one part of the equation. Collecting and cleaning consistent examples, writing a holdout set the model will never see, and then comparing each version with the original model often matters more than the training itself, as the Unsloth guide illustrates with a 60-step run.

That is why Unsloth, as fast as it is for training, does not shorten a project as much as hoped: it removes the technical bottleneck, not the data work.

→
Time the first run
Unsloth's guide suggests max_steps = 60 to move quickly. On a reduced dataset with a few hundred examples, a first complete run—including training and evaluation—provides an estimate of the total time before committing to the full dataset.

#FAQ

Is Unsloth free?+
The core library is licensed under Apache 2.0 and covers LoRA and QLoRA fine-tuning at no cost; the repository has more than 76,000 stars on GitHub. Some optional components, such as the Studio interface, are licensed under AGPL-3.0: the documentation describes this dual license as a way to fund the project while keeping it open. The notebooks run for free in Colab and Kaggle.
Which card should you use to fine-tune a 7-billion-parameter model?+
Unsloth's table lists 5 GB minimum for a 7-billion-parameter model in 4-bit QLoRA, and 19 GB in 16-bit LoRA. A 12 GB card is therefore suitable for QLoRA. The documentation specifies that these are absolute minimums, and that batch size or long examples may require more memory.
How many examples are needed?+
Unsloth’s guide recommends a strict minimum of 100 lines and more than 1,000 lines for better results; if the dataset is too small, you can add synthetic data or a Hugging Face dataset. Consistency matters more than quantity: bad examples teach bad habits just as reliably as good examples teach good ones.
Does the resulting model run in Ollama?+
Yes: export to GGUF with model.save_pretrained_gguf, choose the quantization (q4_k_m is recommended by the documentation), then create the model in Ollama with a Modelfile. Unsloth generates this Modelfile using the training conversation template. If responses become incoherent, first check that the template and end-of-sequence token are the same.
Does it work on AMD, on Mac, or without a GPU?+
Yes for AMD and Intel, with dedicated guides. On Mac, Unsloth Studio supports GGUF training and inference; Core is not yet available for Apple Silicon. Without a GPU, Studio is for chat and data preparation, not training. Core requires a NVIDIA card with compute capability 7.0 or higher.
Will fine-tuning make the model an expert in my data?+
Not sustainably. Unsloth's guide argues that fine-tuning can inject knowledge, but that knowledge becomes outdated as soon as your data changes, whereas a document can be replaced in a minute and cited. Facts and documents belong to information retrieval; tone, format, and register belong to fine-tuning. Combining the two is common.

Recommended hardware: GMKtec EVO-X2 64GB / 1TB (Ryzen AI Max+ 395) — 16 GB NVIDIA card, sufficient for fine-tuning a 7–8B model with QLoRA. All AI hardware →

Did this guide help you?

Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.