Unsloth: fine-tuning an LLM on a large public
Unsloth is a free Python library (Apache 2.0 for the core, with more than 76,000 stars on GitHub) that rewrites the compute-intensive steps of LoRA and QLoRA fine-tuning: up to twice as fast with 70% less VRAM, according to its publisher. Its documentation gives a minimum of 6 GB for an 8-billion-parameter model in 4-bit QLoRA: a 12 GB card is enough, and GGUF export leads to Ollama.
Unsloth makes fine-tuning an open model practical on a consumer graphics card: the same methods used elsewhere, but with rewritten compute kernels that reduce the time and memory required. A 7- or 8-billion-parameter model can therefore be fine-tuned on a 12 GB card, and the result can be exported to GGUF to run in Ollama. The question that comes before everything else remains: do you really need to fine-tune?
#The question to settle before installing anything
Yes, you can fine-tune a 7- or 8-billion-parameter model on a consumer GPU: Unsloth’s documentation gives a minimum of 5 to 6 GB of VRAM in 4-bit QLoRA, so a 12 GB card is suitable, with room for context and the batch. Unsloth rewrites the expensive LoRA and QLoRA fine-tuning steps to achieve, according to the vendor, up to twice the speed with 70% less VRAM, a figure that applies only to certain notebooks. The result can be exported to GGUF to run in Ollama. But the real question comes before installation: fine-tuning changes a model’s behavior (tone, format, register), not what it knows up to date. If you want it to know your documents, build document retrieval first. Also expect dataset preparation to take longer than training.
Fine-tuning changes how a model behaves. Document retrieval changes what it knows when answering. Confusing the two can cost you weeks.
| Your needs | The right answer |
|---|---|
| Answers grounded in your documents | A RAG pipeline: documents can change every day |
| A consistent output format | Fine-tuning or constrained generation |
| A house style, a professional register | Fine-tuning |
| The vocabulary of a specialized field | Fine-tuning, with enough examples |
| Information that changes frequently | RAG, as always: retraining for a price makes no sense |
| Shorter or longer responses | System prompt first; never train for this |
- Fine-tuning or RAG: the detailed decision tree
- LoRA and QLoRA: how they really work
- A complete LoRA fine-tuning example
- Assess whether RAG would do better, with numbers
#What Unsloth changes in practice
Your private ChatGPT, free, on your own machine in an hour — LM Studio, Ollama, Open WebUI, your documents, no cloud.
- Lifetime online access
- PDF + files
- Lifetime updates
The method is not new: the model weights are frozen, and only thin matrices added to each weight are trained—about 1% of the parameters according to Unsloth's guide. LoRA keeps the original model in 16-bit, while QLoRA quantizes it to 4-bit, saving 75% of memory. This is the standard approach; it does not belong to Unsloth.
What the library provides is the implementation, with rewritten compute kernels for the heavy stages. The project exists in three forms: Unsloth Desktop, a native application; Unsloth Studio, a web interface; and Unsloth Core, the code-driven version discussed in this article. It is released under a dual license: Apache 2.0 for the core, and AGPL-3.0 for certain optional components such as the Studio interface. It claims to train language, diffusion, speech-synthesis, and embedding models twice as fast with 70% less VRAM, without loss of accuracy. On mixture-of-experts (MoE) architectures, the published gain is even higher: up to 12 times faster with more than 35% VRAM savings on some recent models. As with all vendor figures, these describe the best case: the README itself shows that the “twice as fast, 70%” claim applies only to certain notebooks.
| Notebook | Advertised speed | Advertised VRAM |
|---|---|---|
| Llama 3.1 (8B) Alpaca | 2 times faster | 70% less |
| gpt-oss (20B) | 2 times faster | 70% less |
| Qwen3.5 (4B), vision | 1.5 times faster | 60% less |
| Gemma 4 (E2B), vision | 1.5 times faster | 50% less |
| Orpheus-TTS (3B) | 1.5 times faster | 50% less |
| embeddinggemma (300M) | 2 times faster | 20% less |
#Which card for which model
| Model size | QLoRA (4-bit) | LoRA (16-bit) | Reading for a consumer-grade card |
|---|---|---|---|
| 3 billion | 3,5 | 8 | Comfortable on any recent card with QLoRA |
| 7 billion | 5 | 19 | QLoRA on 8 GB; 16-bit LoRA requires a 24 GB card |
| 8 billion | 6 | 22 | QLoRA on 8 GB or more; 12 GB leaves some headroom |
| 14 billion | 8,5 | 33 | QLoRA on 12 GB; 16-bit LoRA is out of reach for a 24 GB card |
| 27 billion | 22 | 64 | QLoRA on 24 GB, with no headroom; 16-bit LoRA is impossible on a single card |
| 32 billion | 26 | 76 | Beyond 24 GB, even with QLoRA |
| 70 billion | 41 | 164 | Beyond the reach of a consumer graphics card |
The documentation specifies that these figures are absolute minimums: depending on the model, more may be required. It identifies an excessively large batch size as a frequent cause of exhausted memory, recommending reducing it to 1, 2, or 3, and recommends a context length of 2048 for initial tests.
| Platform | What the documentation says |
|---|---|
| NVIDIA, with Unsloth Core | Linux and Windows; compute capability 7.0 minimum (V100, T4, RTX 20 and later, A100, H100), including Blackwell and DGX Spark. The GTX 1070 and 1080 work, slowly. |
| AMD and Intel | Dedicated guides; training works on these GPUs, with both Core and Studio. |
| Mac | Unsloth Studio supports training, MLX, and GGUF inference (macOS 12 or later). For Core, Apple Silicon (MLX) support is listed as “in preparation.” |
| Without a GPU | Studio works for chatting with GGUF models and for data preparation, not training. |
| Multiple GPUs | Supported by Accelerate and DeepSpeed (FSDP, DDP), with manual setup; the device_map="balanced" option distributes a model that is too large across multiple cards. |
#The dataset is the real work
A fine-tuning run never fails because of the library. It fails because of the data.
- Format
- A two-column dataset—usually question and answer—in the conversation structure expected by the model. The guide recommends starting with an Instruct model: it directly accepts conversation templates (ChatML, ShareGPT) and requires less data than a base model. Consistency matters more than volume.
- Volume
- The Unsloth guide recommends at least 100 lines, and more than 1,000 lines for better results. Tone and format transfer with few examples; a true business behavior requires more, and they must be consistent with one another.
- Quality
- The model imitates what it is shown, including errors. A systematic flaw in the data becomes a systematic flaw in the model.
- Control set
- Examples the model never sees during training. Without them, it is impossible to distinguish learning from simple memorization.
- No code
- Unsloth Studio offers Data Recipes that transform PDFs, CSVs, or DOCX files into datasets through a visual workflow, with a preview before launching the full build.
#From adapter to usable model
The process takes three steps: load a model in 4-bit, train it with one of the official notebooks, then export it. The first code block loads the model and adds the LoRA adapters; the actual training is handled by one of Unsloth’s notebooks, which the guide recommends copying into your local environment.
The name suffix matters. According to the guide, a model ending in unsloth-bnb-4bit is Unsloth’s dynamic 4-bit quantization: it uses slightly more VRAM than standard BitsAndBytes quantization, but offers significantly higher precision. A name ending only in bnb-4bit refers to the standard version. The guide also notes that training and serving benefit from using the same precision: serving in 4-bit means training in 4-bit.
| Parameter | Guide value | Good to know |
|---|---|---|
| per_device_train_batch_size | 2 | Larger: better GPU utilization, but training is slowed by filling; prefer gradient_accumulation_steps |
| gradient_accumulation_steps | 4 | Simulates a larger batch without additional memory |
| max_steps | 60 | Quick trial value; for actual training, replace with num_train_epochs between 1 and 3 |
| learning_rate | 2e-4 | Lower for slower, more precise tuning: try 1e-4, 5e-5, or 2e-5 |
| max_seq_length | 2048 | Recommended length for testing. Setting it higher “to be safe” reserves memory for sequences absent from your data: measure the actual length of your examples first. |
- 01TrainThe result is a LoRA adapter, approximately 100 MB in Unsloth's example, not a complete model.
- 02EvaluateAgainst the control set and against the original model. “Better” must be demonstrated, not assumed. The guide notes that automatic evaluation tools may poorly reflect your criteria.
- 03Merge or keep separateThe adapter can remain separate and be swapped, or it can be merged into the weights with: model.save_pretrained_merged using save_method="merged_16bit".
- 04Export to GGUF and quantizemodel.save_pretrained_gguf produces the file, with q4_k_m, q8_0, or f16 as options. That's what makes the result runnable by Ollama or llama.cpp on ordinary hardware.
#Common pitfalls
- Starting from a base model without realizing it
- The guide recommends Instruct models: they support conversation templates and require less data. A base model requires a different format (Alpaca, Vicuna) and more examples.
- Forget the conversation template
- Examples formatted differently from what the model expects produce training that is technically successful but practically useless.
- Measure on training examples
- This is the mistake that leads to spectacular results being announced and disappointing models being delivered.
- Believing it will replace document research
- What we teach the model becomes outdated when your information changes, and the model does not cite its sources. For changing facts, document search remains safer.
- Underestimating example cleanup
- Near-identical duplicates in the dataset bias training toward those specific cases without anything indicating it in the tracked metrics.
- Change several settings at once
- Changing the learning rate, adapter rank, and sequence length in the same run makes it impossible to know which one produced the observed change.
#Concrete use cases
- Consistent-tone customer support
- A model fine-tuned on real interactions responds in the brand's voice, without repeated style instructions in every prompt.
- Extraction in a strict format
- A fine-tuner trained on input-output examples fixes the output format more reliably than a prompt alone, which is useful before sending the result to a downstream system.
These cases have one thing in common: each can be measured. Before starting a training run, state the success criterion in one verifiable sentence—“the JSON format is respected in at least 95% of the outputs in the control set,” for example—rather than relying on a general impression of improvement. That criterion, not intuition, will tell you whether the result justifies the time invested.
- Source: official Unsloth repository (features, hardware, dual licensing)
- Source: requirements and minimum VRAM by model size
- Source: Unsloth fine-tuning guide
- Source: GGUF export and template issues
- Source: performance figures for MoE models
#The real cost beyond the price of the card
The graphics card is only one part of the equation. Collecting and cleaning consistent examples, writing a holdout set the model will never see, and then comparing each version with the original model often matters more than the training itself, as the Unsloth guide illustrates with a 60-step run.
That is why Unsloth, as fast as it is for training, does not shorten a project as much as hoped: it removes the technical bottleneck, not the data work.
#FAQ
Is Unsloth free?+
Which card should you use to fine-tune a 7-billion-parameter model?+
How many examples are needed?+
Does the resulting model run in Ollama?+
Does it work on AMD, on Mac, or without a GPU?+
Will fine-tuning make the model an expert in my data?+
Recommended hardware: GMKtec EVO-X2 64GB / 1TB (Ryzen AI Max+ 395) — 16 GB NVIDIA card, sufficient for fine-tuning a 7–8B model with QLoRA. All AI hardware →
Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.