Unsloth
Unsloth is an open-source tool for training and running AI language models on your own computer, built to fine-tune them faster and with less graphics-card memory.
Unsloth is an open-source tool for fine-tuning language models. Fine-tuning takes a trained model and customises its behaviour for a domain or a specific task. Unsloth also runs models, and comes as a desktop app, a web interface or a code library. Its core code uses the Apache 2.0 licence, while some extras such as the web interface use AGPL-3.0.
The main problem it tackles is memory. A model like Llama 70B has 70 billion weights, the numbers it learned in training. Unsloth leans on two memory savers. LoRA keeps the original weights frozen. It trains a small set of extra numbers on top instead. QLoRA goes further. It stores the frozen model in 4-bit form, meaning each number is saved in a shorter, rougher way. That cuts memory about four times.
Unsloth’s makers report training about twice as fast with up to 70% less graphics memory. Those are their own benchmarks. Beginners can start in free online notebooks on Colab or Kaggle. When training ends, you can keep a small adapter file, which holds only the changes. Or you can export a GGUF file, a format that Ollama and llama.cpp can run.
Fine-tuning a language model at home runs into two problems.
Follow one small model as Unsloth fine-tunes it.
- 1 · loadYou pick a model and, with one setting, load it in 4-bit form, which uses about four times less memory.
- 2 · attachUnsloth adds LoRA, two thin extra matrices, A and B, to the weights, so the original weights can stay frozen.
- 3 · trainTraining on your question and answer data changes only the small added weights.
- 4 · saveThe result is saved as a small LoRA adapter file, around 100 megabytes in the guide's example.
- 5 · exportYou can export the model to GGUF, a file format that Ollama and llama.cpp can run.
Unsloth trains with LoRA or QLoRA so fine-tuning can run with little graphics memory.
| Who | What they ask | What it works with |
|---|---|---|
| Student | “Can I fine-tune a small model in a free online notebook?” | Unsloth's free Colab or Kaggle notebooks |
| Legal team | “Can a model learn the style of our contracts?” | Question and answer pairs from past contracts |
| Hobbyist | “Can I run my trained model on my laptop afterwards?” | A GGUF export loaded into Ollama or llama.cpp |
- It lets you fine-tune with LoRA or QLoRA locally on a card with little memory.
- Its free Colab and Kaggle notebooks let beginners fine-tune at no cost.
- It saves a trained model as a small adapter or as a GGUF file for local tools.
- It also supports full fine-tuning and reinforcement learning methods.
- You still have to prepare your own question and answer data, and its quality largely decides the result.
- Some models need more memory than the listed minimums, and very old graphics cards are slow.
- Its speed and memory figures are the maker's own benchmarks.
Sources used
This explainer is written in original language. The links below support its factual claims.
- repounslothai/unsloth: README, Unsloth AI · read 28 Sept 2026
- docsUnsloth Docs, Unsloth AI · read 28 Sept 2026
- docsFine-tuning LLMs Guide, Unsloth AI · read 28 Sept 2026
- docsUnsloth Requirements, Unsloth AI · read 28 Sept 2026
- docsSaving to GGUF, Unsloth AI · read 28 Sept 2026
- paperQLoRA: Efficient Finetuning of Quantized LLMs, arXiv · read 28 Sept 2026