Open source

Unsloth

4 min readintermediateUpdated 28 Sept 2026
1 · In one line

Unsloth is an open-source tool for training and running AI language models on your own computer, built to fine-tune them faster and with less graphics-card memory.

1 · What it is

Unsloth is an open-source tool for fine-tuning language models. Fine-tuning takes a trained model and customises its behaviour for a domain or a specific task. Unsloth also runs models, and comes as a desktop app, a web interface or a code library. Its core code uses the Apache 2.0 licence, while some extras such as the web interface use AGPL-3.0.

The main problem it tackles is memory. A model like Llama 70B has 70 billion weights, the numbers it learned in training. Unsloth leans on two memory savers. LoRA keeps the original weights frozen. It trains a small set of extra numbers on top instead. QLoRA goes further. It stores the frozen model in 4-bit form, meaning each number is saved in a shorter, rougher way. That cuts memory about four times.

Unsloth’s makers report training about twice as fast with up to 70% less graphics memory. Those are their own benchmarks. Beginners can start in free online notebooks on Colab or Kaggle. When training ends, you can keep a small adapter file, which holds only the changes. Or you can export a GGUF file, a format that Ollama and llama.cpp can run.

2 · Why it exists

Fine-tuning a language model at home runs into two problems.

Not enough memoryA model such as Llama 70B has 70 billion weights to train.
SpeedUnsloth's makers report about 2x faster training.
3 · How it works

Follow one small model as Unsloth fine-tunes it.

How Unsloth fine-tunes a model Five boxes in a row. A base model is loaded in 4-bit form to save memory. LoRA adds small trainable matrices. The highlighted step trains only those small matrices on a question and answer dataset. The result is saved as a small LoRA adapter file of about 100 megabytes. It can then be exported to GGUF to run in tools like Ollama or llama.cpp. FROM A BASE MODEL TO A FILE YOU CAN RUN Load in 4-bit Add LoRA Train Save adapter Export load_in_4bit about 4x less memory thin matrices A and B; base weights frozen question + answer data; only A and B change LoRA adapter, about 100 MB in the guide GGUF file for Ollama or llama.cpp Freezing the 4-bit base and training only A and B is what keeps memory low.
The highlighted step is training: only the small added weights change, which is why it fits in less memory.
  1. 1 · loadYou pick a model and, with one setting, load it in 4-bit form, which uses about four times less memory.
  2. 2 · attachUnsloth adds LoRA, two thin extra matrices, A and B, to the weights, so the original weights can stay frozen.
  3. 3 · trainTraining on your question and answer data changes only the small added weights.
  4. 4 · saveThe result is saved as a small LoRA adapter file, around 100 megabytes in the guide's example.
  5. 5 · exportYou can export the model to GGUF, a file format that Ollama and llama.cpp can run.

Unsloth trains with LoRA or QLoRA so fine-tuning can run with little graphics memory.

4 · Where it's used
WhoWhat they askWhat it works with
Student“Can I fine-tune a small model in a free online notebook?”Unsloth's free Colab or Kaggle notebooks
Legal team“Can a model learn the style of our contracts?”Question and answer pairs from past contracts
Hobbyist“Can I run my trained model on my laptop afterwards?”A GGUF export loaded into Ollama or llama.cpp
5 · What it solves, and what it doesn't
solves
  • It lets you fine-tune with LoRA or QLoRA locally on a card with little memory.
  • Its free Colab and Kaggle notebooks let beginners fine-tune at no cost.
  • It saves a trained model as a small adapter or as a GGUF file for local tools.
  • It also supports full fine-tuning and reinforcement learning methods.
doesn't solve
  • You still have to prepare your own question and answer data, and its quality largely decides the result.
  • Some models need more memory than the listed minimums, and very old graphics cards are slow.
  • Its speed and memory figures are the maker's own benchmarks.
6 · Go deeper

Sources used

This explainer is written in original language. The links below support its factual claims.

  1. repounslothai/unsloth: README, Unsloth AI · read 28 Sept 2026
  2. docsUnsloth Docs, Unsloth AI · read 28 Sept 2026
  3. docsFine-tuning LLMs Guide, Unsloth AI · read 28 Sept 2026
  4. docsUnsloth Requirements, Unsloth AI · read 28 Sept 2026
  5. docsSaving to GGUF, Unsloth AI · read 28 Sept 2026
  6. paperQLoRA: Efficient Finetuning of Quantized LLMs, arXiv · read 28 Sept 2026