llama.cpp
llama.cpp is an open-source program that runs large language models on your own computer.
llama.cpp is an open-source program that runs large language models on your own computer. It is written in plain C and C++ and needs no other software packages. Its code uses the MIT License, so anyone may use and change it.
A model is usually built in a framework such as PyTorch. To use it with llama.cpp, you first convert it into a file in a format called GGUF. That file holds the model’s weights, which are the numbers it learned in training, plus the details needed to load it.
The big trick is quantization. Each weight is stored with fewer bits, for example 4 instead of 16, so the file shrinks and needs less memory. The model may become slightly less accurate. llama.cpp then runs the file on a processor, a graphics card or both at once. You can run it from the command line, or start a small server that other apps talk to as if it were the OpenAI API.
Running a language model yourself used to run into three walls.
Follow one model from download to a chat reply.
- 1 · convertA script turns a model from a framework such as PyTorch into a GGUF file.
- 2 · shrinkThe quantize tool stores each weight with fewer bits, for example 4 instead of 16.
- 3 · runllama.cpp loads the file and does the model's maths on your processor, graphics card or both.
- 4 · serveYou run it from the command line, or start a local server that apps call like the OpenAI API.
Fewer bits per weight means a smaller file and less memory, with some possible loss of accuracy.
| Who | What they ask | What it works with |
|---|---|---|
| Student | “Can I try an open model on my own laptop?” | A small quantized GGUF model run from the command line |
| App developer | “Can my app call a local model with the same code it uses for a cloud API?” | The OpenAI-compatible routes of llama-server |
| Model publisher | “How do I offer a smaller version of my model for home computers?” | The convert script and the quantize tool |
| Hobbyist | “My model is bigger than my graphics card memory. Can it still run?” | Mixed CPU and GPU inference |
- It runs language models locally with a plain C and C++ program and no extra dependencies.
- It shrinks models by storing each weight in 1.5 to 8 bits, which cuts memory use.
- It supports many chips, from Apple silicon to NVIDIA and AMD graphics cards.
- Its server answers the same kind of requests as the OpenAI API.
- Shrinking a model can lower its accuracy.
- Models from frameworks like PyTorch are converted to GGUF first.
Sources used
This explainer is written in original language. The links below support its factual claims.
- repoggml-org/llama.cpp: LLM inference in C/C++, ggml-org · read 28 Sept 2026
- repollama.cpp LICENSE, ggml-org · read 28 Sept 2026
- docsLLaMA.cpp HTTP Server, ggml-org · read 28 Sept 2026
- docsquantize, ggml-org · read 28 Sept 2026
- officialGGUF, ggml-org · read 28 Sept 2026
- docsGGUF, Hugging Face · read 28 Sept 2026