Self-hosting large language models
Self-hosting means running a language model on a computer you control, so your prompts go to your own machine instead of a company's service.
With a cloud-hosted model, your prompts go to the provider, which processes them to run the service. Self-hosting flips this. You run the model on a computer you control. When it runs locally, the tool maker does not see your prompts. llama.cpp is built to run models on many kinds of hardware, whether on your own computer or in the cloud.
The hard part is size. A model’s weights are the numbers it learned in training. They are usually stored as 32-bit or 16-bit floating-point numbers (numbers with decimals). Quantization stores them with fewer bits (the 0s and 1s a computer uses), so the model needs less memory. Some methods go down to 8-bit or 4-bit integers, and llama.cpp supports levels from 1.5-bit up to 8-bit. Quantization methods try to keep as much accuracy as they can, so it is a trade.
Here is an everyday example. With Ollama installed, typing ollama run gemma4 starts a chat with the Gemma 4 model. The server waits at address 127.0.0.1, on port 11434. A port is like a numbered door. By default, only that computer can reach it. The ollama ps command shows whether the model sits on the GPU (the graphics chip), the CPU (the main chip) or both. Apps written for OpenAI can switch to it by changing the base address (the web address the app sends its requests to).
For engineers, there are dedicated serving tools. vLLM is a library for running and serving models. It batches incoming requests. It also uses a method called PagedAttention to manage the memory the model uses to keep track of the text so far. Hugging Face’s Text Generation Inference server is now in maintenance mode, and its docs recommend vLLM, SGLang, llama.cpp or MLX going forward.
Hosted AI services are easy, but they are not always the right fit.
Follow one model from download to a working app.
- 1 · pickYou download a model file, for example from Hugging Face.
- 2 · shrinkQuantization stores those numbers with fewer bits, so the model needs less memory.
- 3 · loadThe model is loaded into the graphics card, into ordinary memory, or split across both.
- 4 · serveA local server listens for requests, by default only from the same computer.
- 5 · callYour app sends prompts to that server, often with the same client code it would use for OpenAI.
Quantization is a key step: it cuts the memory a model needs.
| Who | What they ask | What it works with |
|---|---|---|
| A startup | “Can one server handle lots of incoming requests?” | A vLLM server that batches requests |
| A developer | “Can my existing OpenAI code run against a local model?” | The OpenAI-compatible endpoint on localhost |
- When the model runs locally, the tool maker does not see your prompts.
- Quantization lets a model load with less memory.
- Apps built for the OpenAI API can switch to a local server by changing the base address (the web address it sends requests to).
- vLLM batches incoming requests.
- Quantization methods try to keep as much accuracy as they can, so accuracy is part of the trade.
- A model can end up partly or fully in system memory instead of the GPU.
- Local servers copy only part of the OpenAI API.
- You configure the server yourself, including its bind address on your network.
Sources used
This explainer is written in original language. The links below support its factual claims.
- docsFAQ, Ollama · read 28 Sept 2026
- repoollama/ollama: Get up and running with open models, Ollama · read 28 Sept 2026
- docsOpenAI compatibility, Ollama · read 28 Sept 2026
- repoggml-org/llama.cpp: LLM inference in C/C++, ggml-org · read 28 Sept 2026
- docsWelcome to vLLM, vLLM project · read 28 Sept 2026
- docsQuantization overview, Hugging Face · read 28 Sept 2026
- docsText Generation Inference, Hugging Face · read 28 Sept 2026