Open source

Ollama

4 min readbeginnerUpdated 28 Sept 2026
1 · In one line

Ollama is an open-source program that downloads open AI models and runs them on your own computer, with a simple command and a local API.

1 · What it is

Ollama is an open-source program for running open AI models, the kind anyone can download, on your own computer. You can install it on Linux, macOS or Windows. Its code is under the MIT License, so anyone may use and change it.

You type one command, such as ollama run with a model’s name. Ollama fetches that model from its online library and saves it in a folder on your machine. Then it loads the model into memory, on the main processor, the graphics card or a mix of both, and you can chat in the terminal, the text window where you type commands. A loaded model stays in memory for five minutes by default.

Behind the scenes, Ollama runs a small local server. It listens only on your own computer, at port 11434, a numbered doorway other programs knock on. Apps can send questions there through Ollama’s REST API, a standard way for programs to talk over the web. They can also use routes that copy part of the OpenAI API. Local requests need no key, and Ollama says it does not see your prompts when you run locally. It also offers optional cloud models, which you can switch off.

2 · Why it exists

Trying an open model on your own computer has three snags.

Finding and fetchingYou have to find a model, download its files and keep track of where they live.
Running it wellThe model has to be loaded into memory on the right chip, the main processor or the graphics card.
Connecting appsOther programs need a steady way to send questions to the model and get answers back.
3 · How it works

Follow one question from your keyboard to a local reply.

How Ollama answers a question on your own computer Five boxes in a row. You type ollama run with a model name. Ollama downloads the model from its library and stores it. The highlighted step, the local Ollama server, loads the model on the CPU, the GPU or both. It listens at localhost port 11434. The reply goes back to the terminal or an app. FROM ONE COMMAND TO A REPLY, ALL ON YOUR MACHINE You ask Pull Stored model Ollama server Reply ollama run + a model name, or an app calls the API download from the Ollama library, then stored saved in a folder on your computer loads it on CPU, GPU or both localhost:11434 in the terminal, or back to the calling app Local requests need no API key. A loaded model stays in memory 5 minutes by default.
The highlighted box is the Ollama server: one program on your computer that loads models and answers every app.
  1. 1 · pickYou choose a model from the Ollama library by name, for example with ollama run.
  2. 2 · pullOllama downloads the model files and keeps them in a folder on your computer.
  3. 3 · loadThe Ollama server loads the model into memory on the processor, the graphics card or both.
  4. 4 · answerYour terminal or any app sends a question to the local API and gets the reply back.

Once a model is downloaded, local requests need no API key and your prompts stay on your machine.

4 · Where it's used
WhoWhat they askWhat it works with
Student“Can I chat with an open model on my laptop without paying for an API?”A model pulled from the Ollama library and run in the terminal
App developer“Can my code talk to a local model the same way it talks to OpenAI?”The OpenAI-compatible routes on localhost port 11434
Hobbyist“Can I make a chatbot with its own personality and settings?”A Modelfile with a SYSTEM message and parameters
Privacy-minded team“Can we keep our prompts off outside servers?”Local-only mode with cloud features turned off
5 · What it solves, and what it doesn't
solves
  • It gets an open model running on your computer with one command.
  • It stores downloaded models locally so they are ready next time.
  • It offers a REST API and OpenAI-compatible routes for other apps.
  • It lets you package a custom model with a Modelfile.
doesn't solve
  • Big models need lots of graphics memory; with too little, replies can be slower.
  • Cloud models send your prompts to Ollama's servers, so they are not fully local.
6 · Go deeper

Sources used

This explainer is written in original language. The links below support its factual claims.

  1. repoollama/ollama: Start building with open models, Ollama · read 28 Sept 2026
  2. repoOllama LICENSE, Ollama · read 28 Sept 2026
  3. docsAPI introduction, Ollama · read 28 Sept 2026
  4. docsOpenAI compatibility, Ollama · read 28 Sept 2026
  5. docsModelfile Reference, Ollama · read 28 Sept 2026
  6. docsFAQ, Ollama · read 28 Sept 2026
  7. docsQuickstart, Ollama · read 28 Sept 2026