Large language model operations
The day-to-day work of running an app built on a language model, from managing prompts and testing answers to watching speed and cost.
LLMOps is short for “large language model operations”. LLMOps is the set of habits and tools for running an app built on one once real people depend on it. It covers the everyday work around the model: keeping prompts under control, testing answers, and watching speed and cost. Microsoft calls the same work generative AI operations, or GenAIOps, and treats it as an extension of MLOps.
Google Cloud describes LLMOps as a specialized subset of MLOps, short for machine learning operations. Much stays the same, but the focus moves. Many language models start as a general foundation model, a big model trained for many jobs. It is then fine-tuned, meaning trained a little more for one job, instead of being built from zero. Often, an app team never trains anything. They work on the pipeline around the model. That means the steps that prepare a question, send it and handle the answer.
That shifts what needs managing. In classic MLOps, the app that calls the model is often left outside the process. With a language model app, you manage that part too, because it writes the prompt, the instructions sent to the model. Prompt templates matter for accurate, reliable answers.
Picture a school that runs a homework-help chatbot. A teacher asks for friendlier replies, and a developer rewrites one sentence of the prompt. With LLMOps, the new wording is saved as version 3 in a prompt registry, a store that lets teams version, track and reuse prompts. Versions cannot be changed after saving, so version 2 always means the same text.
Next comes evaluation. Evals, short for evaluations, test the model’s answers against rules the team chooses. OpenAI calls them essential, especially when you upgrade or try a new model. The team runs version 2 and version 3 on the same set of real homework questions and compares the scores. Only if version 3 holds up does the production alias, a movable label saying which version is live, point to it. If something goes wrong later, moving the label back is a rollback.
Measuring quality also works differently. Classic models have clear scores, such as accuracy (how often they are right). A chatbot that answers from school notes needs measures like groundedness, meaning whether the answer sticks to its source, and relevancy. These tasks have no single right answer, so feedback from real users is often critical. Feeding it back into the pipeline helps both testing and later fine-tuning.
Once live, the app needs watching. Microsoft suggests watching response time, how many tokens (small chunks of text) are used, and 429 (too many requests) errors. Tracing tools such as MLflow record the time and tokens each step used.
Cost is part of the job as well. If you use a model as a hosted service, someone else runs the servers. Instead you care about how much work the service handles, your quota (your usage allowance), and being slowed down when you hit limits. Azure OpenAI bills by tokens, so watching quota usage is how a team keeps spending in check.
How much of this a project needs depends on the project.
| MLOps | LLMOps | |
|---|---|---|
| Typical quality measures | Accuracy, F1 score | Groundedness, relevancy, human feedback |
A language model app that works in a demo can still fail once real people use it.
Follow one prompt change from edit to production.
- 1 · versionSave each prompt edit as a new, unchangeable version in a registry.
- 2 · evaluateRun evals, tests of model output against criteria you choose, before the new version goes live.
- 3 · releasePoint a production alias at the new version, which also makes rolling back easy.
- 4 · monitorTrack latency, token usage and errors while real users send requests.
- 5 · learnFeed what you learn from real traffic back into the next round of changes.
In LLMOps, a prompt edit is a release, so it gets tested like one.
| Who | What they ask | What it works with |
|---|---|---|
| App developer | “Did my new prompt make answers worse on the questions we already handle?” | Eval results for the old and new prompt versions |
| Operations engineer | “Why are some users waiting so long for a reply?” | Traces with latency and token usage per step |
| Product owner | “Which prompt version is live right now, and can we go back?” | The production alias in the prompt registry |
| Finance team | “How much are model calls costing us each week?” | Token usage and quota logs |
- It keeps a clear record of which prompt version is running.
- It catches quality drops from a prompt or model change before release.
- It shows where time and tokens go in each request.
- It makes rolling back a bad change straightforward.
- It does not make the model itself smarter or more truthful.
- It does not remove the need for human judgement on open-ended answers.
- It does not set how much a model provider charges per token.
Sources used
This explainer is written in original language. The links below support its factual claims.
- docsGenerative AI operations for organizations with MLOps investments, Microsoft Learn · read 28 Sept 2026
- officialWhat is LLMOps (large language model operations)?, Google Cloud · read 28 Sept 2026
- officialWhat is LLMOps?, Databricks · read 28 Sept 2026
- docsPrompt Registry, MLflow · read 28 Sept 2026
- docsWorking with evals, OpenAI · read 28 Sept 2026
- docsLLM Tracing and Agent Observability, MLflow · read 28 Sept 2026