MLOpsBuilding with AI

Machine learning operations

5 min readintermediateUpdated 28 Sept 2026
1 · In one line

MLOps is a way of working that makes building and releasing machine learning models simpler and more automatic.

1 · What it is

MLOps, short for machine learning operations, covers the practices teams use to make building and releasing machine learning models more automatic and simpler. Google Cloud describes MLOps as a culture and practice. It joins building ML (machine learning) systems with running them. It borrows ideas from DevOps, a set of software habits such as continuous integration and continuous delivery.

Why is a separate name needed? A 2015 paper at NeurIPS, a major AI research conference, warned that machine learning systems often carry large ongoing maintenance costs. A model can also get worse because the data it sees keeps changing, not only because of bugs in its code.

Picture a model that predicts how many ice creams a shop will sell each day. Then the shop moves to a busier street. The model’s inputs look different, and its guesses slowly drift off. MLOps is the set of routines that notices this and fixes it safely.

The first routine is testing more than code. A feature here means one input column, such as temperature or day of the week. A pipeline is the chain of steps that turns data into a model. A schema is the layout the data should follow. Missing features, surprise features or features with odd values all break that layout. When this happens, the pipeline should halt so people on the team can look into it.

The second routine is automation. Google Cloud’s guide describes three levels, from no automation up to fully automated pipelines. At level 0, a team runs every step by hand. At level 1, a team ships the whole training pipeline, not just one trained model. That pipeline can retrain and redeploy models by itself. A new run can start in several ways: by hand, on a timetable, or when fresh data arrives. It can also start when results get noticeably worse, or when the inputs look very different from before.

The most advanced stage, level 2, adds automated CI/CD (continuous integration and delivery). Here it means new pipeline parts are built, tested and deployed without manual steps. The finished model goes into the registry and is served to apps.

The third routine is a gate before release. A test set is data the model did not learn from, kept aside for marking its work. The new model is scored on a test set, and its scores are compared with the current model’s. The new model replaces the live one only if it does better.

The fourth routine is keeping records. A model registry is a catalogue of trained models. It links each version to the run that produced it, so a team can trace how it was trained. Azure Machine Learning can also log who published a model, why it changed and when it was deployed. MLflow is another tool with a model registry that tracks versions of each model. In MLflow, an alias is a movable name, such as champion, that points to one model version. Apps can ask for the champion instead of a version number, so the team switches the live model by moving the label.

The fifth routine is careful release and watching. An endpoint is the web address other apps call to get predictions. Azure Machine Learning can split traffic between two versions behind one endpoint. Each version gets only a share of the traffic. This is called A/B testing. Azure Machine Learning can also raise an alert when drift shows up in the data.

MLOps is a way of working, not one product. It does not replace people who check a model for fairness and bias before it goes live.

2 · Why it exists

Training a good model once is the easy part; keeping it useful in production is harder.

Model code is a small partIn a real system, the model code is only a small part of the whole system.
Models go staleWhen the data a model sees drifts away from its training data, its predictions can quietly get worse.
Hard to repeatExperiments try many settings, so without tracking it is hard to know which version worked or to rebuild it.
3 · How it works

Follow one model around the loop, from new data back to retraining.

The key gate: a newly trained model is compared with the current one before it can replace it.
  1. 1 · validateCheck the new data before training, and stop if it looks wrong.
  2. 2 · trainAn automated pipeline retrains the model on the fresh data.
  3. 3 · compareTest the new model and promote it only if it beats the current one.
  4. 4 · deployPush the approved model to the registry and serve it as a prediction service.
  5. 5 · monitorWatch live data and model quality, and trigger retraining when they slip.

MLOps adds continuous training to the usual software habits of testing and shipping.

4 · Where it's used
WhoWhat they askWhat it works with
Data scientist“Which experiment produced the model that is live right now?”Model registry lineage linking each version to its training run
ML engineer“Can the same pipeline run in testing and in production?”A versioned, automated training pipeline
Operations team“Has the incoming data changed enough to hurt predictions?”Data drift and model quality alerts
Reviewer or auditor“Who published this model and when did it go live?”Lifecycle metadata with publish and deployment records
5 · What it solves, and what it doesn't
solves
  • It turns manual, one-off training steps into a repeatable pipeline.
  • It checks data and models before a new version reaches users.
  • It keeps versions so a team can compare models or roll back.
  • It spots drifting data and can trigger retraining.
doesn't solve
  • It is mostly about running a model well, not about designing a better one.
  • It doesn't remove the need for human review of fairness and bias.
6 · Go deeper

Sources used

This explainer is written in original language. The links below support its factual claims.

  1. docsMLOps: Continuous delivery and automation pipelines in machine learning, Google Cloud · read 28 Sept 2026
  2. officialWhat is MLOps?, Amazon Web Services · read 28 Sept 2026
  3. docsMLOps model management with Azure Machine Learning, Microsoft Learn · read 28 Sept 2026
  4. docsMLflow Model Registry, MLflow · read 28 Sept 2026
  5. paperHidden Technical Debt in Machine Learning Systems, NeurIPS · read 28 Sept 2026