Products

Vertex AI (now Gemini Enterprise Agent Platform)

5 min readintermediateUpdated 28 Sept 2026
1 · In one line

Vertex AI is Google Cloud’s platform for choosing AI models, tuning them on your own examples and running models and agents for real users.

1 · What it is

Vertex AI is Google Cloud’s service for building with AI models. Google Cloud now calls it Gemini Enterprise Agent Platform and presents it as the next step for Vertex AI.

The platform does four jobs. First, it helps you pick a model. Google says it offers more than 200 foundation models (large general-purpose models that others build on). The catalogue, called Model Garden, lists Google’s own models (Gemini, Gemma, Veo, Lyria). Partner models on the list include Anthropic’s Claude, Grok and Mistral AI models, and there are open-weights options (models whose trained numbers anyone can download) such as Qwen, Llama and DeepSeek. Companies can also set a rule that allows only models they have checked.

Second, it helps you tune a model. Tuning means training an existing model a little more on your own examples, so it gets better at one narrow job. Google’s advice is to try prompting first. Fine-tuning works best with a sizeable labelled set, ideally 100 examples or more. A labelled example is an input paired with the answer you want, such as a review paired with “unhappy”. Parameter-efficient tuning changes only a fairly small part of the model’s parameters, the numbers that shape its answers. Full fine-tuning changes every parameter.

Third, it runs models for real traffic. There are two ways to reach a model. You can use a managed API (Google calls this Model as a Service), or deploy the model yourself. With a managed API, you send a request and get an answer back, and you never run a server. Deploying means linking a model to its own machines so it can answer quickly, and giving it an endpoint, a web address your app sends requests to. You are charged for the machines behind a deployed open model. In return you get controls. Autoscaling adds or removes serving machines based on how many requests arrive at once. A new model version can share an endpoint with the old one. It starts with a small share of traffic and gets more until it handles all of it.

Fourth, it runs agents. Agent Runtime can run agents built with Agent Development Kit, a ready-made toolkit for building agents. Agents on Agent Runtime can use Sessions and Memory Bank to keep track of ongoing conversations and long-term memories.

Before adding more training examples, Google advises checking where the model goes wrong. For agents, Google offers an Example Store and an Evaluation Service. They help you test, monitor and trace what an agent does, and improve it over time.

Picture a small online shop that wants every review tagged with its own labels. The team tries a Gemini model from Model Garden with a clear prompt. If the labels keep drifting, they collect a few hundred reviews that staff have already tagged and tune the model on them. They put the tuned version on an endpoint, and their website sends each new review there. Later, they give a better version a small slice of reviews first.

2 · Why it exists

Turning a model into a working product raises three problems.

Too many choicesModels come from Google, from partner companies and from open-weights projects.
General, not specificA general model may not follow your format, rules or vocabulary well enough on its own.
Running it is hardA model or agent needs servers, a web address to call, and care for ongoing conversations before real users can rely on it.
3 · How it works

Follow one model from the shelf to a live app.

How a model moves through Vertex AI A team picks a model in Model Garden and tests prompts. The highlighted step is tuning on labelled examples, ideally a hundred or more, which makes a new model version. The model is deployed to an endpoint that apps call. Instead, a team can build an agent that runs on Agent Runtime and can use Sessions and Memory Bank. An evaluation service tests, monitors and traces what agents do. FROM THE MODEL SHELF TO A LIVE APP Model Garden Try prompts Tune Endpoint Your app Agent Runtime Evaluate Google, partner and open models clear instructions come first 100+ labelled examples, ideally → new version a web address apps can call sends a request, gets an answer runs agents with Sessions and Memory Bank test, monitor and trace agents or an agent Tuning is optional: Google suggests it only if prompting is not enough.
Tuning is optional; Google suggests moving on to it only if prompting is not enough.
  1. 1 · chooseYou browse Model Garden, a catalogue of Google, partner and open models, and test a few with sample prompts.
  2. 2 · promptYou first try to get good results with clear instructions alone, because Google recommends starting there.
  3. 3 · tuneIf prompting falls short, you tune the model on labelled examples so it learns your task.
  4. 4 · deployYou put the model on an endpoint, which gives your app a web address to send requests to.
  5. 5 · runFor agents, a managed runtime runs them, and they can use Sessions and Memory Bank to keep track of conversations.

Tuning makes a new version of the model.

4 · Where it's used
WhoWhat they askWhat it works with
Online shop“Can we tag each customer review as happy, unhappy or mixed, in our own labels?”A Gemini model tuned on a few hundred labelled reviews
Research lab“Can we try an open model without buying our own servers?”An open model deployed from Model Garden to an endpoint
Support team“Can our help agent remember what a customer said last week?”An agent running on Agent Runtime with Memory Bank
Security team“Can we allow only the models we have reviewed?”A Model Garden organisation policy
5 · What it solves, and what it doesn't
solves
  • It gathers Google, partner and open models in one catalogue with one way to deploy them.
  • It lets you tune Gemini models on your own labelled examples.
  • It runs deployed models behind an endpoint so apps can request answers.
  • It gives agents a managed place to run, plus Sessions and Memory Bank services they can use.
doesn't solve
  • Tuning needs good, well-labelled data; quality matters more than quantity.
  • Full fine-tuning needs more computing power to tune and to serve, so it costs more overall.
  • Some open models are flagged as suspicious in Model Garden but can still be deployed, so review them yourself.
  • If you deploy an open model yourself, you pay for the machines that serve it.
6 · Go deeper

Sources used

This explainer is written in original language. The links below support its factual claims.

  1. docsAgent Platform overview, Google Cloud · read 28 Sept 2026
  2. officialGemini Enterprise Agent Platform, Google Cloud · read 28 Sept 2026
  3. docsOverview of models on Agent Platform, Google Cloud · read 28 Sept 2026
  4. docsOverview of Model Garden, Google Cloud · read 28 Sept 2026
  5. docsIntroduction to tuning, Google Cloud · read 28 Sept 2026
  6. docsDeploy a model to an endpoint, Google Cloud · read 28 Sept 2026
  7. docsScale your agents, Google Cloud · read 28 Sept 2026