Vertex AI (now Gemini Enterprise Agent Platform)
Vertex AI is Google Cloud’s platform for choosing AI models, tuning them on your own examples and running models and agents for real users.
Vertex AI is Google Cloud’s service for building with AI models. Google Cloud now calls it Gemini Enterprise Agent Platform and presents it as the next step for Vertex AI.
The platform does four jobs. First, it helps you pick a model. Google says it offers more than 200 foundation models (large general-purpose models that others build on). The catalogue, called Model Garden, lists Google’s own models (Gemini, Gemma, Veo, Lyria). Partner models on the list include Anthropic’s Claude, Grok and Mistral AI models, and there are open-weights options (models whose trained numbers anyone can download) such as Qwen, Llama and DeepSeek. Companies can also set a rule that allows only models they have checked.
Second, it helps you tune a model. Tuning means training an existing model a little more on your own examples, so it gets better at one narrow job. Google’s advice is to try prompting first. Fine-tuning works best with a sizeable labelled set, ideally 100 examples or more. A labelled example is an input paired with the answer you want, such as a review paired with “unhappy”. Parameter-efficient tuning changes only a fairly small part of the model’s parameters, the numbers that shape its answers. Full fine-tuning changes every parameter.
Third, it runs models for real traffic. There are two ways to reach a model. You can use a managed API (Google calls this Model as a Service), or deploy the model yourself. With a managed API, you send a request and get an answer back, and you never run a server. Deploying means linking a model to its own machines so it can answer quickly, and giving it an endpoint, a web address your app sends requests to. You are charged for the machines behind a deployed open model. In return you get controls. Autoscaling adds or removes serving machines based on how many requests arrive at once. A new model version can share an endpoint with the old one. It starts with a small share of traffic and gets more until it handles all of it.
Fourth, it runs agents. Agent Runtime can run agents built with Agent Development Kit, a ready-made toolkit for building agents. Agents on Agent Runtime can use Sessions and Memory Bank to keep track of ongoing conversations and long-term memories.
Before adding more training examples, Google advises checking where the model goes wrong. For agents, Google offers an Example Store and an Evaluation Service. They help you test, monitor and trace what an agent does, and improve it over time.
Picture a small online shop that wants every review tagged with its own labels. The team tries a Gemini model from Model Garden with a clear prompt. If the labels keep drifting, they collect a few hundred reviews that staff have already tagged and tune the model on them. They put the tuned version on an endpoint, and their website sends each new review there. Later, they give a better version a small slice of reviews first.
Turning a model into a working product raises three problems.
Follow one model from the shelf to a live app.
- 1 · chooseYou browse Model Garden, a catalogue of Google, partner and open models, and test a few with sample prompts.
- 2 · promptYou first try to get good results with clear instructions alone, because Google recommends starting there.
- 3 · tuneIf prompting falls short, you tune the model on labelled examples so it learns your task.
- 4 · deployYou put the model on an endpoint, which gives your app a web address to send requests to.
- 5 · runFor agents, a managed runtime runs them, and they can use Sessions and Memory Bank to keep track of conversations.
Tuning makes a new version of the model.
| Who | What they ask | What it works with |
|---|---|---|
| Online shop | “Can we tag each customer review as happy, unhappy or mixed, in our own labels?” | A Gemini model tuned on a few hundred labelled reviews |
| Research lab | “Can we try an open model without buying our own servers?” | An open model deployed from Model Garden to an endpoint |
| Support team | “Can our help agent remember what a customer said last week?” | An agent running on Agent Runtime with Memory Bank |
| Security team | “Can we allow only the models we have reviewed?” | A Model Garden organisation policy |
- It gathers Google, partner and open models in one catalogue with one way to deploy them.
- It lets you tune Gemini models on your own labelled examples.
- It runs deployed models behind an endpoint so apps can request answers.
- It gives agents a managed place to run, plus Sessions and Memory Bank services they can use.
- Tuning needs good, well-labelled data; quality matters more than quantity.
- Full fine-tuning needs more computing power to tune and to serve, so it costs more overall.
- Some open models are flagged as suspicious in Model Garden but can still be deployed, so review them yourself.
- If you deploy an open model yourself, you pay for the machines that serve it.
Sources used
This explainer is written in original language. The links below support its factual claims.
- docsAgent Platform overview, Google Cloud · read 28 Sept 2026
- officialGemini Enterprise Agent Platform, Google Cloud · read 28 Sept 2026
- docsOverview of models on Agent Platform, Google Cloud · read 28 Sept 2026
- docsOverview of Model Garden, Google Cloud · read 28 Sept 2026
- docsIntroduction to tuning, Google Cloud · read 28 Sept 2026
- docsDeploy a model to an endpoint, Google Cloud · read 28 Sept 2026
- docsScale your agents, Google Cloud · read 28 Sept 2026