# When to Use Gemini vs OpenMaaS Models on Agent Platform: A Complete Decision Guide

> Decide between Gemini and Open MaaS models for your Agent Platform. Learn when to use Gemini for security and Google Cloud integration or Open MaaS for model variety and pricing.

- Repository: [Google/skills](https://github.com/google/skills)
- Tags: decision-guide
- Published: 2026-08-09

---

**Choose Gemini for enterprise security, compliance, and deep Google Cloud integration; choose Open MaaS for third-party model variety, competitive pricing, and OpenAI SDK compatibility.**

The Agent Platform (formerly Gemini Enterprise Agent Platform) supports two distinct model families for inference workloads. According to the source code in [`skills/cloud/agent-platform-inference/SKILL.md`](https://github.com/google/skills/blob/main/skills/cloud/agent-platform-inference/SKILL.md), selecting the right approach depends on your security requirements, existing codebase, and specific model capabilities. This guide breaks down the decision framework and implementation patterns used by the platform's inference skill.

## Key Differences Between Gemini and Open MaaS Models

Understanding the fundamental distinctions between these model families is critical for architectural decisions.

### Gemini (First-Party) Models

**Gemini models** are Google's native, first-party offerings including `gemini-1.5-flash` and `gemini-1.5-pro`. These models provide **tight integration with Google Cloud IAM**, audit logs, and VPC Service Controls. They support advanced features such as `generateContent` tool calls, system instructions, and fine-tuned LoRA adapters.

Use Gemini when you need consistent **regional availability and quota management**, or when your workload requires built-in compliance and security controls that align with existing Google Cloud infrastructure.

### Open MaaS (Third-Party) Models

**Open MaaS (Model-as-a-Service)** hosts third-party models from publishers like Meta (Llama 3), DeepSeek, and Qwen. These models are served through a **standard OpenAI-compatible endpoint**, making them ideal for teams with existing OpenAI client codebases.

This family offers broad model variety and competitive pricing, though you must handle model-specific licensing and security considerations yourself. The global endpoint architecture means no regional availability checks are required.

## Decision Framework from the Agent Platform Inference Skill

The [`skills/cloud/agent-platform-inference/SKILL.md`](https://github.com/google/skills/blob/main/skills/cloud/agent-platform-inference/SKILL.md) file defines a four-step decision tree for model selection:

1. **Identify the model family** – Determine if the user explicitly requests a Gemini model name (`gemini-*`) or a third-party model name (`meta/llama-*`, `deepseek-*`, etc.). If unclear, prompt the user for clarification.
2. **Choose the appropriate SDK** – For Gemini, use the **GenAI SDK** (preferred) or the legacy Vertex AI SDK. For Open MaaS, use the **OpenAI SDK** (highly recommended) with the GenAI SDK as a fallback option.
3. **Validate region and availability** – Gemini models require region-specific publisher endpoints; you must probe the exact model and region before generating code. Open MaaS models use a global endpoint, eliminating regional validation.
4. **Apply troubleshooting logic** – Handle specific error codes such as `429` (Resource Exhausted) or `400` (User Validation) using the skill's error-handling guidance.

## Implementation Examples

The following patterns demonstrate the correct SDK usage for each model family as specified in the Agent Platform Inference skill.

### Calling Gemini Models with the GenAI SDK

For first-party models, initialize the Vertex AI SDK and use the `GenerativeModel` class:

```python
from vertexai.preview import GenerativeModel
import vertexai

vertexai.init(project="my-gcp-project", location="us-central1")

model = GenerativeModel("gemini-1.5-flash")
response = model.generate_content("Explain quantum computing in simple terms.")
print(response.text)

```

This approach leverages the preferred SDK for Gemini models, providing access to multimodal inputs and function calling capabilities.

### Calling Open MaaS Models with the OpenAI SDK

For third-party models, configure the OpenAI client with the Agent Platform base URL:

```python
from openai import OpenAI

client = OpenAI(
    base_url="https://us-central1-aiplatform.googleapis.com/v1beta1/projects/my-project/locations/us-central1/publishers/openai",
    api_key="YOUR_GCLOUD_OAUTH_TOKEN"
)

resp = client.chat.completions.create(
    model="meta/llama-3.2-70b-instruct-maas",
    messages=[{"role": "user", "content": "Summarize the plot of 'The Matrix'"}],
)
print(resp.choices[0].message.content)

```

The skill explicitly recommends the OpenAI SDK for Open MaaS to maximize compatibility with existing codebases.

### Accessing Custom or Fine-Tuned Endpoints

For custom LoRA adapters or private endpoints, call the numeric endpoint directly using REST:

```bash
curl -sS -H "Authorization: Bearer $(gcloud auth print-access-token)" \
     -H "Content-Type: application/json" \
     "https://us-central1-aiplatform.googleapis.com/v1/projects/my-project/locations/us-central1/endpoints/1234567890123456789:predict" \
     -d '{"instances": [{"prompt": "Translate to French: Hello world"}]}'

```

This pattern applies when working with fine-tuned Gemini models or specialized deployments that require direct endpoint access.

## Summary

- **Gemini models** provide enterprise-grade security, Google Cloud IAM integration, and advanced features like tool calling, but require region-specific endpoint validation.
- **Open MaaS models** offer third-party variety and OpenAI SDK compatibility through a global endpoint, ideal for migrating existing OpenAI client code.
- The **GenAI SDK** is the preferred interface for Gemini, while the **OpenAI SDK** is recommended for Open MaaS.
- Always check **regional availability** for Gemini models using the probe logic defined in [`skills/cloud/agent-platform-inference/SKILL.md`](https://github.com/google/skills/blob/main/skills/cloud/agent-platform-inference/SKILL.md), while Open MaaS models require no such validation.

## Frequently Asked Questions

### What is the main SDK difference between Gemini and Open MaaS on Agent Platform?

Gemini models are optimized for the **GenAI SDK** (or legacy Vertex AI SDK), providing native access to multimodal capabilities and Google Cloud security features. Open MaaS models are designed for the **OpenAI SDK**, offering a drop-in replacement for existing OpenAI client implementations without code migration.

### Do I need to check region availability for Open MaaS models?

No. Open MaaS models are served through a **global endpoint**, eliminating the need for region-specific availability checks. Gemini models, however, require explicit validation of regional endpoints and quota availability before making API calls.

### Can I use the GenAI SDK with Open MaaS models?

Yes, though it is considered a **fallback option**. The Agent Platform Inference skill highly recommends using the **OpenAI SDK** for Open MaaS models to ensure maximum compatibility and leverage existing client code patterns.

### How do I handle authentication for custom model endpoints?

Custom endpoints for fine-tuned models require **OAuth 2.0 bearer tokens** obtained via `gcloud auth print-access-token`. Include this token in the `Authorization` header when making direct REST calls to numeric endpoint IDs, as shown in the custom endpoint implementation example.