When to Use Gemini vs OpenMaaS Models on Agent Platform: A Complete Decision Guide

Choose Gemini for enterprise security, compliance, and deep Google Cloud integration; choose Open MaaS for third-party model variety, competitive pricing, and OpenAI SDK compatibility.

The Agent Platform (formerly Gemini Enterprise Agent Platform) supports two distinct model families for inference workloads. According to the source code in skills/cloud/agent-platform-inference/SKILL.md, selecting the right approach depends on your security requirements, existing codebase, and specific model capabilities. This guide breaks down the decision framework and implementation patterns used by the platform's inference skill.

Key Differences Between Gemini and Open MaaS Models

Understanding the fundamental distinctions between these model families is critical for architectural decisions.

Gemini (First-Party) Models

Gemini models are Google's native, first-party offerings including gemini-1.5-flash and gemini-1.5-pro. These models provide tight integration with Google Cloud IAM, audit logs, and VPC Service Controls. They support advanced features such as generateContent tool calls, system instructions, and fine-tuned LoRA adapters.

Use Gemini when you need consistent regional availability and quota management, or when your workload requires built-in compliance and security controls that align with existing Google Cloud infrastructure.

Open MaaS (Third-Party) Models

Open MaaS (Model-as-a-Service) hosts third-party models from publishers like Meta (Llama 3), DeepSeek, and Qwen. These models are served through a standard OpenAI-compatible endpoint, making them ideal for teams with existing OpenAI client codebases.

This family offers broad model variety and competitive pricing, though you must handle model-specific licensing and security considerations yourself. The global endpoint architecture means no regional availability checks are required.

Decision Framework from the Agent Platform Inference Skill

The skills/cloud/agent-platform-inference/SKILL.md file defines a four-step decision tree for model selection:

  1. Identify the model family – Determine if the user explicitly requests a Gemini model name (gemini-*) or a third-party model name (meta/llama-*, deepseek-*, etc.). If unclear, prompt the user for clarification.
  2. Choose the appropriate SDK – For Gemini, use the GenAI SDK (preferred) or the legacy Vertex AI SDK. For Open MaaS, use the OpenAI SDK (highly recommended) with the GenAI SDK as a fallback option.
  3. Validate region and availability – Gemini models require region-specific publisher endpoints; you must probe the exact model and region before generating code. Open MaaS models use a global endpoint, eliminating regional validation.
  4. Apply troubleshooting logic – Handle specific error codes such as 429 (Resource Exhausted) or 400 (User Validation) using the skill's error-handling guidance.

Implementation Examples

The following patterns demonstrate the correct SDK usage for each model family as specified in the Agent Platform Inference skill.

Calling Gemini Models with the GenAI SDK

For first-party models, initialize the Vertex AI SDK and use the GenerativeModel class:

from vertexai.preview import GenerativeModel
import vertexai

vertexai.init(project="my-gcp-project", location="us-central1")

model = GenerativeModel("gemini-1.5-flash")
response = model.generate_content("Explain quantum computing in simple terms.")
print(response.text)

This approach leverages the preferred SDK for Gemini models, providing access to multimodal inputs and function calling capabilities.

Calling Open MaaS Models with the OpenAI SDK

For third-party models, configure the OpenAI client with the Agent Platform base URL:

from openai import OpenAI

client = OpenAI(
    base_url="https://us-central1-aiplatform.googleapis.com/v1beta1/projects/my-project/locations/us-central1/publishers/openai",
    api_key="YOUR_GCLOUD_OAUTH_TOKEN"
)

resp = client.chat.completions.create(
    model="meta/llama-3.2-70b-instruct-maas",
    messages=[{"role": "user", "content": "Summarize the plot of 'The Matrix'"}],
)
print(resp.choices[0].message.content)

The skill explicitly recommends the OpenAI SDK for Open MaaS to maximize compatibility with existing codebases.

Accessing Custom or Fine-Tuned Endpoints

For custom LoRA adapters or private endpoints, call the numeric endpoint directly using REST:

curl -sS -H "Authorization: Bearer $(gcloud auth print-access-token)" \
     -H "Content-Type: application/json" \
     "https://us-central1-aiplatform.googleapis.com/v1/projects/my-project/locations/us-central1/endpoints/1234567890123456789:predict" \
     -d '{"instances": [{"prompt": "Translate to French: Hello world"}]}'

This pattern applies when working with fine-tuned Gemini models or specialized deployments that require direct endpoint access.

Summary

  • Gemini models provide enterprise-grade security, Google Cloud IAM integration, and advanced features like tool calling, but require region-specific endpoint validation.
  • Open MaaS models offer third-party variety and OpenAI SDK compatibility through a global endpoint, ideal for migrating existing OpenAI client code.
  • The GenAI SDK is the preferred interface for Gemini, while the OpenAI SDK is recommended for Open MaaS.
  • Always check regional availability for Gemini models using the probe logic defined in skills/cloud/agent-platform-inference/SKILL.md, while Open MaaS models require no such validation.

Frequently Asked Questions

What is the main SDK difference between Gemini and Open MaaS on Agent Platform?

Gemini models are optimized for the GenAI SDK (or legacy Vertex AI SDK), providing native access to multimodal capabilities and Google Cloud security features. Open MaaS models are designed for the OpenAI SDK, offering a drop-in replacement for existing OpenAI client implementations without code migration.

Do I need to check region availability for Open MaaS models?

No. Open MaaS models are served through a global endpoint, eliminating the need for region-specific availability checks. Gemini models, however, require explicit validation of regional endpoints and quota availability before making API calls.

Can I use the GenAI SDK with Open MaaS models?

Yes, though it is considered a fallback option. The Agent Platform Inference skill highly recommends using the OpenAI SDK for Open MaaS models to ensure maximum compatibility and leverage existing client code patterns.

How do I handle authentication for custom model endpoints?

Custom endpoints for fine-tuned models require OAuth 2.0 bearer tokens obtained via gcloud auth print-access-token. Include this token in the Authorization header when making direct REST calls to numeric endpoint IDs, as shown in the custom endpoint implementation example.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →