How to Use OpenAI SDK with OpenMaaS Models on Vertex AI Agent Platform

The Vertex AI Agent Platform exposes OpenMaaS (Model-as-a-Service) models through a standard OpenAI-compatible endpoint, allowing you to use the official openai Python SDK with only a modified base_url and Google Cloud access token authentication.

The google/skills repository provides working examples that demonstrate how to integrate the OpenAI SDK with third-party models hosted on Google Cloud's Agent Platform. By leveraging the OpenAI-compatible API surface, you can deploy cutting-edge open models like DeepSeek and GLM using familiar SDK patterns without custom client implementations.

Authentication and Client Configuration

To connect the OpenAI SDK to Vertex AI, you must authenticate using Google Cloud credentials and configure the client to point to the Vertex AI endpoint instead of the default OpenAI API.

Obtaining Google Cloud Access Tokens

Unlike standard OpenAI API keys, the Agent Platform requires a valid Google Cloud OAuth access token. As implemented in openmaas_openai_sdk.py (lines 8-12), you retrieve this token from the default application credentials:

import google.auth
import google.auth.transport.requests

def get_gcp_access_token():
    creds, _ = google.auth.default()
    # Refresh to guarantee a valid token

    creds.refresh(google.auth.transport.requests.Request())
    return creds.token

This function acquires credentials from your environment—whether from a service account, gcloud CLI login, or workload identity—and refreshes them to ensure a valid token for the API request.

Configuring the Base URL

The OpenAI client must target the Vertex AI REST endpoint rather than api.openai.com. According to openmaas_openai_sdk.py (lines 17-20), construct the base_url using your Google Cloud project ID and desired region:

_, project_id = google.auth.default()

client = openai.OpenAI(
    base_url=(
        f"https://aiplatform.googleapis.com/v1/projects/{project_id}"
        "/locations/global/endpoints/openapi"
    ),
    api_key=get_gcp_access_token(),
)

The endpoint follows the pattern https://aiplatform.googleapis.com/v1/projects/<PROJECT_ID>/locations/<REGION>/endpoints/openapi. The example above uses the global location, but you can substitute specific regions like us-central1 for latency-sensitive workloads.

Model Naming Conventions

OpenMaaS models on the Agent Platform use a publisher/model identifier format. As shown in openmaas_openai_sdk.py (lines 22-24), you specify models using this syntax:

  • deepseek-ai/deepseek-v3.2-maas
  • zai-org/glm-5-maas

This string is passed directly to the model parameter in your chat completion requests, replacing the standard OpenAI model names like gpt-4 or gpt-3.5-turbo.

Making Chat Completion Requests

Once configured, the request flow mirrors standard OpenAI SDK usage. The example in openmaas_openai_sdk.py (lines 25-33) demonstrates a complete chat completion:

model_name = "zai-org/glm-5-maas"

response = client.chat.completions.create(
    model=model_name,
    messages=[
        {"role": "user", "content": "Explain quantum computing in simple terms."}
    ],
)

print(response.choices[0].message.content)

The SDK serializes the payload, sends it to the Vertex AI endpoint, and returns a response object compatible with the OpenAI schema. You access the generated content via response.choices[0].message.content, exactly as you would with native OpenAI models.

Regional Endpoint Configuration

While the default configuration uses the global endpoint, you can optimize for latency by targeting specific regions. As demonstrated in gemini_openai_sdk.py (lines 19-21), modify the location segment of the URL:


# Target a specific region instead of global

base_url=(
    f"https://aiplatform.googleapis.com/v1/projects/{project_id}"
    "/locations/us-central1/endpoints/openapi"
)

Regional endpoints reduce network latency by processing requests within specific geographic boundaries, which is critical for production applications requiring sub-second response times.

Complete Working Example

Below is the full implementation from openmaas_openai_sdk.py, combining authentication, client configuration, and inference:

"""Use the OpenAI SDK with an OpenMaaS model on Vertex AI."""

import google.auth
import google.auth.transport.requests
import openai

# Obtain a fresh GCP access token from default credentials

def get_gcp_access_token():
    creds, _ = google.auth.default()
    creds.refresh(google.auth.transport.requests.Request())
    return creds.token

# Resolve the current Google Cloud project ID

_, project_id = google.auth.default()

# Build the OpenAI client pointing to the Vertex AI "openapi" endpoint

client = openai.OpenAI(
    base_url=(
        f"https://aiplatform.googleapis.com/v1/projects/{project_id}"
        "/locations/global/endpoints/openapi"
    ),
    api_key=get_gcp_access_token(),
)

# Choose an OpenMaaS model (publisher/model format)

model_name = "zai-org/glm-5-maas"

# Send a chat request

response = client.chat.completions.create(
    model=model_name,
    messages=[
        {"role": "user", "content": "Explain quantum computing in simple terms."}
    ],
)

print(response.choices[0].message.content)

Summary

  • Authentication: Use google.auth.default() to obtain an OAuth access token and pass it as the api_key parameter when instantiating the OpenAI client.
  • Endpoint: Set base_url to https://aiplatform.googleapis.com/v1/projects/{project_id}/locations/{region}/endpoints/openapi to route requests through Vertex AI Agent Platform.
  • Model Format: Specify OpenMaaS models using the publisher/model syntax (e.g., deepseek-ai/deepseek-v3.2-maas).
  • Compatibility: The response schema matches standard OpenAI SDK expectations, allowing drop-in replacement for existing applications.
  • Regional Optimization: Replace global with specific regions like us-central1 in the endpoint URL to reduce latency.

Frequently Asked Questions

Do I need to modify the OpenAI SDK to work with Vertex AI Agent Platform?

No modification is required. The official openai Python SDK works out of the box when you configure the base_url to point to the Vertex AI endpoint and provide a valid Google Cloud access token as the api_key. This OpenAI-compatible interface is maintained by the Agent Platform to ensure seamless integration with existing codebases.

What model names are available for OpenMaaS on Agent Platform?

OpenMaaS models use a publisher/model naming convention. Examples include deepseek-ai/deepseek-v3.2-maas and zai-org/glm-5-maas. Refer to the SKILL.md file in the google/skills repository for the current list of supported models and their specific capabilities.

How does authentication differ from standard OpenAI API usage?

Instead of using a static API key from the OpenAI dashboard, you authenticate using Google Cloud's OAuth 2.0 access tokens obtained via google.auth.default(). These tokens are temporary and must be refreshed periodically, which the SDK handles automatically when you implement the credential refresh pattern shown in openmaas_openai_sdk.py.

Can I use region-specific endpoints for better performance?

Yes. While the examples default to the global location, you can specify regional endpoints by changing the location segment in the URL to regions like us-central1, europe-west4, or asia-northeast1. This is documented in gemini_openai_sdk.py and allows you to minimize network latency by processing requests closer to your application deployment.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →