# How to Use OpenAI SDK with OpenMaaS Models on Vertex AI Agent Platform

> Leverage OpenAI SDK with OpenMaaS models on Vertex AI Agent Platform. Connect easily using an OpenAI-compatible endpoint and Google Cloud authentication for seamless integration.

- Repository: [Google/skills](https://github.com/google/skills)
- Tags: how-to-guide
- Published: 2026-08-09

---

**The Vertex AI Agent Platform exposes OpenMaaS (Model-as-a-Service) models through a standard OpenAI-compatible endpoint, allowing you to use the official `openai` Python SDK with only a modified `base_url` and Google Cloud access token authentication.**

The `google/skills` repository provides working examples that demonstrate how to integrate the OpenAI SDK with third-party models hosted on Google Cloud's Agent Platform. By leveraging the OpenAI-compatible API surface, you can deploy cutting-edge open models like DeepSeek and GLM using familiar SDK patterns without custom client implementations.

## Authentication and Client Configuration

To connect the OpenAI SDK to Vertex AI, you must authenticate using Google Cloud credentials and configure the client to point to the Vertex AI endpoint instead of the default OpenAI API.

### Obtaining Google Cloud Access Tokens

Unlike standard OpenAI API keys, the Agent Platform requires a valid Google Cloud OAuth access token. As implemented in [`openmaas_openai_sdk.py`](https://github.com/google/skills/blob/main/openmaas_openai_sdk.py) (lines 8-12), you retrieve this token from the default application credentials:

```python
import google.auth
import google.auth.transport.requests

def get_gcp_access_token():
    creds, _ = google.auth.default()
    # Refresh to guarantee a valid token

    creds.refresh(google.auth.transport.requests.Request())
    return creds.token

```

This function acquires credentials from your environment—whether from a service account, `gcloud` CLI login, or workload identity—and refreshes them to ensure a valid token for the API request.

### Configuring the Base URL

The OpenAI client must target the Vertex AI REST endpoint rather than `api.openai.com`. According to [`openmaas_openai_sdk.py`](https://github.com/google/skills/blob/main/openmaas_openai_sdk.py) (lines 17-20), construct the `base_url` using your Google Cloud project ID and desired region:

```python
_, project_id = google.auth.default()

client = openai.OpenAI(
    base_url=(
        f"https://aiplatform.googleapis.com/v1/projects/{project_id}"
        "/locations/global/endpoints/openapi"
    ),
    api_key=get_gcp_access_token(),
)

```

The endpoint follows the pattern `https://aiplatform.googleapis.com/v1/projects/<PROJECT_ID>/locations/<REGION>/endpoints/openapi`. The example above uses the `global` location, but you can substitute specific regions like `us-central1` for latency-sensitive workloads.

## Model Naming Conventions

OpenMaaS models on the Agent Platform use a `publisher/model` identifier format. As shown in [`openmaas_openai_sdk.py`](https://github.com/google/skills/blob/main/openmaas_openai_sdk.py) (lines 22-24), you specify models using this syntax:

- `deepseek-ai/deepseek-v3.2-maas`
- `zai-org/glm-5-maas`

This string is passed directly to the `model` parameter in your chat completion requests, replacing the standard OpenAI model names like `gpt-4` or `gpt-3.5-turbo`.

## Making Chat Completion Requests

Once configured, the request flow mirrors standard OpenAI SDK usage. The example in [`openmaas_openai_sdk.py`](https://github.com/google/skills/blob/main/openmaas_openai_sdk.py) (lines 25-33) demonstrates a complete chat completion:

```python
model_name = "zai-org/glm-5-maas"

response = client.chat.completions.create(
    model=model_name,
    messages=[
        {"role": "user", "content": "Explain quantum computing in simple terms."}
    ],
)

print(response.choices[0].message.content)

```

The SDK serializes the payload, sends it to the Vertex AI endpoint, and returns a response object compatible with the OpenAI schema. You access the generated content via `response.choices[0].message.content`, exactly as you would with native OpenAI models.

## Regional Endpoint Configuration

While the default configuration uses the `global` endpoint, you can optimize for latency by targeting specific regions. As demonstrated in [`gemini_openai_sdk.py`](https://github.com/google/skills/blob/main/gemini_openai_sdk.py) (lines 19-21), modify the location segment of the URL:

```python

# Target a specific region instead of global

base_url=(
    f"https://aiplatform.googleapis.com/v1/projects/{project_id}"
    "/locations/us-central1/endpoints/openapi"
)

```

Regional endpoints reduce network latency by processing requests within specific geographic boundaries, which is critical for production applications requiring sub-second response times.

## Complete Working Example

Below is the full implementation from [`openmaas_openai_sdk.py`](https://github.com/google/skills/blob/main/openmaas_openai_sdk.py), combining authentication, client configuration, and inference:

```python
"""Use the OpenAI SDK with an OpenMaaS model on Vertex AI."""

import google.auth
import google.auth.transport.requests
import openai

# Obtain a fresh GCP access token from default credentials

def get_gcp_access_token():
    creds, _ = google.auth.default()
    creds.refresh(google.auth.transport.requests.Request())
    return creds.token

# Resolve the current Google Cloud project ID

_, project_id = google.auth.default()

# Build the OpenAI client pointing to the Vertex AI "openapi" endpoint

client = openai.OpenAI(
    base_url=(
        f"https://aiplatform.googleapis.com/v1/projects/{project_id}"
        "/locations/global/endpoints/openapi"
    ),
    api_key=get_gcp_access_token(),
)

# Choose an OpenMaaS model (publisher/model format)

model_name = "zai-org/glm-5-maas"

# Send a chat request

response = client.chat.completions.create(
    model=model_name,
    messages=[
        {"role": "user", "content": "Explain quantum computing in simple terms."}
    ],
)

print(response.choices[0].message.content)

```

## Summary

- **Authentication**: Use `google.auth.default()` to obtain an OAuth access token and pass it as the `api_key` parameter when instantiating the OpenAI client.
- **Endpoint**: Set `base_url` to `https://aiplatform.googleapis.com/v1/projects/{project_id}/locations/{region}/endpoints/openapi` to route requests through Vertex AI Agent Platform.
- **Model Format**: Specify OpenMaaS models using the `publisher/model` syntax (e.g., `deepseek-ai/deepseek-v3.2-maas`).
- **Compatibility**: The response schema matches standard OpenAI SDK expectations, allowing drop-in replacement for existing applications.
- **Regional Optimization**: Replace `global` with specific regions like `us-central1` in the endpoint URL to reduce latency.

## Frequently Asked Questions

### Do I need to modify the OpenAI SDK to work with Vertex AI Agent Platform?

No modification is required. The official `openai` Python SDK works out of the box when you configure the `base_url` to point to the Vertex AI endpoint and provide a valid Google Cloud access token as the `api_key`. This OpenAI-compatible interface is maintained by the Agent Platform to ensure seamless integration with existing codebases.

### What model names are available for OpenMaaS on Agent Platform?

OpenMaaS models use a `publisher/model` naming convention. Examples include `deepseek-ai/deepseek-v3.2-maas` and `zai-org/glm-5-maas`. Refer to the [`SKILL.md`](https://github.com/google/skills/blob/main/SKILL.md) file in the `google/skills` repository for the current list of supported models and their specific capabilities.

### How does authentication differ from standard OpenAI API usage?

Instead of using a static API key from the OpenAI dashboard, you authenticate using Google Cloud's OAuth 2.0 access tokens obtained via `google.auth.default()`. These tokens are temporary and must be refreshed periodically, which the SDK handles automatically when you implement the credential refresh pattern shown in [`openmaas_openai_sdk.py`](https://github.com/google/skills/blob/main/openmaas_openai_sdk.py).

### Can I use region-specific endpoints for better performance?

Yes. While the examples default to the `global` location, you can specify regional endpoints by changing the location segment in the URL to regions like `us-central1`, `europe-west4`, or `asia-northeast1`. This is documented in [`gemini_openai_sdk.py`](https://github.com/google/skills/blob/main/gemini_openai_sdk.py) and allows you to minimize network latency by processing requests closer to your application deployment.