How to Use OpenAI SDK with OpenMaaS Models on Vertex AI Agent Platform
The Vertex AI Agent Platform exposes OpenMaaS (Model-as-a-Service) models through a standard OpenAI-compatible endpoint, allowing you to use the official openai Python SDK with only a modified base_url and Google Cloud access token authentication.
The google/skills repository provides working examples that demonstrate how to integrate the OpenAI SDK with third-party models hosted on Google Cloud's Agent Platform. By leveraging the OpenAI-compatible API surface, you can deploy cutting-edge open models like DeepSeek and GLM using familiar SDK patterns without custom client implementations.
Authentication and Client Configuration
To connect the OpenAI SDK to Vertex AI, you must authenticate using Google Cloud credentials and configure the client to point to the Vertex AI endpoint instead of the default OpenAI API.
Obtaining Google Cloud Access Tokens
Unlike standard OpenAI API keys, the Agent Platform requires a valid Google Cloud OAuth access token. As implemented in openmaas_openai_sdk.py (lines 8-12), you retrieve this token from the default application credentials:
import google.auth
import google.auth.transport.requests
def get_gcp_access_token():
creds, _ = google.auth.default()
# Refresh to guarantee a valid token
creds.refresh(google.auth.transport.requests.Request())
return creds.token
This function acquires credentials from your environment—whether from a service account, gcloud CLI login, or workload identity—and refreshes them to ensure a valid token for the API request.
Configuring the Base URL
The OpenAI client must target the Vertex AI REST endpoint rather than api.openai.com. According to openmaas_openai_sdk.py (lines 17-20), construct the base_url using your Google Cloud project ID and desired region:
_, project_id = google.auth.default()
client = openai.OpenAI(
base_url=(
f"https://aiplatform.googleapis.com/v1/projects/{project_id}"
"/locations/global/endpoints/openapi"
),
api_key=get_gcp_access_token(),
)
The endpoint follows the pattern https://aiplatform.googleapis.com/v1/projects/<PROJECT_ID>/locations/<REGION>/endpoints/openapi. The example above uses the global location, but you can substitute specific regions like us-central1 for latency-sensitive workloads.
Model Naming Conventions
OpenMaaS models on the Agent Platform use a publisher/model identifier format. As shown in openmaas_openai_sdk.py (lines 22-24), you specify models using this syntax:
deepseek-ai/deepseek-v3.2-maaszai-org/glm-5-maas
This string is passed directly to the model parameter in your chat completion requests, replacing the standard OpenAI model names like gpt-4 or gpt-3.5-turbo.
Making Chat Completion Requests
Once configured, the request flow mirrors standard OpenAI SDK usage. The example in openmaas_openai_sdk.py (lines 25-33) demonstrates a complete chat completion:
model_name = "zai-org/glm-5-maas"
response = client.chat.completions.create(
model=model_name,
messages=[
{"role": "user", "content": "Explain quantum computing in simple terms."}
],
)
print(response.choices[0].message.content)
The SDK serializes the payload, sends it to the Vertex AI endpoint, and returns a response object compatible with the OpenAI schema. You access the generated content via response.choices[0].message.content, exactly as you would with native OpenAI models.
Regional Endpoint Configuration
While the default configuration uses the global endpoint, you can optimize for latency by targeting specific regions. As demonstrated in gemini_openai_sdk.py (lines 19-21), modify the location segment of the URL:
# Target a specific region instead of global
base_url=(
f"https://aiplatform.googleapis.com/v1/projects/{project_id}"
"/locations/us-central1/endpoints/openapi"
)
Regional endpoints reduce network latency by processing requests within specific geographic boundaries, which is critical for production applications requiring sub-second response times.
Complete Working Example
Below is the full implementation from openmaas_openai_sdk.py, combining authentication, client configuration, and inference:
"""Use the OpenAI SDK with an OpenMaaS model on Vertex AI."""
import google.auth
import google.auth.transport.requests
import openai
# Obtain a fresh GCP access token from default credentials
def get_gcp_access_token():
creds, _ = google.auth.default()
creds.refresh(google.auth.transport.requests.Request())
return creds.token
# Resolve the current Google Cloud project ID
_, project_id = google.auth.default()
# Build the OpenAI client pointing to the Vertex AI "openapi" endpoint
client = openai.OpenAI(
base_url=(
f"https://aiplatform.googleapis.com/v1/projects/{project_id}"
"/locations/global/endpoints/openapi"
),
api_key=get_gcp_access_token(),
)
# Choose an OpenMaaS model (publisher/model format)
model_name = "zai-org/glm-5-maas"
# Send a chat request
response = client.chat.completions.create(
model=model_name,
messages=[
{"role": "user", "content": "Explain quantum computing in simple terms."}
],
)
print(response.choices[0].message.content)
Summary
- Authentication: Use
google.auth.default()to obtain an OAuth access token and pass it as theapi_keyparameter when instantiating the OpenAI client. - Endpoint: Set
base_urltohttps://aiplatform.googleapis.com/v1/projects/{project_id}/locations/{region}/endpoints/openapito route requests through Vertex AI Agent Platform. - Model Format: Specify OpenMaaS models using the
publisher/modelsyntax (e.g.,deepseek-ai/deepseek-v3.2-maas). - Compatibility: The response schema matches standard OpenAI SDK expectations, allowing drop-in replacement for existing applications.
- Regional Optimization: Replace
globalwith specific regions likeus-central1in the endpoint URL to reduce latency.
Frequently Asked Questions
Do I need to modify the OpenAI SDK to work with Vertex AI Agent Platform?
No modification is required. The official openai Python SDK works out of the box when you configure the base_url to point to the Vertex AI endpoint and provide a valid Google Cloud access token as the api_key. This OpenAI-compatible interface is maintained by the Agent Platform to ensure seamless integration with existing codebases.
What model names are available for OpenMaaS on Agent Platform?
OpenMaaS models use a publisher/model naming convention. Examples include deepseek-ai/deepseek-v3.2-maas and zai-org/glm-5-maas. Refer to the SKILL.md file in the google/skills repository for the current list of supported models and their specific capabilities.
How does authentication differ from standard OpenAI API usage?
Instead of using a static API key from the OpenAI dashboard, you authenticate using Google Cloud's OAuth 2.0 access tokens obtained via google.auth.default(). These tokens are temporary and must be refreshed periodically, which the SDK handles automatically when you implement the credential refresh pattern shown in openmaas_openai_sdk.py.
Can I use region-specific endpoints for better performance?
Yes. While the examples default to the global location, you can specify regional endpoints by changing the location segment in the URL to regions like us-central1, europe-west4, or asia-northeast1. This is documented in gemini_openai_sdk.py and allows you to minimize network latency by processing requests closer to your application deployment.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →