Difference Between Custom Endpoints and Publisher Endpoints in Vertex AI Agent Platform

Custom endpoints are dedicated, region-locked resources you deploy for tuned models or self-hosted LLMs, while publisher endpoints are shared routing services for first-party Gemini and OpenMaaS models that require runtime regional availability checks.

Understanding the distinction between these two invocation patterns is critical for architecting inference workloads on Google Cloud. According to the source code in the google/skills repository, specifically skills/cloud/agent-platform-inference/SKILL.md, the Vertex AI Agent Platform handles routing, regional validation, and SDK compatibility differently depending on whether you target a specific endpoint resource or the shared publisher service.

What Are Custom Endpoints?

A custom endpoint is a concrete, immutable resource that you create (or that is created automatically when you tune a Gemini model or deploy an open-source LLM). It represents a specific deployment with a fixed regional location.

Resource Structure and Routing

Custom endpoints follow a strict resource naming convention: projects/<PROJECT>/locations/<REGION>/endpoints/<ENDPOINT_ID>. When you send a request, the routing is direct to that specific resource. As implemented in the google/skills source, if the caller's region does not match the endpoint's deployment region, the lookup returns a clean 404 error without incurring inference costs. This behavior is documented in skills/cloud/agent-platform-inference/SKILL.md at lines 44-48.

When to Use Custom Endpoints

Use a custom endpoint in the following scenarios:

  • Tuned Gemini models served on a numeric endpoint ID (e.g., projects/my-project/locations/us-central1/endpoints/1234567890)
  • Self-deployed OSS LLMs such as Llama, DeepSeek, Qwen, or Gemma running on your own infrastructure
  • Legacy custom models requiring specific resource isolation

What Are Publisher Endpoints?

A publisher endpoint is a shared "publisher" service that acts as a façade, routing requests to first-party Gemini models or OpenMaaS (Open Model-as-a-Service) offerings through a global or regional base URL.

Shared Service Architecture

Unlike custom endpoints, publisher endpoints do not represent a specific resource you own. Instead, they use base URLs such as https://aiplatform.googleapis.com/v1/projects/<PROJECT>/locations/global/endpoints/openapi or regional variants. The service handles routing to the underlying model, which may vary in availability by region. This architectural distinction is detailed in skills/cloud/agent-platform-inference/SKILL.md at lines 55-58.

Regional Availability Considerations

Publisher endpoints require runtime availability verification. You must probe the model in the target region (for example, via a :generateContent call) to confirm support. A 200 response indicates availability, while a 404 means the model is not supported in that region, and you must report the supported regions accordingly. This check is mandatory because, unlike custom endpoints, the publisher service is not tied to a single immutable region.

Key Differences Between Custom and Publisher Endpoints

Aspect Custom Endpoint Publisher Endpoint
Resource Type Specific endpoint resource you own Shared routing service
Region Behavior Fixed at deployment time; mismatched regions return immediate 404 Variable availability; requires runtime probing
Model Types Tuned Gemini, self-hosted OSS LLMs First-party Gemini, OpenMaaS models
Typical SDK Vertex AI SDK GenAI SDK (google-genai) or OpenAI SDK
URL Pattern .../locations/<REGION>/endpoints/<ENDPOINT_ID> .../locations/<REGION>/endpoints/openapi

Implementation Examples

The google/skills repository provides concrete implementation patterns for both endpoint types in the skills/cloud/agent-platform-inference/scripts/ directory.

Calling a Custom Endpoint with Vertex AI SDK

When targeting a custom endpoint, initialize the Vertex AI client for the specific region and reference the exact endpoint ID:


# scripts/gemini_vertexai_sdk.py

import google.auth
import vertexai
from vertexai.generative_models import GenerativeModel

# Obtain default project and token

_, project_id = google.auth.default()

# Initialise the client for the region where the endpoint lives

vertexai.init(project=project_id, location="us-central1")

# Directly reference the custom endpoint ID

custom_endpoint = "1234567890"          # <-- replace with your endpoint ID

model = GenerativeModel(
    "gemini-2.5-pro",
    # Tell the SDK to use the explicit endpoint resource

    endpoint=f"projects/{project_id}/locations/us-central1/endpoints/{custom_endpoint}",
)

resp = model.generate_content("Why is the sky blue?")
print(resp.text)

Calling a Publisher Endpoint with OpenAI SDK

For OpenMaaS models accessed via the publisher endpoint, use the OpenAI SDK with the global or regional openapi base URL:


# scripts/openmaas_openai_sdk.py

import google.auth
from google.auth.transport.requests import Request
from openai import OpenAI

def get_gcp_token():
    creds, _ = google.auth.default()
    creds.refresh(Request())
    return creds.token

# Authentication – use the GCP token as the API key

openai_client = OpenAI(
    api_key=get_gcp_token(),
    base_url="https://aiplatform.googleapis.com/v1/projects/YOUR_PROJECT/locations/global/endpoints/openapi",
)

# Call an OpenMaaS model (e.g., DeepSeek) via the publisher endpoint

completion = openai_client.chat.completions.create(
    model="deepseek-chat",
    messages=[{"role": "user", "content": "Explain reinforcement learning"}],
)
print(completion.choices[0].message.content)

Summary

  • Custom endpoints are immutable, region-specific resources ideal for tuned models and self-hosted LLMs, returning immediate 404s for region mismatches.
  • Publisher endpoints are shared services for first-party and OpenMaaS models that require runtime regional availability checks before invocation.
  • Vertex AI SDK is the preferred client for custom endpoints, while GenAI SDK or OpenAI SDK are recommended for publisher endpoints.
  • Always verify regional availability programmatically when using publisher endpoints, as documented in skills/cloud/agent-platform-inference/SKILL.md.

Frequently Asked Questions

Can I use a publisher endpoint for a fine-tuned Gemini model?

Yes, but only for LoRA adapters. According to the google/skills source, fine-tuned Gemini LoRA adapters still route via the publisher endpoint rather than requiring a custom endpoint resource. However, fully tuned base models require a custom endpoint with a specific numeric endpoint ID.

Why do I get a 404 error when calling a custom endpoint?

A 404 error on a custom endpoint indicates a region mismatch. Because custom endpoints are immutable resources tied to a specific region at deployment time, calling them from a different region results in an immediate 404 without inference charges. Ensure your client initialization location matches the endpoint's region exactly.

Which SDK should I use for OpenMaaS models?

Use the OpenAI SDK with the publisher endpoint base URL for OpenMaaS models like DeepSeek, Llama, or Qwen. Set the base_url to https://aiplatform.googleapis.com/v1/projects/<PROJECT>/locations/global/endpoints/openapi (or the appropriate regional variant) and authenticate using a GCP access token.

Are publisher endpoints available in all regions?

No. Publisher endpoints have variable regional availability depending on the specific model. First-party Gemini models are only available in certain regions, while OpenMaaS models may use the global openapi endpoint. You must probe the model with a test request to verify availability before deploying production workloads.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →