# Difference Between Custom Endpoints and Publisher Endpoints in Vertex AI Agent Platform

> Understand the difference between custom endpoints and publisher endpoints in Vertex AI. Learn how to deploy tuned models or use shared services for Gemini and OpenMaaS.

- Repository: [Google/skills](https://github.com/google/skills)
- Tags: deep-dive
- Published: 2026-08-09

---

**Custom endpoints are dedicated, region-locked resources you deploy for tuned models or self-hosted LLMs, while publisher endpoints are shared routing services for first-party Gemini and OpenMaaS models that require runtime regional availability checks.**

Understanding the distinction between these two invocation patterns is critical for architecting inference workloads on Google Cloud. According to the source code in the `google/skills` repository, specifically [`skills/cloud/agent-platform-inference/SKILL.md`](https://github.com/google/skills/blob/main/skills/cloud/agent-platform-inference/SKILL.md), the Vertex AI Agent Platform handles routing, regional validation, and SDK compatibility differently depending on whether you target a specific endpoint resource or the shared publisher service.

## What Are Custom Endpoints?

A **custom endpoint** is a concrete, immutable resource that you create (or that is created automatically when you tune a Gemini model or deploy an open-source LLM). It represents a specific deployment with a fixed regional location.

### Resource Structure and Routing

Custom endpoints follow a strict resource naming convention: `projects/<PROJECT>/locations/<REGION>/endpoints/<ENDPOINT_ID>`. When you send a request, the routing is direct to that specific resource. As implemented in the `google/skills` source, if the caller's region does not match the endpoint's deployment region, the lookup returns a clean **404 error** without incurring inference costs. This behavior is documented in [`skills/cloud/agent-platform-inference/SKILL.md`](https://github.com/google/skills/blob/main/skills/cloud/agent-platform-inference/SKILL.md) at lines 44-48.

### When to Use Custom Endpoints

Use a custom endpoint in the following scenarios:

- **Tuned Gemini models** served on a numeric endpoint ID (e.g., `projects/my-project/locations/us-central1/endpoints/1234567890`)
- **Self-deployed OSS LLMs** such as Llama, DeepSeek, Qwen, or Gemma running on your own infrastructure
- **Legacy custom models** requiring specific resource isolation

## What Are Publisher Endpoints?

A **publisher endpoint** is a shared "publisher" service that acts as a façade, routing requests to first-party Gemini models or OpenMaaS (Open Model-as-a-Service) offerings through a global or regional base URL.

### Shared Service Architecture

Unlike custom endpoints, publisher endpoints do not represent a specific resource you own. Instead, they use base URLs such as `https://aiplatform.googleapis.com/v1/projects/<PROJECT>/locations/global/endpoints/openapi` or regional variants. The service handles routing to the underlying model, which may vary in availability by region. This architectural distinction is detailed in [`skills/cloud/agent-platform-inference/SKILL.md`](https://github.com/google/skills/blob/main/skills/cloud/agent-platform-inference/SKILL.md) at lines 55-58.

### Regional Availability Considerations

Publisher endpoints require **runtime availability verification**. You must probe the model in the target region (for example, via a `:generateContent` call) to confirm support. A 200 response indicates availability, while a 404 means the model is not supported in that region, and you must report the supported regions accordingly. This check is mandatory because, unlike custom endpoints, the publisher service is not tied to a single immutable region.

## Key Differences Between Custom and Publisher Endpoints

| Aspect | Custom Endpoint | Publisher Endpoint |
|--------|----------------|-------------------|
| **Resource Type** | Specific endpoint resource you own | Shared routing service |
| **Region Behavior** | Fixed at deployment time; mismatched regions return immediate 404 | Variable availability; requires runtime probing |
| **Model Types** | Tuned Gemini, self-hosted OSS LLMs | First-party Gemini, OpenMaaS models |
| **Typical SDK** | Vertex AI SDK | GenAI SDK (`google-genai`) or OpenAI SDK |
| **URL Pattern** | `.../locations/<REGION>/endpoints/<ENDPOINT_ID>` | `.../locations/<REGION>/endpoints/openapi` |

## Implementation Examples

The `google/skills` repository provides concrete implementation patterns for both endpoint types in the `skills/cloud/agent-platform-inference/scripts/` directory.

### Calling a Custom Endpoint with Vertex AI SDK

When targeting a custom endpoint, initialize the Vertex AI client for the specific region and reference the exact endpoint ID:

```python

# scripts/gemini_vertexai_sdk.py

import google.auth
import vertexai
from vertexai.generative_models import GenerativeModel

# Obtain default project and token

_, project_id = google.auth.default()

# Initialise the client for the region where the endpoint lives

vertexai.init(project=project_id, location="us-central1")

# Directly reference the custom endpoint ID

custom_endpoint = "1234567890"          # <-- replace with your endpoint ID

model = GenerativeModel(
    "gemini-2.5-pro",
    # Tell the SDK to use the explicit endpoint resource

    endpoint=f"projects/{project_id}/locations/us-central1/endpoints/{custom_endpoint}",
)

resp = model.generate_content("Why is the sky blue?")
print(resp.text)

```

### Calling a Publisher Endpoint with OpenAI SDK

For OpenMaaS models accessed via the publisher endpoint, use the OpenAI SDK with the global or regional `openapi` base URL:

```python

# scripts/openmaas_openai_sdk.py

import google.auth
from google.auth.transport.requests import Request
from openai import OpenAI

def get_gcp_token():
    creds, _ = google.auth.default()
    creds.refresh(Request())
    return creds.token

# Authentication – use the GCP token as the API key

openai_client = OpenAI(
    api_key=get_gcp_token(),
    base_url="https://aiplatform.googleapis.com/v1/projects/YOUR_PROJECT/locations/global/endpoints/openapi",
)

# Call an OpenMaaS model (e.g., DeepSeek) via the publisher endpoint

completion = openai_client.chat.completions.create(
    model="deepseek-chat",
    messages=[{"role": "user", "content": "Explain reinforcement learning"}],
)
print(completion.choices[0].message.content)

```

## Summary

- **Custom endpoints** are immutable, region-specific resources ideal for tuned models and self-hosted LLMs, returning immediate 404s for region mismatches.
- **Publisher endpoints** are shared services for first-party and OpenMaaS models that require runtime regional availability checks before invocation.
- **Vertex AI SDK** is the preferred client for custom endpoints, while **GenAI SDK** or **OpenAI SDK** are recommended for publisher endpoints.
- Always verify regional availability programmatically when using publisher endpoints, as documented in [`skills/cloud/agent-platform-inference/SKILL.md`](https://github.com/google/skills/blob/main/skills/cloud/agent-platform-inference/SKILL.md).

## Frequently Asked Questions

### Can I use a publisher endpoint for a fine-tuned Gemini model?

Yes, but only for LoRA adapters. According to the `google/skills` source, fine-tuned Gemini LoRA adapters still route via the publisher endpoint rather than requiring a custom endpoint resource. However, fully tuned base models require a custom endpoint with a specific numeric endpoint ID.

### Why do I get a 404 error when calling a custom endpoint?

A 404 error on a custom endpoint indicates a **region mismatch**. Because custom endpoints are immutable resources tied to a specific region at deployment time, calling them from a different region results in an immediate 404 without inference charges. Ensure your client initialization location matches the endpoint's region exactly.

### Which SDK should I use for OpenMaaS models?

Use the **OpenAI SDK** with the publisher endpoint base URL for OpenMaaS models like DeepSeek, Llama, or Qwen. Set the `base_url` to `https://aiplatform.googleapis.com/v1/projects/<PROJECT>/locations/global/endpoints/openapi` (or the appropriate regional variant) and authenticate using a GCP access token.

### Are publisher endpoints available in all regions?

No. Publisher endpoints have **variable regional availability** depending on the specific model. First-party Gemini models are only available in certain regions, while OpenMaaS models may use the global `openapi` endpoint. You must probe the model with a test request to verify availability before deploying production workloads.