Vertex AI Dedicated vs Shared Endpoints: Architecture, DNS, and Deployment Differences
Dedicated endpoints provide isolated DNS names and guaranteed traffic isolation for production workloads, while shared endpoints use regional DNS infrastructure for simpler prototyping and reduced management overhead.
Vertex AI supports two distinct endpoint architectures for model deployment, each optimized for different operational requirements. Understanding the differences between dedicated vs shared endpoints in Vertex AI deployment is critical for optimizing latency, ensuring traffic isolation, and maintaining correct request routing. This guide examines the implementation details found in the google/skills repository to help you choose the right configuration for your machine learning workloads.
DNS and Network Architecture
The most fundamental difference lies in how each endpoint type handles DNS resolution and network addressing.
Dedicated Endpoint DNS Structure
A dedicated endpoint receives its own unique DNS name following the pattern:
<ENDPOINT_ID>.<REGION>-<PROJECT_NUM>.prediction.vertexai.goog
This endpoint is completely isolated at the network level. According to the source documentation in skills/cloud/agent-platform-inference/SKILL.md (lines 459-464), a dedicated endpoint "has its own DNS" and critically cannot be reached via the shared regional DNS. Attempting to call a dedicated endpoint using the shared base URL results in a "cannot resolve DNS" error.
Shared Endpoint DNS Structure
Shared endpoints utilize the common regional DNS format:
<REGION>-aiplatform.googleapis.com
Multiple models deployed on shared endpoints all receive traffic through this single regional URL. The system routes requests to the appropriate model based on the endpoint ID specified in the request payload.
Traffic Isolation and Performance
Isolation Guarantees
Dedicated endpoints offer complete traffic isolation. Only the specific model attached to the endpoint receives traffic, eliminating cross-model interference. This isolation is mandatory for private-preview features such as custom model routing and high-throughput workloads that require predictable performance.
Shared endpoints route traffic through a shared load balancer. While this reduces management complexity, multiple models share the same infrastructure, which can introduce variable latency depending on overall platform load.
Latency and Throughput Characteristics
Dedicated endpoints typically deliver lower latency and higher throughput because the request path is shorter and backend resources are tuned for a single model. The evaluation script in skills/cloud/agent-platform-eval-flywheel/scripts/endpoint_evaluation.py (lines 102-105) demonstrates how production workloads switch to dedicated URLs when dedicated_endpoint_dns is provided.
Shared endpoints incur slightly higher latency because requests pass through a shared proxy layer that must resolve which model to invoke before processing.
Configuration and Implementation
The dedicatedEndpointDns Flag
Vertex AI uses the boolean field dedicatedEndpointDns to distinguish between deployment types:
- Dedicated: Set
dedicatedEndpointDnstotrueand populate with the dedicated DNS string - Shared: Omit
dedicatedEndpointDnsor set tofalse
As noted in skills/cloud/agent-platform-deploy/SKILL.md (line 139), certain advanced models are only deployable to dedicated endpoints due to their infrastructure requirements.
URL Construction Patterns
The request URL construction differs fundamentally between the two types. The endpoint_evaluation.py script shows that your application logic must branch based on whether a dedicated DNS value is supplied.
Practical Implementation Examples
Accessing a Shared Endpoint
When calling a shared endpoint, construct the URL using the regional DNS:
project_id = "my-project"
location = "us-central1"
endpoint_id = "1234567890123456789"
# Shared-endpoint URL format
url = f"https://{location}-aiplatform.googleapis.com/v1/projects/{project_id}/locations/{location}/endpoints/{endpoint_id}:predict"
Accessing a Dedicated Endpoint
For dedicated endpoints, you must use the unique DNS name assigned to that specific endpoint:
project_id = "my-project"
location = "us-central1"
endpoint_id = "1234567890123456789"
# Construct dedicated DNS from endpoint metadata
dedicated_dns = f"{endpoint_id}.{location}-{project_id}.prediction.vertexai.goog"
# Dedicated-endpoint URL (must use dedicated DNS)
url = f"https://{dedicated_dns}/v1/projects/{project_id}/locations/{location}/endpoints/{endpoint_id}:predict"
Dynamic Endpoint Selection Script
Use this pattern to handle both endpoint types in your deployment scripts:
#!/usr/bin/env python3
import argparse
parser = argparse.ArgumentParser()
parser.add_argument("--endpoint_id", required=True)
parser.add_argument("--dedicated_endpoint_dns", default=None,
help="Dedicated DNS for the endpoint (if any)")
args = parser.parse_args()
if args.dedicated_endpoint_dns:
endpoint_url = f"https://{args.dedicated_endpoint_dns}/v1/projects/$PROJECT/locations/$LOCATION/endpoints/{args.endpoint_id}:predict"
else:
endpoint_url = f"https://$LOCATION-aiplatform.googleapis.com/v1/projects/$PROJECT/locations/$LOCATION/endpoints/{args.endpoint_id}:predict"
print(f"Using endpoint URL: {endpoint_url}")
Selection Criteria for Production Workloads
Choose your endpoint architecture based on these operational requirements:
- Use dedicated endpoints when you need guaranteed isolation, custom routing logic, VPC Service Controls (VPC-SC), or high-throughput production workloads that cannot tolerate shared infrastructure variability.
- Use shared endpoints for rapid prototyping, experimentation, or when minimizing management overhead takes priority over latency guarantees.
Summary
- Dedicated endpoints use unique DNS names (
<ENDPOINT_ID>.<REGION>-<PROJECT_NUM>.prediction.vertexai.goog) and provide complete traffic isolation, lower latency, and support for advanced features like custom routing. - Shared endpoints use regional DNS (
<REGION>-aiplatform.googleapis.com) and share infrastructure across models, trading isolation for simplified management. - The
dedicatedEndpointDnsconfiguration flag determines which architecture Vertex AI provisions for your model. - You cannot reach a dedicated endpoint via the shared DNS, and conversely, shared endpoints reject requests using the dedicated DNS format.
- Production workloads requiring consistent performance should use dedicated endpoints according to the implementation patterns in
skills/cloud/agent-platform-eval-flywheel/scripts/endpoint_evaluation.py.
Frequently Asked Questions
Can I convert an existing shared endpoint to a dedicated endpoint?
No, you cannot convert between endpoint types after deployment. You must deploy a new model version to a dedicated endpoint and migrate your traffic. The endpoint type is determined at creation time via the dedicatedEndpointDns configuration and remains immutable for the lifecycle of that endpoint resource.
Why does my dedicated endpoint return DNS resolution errors?
You are likely using the shared regional DNS (<REGION>-aiplatform.googleapis.com) instead of the dedicated DNS name. As documented in skills/cloud/agent-platform-inference/SKILL.md, dedicated endpoints require the specific DNS format <ENDPOINT_ID>.<REGION>-<PROJECT_NUM>.prediction.vertexai.goog. Verify that your client code constructs the URL using the dedicated endpoint's specific DNS property.
Do dedicated endpoints cost more than shared endpoints?
While dedicated endpoints provide isolated infrastructure that may incur different pricing based on resource guarantees, the primary cost difference stems from provisioned throughput requirements. Dedicated endpoints are required for high-throughput or custom routing scenarios where shared infrastructure would be insufficient. Consult the Vertex AI pricing documentation for specific rates based on your deployment configuration and traffic patterns.
Which endpoint type should I use for VPC Service Controls (VPC-SC)?
You must use dedicated endpoints for VPC-SC compliance. Shared endpoints utilize multi-tenant infrastructure that cannot be fully isolated within your VPC perimeter. The dedicated endpoint architecture provides the network isolation required to enforce VPC-SC policies, as indicated in the deployment notes within the google/skills repository.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →