How to Configure the Agent Platform RAG Engine: A Complete Setup Guide
To configure the Agent Platform RAG engine, authenticate to Google Cloud, initialize the Vertex AI SDK, discover your corpora, and configure a VertexRagStore tool to retrieve and ground context in Gemini model requests.
Setting up the Agent Platform RAG engine enables your LLM applications to generate responses grounded in proprietary document corpora. According to the google/skills repository, the complete configuration workflow is documented in skills/cloud/agent-platform-rag-engine-management/SKILL.md and consists of three logical layers: environment setup, corpus discovery, and retrieval-augmented generation tooling. This guide walks through each layer using the exact method signatures and parameters defined in the source code.
Prerequisites and Environment Setup
Before interacting with the Agent Platform RAG engine, you must establish Google Cloud credentials and install the required SDKs. As specified in lines 47-71 of the skill definition, this involves authenticating via the gcloud CLI and installing google-cloud-aiplatform for corpus operations and google-genai for model interactions.
Run the following commands once to prepare your environment:
# Authenticate with Google Cloud
gcloud auth login
gcloud auth application-default login
# Create and activate a virtual environment
python3 -m venv ~/rag_agent_venv
source ~/rag_agent_venv/bin/activate
# Install the required SDKs
pip install google-cloud-aiplatform google-genai
Initializing Vertex AI and Discovering Corpora
Once authenticated, initialize the Vertex AI client with your project ID and region to enable API calls. The SDK automatically handles pagination when listing corpora, though manual pagination controls are available for large-scale projects (lines 107-131).
import vertexai
from vertexai.preview import rag
# Replace with your values
project_id = "my-project"
region = "us-central1"
vertexai.init(project=project_id, location=region)
# List all corpora (automatic pagination)
all_corpora = list(rag.list_corpora())
print(f"Found {len(all_corpora)} corpora:")
for c in all_corpora:
print(f"- {c.display_name} ({c.name})")
Listing Files Within a Corpus
To verify indexed content before retrieval, enumerate files inside a specific corpus using its resource name. This step is optional but recommended for confirming that your documents are properly ingested.
corpus_name = f"projects/{project_id}/locations/{region}/ragCorpora/your-corpus-id"
files = list(rag.list_files(corpus_name=corpus_name))
print(f"Found {len(files)} files in the corpus:")
for f in files:
print(f"- {f.display_name} ({f.name})")
Configuring RAG Retrieval and Generation
The final configuration layer connects your corpus to a Gemini model through the VertexRagStore tool. This process involves executing a retrieval query to fetch relevant passages, then supplying those passages to the model via a tool configuration that includes retrieval parameters like top_k and vector_similarity_threshold (lines 190-236).
Retrieving Context with rag.retrieval_query
Use the rag.retrieval_query function to fetch the most relevant document fragments for a user query. The similarity_top_k parameter controls how many passages are returned for grounding.
query = "What is the speed of light?"
response = rag.retrieval_query(
rag_corpora=[corpus_name],
text=query,
similarity_top_k=3 # Return the 3 most similar passages
)
for ctx in response.contexts.contexts:
print("Context:", ctx.text)
print("Source :", ctx.source_uri)
Grounding Gemini with VertexRagStore
Construct a VertexRagStore tool definition that points to your corpus and specifies retrieval behavior. Pass this tool to the model's generate_content method to produce answers grounded in your documents.
from google import genai
from google.genai import types
client = genai.Client(enterprise=True, project=project_id, location=region)
# Define the RAG tool that points to the corpus
rag_tool = types.Tool(
retrieval=types.Retrieval(
vertex_rag_store=types.VertexRagStore(
rag_resources=[
types.VertexRagStoreRagResource(rag_corpus=corpus_name)
],
rag_retrieval_config=types.RagRetrievalConfig(
top_k=3,
filter=types.RagRetrievalConfigFilter(
vector_similarity_threshold=0.5
),
),
)
)
)
# Generate a grounded answer
response = client.models.generate_content(
model="gemini-2.5-flash",
contents="What is the speed of light?",
config=types.GenerateContentConfig(tools=[rag_tool])
)
print("Grounded answer:", response.text)
Summary
- Authenticate to Google Cloud and install
google-cloud-aiplatformandgoogle-genaito establish API credentials and dependencies. - Initialize Vertex AI with
vertexai.init()using your project ID and region, then discover available corpora viarag.list_corpora(). - Retrieve relevant context by calling
rag.retrieval_query()withrag_corporaandsimilarity_top_kparameters to fetch specific document passages. - Ground Gemini model responses by configuring a
VertexRagStoretool withrag_retrieval_configsettings includingtop_kandvector_similarity_threshold, enabling traceable, source-cited generation.
Frequently Asked Questions
What Python SDKs are required to configure the Agent Platform RAG engine?
You need google-cloud-aiplatform for corpus and retrieval operations, and google-genai for model inference. Install both via pip after running gcloud auth application-default login to ensure your environment has proper credentials.
How do I locate the resource name for my RAG corpus?
After initializing Vertex AI with vertexai.init(), call list(rag.list_corpora()) to enumerate all corpora. Each returned object contains a name attribute formatted as projects/{project_id}/locations/{region}/ragCorpora/{corpus_id}, which serves as the unique identifier for subsequent operations.
Which parameters control the quality and quantity of retrieved context?
The similarity_top_k parameter in rag.retrieval_query() and the top_k field within RagRetrievalConfig define how many passages are fetched. Additionally, setting vector_similarity_threshold in RagRetrievalConfigFilter excludes low-relevance chunks below the specified semantic similarity score.
Can I query multiple corpora in a single generation request?
Yes, the rag_resources list inside VertexRagStore accepts multiple VertexRagStoreRagResource objects. By appending additional corpus resource names to this list, you can retrieve context across several document collections simultaneously for broader knowledge coverage.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →