How to Integrate Google's Generative Models Using the generative-ai-gcp Directory
You can integrate Google's generative models by using the three helper modules in tools/llmevalkit/src/—gcp_prompt.py, gcp_dataset.py, and gcp_evaluation.py—to build prompts, batch process data, and evaluate outputs via the Vertex AI API.
The generative-ai-gcp directory within the GoogleCloudPlatform/generative-ai repository provides production-ready Python utilities for calling Gemini and Vertex AI large language models. These lightweight helpers abstract the boilerplate of authentication, prompt construction, and batch inference while maintaining full access to the underlying Vertex AI SDK.
Overview of the generative-ai-gcp Toolkit
The integration logic is encapsulated in three thin helper modules under tools/llmevalkit/src/. Each module handles a distinct stage of the LLM workflow, allowing you to compose solutions without importing the entire Vertex AI SDK boilerplate into your application code.
Prompt Construction with gcp_prompt.py
The gcp_prompt.py module standardizes how you build inputs for Gemini and Vertex AI models. It exposes create_prompt(), add_context(), and to_message_list() to assemble system instructions, few-shot examples, and user queries into the message-list format required by the Gemini API.
Data Loading with gcp_dataset.py
The gcp_dataset.py module handles ingestion from CSV, BigQuery, or JSON sources and prepares data for batch inference. Key functions include load_dataset() for reading files from Cloud Storage or local paths, and batch_iterable() for chunking records into batches that respect model token limits.
Evaluation with gcp_evaluation.py
The gcp_evaluation.py module wraps the Vertex AI Evaluation API to score model outputs. The evaluate_responses() function computes metrics such as BLEU and Rouge-L, while write_to_bq() persists results to a BigQuery table defined by the EVAL_BQ_DATASET environment variable.
Authentication and Environment Setup
Before invoking any helper functions, you must authenticate to GCP and configure your environment.
Application Default Credentials (ADC) is the recommended approach. Run the following command locally:
gcloud auth application-default login
For production deployments, mount a service-account key and set the GOOGLE_APPLICATION_CREDENTIALS environment variable. The Quickbot sample application demonstrates this pattern in gemini/sample-apps/quickbot/backend/Dockerfile.local:
# The secret containing the service-account JSON is mounted at /tmp/gcp_adc.json
ENV GOOGLE_APPLICATION_CREDENTIALS=/tmp/gcp_adc.json
Define these environment variables in your shell or deployment manifest:
export GCP_PROJECT_ID="my-gcp-project"
export GCP_REGION="us-central1"
export EVAL_BQ_DATASET="my_eval_dataset"
Install the required dependencies:
pip install -r tools/llmevalkit/requirements.txt
Step-by-Step Integration Workflow
Building and Sending Prompts
Use create_prompt() to assemble your input and to_message_list() to convert it into the format expected by the Gemini API.
from tools.llmevalkit.src.gcp_prompt import create_prompt
from vertexai.language_models import TextGenerationModel
import os
project = os.getenv("GCP_PROJECT_ID")
region = os.getenv("GCP_REGION")
# Build a chat-style prompt
prompt = create_prompt(
system="You are a helpful assistant that always speaks in a friendly tone.",
user="Explain quantum computing in two sentences."
)
# Initialize the Gemini model
model = TextGenerationModel.from_pretrained(
"gemini-1.5-pro",
project=project,
location=region,
)
# Generate the response
response = model.predict(prompt.to_message_list())
print(response.text)
Processing Large Datasets in Batches
For high-throughput scenarios, use load_dataset() to read from Cloud Storage and batch_iterable() to chunk inputs.
import pandas as pd
from tools.llmevalkit.src.gcp_dataset import load_dataset, batch_iterable
from tools.llmevalkit.src.gcp_prompt import create_prompt
from vertexai.language_models import TextGenerationModel
import os
df = load_dataset("gs://my-bucket/queries.csv") # CSV with a column 'question'
batches = batch_iterable(df["question"], batch_size=8)
model = TextGenerationModel.from_pretrained(
"gemini-1.5-flash",
project=os.getenv("GCP_PROJECT_ID"),
location=os.getenv("GCP_REGION"),
)
answers = []
for batch in batches:
prompts = [create_prompt(user=q).to_message_list() for q in batch]
resp = model.batch_predict(prompts)
answers.extend([r.text for r in resp])
df["answer"] = answers
df.to_csv("gs://my-bucket/answers.csv", index=False)
Evaluating Model Outputs
After generating responses, call evaluate_responses() to compute metrics and persist them to BigQuery.
from tools.llmevalkit.src.gcp_evaluation import evaluate_responses
import os
import pandas as pd
# Assume df has columns: prompt, model_output, ground_truth
metrics = evaluate_responses(
dataset=df,
project=os.getenv("GCP_PROJECT_ID"),
location=os.getenv("GCP_REGION"),
bq_dataset=os.getenv("EVAL_BQ_DATASET")
)
print("Overall BLEU:", metrics["bleu"])
print("Average Rouge-L:", metrics["rouge_l"])
Deployment Patterns
Cloud Functions and Cloud Run
The helper modules are stateless and fit naturally into serverless containers. The Quickbot sample application demonstrates a production-ready Cloud Run deployment. In gemini/sample-apps/quickbot/backend/Dockerfile.local, the container mounts a service-account key at /tmp/gcp_adc.json and sets GOOGLE_APPLICATION_CREDENTIALS to enable ADC inside the container.
Your Cloud Function handler can import the same utilities and invoke the model in response to HTTP triggers, reusing the authentication flow described above.
Vertex AI Notebooks
For interactive development, install the toolkit in a Vertex AI Workbench notebook. The workshops/rag-ops/2.5_mvp_evaluation_vertexai_eval.ipynb notebook demonstrates how to evaluate RAG pipelines using the gcp_evaluation module directly within a managed notebook environment.
Summary
- The
generative-ai-gcpdirectory provides three minimal helper modules—gcp_prompt.py,gcp_dataset.py, andgcp_evaluation.py—that wrap the Vertex AI SDK for common LLM workflows. - Authentication relies on Application Default Credentials (ADC) via
gcloud auth application-default loginor a mounted service-account key set viaGOOGLE_APPLICATION_CREDENTIALS. - Prompt construction uses
create_prompt()andto_message_list()to format inputs for Gemini models likegemini-1.5-proandgemini-1.5-flash. - Batch processing leverages
load_dataset()andbatch_iterable()to stream large datasets from Cloud Storage and respect token limits. - Evaluation is handled by
evaluate_responses(), which computes BLEU, Rouge-L, and other metrics and writes results to BigQuery using theEVAL_BQ_DATASETenvironment variable.
Frequently Asked Questions
What authentication method does the generative-ai-gcp directory use?
The helper modules rely on Application Default Credentials (ADC). When running locally, execute gcloud auth application-default login to populate credentials. In production deployments such as Cloud Run or Cloud Functions, mount a service-account JSON key and set the GOOGLE_APPLICATION_CREDENTIALS environment variable to the key path, as demonstrated in the Quickbot Dockerfile.
Can I use generative-ai-gcp with custom Vertex AI models instead of Gemini?
Yes. While the examples highlight Gemini models (gemini-1.5-pro, gemini-1.5-flash), the TextGenerationModel.from_pretrained() method accepts any model ID available in your Vertex AI Model Garden. The prompt helpers in gcp_prompt.py generate standard message lists compatible with both Gemini and other Vertex AI LLM endpoints.
How does batch processing handle model token limits?
The batch_iterable() function in gcp_dataset.py chunks input sequences into batches that respect the maximum token limit for the target model. When iterating over a dataset, the helper yields subsets sized to stay within the model's context window, preventing request failures due to payload size while maximizing throughput.
Where are evaluation metrics stored after running evaluate_responses?
By default, evaluate_responses() persists metrics to a BigQuery dataset specified by the EVAL_BQ_DATASET environment variable. The function writes aggregated scores—such as BLEU and Rouge-L—alongside individual response records to a table within that dataset, enabling downstream analysis and dashboarding in Looker or BigQuery Studio.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →