Deploy OpenViking on Kubernetes Using Helm Charts: Complete Production Guide
The OpenViking repository provides an official Helm chart under examples/k8s-helm/ that packages the vector database server, Kubernetes Service, configuration secrets, and optional persistent storage into a single declarative release, enabling production-ready deployments on any cloud provider with autoscaling support.
OpenViking is an open-source multimodal vector database developed by VolcEngine. To deploy OpenViking on Kubernetes using Helm charts, you use the chart defined in the examples/k8s-helm/ directory, which templates the complete application stack—including the container image, HTTP service exposure, and runtime configuration—into reusable Kubernetes manifests.
Helm Chart Architecture and Components
The chart abstracts the entire OpenViking deployment into four core Kubernetes resources defined in the templates/ directory and configured via values.yaml.
Chart Metadata and Structure
The chart definition resides in examples/k8s-helm/Chart.yaml, which declares the application name, version, and description. The template files in examples/k8s-helm/templates/ generate the actual cluster resources:
deployment.yaml— Creates the OpenViking server pod, mounts the configuration secret, and attaches the optional persistent volume.service.yaml— Exposes the HTTP API on port 1933 with configurable type (ClusterIPorLoadBalancer).secret.yaml— Renders sensitive configuration fromopenviking.configinto a Kubernetes Secret (restricted to embedding and VLM sections only)._helpers.tpl— Contains reusable Helm helper functions for consistent resource naming.
Container and Runtime Configuration
The container image reference is configurable via values.yaml under the image stanza. The default image is ghcr.io/astral-sh/uv:python3.12-bookworm, though you should override this with the official OpenViking server image for production deployments. Runtime configuration—including API keys for embedding services—is injected through the secret mounted at ov.conf.
Storage and Persistence
By default, the chart uses an emptyDir volume for ephemeral storage, meaning vector data and AGFS files are lost when pods restart. For production workloads, enable persistent storage by setting openviking.dataVolume.enabled=true in your values file, which provisions a PersistentVolumeClaim (PVC) defined in templates/deployment.yaml.
Autoscaling and Cloud Provider Integration
The chart supports Horizontal Pod Autoscaling (HPA) via the autoscaling block in values.yaml (disabled by default). When enabled, Kubernetes automatically scales the openviking-server pods between minReplicas and maxReplicas based on CPU and memory metrics.
For external access, the cloudProvider value (accepting gcp, aws, or generic) injects cloud-specific annotations into the Service manifest in templates/service.yaml, allowing the cloud controller to provision external load balancer IPs automatically.
Configuring the Helm Deployment
Before installing, prepare a custom values file to override defaults in examples/k8s-helm/values.yaml.
Cloud Provider and Networking
Set the service type and cloud annotations to expose OpenViking externally:
cloudProvider: gcp # or "aws"
service:
type: LoadBalancer
port: 1933
API Keys and Secrets
Store sensitive configuration in the openviking.config section. The chart restricts this to embedding and VLM configurations only, rendering them into a Kubernetes Secret:
openviking:
config:
embedding:
dense:
api_key: "YOUR_VOLCENGINE_API_KEY"
Persistent Volume Configuration
Enable PVC-based storage for vector data persistence:
openviking:
dataVolume:
enabled: true
usePVC: true
size: 50Gi
storageClassName: standard
accessModes:
- ReadWriteOnce
Autoscaling Parameters
Configure HPA for dynamic scaling:
autoscaling:
enabled: true
minReplicas: 2
maxReplicas: 10
targetCPUUtilizationPercentage: 80
Step-by-Step Deployment Guide
Follow these commands to deploy OpenViking on Kubernetes using Helm charts from the source repository.
1. Clone the Repository
Download the chart source to your local machine:
git clone https://github.com/volcengine/OpenViking.git
cd OpenViking/examples/k8s-helm
2. Prepare Custom Values
Create my-values.yaml with your specific configuration:
cloudProvider: gcp
replicaCount: 2
openviking:
config:
embedding:
dense:
api_key: YOUR_VOLCENGINE_API_KEY
dataVolume:
enabled: true
usePVC: true
size: 50Gi
storageClassName: standard
autoscaling:
enabled: true
minReplicas: 2
maxReplicas: 10
3. Install the Helm Release
Deploy OpenViking using your custom values:
helm install openviking ./openviking -f my-values.yaml
For a quick test with inline overrides instead of a file:
helm install openviking ./openviking \
--set cloudProvider=aws \
--set openviking.config.embedding.dense.api_key=YOUR_API_KEY \
--set replicaCount=3
4. Verify the Deployment
Check that pods are running and the service has been assigned an external IP:
kubectl get pods -l app.kubernetes.io/name=openviking
kubectl get svc openviking
5. Connect with the OpenViking Client
Configure the CLI to connect to your Kubernetes service. First, capture the load balancer IP:
export OPENVIKING_IP=$(kubectl get svc openviking -o jsonpath='{.status.loadBalancer.ingress[0].ip}')
Create the client configuration file:
mkdir -p ~/.openviking
cat > ~/.openviking/ovcli.conf <<EOF
{
"url": "http://$OPENVIKING_IP:1933",
"api_key": null,
"output": "table"
}
EOF
Verify connectivity:
openviking health
6. Upgrade and Maintenance
To change configuration or scale replicas, modify my-values.yaml and run:
helm upgrade openviking ./openviking -f my-values.yaml
7. Uninstall the Deployment
Remove the release and clean up resources:
helm uninstall openviking
kubectl delete pvc openviking-data # If persistent storage was enabled
Connecting via Python SDK
After deploying OpenViking on Kubernetes using Helm charts, connect your application using the Python client:
import openviking as ov
# Replace <lb-ip> with the external IP from `kubectl get svc`
client = ov.OpenViking(url="http://<lb-ip>:1933", api_key=None)
client.initialize()
# Add a PDF document
client.add_resource(path="./doc.pdf")
client.wait_processed()
# Query the vector database
results = client.find("search term")
print(results)
client.close()
Summary
- The official Helm chart in
examples/k8s-helm/packages OpenViking into a complete Kubernetes application with configurable storage, networking, and autoscaling. - Configuration management uses
values.yamlto generate Kubernetes Secrets for API keys and ConfigMaps for runtime settings, stored intemplates/secret.yaml. - Storage defaults to ephemeral
emptyDirvolumes, but production deployments should enable PVCs viaopenviking.dataVolume.enabled=trueintemplates/deployment.yaml. - Cloud provider integration supports GCP and AWS load balancer annotations through the
cloudProvidervalue intemplates/service.yaml. - Horizontal Pod Autoscaling can be enabled via the
autoscalingblock to automatically scale the OpenViking server based on resource utilization.
Frequently Asked Questions
What is the default service type for OpenViking in Kubernetes?
The default service type is ClusterIP, which exposes the service only within the cluster on port 1933. To enable external access, set service.type to LoadBalancer in your values file and specify a cloudProvider (GCP or AWS) to inject the appropriate cloud controller annotations in templates/service.yaml.
How do I configure cloud-specific load balancers for OpenViking?
Set the cloudProvider value to gcp or aws in your values.yaml file. The chart automatically merges the corresponding cloud provider annotations into the Service manifest, allowing the Kubernetes cloud controller to provision an external load balancer IP automatically without manual annotation editing.
Can I use ephemeral storage instead of a PersistentVolumeClaim?
Yes. By default, the chart uses an emptyDir volume for the OpenViking data directory, making storage ephemeral and suitable for stateless testing. For production workloads requiring data persistence across pod restarts, set openviking.dataVolume.enabled=true and usePVC=true to provision a PersistentVolumeClaim as defined in templates/deployment.yaml.
How do I update the OpenViking configuration after deployment?
Modify your custom values file and run helm upgrade openviking ./openviking -f my-values.yaml. The upgrade process updates the Kubernetes Secret in templates/secret.yaml and triggers a rolling restart of the Deployment to load the new ov.conf configuration. Note that only the embedding and VLM sections are permitted in the secret configuration according to the chart's validation logic in values.yaml.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →