How to Customize the Vector Store in WeKnora: Configuration and Deployment Guide

Customizing a vector store in WeKnora involves defining a VectorStore record with engine-specific connection and index configurations, or setting RETRIEVE_DRIVER environment variables for ephemeral environment stores, then registering them via the EngineFactory to enable hybrid retrieval for knowledge bases.

WeKnora, Tencent’s open-source knowledge management platform, abstracts vector storage through pluggable backends including Elasticsearch, Qdrant, Milvus, Weaviate, and Tencent VectorDB. To customize the vector store in WeKnora, administrators must understand the dual configuration model: persistent database records managed via REST API, and transient environment stores derived from runtime variables. Both methods integrate with the retrieval engine registry to power semantic search across knowledge bases.

Understanding the VectorStore Data Model

The core abstraction resides in internal/types/vectorstore.go, where the VectorStore struct defines the schema persisted to the database.

Core Struct Definition

According to the source in internal/types/vectorstore.go (lines 40-48), the struct captures metadata, engine type, and configuration blobs:

type VectorStore struct {
    ID               string          `json:"id" gorm:"type:varchar(36);primaryKey"`
    TenantID         uint64          `json:"tenant_id"`
    Name             string          `json:"name" gorm:"type:varchar(255);not null"`
    EngineType       RetrieverEngineType `json:"engine_type" gorm:"type:varchar(50);not null"`
    ConnectionConfig ConnectionConfig `json:"connection_config" gorm:"type:json"`
    IndexConfig      IndexConfig      `json:"index_config" gorm:"type:json"`
    CreatedAt        time.Time
    UpdatedAt        time.Time
    DeletedAt        gorm.DeletedAt `gorm:"index"`
}

Supported Engine Types

The validEngineTypes slice (lines 72-95) enumerates persistable engines: elasticsearch, qdrant, milvus, weaviate, doris, opensearch, and tencent_vectordb. Note that postgres and sqlite are reserved exclusively for environment stores and excluded from the UI selection list.

Configuring Persistent Vector Stores

Persistent stores are rows in the vector_stores table created via the REST API or administration interface.

Connection and Index Configuration

The ConnectionConfig struct stores driver-specific parameters including host, port, API keys, and TLS settings. Sensitive fields are encrypted at rest through custom Value and Scan methods defined on the type. The IndexConfig struct defines engine-specific collection settings such as shard count, replication factor, and HNSW parameters, validated by ValidateIndexConfig.

Validation Rules

Before persistence, the Validate() method (lines 102-114) enforces:

  • Non-empty Name field
  • EngineType must exist in validEngineTypes
  • TenantID must be non-zero
  • IndexConfig numeric ranges and name patterns must comply with engine constraints

Creating a Store Programmatically

When interacting with the Go SDK directly, instantiate and validate the struct before handing it to the service layer:

store := types.VectorStore{
    TenantID:  42,
    Name:      "Production-Qdrant",
    EngineType: types.QdrantRetrieverEngineType,
    ConnectionConfig: types.ConnectionConfig{
        Host:   "qdrant.internal",
        Port:   6334,
        APIKey: "sk-secret-key",
    },
    IndexConfig: types.IndexConfig{
        CollectionPrefix: "kb_embeddings",
        ShardNumber:      4,
        ReplicationFactor: 2,
    },
}

if err := store.Validate(); err != nil {
    log.Fatal(err)
}
// The VectorStoreService calls BeforeCreate to encrypt the API key before DB insertion

Implementing Environment Vector Stores

For dynamic, infrastructure-as-code deployments, WeKnora supports environment vector stores that require no database persistence. When the RETRIEVE_DRIVER environment variable contains comma-separated entries like qdrant,elasticsearch_v8, the BuildEnvVectorStores function (lines 89-120) generates read-only VectorStore objects with IDs prefixed by __env_.

Configuration example:

export RETRIEVE_DRIVER="qdrant,elasticsearch_v8"
export QDRANT_HOST="qdrant.mycompany.com"
export QDRANT_API_KEY="s3cr3t-apikey"
export ELASTICSEARCH_ADDR="http://es.mycompany.com:9200"
export ELASTICSEARCH_USERNAME="elastic"
export ELASTICSEARCH_PASSWORD="elastic-pwd"

These virtual stores appear in the registry as __env_qdrant__ and __env_elasticsearch_v8__. They are read-only and never persisted to the database, making them ideal for containerized deployments.

Registry Integration and Knowledge Base Binding

EngineFactory Initialization

During server startup, internal/container/engine_factory.go loads all rows from the vector_stores table. For each record, createEngineServiceFromStore instantiates the appropriate engine client and registers it under the store ID via registry.RegisterWithStoreID. This registration enables hybrid retrieval capabilities that combine vector similarity with keyword search.

Binding to Knowledge Bases

Knowledge bases specify their target store via the vector_store_id foreign key. If omitted, WeKnora defaults to the tenant’s environment store. This binding logic is enforced during KB creation in internal/handler/knowledgebase.go.

REST API for Vector Store Management

The platform exposes CRUD operations under /api/v1/vector-stores, registered in internal/container/container.go:

  • GET /api/v1/vector-stores – List all stores (responses redact credentials)
  • POST /api/v1/vector-stores – Create store (requires manage_vector_stores capability)
  • PATCH /api/v1/vector-stores/{id} – Update mutable fields
  • DELETE /api/v1/vector-stores/{id} – Soft delete
  • GET /api/v1/vector-stores/types – Retrieve VectorStoreTypeInfo metadata detailing engine configurations and field schemas

Example API call to create a Milvus store:

curl -X POST https://weknora.example.com/api/v1/vector-stores \
  -H "Authorization: Bearer $TOKEN" \
  -H "Content-Type: application/json" \
  -d '{
    "tenant_id": 42,
    "name": "Milvus-Prod",
    "engine_type": "milvus",
    "connection_config": {
      "addr": "milvus.internal:19530",
      "username": "root",
      "password": "milvus-pwd"
    },
    "index_config": {
      "collection_name": "embeddings",
      "shards_num": 4,
      "replica_number": 2
    }
  }'

The response returns a VectorStoreResponse with passwords and API keys redacted for security.

Summary

  • Define vector stores either as persistent database records or transient environment stores via RETRIEVE_DRIVER.
  • Configure engine-specific parameters through ConnectionConfig (authentication, TLS) and IndexConfig (sharding, replication, HNSW settings).
  • Validate configurations using the Validate() method before persistence to ensure TenantID, Name, and EngineType constraints are satisfied.
  • Register stores automatically via EngineFactory at startup, making them available for hybrid retrieval operations.
  • Bind knowledge bases to specific stores using vector_store_id or fallback to default environment stores when omitted.
  • Manage stores programmatically via the /api/v1/vector-stores endpoints with proper capability checks.

Frequently Asked Questions

What vector database engines does WeKnora support?

WeKnora supports Elasticsearch, Qdrant, Milvus, Weaviate, Apache Doris, OpenSearch, and Tencent VectorDB for persistent storage. PostgreSQL and SQLite are available exclusively as environment stores for local development scenarios and cannot be persisted through the UI.

How does WeKnora protect sensitive connection credentials?

The VectorStore struct encrypts sensitive fields within ConnectionConfig at rest using GORM's Value and Scan hooks (defined in internal/types/vectorstore.go). API responses automatically redact passwords and API keys through the VectorStoreResponse wrapper, ensuring credentials never traverse the network in plaintext.

Can I use multiple vector stores simultaneously?

Yes. The EngineFactory registers every configured store—both persistent and environment—under unique IDs during startup. Knowledge bases can bind to different stores, enabling hybrid architectures where, for example, one KB uses Qdrant for high-frequency queries while another uses Elasticsearch for full-text hybrid search.

What happens if I don't specify a vector_store_id for my knowledge base?

If vector_store_id is omitted during knowledge base creation, WeKnora automatically assigns the tenant’s default environment vector store derived from RETRIEVE_DRIVER. This fallback mechanism ensures retrieval operations function immediately without explicit store configuration, though production deployments should specify dedicated stores for isolation and performance tuning.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →