How to Wire a Custom LLM Provider Using the OpenAI-Compatible API in WeKnora

You can wire any OpenAI-compatible LLM into WeKnora by configuring a custom base_url in your YAML configuration or instantiating the generic LiteLLM adapter programmatically, leveraging the Provider interface without modifying core source code.

WeKnora is an open-source knowledge retrieval framework that standardizes LLM interactions through a unified provider abstraction. By leveraging the Provider interface defined in internal/models/provider/provider.go, you can integrate self-hosted models, third-party SaaS APIs, or proxy services that expose the standard OpenAI HTTP specification. This guide details the exact file paths, configuration patterns, and implementation strategies required to connect your custom backend.

Understanding the Provider Architecture

WeKnora decouples LLM access from specific vendors through the Provider interface located at internal/models/provider/provider.go. This interface defines standardized methods for Chat, Embeddings, and other model capabilities, while concrete implementations handle the underlying HTTP transport.

The repository ships with built-in providers including openai, azure_openai, gemini, and a generic litellm adapter. Because the OpenAI-compatible provider in internal/models/provider/openai.go parameterizes the base URL, you can redirect it to any service implementing the /v1/chat/completions endpoint without writing new provider code.

Step 1: Choose Your Implementation Strategy

When wiring a custom LLM, you have two implementation paths depending on your complexity requirements.

For most custom backends, use the LiteLLM adapter in internal/models/provider/litellm.go. This generic implementation accepts arbitrary base URLs and handles request serialization for any OpenAI-compatible endpoint.

import "github.com/tencent/WeKnora/internal/models/provider"

func main() {
    lite := provider.NewLiteLLM(
        "https://api.my-llm.com/v1",
        "sk-********",
        map[string]string{"X-Tenant-Id": "my-tenant-123"},
        "gpt-4o-mini",
    )
    resp, err := lite.Chat(context.Background(), messages)
}

Option B: Create a Custom Provider Implementation

If you require specialized authentication flows or request shaping, copy internal/models/provider/openai.go as a template. Replace the hard-coded OpenAIBaseURL constant with your endpoint logic, then register your new provider identifier in internal/models/provider/provider.go.

Step 2: Configure Your Custom LLM in YAML

The simplest integration method requires only YAML configuration. Add your custom model to configs/models.yaml (or your runtime configuration file):

models:
  - name: my-custom-gpt
    provider: openai
    base_url: https://api.my-llm.com/v1
    api_key: $MY_LLM_API_KEY
    model: gpt-4o-mini

The provider: openai field tells WeKnora to instantiate the OpenAI-compatible provider, while base_url redirects requests to your custom endpoint. The system matches the provider value against constants defined in provider.go.

Step 3: Add Custom HTTP Headers

For services requiring additional authentication headers (such as tenant IDs or custom signatures), WeKnora provides the extraheaders.go utility. You can inject headers directly in your YAML configuration:

models:
  - name: my-custom-gpt
    provider: openai
    base_url: https://api.my-llm.com/v1
    api_key: $MY_LLM_API_KEY
    model: gpt-4o-mini
    custom_headers:
      X-Tenant-Id: my-tenant-123
      X-Custom-Auth: bearer-token

The provider automatically merges these headers into every HTTP request via the mechanism implemented in internal/utils/extraheaders.go.

Step 4: Invoke the Model Programmatically

Once configured, use the standard provider factory to instantiate your custom LLM in Go code:

import (
    "github.com/tencent/WeKnora/internal/models/provider"
    "github.com/tencent/WeKnora/internal/types"
)

func queryCustomLLM() {
    cfg := types.ModelConfig{
        Name:     "my-custom-gpt",
        Provider: "openai",
        BaseURL:  "https://api.my-llm.com/v1",
        APIKey:   os.Getenv("MY_LLM_API_KEY"),
        Model:    "gpt-4o-mini",
    }

    prov, err := provider.NewProvider(cfg)
    if err != nil {
        panic(err)
    }

    resp, err := prov.Chat(context.Background(), []types.ChatMessage{
        {Role: "user", Content: "Explain quantum computing"},
    })
}

The factory method NewProvider returns the concrete implementation based on the Provider field, allowing the rest of the WeKnora codebase (including VLM, reranker, and embedding modules) to interact with your custom LLM through the unified interface.

Key Implementation Files

Reference these source files when implementing your custom provider:

Summary

  • WeKnora uses a Provider interface (internal/models/provider/provider.go) to abstract LLM implementations, allowing you to swap backends without changing business logic.
  • Reuse existing adapters by pointing the openai or litellm provider at your custom base_url in YAML configuration.
  • Configure via YAML in configs/models.yaml using the provider, base_url, and custom_headers fields.
  • Extend programmatically using provider.NewLiteLLM() for direct instantiation or provider.NewProvider() for factory-based creation.
  • Add custom headers through the custom_headers YAML map or the extraheaders.go utility for authentication requirements.

Frequently Asked Questions

Can I use WeKnora with a self-hosted LLM like Ollama or vLLM?

Yes. Both Ollama and vLLM expose OpenAI-compatible HTTP APIs. Configure your base_url to point to their respective endpoints (e.g., http://localhost:11434/v1 for Ollama) and set provider: openai or provider: litellm in your YAML. The LiteLLM adapter in internal/models/provider/litellm.go handles the protocol translation automatically.

How do I handle authentication schemes beyond the standard API key?

Use the custom_headers field in your model configuration to inject additional authentication tokens, tenant IDs, or signing headers. These headers are merged into every request by the mechanism defined in internal/utils/extraheaders.go. For complex authentication flows requiring request signing, implement a custom provider by copying internal/models/provider/openai.go and modifying the request preparation logic.

Do I need to modify the WeKnora source code to add a new provider?

No. For standard OpenAI-compatible APIs, you only need to modify YAML configuration. The existing openai provider accepts arbitrary base_url values at runtime. You only need to write custom Go code if your LLM uses a non-standard protocol or requires specialized request/response handling that the LiteLLM adapter cannot accommodate.

Which OpenAI endpoints must my custom LLM implement?

At minimum, your service must implement /v1/chat/completions for chat functionality. If you use embedding features, implement /v1/embeddings following the EmbeddingRequest structure defined in internal/models/provider/requesty.go. The WeKnora provider automatically selects the appropriate endpoint based on the model capability flags configured in your YAML.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →