# How to Wire a Custom LLM Provider Using the OpenAI-Compatible API in WeKnora

> Effortlessly integrate custom LLM providers with WeKnora using the OpenAI-compatible API. Configure base_url or use the LiteLLM adapter for seamless customization.

- Repository: [Tencent/WeKnora](https://github.com/tencent/WeKnora)
- Tags: how-to-guide
- Published: 2026-09-12

---

**You can wire any OpenAI-compatible LLM into WeKnora by configuring a custom `base_url` in your YAML configuration or instantiating the generic LiteLLM adapter programmatically, leveraging the Provider interface without modifying core source code.**

WeKnora is an open-source knowledge retrieval framework that standardizes LLM interactions through a unified provider abstraction. By leveraging the `Provider` interface defined in [`internal/models/provider/provider.go`](https://github.com/Tencent/WeKnora/blob/main/internal/models/provider/provider.go), you can integrate self-hosted models, third-party SaaS APIs, or proxy services that expose the standard OpenAI HTTP specification. This guide details the exact file paths, configuration patterns, and implementation strategies required to connect your custom backend.

## Understanding the Provider Architecture

WeKnora decouples LLM access from specific vendors through the **Provider** interface located at [`internal/models/provider/provider.go`](https://github.com/Tencent/WeKnora/blob/main/internal/models/provider/provider.go). This interface defines standardized methods for **Chat**, **Embeddings**, and other model capabilities, while concrete implementations handle the underlying HTTP transport.

The repository ships with built-in providers including `openai`, `azure_openai`, `gemini`, and a generic `litellm` adapter. Because the OpenAI-compatible provider in [`internal/models/provider/openai.go`](https://github.com/Tencent/WeKnora/blob/main/internal/models/provider/openai.go) parameterizes the base URL, you can redirect it to any service implementing the `/v1/chat/completions` endpoint without writing new provider code.

## Step 1: Choose Your Implementation Strategy

When wiring a custom LLM, you have two implementation paths depending on your complexity requirements.

### Option A: Reuse the LiteLLM Adapter (Recommended)

For most custom backends, use the **LiteLLM** adapter in [`internal/models/provider/litellm.go`](https://github.com/Tencent/WeKnora/blob/main/internal/models/provider/litellm.go). This generic implementation accepts arbitrary base URLs and handles request serialization for any OpenAI-compatible endpoint.

```go
import "github.com/tencent/WeKnora/internal/models/provider"

func main() {
    lite := provider.NewLiteLLM(
        "https://api.my-llm.com/v1",
        "sk-********",
        map[string]string{"X-Tenant-Id": "my-tenant-123"},
        "gpt-4o-mini",
    )
    resp, err := lite.Chat(context.Background(), messages)
}

```

### Option B: Create a Custom Provider Implementation

If you require specialized authentication flows or request shaping, copy [`internal/models/provider/openai.go`](https://github.com/Tencent/WeKnora/blob/main/internal/models/provider/openai.go) as a template. Replace the hard-coded `OpenAIBaseURL` constant with your endpoint logic, then register your new provider identifier in [`internal/models/provider/provider.go`](https://github.com/Tencent/WeKnora/blob/main/internal/models/provider/provider.go).

## Step 2: Configure Your Custom LLM in YAML

The simplest integration method requires only YAML configuration. Add your custom model to [`configs/models.yaml`](https://github.com/Tencent/WeKnora/blob/main/configs/models.yaml) (or your runtime configuration file):

```yaml
models:
  - name: my-custom-gpt
    provider: openai
    base_url: https://api.my-llm.com/v1
    api_key: $MY_LLM_API_KEY
    model: gpt-4o-mini

```

The `provider: openai` field tells WeKnora to instantiate the OpenAI-compatible provider, while `base_url` redirects requests to your custom endpoint. The system matches the `provider` value against constants defined in [`provider.go`](https://github.com/Tencent/WeKnora/blob/main/provider.go).

## Step 3: Add Custom HTTP Headers

For services requiring additional authentication headers (such as tenant IDs or custom signatures), WeKnora provides the [`extraheaders.go`](https://github.com/Tencent/WeKnora/blob/main/extraheaders.go) utility. You can inject headers directly in your YAML configuration:

```yaml
models:
  - name: my-custom-gpt
    provider: openai
    base_url: https://api.my-llm.com/v1
    api_key: $MY_LLM_API_KEY
    model: gpt-4o-mini
    custom_headers:
      X-Tenant-Id: my-tenant-123
      X-Custom-Auth: bearer-token

```

The provider automatically merges these headers into every HTTP request via the mechanism implemented in [`internal/utils/extraheaders.go`](https://github.com/Tencent/WeKnora/blob/main/internal/utils/extraheaders.go).

## Step 4: Invoke the Model Programmatically

Once configured, use the standard provider factory to instantiate your custom LLM in Go code:

```go
import (
    "github.com/tencent/WeKnora/internal/models/provider"
    "github.com/tencent/WeKnora/internal/types"
)

func queryCustomLLM() {
    cfg := types.ModelConfig{
        Name:     "my-custom-gpt",
        Provider: "openai",
        BaseURL:  "https://api.my-llm.com/v1",
        APIKey:   os.Getenv("MY_LLM_API_KEY"),
        Model:    "gpt-4o-mini",
    }

    prov, err := provider.NewProvider(cfg)
    if err != nil {
        panic(err)
    }

    resp, err := prov.Chat(context.Background(), []types.ChatMessage{
        {Role: "user", Content: "Explain quantum computing"},
    })
}

```

The factory method `NewProvider` returns the concrete implementation based on the `Provider` field, allowing the rest of the WeKnora codebase (including VLM, reranker, and embedding modules) to interact with your custom LLM through the unified interface.

## Key Implementation Files

Reference these source files when implementing your custom provider:

- [`internal/models/provider/provider.go`](https://github.com/Tencent/WeKnora/blob/main/internal/models/provider/provider.go) — Defines the `Provider` interface and built-in provider constants
- [`internal/models/provider/openai.go`](https://github.com/Tencent/WeKnora/blob/main/internal/models/provider/openai.go) — Concrete OpenAI-compatible implementation handling base URL configuration and request shaping
- [`internal/models/provider/litellm.go`](https://github.com/Tencent/WeKnora/blob/main/internal/models/provider/litellm.go) — Generic adapter for arbitrary OpenAI-compatible endpoints
- [`internal/models/provider/requesty.go`](https://github.com/Tencent/WeKnora/blob/main/internal/models/provider/requesty.go) — Helper structs describing model capabilities (chat, embedding, rerank)
- [`internal/utils/extraheaders.go`](https://github.com/Tencent/WeKnora/blob/main/internal/utils/extraheaders.go) — Mechanism for attaching custom HTTP headers to provider requests
- [`configs/models.yaml`](https://github.com/Tencent/WeKnora/blob/main/configs/models.yaml) — Example configuration file where you register custom model entries

## Summary

- **WeKnora uses a Provider interface** ([`internal/models/provider/provider.go`](https://github.com/Tencent/WeKnora/blob/main/internal/models/provider/provider.go)) to abstract LLM implementations, allowing you to swap backends without changing business logic.
- **Reuse existing adapters** by pointing the `openai` or `litellm` provider at your custom `base_url` in YAML configuration.
- **Configure via YAML** in [`configs/models.yaml`](https://github.com/Tencent/WeKnora/blob/main/configs/models.yaml) using the `provider`, `base_url`, and `custom_headers` fields.
- **Extend programmatically** using `provider.NewLiteLLM()` for direct instantiation or `provider.NewProvider()` for factory-based creation.
- **Add custom headers** through the `custom_headers` YAML map or the [`extraheaders.go`](https://github.com/Tencent/WeKnora/blob/main/extraheaders.go) utility for authentication requirements.

## Frequently Asked Questions

### Can I use WeKnora with a self-hosted LLM like Ollama or vLLM?

Yes. Both Ollama and vLLM expose OpenAI-compatible HTTP APIs. Configure your `base_url` to point to their respective endpoints (e.g., `http://localhost:11434/v1` for Ollama) and set `provider: openai` or `provider: litellm` in your YAML. The LiteLLM adapter in [`internal/models/provider/litellm.go`](https://github.com/Tencent/WeKnora/blob/main/internal/models/provider/litellm.go) handles the protocol translation automatically.

### How do I handle authentication schemes beyond the standard API key?

Use the `custom_headers` field in your model configuration to inject additional authentication tokens, tenant IDs, or signing headers. These headers are merged into every request by the mechanism defined in [`internal/utils/extraheaders.go`](https://github.com/Tencent/WeKnora/blob/main/internal/utils/extraheaders.go). For complex authentication flows requiring request signing, implement a custom provider by copying [`internal/models/provider/openai.go`](https://github.com/Tencent/WeKnora/blob/main/internal/models/provider/openai.go) and modifying the request preparation logic.

### Do I need to modify the WeKnora source code to add a new provider?

No. For standard OpenAI-compatible APIs, you only need to modify YAML configuration. The existing `openai` provider accepts arbitrary `base_url` values at runtime. You only need to write custom Go code if your LLM uses a non-standard protocol or requires specialized request/response handling that the LiteLLM adapter cannot accommodate.

### Which OpenAI endpoints must my custom LLM implement?

At minimum, your service must implement `/v1/chat/completions` for chat functionality. If you use embedding features, implement `/v1/embeddings` following the `EmbeddingRequest` structure defined in [`internal/models/provider/requesty.go`](https://github.com/Tencent/WeKnora/blob/main/internal/models/provider/requesty.go). The WeKnora provider automatically selects the appropriate endpoint based on the model capability flags configured in your YAML.