# How to Integrate Different LLM Providers with WeKnora: A Complete Guide

> Easily integrate diverse LLM providers like OpenAI, Azure, and HuggingFace with WeKnora. Learn how to switch between models seamlessly without code changes using our complete guide.

- Repository: [Tencent/WeKnora](https://github.com/tencent/WeKnora)
- Tags: how-to-guide
- Published: 2026-09-13

---

**WeKnora abstracts LLM interactions behind a Model entity that stores provider credentials and configuration, allowing you to switch between OpenAI, Azure, HuggingFace, and other providers without code changes by simply registering new model records via the MCP client.**

Integrating different LLM providers with WeKnora requires no modifications to the core codebase. The Tencent/WeKnora platform decouples vendor-specific logic from the chat engine through a unified **Model** abstraction, letting you register OpenAI, Azure OpenAI, HuggingFace inference endpoints, or local models via JSON configuration alone.

## Understanding WeKnora's LLM Abstraction Layer

WeKnora treats every LLM integration as a database record rather than hardcoded logic. When a chat or agent session initiates, the server looks up the default model (or the specific one defined in the session strategy) and forwards requests to the corresponding provider implementation.

### The Model Entity Structure

Each model record in WeKnora defines five critical fields:

- **name** – Human-readable identifier (e.g., *OpenAI-GPT-4*).
- **type** – Model purpose: `KnowledgeQA` for chat completion, `Embedding` for vector generation, or `Rerank` for result reordering.
- **source** – Provider identifier: `openai`, `azure`, `huggingface`, `local`, etc.
- **parameters** – Provider-specific configuration object containing API keys, endpoint URLs, deployment names, and model identifiers.
- **is_default** – Boolean flag that automatically assigns this model to new sessions of its designated `type`.

This architecture means you can integrate a new LLM provider by inserting a single JSON record through the MCP client, without touching [`weknora_mcp_server.py`](https://github.com/Tencent/WeKnora/blob/main/weknora_mcp_server.py) or any server-side Python code.

## Step-by-Step Integration Workflow

The standard integration flow follows five sequential stages:

1. **Prepare provider credentials** – Obtain API keys, endpoint URLs, and deployment names from your LLM vendor.
2. **Register the provider** – Issue a `POST /models` request via the MCP client with `source` set to the provider name and `parameters` containing authentication details.
3. **Set default status (optional)** – Mark `is_default: true` to make this LLM the automatic choice for new sessions.
4. **Create a chat session** – Reference the model explicitly via `summary_model_id` or rely on the default selection.
5. **Invoke chat** – Calls to `chat()` or `agent_chat()` route automatically to the registered provider through `_consume_sse_stream`.

## Practical Implementation: Registering LLM Providers

All provider registration happens through the `WeKnoraMCPClient` class in [`mcp-server/weknora_mcp_server.py`](https://github.com/Tencent/WeKnora/blob/main/mcp-server/weknora_mcp_server.py). The `create_model` method (lines 998–1018) handles the underlying HTTP request to the models endpoint.

### Registering an OpenAI Model (GPT-4)

The following snippet registers GPT-4 as a `KnowledgeQA` model and sets it as the default for all new chat sessions:

```python
from mcp_server.weknora_mcp_server import WeKnoraMCPClient

client = WeKnoraMCPClient(base_url="https://your-weknora-instance.com/api")

model = client.create_model(
    name="OpenAI-GPT-4",
    model_type="KnowledgeQA",
    source="openai",
    description="OpenAI GPT-4 for QA",
    parameters={
        "api_key": "<YOUR_OPENAI_API_KEY>",
        "base_url": "https://api.openai.com/v1",
        "model_name": "gpt-4"
    },
    is_default=True
)
print("Created model ID:", model["id"])

```

*Implementation reference:* `create_model` method, lines 998–1018 in [`mcp-server/weknora_mcp_server.py`](https://github.com/Tencent/WeKnora/blob/main/mcp-server/weknora_mcp_server.py).

### Registering an Azure OpenAI Model

Azure OpenAI requires additional parameters for the deployment name and API version:

```python
model = client.create_model(
    name="Azure-GPT-35-Turbo",
    model_type="KnowledgeQA",
    source="azure",
    description="Azure OpenAI deployment",
    parameters={
        "api_key": "<YOUR_AZURE_API_KEY>",
        "base_url": "https://YOUR_RESOURCE_NAME.openai.azure.com",
        "deployment_name": "gpt-35-turbo",
        "api_version": "2023-05-15"
    },
    is_default=False
)
print("Azure model ID:", model["id"])

```

### Registering a HuggingFace Inference Endpoint

For self-hosted models or HuggingFace Inference API:

```python
model = client.create_model(
    name="HF-Llama-2-7B",
    model_type="KnowledgeQA",
    source="huggingface",
    description="Self-hosted Llama-2-7B via HF Inference API",
    parameters={
        "api_key": "<HF_TOKEN>",
        "base_url": "http://your-hf-server.com",
        "model_name": "meta-llama/Llama-2-7b-chat-hf"
    },
    is_default=False
)
print("HF model ID:", model["id"])

```

## Creating Chat Sessions with Custom Providers

After registration, bind the model to a session using `create_session` (lines 1034–1060). The `summary_model_id` parameter explicitly routes requests to your chosen provider:

```python
session = client.create_session(
    kb_id="my_knowledge_base",
    max_rounds=8,
    enable_rewrite=True,
    fallback_response="I'm not sure how to answer that.",
    summary_model_id=model["id"],
    title="Demo session",
    description="Chat using the newly added LLM"
)
print("Session ID:", session["id"])

```

Send messages through the provider-agnostic `chat` method, which internally calls `_consume_sse_stream` (lines 1072–1105) to handle Server-Sent Events:

```python
response = client.chat(
    session_id=session["id"],
    messages=[
        {"role": "user", "content": "What is the capital of France?"}
    ]
)
print("Answer:", response["answer"])
print("References:", response.get("references", []))

```

## Managing Default Providers and Runtime Switching

Because models persist in the platform database, you can swap default providers without restarting services. Update the `is_default` flag via the `PATCH /models/{model_id}` endpoint:

```python
client.update_model(
    model_id=azure_model_id,
    data={"is_default": True}
)

```

When `is_default` is `True`, all new sessions automatically use that model unless overridden by an explicit `summary_model_id` in the session strategy.

## Key Implementation Files in Tencent/WeKnora

The following source files contain the core logic for LLM provider integration:

- **[`mcp-server/weknora_mcp_server.py`](https://github.com/Tencent/WeKnora/blob/main/mcp-server/weknora_mcp_server.py)** – Contains `create_model` (lines 998–1018) for provider registration, `create_session` (lines 1034–1060) for session binding, and `_consume_sse_stream` (lines 1072–1105) for response streaming.
- **[`docreader/main.py`](https://github.com/Tencent/WeKnora/blob/main/docreader/main.py)** – Implements the document processing pipeline that prepares knowledge bases for LLM consumption.
- **[`examples/mcp-demo/server.py`](https://github.com/Tencent/WeKnora/blob/main/examples/mcp-demo/server.py)** – Provides a runnable reference implementation demonstrating the Model abstraction in action.

## Summary

- WeKnora uses a **Model** entity to abstract LLM providers, storing configuration in database records rather than code.
- The **MCP client** (`WeKnoraMCPClient`) provides `create_model`, `create_session`, and `chat` methods to register and invoke providers.
- Supported `source` values include `openai`, `azure`, `huggingface`, and `local`, each requiring specific `parameters` JSON objects.
- The `is_default` flag enables zero-configuration provider switching for new sessions.
- No source code changes are required to integrate new LLM vendors; all configuration happens through the `POST /models` API.

## Frequently Asked Questions

### What LLM providers does WeKnora support?

WeKnora supports any provider that implements a compatible HTTP interface. The `source` field accepts values like `openai`, `azure`, `huggingface`, and `local`. You can register multiple models from different vendors simultaneously and select between them at the session level using `summary_model_id`.

### Do I need to modify WeKnora source code to add a new provider?

No. The provider logic lives entirely in the server-side implementation of [`weknora_mcp_server.py`](https://github.com/Tencent/WeKnora/blob/main/weknora_mcp_server.py). As a user, you only interact with the MCP client's `create_model` method to register new providers via JSON configuration. The core platform handles request routing automatically based on the model's `source` and `parameters` fields.

### How does WeKnora route requests to the correct LLM provider?

When you call `chat()`, the server looks up the session's `summary_model_id` (or the default model for the `KnowledgeQA` type) and extracts the `source` and `parameters` from the model record. The `_consume_sse_stream` method then constructs the appropriate HTTP request to the provider's endpoint and streams back the response, translating vendor-specific SSE formats into WeKnora's unified response schema.

### Can I use different LLM providers for different knowledge bases?

Yes. Since the model selection happens at the session level via `create_session`, you can create separate sessions for different knowledge bases, each pointing to a different `summary_model_id`. This allows one knowledge base to use Azure OpenAI while another uses a local HuggingFace model, with both running concurrently on the same WeKnora instance.