How to Integrate Different LLM Providers with WeKnora: A Complete Guide

WeKnora abstracts LLM interactions behind a Model entity that stores provider credentials and configuration, allowing you to switch between OpenAI, Azure, HuggingFace, and other providers without code changes by simply registering new model records via the MCP client.

Integrating different LLM providers with WeKnora requires no modifications to the core codebase. The Tencent/WeKnora platform decouples vendor-specific logic from the chat engine through a unified Model abstraction, letting you register OpenAI, Azure OpenAI, HuggingFace inference endpoints, or local models via JSON configuration alone.

Understanding WeKnora's LLM Abstraction Layer

WeKnora treats every LLM integration as a database record rather than hardcoded logic. When a chat or agent session initiates, the server looks up the default model (or the specific one defined in the session strategy) and forwards requests to the corresponding provider implementation.

The Model Entity Structure

Each model record in WeKnora defines five critical fields:

  • name – Human-readable identifier (e.g., OpenAI-GPT-4).
  • type – Model purpose: KnowledgeQA for chat completion, Embedding for vector generation, or Rerank for result reordering.
  • source – Provider identifier: openai, azure, huggingface, local, etc.
  • parameters – Provider-specific configuration object containing API keys, endpoint URLs, deployment names, and model identifiers.
  • is_default – Boolean flag that automatically assigns this model to new sessions of its designated type.

This architecture means you can integrate a new LLM provider by inserting a single JSON record through the MCP client, without touching weknora_mcp_server.py or any server-side Python code.

Step-by-Step Integration Workflow

The standard integration flow follows five sequential stages:

  1. Prepare provider credentials – Obtain API keys, endpoint URLs, and deployment names from your LLM vendor.
  2. Register the provider – Issue a POST /models request via the MCP client with source set to the provider name and parameters containing authentication details.
  3. Set default status (optional) – Mark is_default: true to make this LLM the automatic choice for new sessions.
  4. Create a chat session – Reference the model explicitly via summary_model_id or rely on the default selection.
  5. Invoke chat – Calls to chat() or agent_chat() route automatically to the registered provider through _consume_sse_stream.

Practical Implementation: Registering LLM Providers

All provider registration happens through the WeKnoraMCPClient class in mcp-server/weknora_mcp_server.py. The create_model method (lines 998–1018) handles the underlying HTTP request to the models endpoint.

Registering an OpenAI Model (GPT-4)

The following snippet registers GPT-4 as a KnowledgeQA model and sets it as the default for all new chat sessions:

from mcp_server.weknora_mcp_server import WeKnoraMCPClient

client = WeKnoraMCPClient(base_url="https://your-weknora-instance.com/api")

model = client.create_model(
    name="OpenAI-GPT-4",
    model_type="KnowledgeQA",
    source="openai",
    description="OpenAI GPT-4 for QA",
    parameters={
        "api_key": "<YOUR_OPENAI_API_KEY>",
        "base_url": "https://api.openai.com/v1",
        "model_name": "gpt-4"
    },
    is_default=True
)
print("Created model ID:", model["id"])

Implementation reference: create_model method, lines 998–1018 in mcp-server/weknora_mcp_server.py.

Registering an Azure OpenAI Model

Azure OpenAI requires additional parameters for the deployment name and API version:

model = client.create_model(
    name="Azure-GPT-35-Turbo",
    model_type="KnowledgeQA",
    source="azure",
    description="Azure OpenAI deployment",
    parameters={
        "api_key": "<YOUR_AZURE_API_KEY>",
        "base_url": "https://YOUR_RESOURCE_NAME.openai.azure.com",
        "deployment_name": "gpt-35-turbo",
        "api_version": "2023-05-15"
    },
    is_default=False
)
print("Azure model ID:", model["id"])

Registering a HuggingFace Inference Endpoint

For self-hosted models or HuggingFace Inference API:

model = client.create_model(
    name="HF-Llama-2-7B",
    model_type="KnowledgeQA",
    source="huggingface",
    description="Self-hosted Llama-2-7B via HF Inference API",
    parameters={
        "api_key": "<HF_TOKEN>",
        "base_url": "http://your-hf-server.com",
        "model_name": "meta-llama/Llama-2-7b-chat-hf"
    },
    is_default=False
)
print("HF model ID:", model["id"])

Creating Chat Sessions with Custom Providers

After registration, bind the model to a session using create_session (lines 1034–1060). The summary_model_id parameter explicitly routes requests to your chosen provider:

session = client.create_session(
    kb_id="my_knowledge_base",
    max_rounds=8,
    enable_rewrite=True,
    fallback_response="I'm not sure how to answer that.",
    summary_model_id=model["id"],
    title="Demo session",
    description="Chat using the newly added LLM"
)
print("Session ID:", session["id"])

Send messages through the provider-agnostic chat method, which internally calls _consume_sse_stream (lines 1072–1105) to handle Server-Sent Events:

response = client.chat(
    session_id=session["id"],
    messages=[
        {"role": "user", "content": "What is the capital of France?"}
    ]
)
print("Answer:", response["answer"])
print("References:", response.get("references", []))

Managing Default Providers and Runtime Switching

Because models persist in the platform database, you can swap default providers without restarting services. Update the is_default flag via the PATCH /models/{model_id} endpoint:

client.update_model(
    model_id=azure_model_id,
    data={"is_default": True}
)

When is_default is True, all new sessions automatically use that model unless overridden by an explicit summary_model_id in the session strategy.

Key Implementation Files in Tencent/WeKnora

The following source files contain the core logic for LLM provider integration:

  • mcp-server/weknora_mcp_server.py – Contains create_model (lines 998–1018) for provider registration, create_session (lines 1034–1060) for session binding, and _consume_sse_stream (lines 1072–1105) for response streaming.
  • docreader/main.py – Implements the document processing pipeline that prepares knowledge bases for LLM consumption.
  • examples/mcp-demo/server.py – Provides a runnable reference implementation demonstrating the Model abstraction in action.

Summary

  • WeKnora uses a Model entity to abstract LLM providers, storing configuration in database records rather than code.
  • The MCP client (WeKnoraMCPClient) provides create_model, create_session, and chat methods to register and invoke providers.
  • Supported source values include openai, azure, huggingface, and local, each requiring specific parameters JSON objects.
  • The is_default flag enables zero-configuration provider switching for new sessions.
  • No source code changes are required to integrate new LLM vendors; all configuration happens through the POST /models API.

Frequently Asked Questions

What LLM providers does WeKnora support?

WeKnora supports any provider that implements a compatible HTTP interface. The source field accepts values like openai, azure, huggingface, and local. You can register multiple models from different vendors simultaneously and select between them at the session level using summary_model_id.

Do I need to modify WeKnora source code to add a new provider?

No. The provider logic lives entirely in the server-side implementation of weknora_mcp_server.py. As a user, you only interact with the MCP client's create_model method to register new providers via JSON configuration. The core platform handles request routing automatically based on the model's source and parameters fields.

How does WeKnora route requests to the correct LLM provider?

When you call chat(), the server looks up the session's summary_model_id (or the default model for the KnowledgeQA type) and extracts the source and parameters from the model record. The _consume_sse_stream method then constructs the appropriate HTTP request to the provider's endpoint and streams back the response, translating vendor-specific SSE formats into WeKnora's unified response schema.

Can I use different LLM providers for different knowledge bases?

Yes. Since the model selection happens at the session level via create_session, you can create separate sessions for different knowledge bases, each pointing to a different summary_model_id. This allows one knowledge base to use Azure OpenAI while another uses a local HuggingFace model, with both running concurrently on the same WeKnora instance.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →