How to Integrate Different LLM Providers with WeKnora: A Complete Guide
WeKnora abstracts LLM interactions behind a Model entity that stores provider credentials and configuration, allowing you to switch between OpenAI, Azure, HuggingFace, and other providers without code changes by simply registering new model records via the MCP client.
Integrating different LLM providers with WeKnora requires no modifications to the core codebase. The Tencent/WeKnora platform decouples vendor-specific logic from the chat engine through a unified Model abstraction, letting you register OpenAI, Azure OpenAI, HuggingFace inference endpoints, or local models via JSON configuration alone.
Understanding WeKnora's LLM Abstraction Layer
WeKnora treats every LLM integration as a database record rather than hardcoded logic. When a chat or agent session initiates, the server looks up the default model (or the specific one defined in the session strategy) and forwards requests to the corresponding provider implementation.
The Model Entity Structure
Each model record in WeKnora defines five critical fields:
- name – Human-readable identifier (e.g., OpenAI-GPT-4).
- type – Model purpose:
KnowledgeQAfor chat completion,Embeddingfor vector generation, orRerankfor result reordering. - source – Provider identifier:
openai,azure,huggingface,local, etc. - parameters – Provider-specific configuration object containing API keys, endpoint URLs, deployment names, and model identifiers.
- is_default – Boolean flag that automatically assigns this model to new sessions of its designated
type.
This architecture means you can integrate a new LLM provider by inserting a single JSON record through the MCP client, without touching weknora_mcp_server.py or any server-side Python code.
Step-by-Step Integration Workflow
The standard integration flow follows five sequential stages:
- Prepare provider credentials – Obtain API keys, endpoint URLs, and deployment names from your LLM vendor.
- Register the provider – Issue a
POST /modelsrequest via the MCP client withsourceset to the provider name andparameterscontaining authentication details. - Set default status (optional) – Mark
is_default: trueto make this LLM the automatic choice for new sessions. - Create a chat session – Reference the model explicitly via
summary_model_idor rely on the default selection. - Invoke chat – Calls to
chat()oragent_chat()route automatically to the registered provider through_consume_sse_stream.
Practical Implementation: Registering LLM Providers
All provider registration happens through the WeKnoraMCPClient class in mcp-server/weknora_mcp_server.py. The create_model method (lines 998–1018) handles the underlying HTTP request to the models endpoint.
Registering an OpenAI Model (GPT-4)
The following snippet registers GPT-4 as a KnowledgeQA model and sets it as the default for all new chat sessions:
from mcp_server.weknora_mcp_server import WeKnoraMCPClient
client = WeKnoraMCPClient(base_url="https://your-weknora-instance.com/api")
model = client.create_model(
name="OpenAI-GPT-4",
model_type="KnowledgeQA",
source="openai",
description="OpenAI GPT-4 for QA",
parameters={
"api_key": "<YOUR_OPENAI_API_KEY>",
"base_url": "https://api.openai.com/v1",
"model_name": "gpt-4"
},
is_default=True
)
print("Created model ID:", model["id"])
Implementation reference: create_model method, lines 998–1018 in mcp-server/weknora_mcp_server.py.
Registering an Azure OpenAI Model
Azure OpenAI requires additional parameters for the deployment name and API version:
model = client.create_model(
name="Azure-GPT-35-Turbo",
model_type="KnowledgeQA",
source="azure",
description="Azure OpenAI deployment",
parameters={
"api_key": "<YOUR_AZURE_API_KEY>",
"base_url": "https://YOUR_RESOURCE_NAME.openai.azure.com",
"deployment_name": "gpt-35-turbo",
"api_version": "2023-05-15"
},
is_default=False
)
print("Azure model ID:", model["id"])
Registering a HuggingFace Inference Endpoint
For self-hosted models or HuggingFace Inference API:
model = client.create_model(
name="HF-Llama-2-7B",
model_type="KnowledgeQA",
source="huggingface",
description="Self-hosted Llama-2-7B via HF Inference API",
parameters={
"api_key": "<HF_TOKEN>",
"base_url": "http://your-hf-server.com",
"model_name": "meta-llama/Llama-2-7b-chat-hf"
},
is_default=False
)
print("HF model ID:", model["id"])
Creating Chat Sessions with Custom Providers
After registration, bind the model to a session using create_session (lines 1034–1060). The summary_model_id parameter explicitly routes requests to your chosen provider:
session = client.create_session(
kb_id="my_knowledge_base",
max_rounds=8,
enable_rewrite=True,
fallback_response="I'm not sure how to answer that.",
summary_model_id=model["id"],
title="Demo session",
description="Chat using the newly added LLM"
)
print("Session ID:", session["id"])
Send messages through the provider-agnostic chat method, which internally calls _consume_sse_stream (lines 1072–1105) to handle Server-Sent Events:
response = client.chat(
session_id=session["id"],
messages=[
{"role": "user", "content": "What is the capital of France?"}
]
)
print("Answer:", response["answer"])
print("References:", response.get("references", []))
Managing Default Providers and Runtime Switching
Because models persist in the platform database, you can swap default providers without restarting services. Update the is_default flag via the PATCH /models/{model_id} endpoint:
client.update_model(
model_id=azure_model_id,
data={"is_default": True}
)
When is_default is True, all new sessions automatically use that model unless overridden by an explicit summary_model_id in the session strategy.
Key Implementation Files in Tencent/WeKnora
The following source files contain the core logic for LLM provider integration:
mcp-server/weknora_mcp_server.py– Containscreate_model(lines 998–1018) for provider registration,create_session(lines 1034–1060) for session binding, and_consume_sse_stream(lines 1072–1105) for response streaming.docreader/main.py– Implements the document processing pipeline that prepares knowledge bases for LLM consumption.examples/mcp-demo/server.py– Provides a runnable reference implementation demonstrating the Model abstraction in action.
Summary
- WeKnora uses a Model entity to abstract LLM providers, storing configuration in database records rather than code.
- The MCP client (
WeKnoraMCPClient) providescreate_model,create_session, andchatmethods to register and invoke providers. - Supported
sourcevalues includeopenai,azure,huggingface, andlocal, each requiring specificparametersJSON objects. - The
is_defaultflag enables zero-configuration provider switching for new sessions. - No source code changes are required to integrate new LLM vendors; all configuration happens through the
POST /modelsAPI.
Frequently Asked Questions
What LLM providers does WeKnora support?
WeKnora supports any provider that implements a compatible HTTP interface. The source field accepts values like openai, azure, huggingface, and local. You can register multiple models from different vendors simultaneously and select between them at the session level using summary_model_id.
Do I need to modify WeKnora source code to add a new provider?
No. The provider logic lives entirely in the server-side implementation of weknora_mcp_server.py. As a user, you only interact with the MCP client's create_model method to register new providers via JSON configuration. The core platform handles request routing automatically based on the model's source and parameters fields.
How does WeKnora route requests to the correct LLM provider?
When you call chat(), the server looks up the session's summary_model_id (or the default model for the KnowledgeQA type) and extracts the source and parameters from the model record. The _consume_sse_stream method then constructs the appropriate HTTP request to the provider's endpoint and streams back the response, translating vendor-specific SSE formats into WeKnora's unified response schema.
Can I use different LLM providers for different knowledge bases?
Yes. Since the model selection happens at the session level via create_session, you can create separate sessions for different knowledge bases, each pointing to a different summary_model_id. This allows one knowledge base to use Azure OpenAI while another uses a local HuggingFace model, with both running concurrently on the same WeKnora instance.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →