How WeKnora Manages Per-Knowledge-Base Model Selection and Thinking-Mode Override

WeKnora determines which LLM to use by inspecting knowledge‑base configurations and allows callers to override the reasoning style through explicit API flags or custom agent settings.

In the Tencent/WeKnora codebase, routing a query to the correct model requires balancing per-knowledge-base preferences with request-specific overrides. The system implements a deterministic selection algorithm in internal/application/service/session_knowledge_qa.go and a pluggable thinking-mode strategy in internal/models/chat/thinking.go. This article explains exactly how per-knowledge-base model selection works and how the framework handles model thinking-mode overrides across different providers.

Understanding the Model Selection Pipeline

Every QA request begins by identifying which knowledge bases (KBs) are relevant. The service then applies a strict priority chain to pick the chat model that will process the query.

Resolving Knowledge Bases

When a request arrives, sessionService.resolveKnowledgeBases (or resolveKnowledgeBasesFromAgent) builds the definitive list of KB IDs that the pipeline will search. This step establishes the scope for downstream model selection.

// From session_knowledge_qa.go
kbIDs, err := svc.resolveKnowledgeBases(ctx, session, req.KnowledgeBaseIDs)

The Selection Priority Chain

The helper selectChatModelID implements a three-tier priority system:

  1. Remote Summary Model – If any resolved KB has a SummaryModelID with Source == ModelSourceRemote, that model is selected immediately.
  2. First KB's Local Model – If no Remote model exists, the system uses the first KB's configured summary model.
  3. Global Fallback – When no KB supplies a model, the pipeline falls back to the first globally available ModelTypeKnowledgeQA model.

This logic is located in internal/application/service/session_knowledge_qa.go (lines 40–46).

Implementing Per-Knowledge-Base Model Selection

The selected chatModelID is stored in ChatManage.PipelineRequest.ChatModelID and later fed to the chat provider via modelService.GetChatModel. This guarantees that the LLM matches the KB's capabilities—such as vision support or tool calling—especially when a Remote model is specified.

// Example: QA request using a Remote model from KB "kb-42"
kbIDs := []string{"kb-42"}
modelID, _ := svc.selectChatModelID(ctx, session, kbIDs, nil)
// modelID == kb42.SummaryModelID (Remote)

If the request involves multiple KBs, the first one with a Remote model wins, ensuring consistent behavior across distributed knowledge sources.

Managing Thinking-Mode Overrides

Thinking mode—extended reasoning before answering—varies by provider. WeKnora abstracts this through the Thinking() interface and provides two override mechanisms.

Provider-Specific Thinking Strategies

Each provider adapter implements Thinking() to declare its default strategy:

  • Qwen models – Use the boolean field enable_thinking
  • Other providers – Use a nested thinking object with a type field (e.g., "enabled")

The parseThinkingOverride function in internal/models/chat/thinking.go (lines 120–159) decodes the request's ExtraConfig to determine which strategy to apply.

Request-Level Overrides via ExtraConfig

Callers can override the default thinking behavior by supplying extra_config.thinking_control in the API request. Valid values are "enable_thinking" or "thinking_type", mapping to the concrete ThinkingStrategy implementations (thinkingTypeField or enableThinking).

// Override to use legacy "enable_thinking" flag for Qwen
req.ExtraConfig = map[string]string{
    "thinking_control": "enable_thinking",
}

The Apply method of the selected strategy then injects the correct fields into the outbound request struct:

  • For Qwen: Sets EnableThinking *bool (always sent, forced off for non-streaming requests)
  • For others: Adds Thinking *ThinkingConfig with {"type": "enabled"}

This logic resides in internal/models/chat/thinking.go around lines 61–106.

Agent-Level Overrides

Custom agents provide a blunt instrument for thinking control. The CustomAgent.Config.Thinking field (*bool) can explicitly enable or disable reasoning for the entire request.

In internal/types/custom_agent.go (lines 130–132), applyAgentOverridesToChatManage copies this boolean into ChatManage.PipelineRequest.Thinking. When the provider adapter builds the final request in internal/models/chat/provider.go (lines 136–144), an explicit thinkingOverride takes precedence over the provider's default Thinking() implementation.


# In a custom agent definition (YAML)

thinking: true   # Forces thinking on for this request

Summary

Per-knowledge-base model selection and thinking-mode override in WeKnora follow a clear, deterministic flow:

  • KB Resolution – resolveKnowledgeBases establishes the search scope.
  • Model Priority – Remote summary models take precedence over local KB models and global fallbacks.
  • Strategy Pattern – Thinking-mode serialization uses provider-specific strategies (enable_thinking vs. thinking.type).
  • Dual Override Paths – Callers can tune reasoning via extra_config.thinking_control or enforce it through CustomAgent.Config.Thinking.
  • Adapter Enforcement – The final ChatAdapter respects explicit overrides before falling back to provider defaults.

Frequently Asked Questions

How does WeKnora handle model selection when multiple knowledge bases are involved?

The system evaluates the resolved KB list in order. The first KB with a Remote summary model (ModelSourceRemote) dictates the model ID. If none exist, the first KB's local summary model is used. Only when no KB specifies a model does the pipeline fall back to a generic ModelTypeKnowledgeQA instance.

What is the difference between thinking_control and the CustomAgent Thinking field?

The extra_config.thinking_control string selects the wire-format strategy (enable_thinking boolean vs. thinking object) but respects the provider's default value. The CustomAgent Thinking boolean (*bool) is a hard override that forces reasoning on or off regardless of provider defaults, stored in PipelineRequest.Thinking.

Which file contains the logic for applying thinking-mode overrides to the final request?

The Apply method implementations in internal/models/chat/thinking.go (lines 61–106) encode the strategy into provider-specific structs. However, the ultimate precedence logic—checking for thinkingOverride before defaulting to provider.Thinking()—lives in internal/models/chat/provider.go (lines 136–144).

Can thinking mode be disabled for non-streaming requests?

Yes. The Qwen adapter explicitly forces EnableThinking to false when streaming is disabled, regardless of the override strategy. This behavior is hardcoded in the Apply method for the enableThinking strategy to comply with API constraints.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →