# How WeKnora Manages Per-Knowledge-Base Model Selection and Thinking-Mode Override

> Discover how WeKnora manages per knowledge base model selection and allows API override of thinking modes. Learn to customize your LLM reasoning for Tencent WeKnora.

- Repository: [Tencent/WeKnora](https://github.com/tencent/WeKnora)
- Tags: internals
- Published: 2026-09-12

---

**WeKnora determines which LLM to use by inspecting knowledge‑base configurations and allows callers to override the reasoning style through explicit API flags or custom agent settings.**

In the Tencent/WeKnora codebase, routing a query to the correct model requires balancing per-knowledge-base preferences with request-specific overrides. The system implements a deterministic selection algorithm in [`internal/application/service/session_knowledge_qa.go`](https://github.com/Tencent/WeKnora/blob/main/internal/application/service/session_knowledge_qa.go) and a pluggable thinking-mode strategy in [`internal/models/chat/thinking.go`](https://github.com/Tencent/WeKnora/blob/main/internal/models/chat/thinking.go). This article explains exactly how per-knowledge-base model selection works and how the framework handles model thinking-mode overrides across different providers.

## Understanding the Model Selection Pipeline

Every QA request begins by identifying which knowledge bases (KBs) are relevant. The service then applies a strict priority chain to pick the chat model that will process the query.

### Resolving Knowledge Bases

When a request arrives, `sessionService.resolveKnowledgeBases` (or `resolveKnowledgeBasesFromAgent`) builds the definitive list of KB IDs that the pipeline will search. This step establishes the scope for downstream model selection.

```go
// From session_knowledge_qa.go
kbIDs, err := svc.resolveKnowledgeBases(ctx, session, req.KnowledgeBaseIDs)

```

### The Selection Priority Chain

The helper `selectChatModelID` implements a three-tier priority system:

1. **Remote Summary Model** – If any resolved KB has a `SummaryModelID` with `Source == ModelSourceRemote`, that model is selected immediately.
2. **First KB's Local Model** – If no Remote model exists, the system uses the first KB's configured summary model.
3. **Global Fallback** – When no KB supplies a model, the pipeline falls back to the first globally available `ModelTypeKnowledgeQA` model.

This logic is located in [`internal/application/service/session_knowledge_qa.go`](https://github.com/Tencent/WeKnora/blob/main/internal/application/service/session_knowledge_qa.go) (lines 40–46).

## Implementing Per-Knowledge-Base Model Selection

The selected `chatModelID` is stored in `ChatManage.PipelineRequest.ChatModelID` and later fed to the chat provider via `modelService.GetChatModel`. This guarantees that the LLM matches the KB's capabilities—such as vision support or tool calling—especially when a Remote model is specified.

```go
// Example: QA request using a Remote model from KB "kb-42"
kbIDs := []string{"kb-42"}
modelID, _ := svc.selectChatModelID(ctx, session, kbIDs, nil)
// modelID == kb42.SummaryModelID (Remote)

```

If the request involves multiple KBs, the first one with a Remote model wins, ensuring consistent behavior across distributed knowledge sources.

## Managing Thinking-Mode Overrides

Thinking mode—extended reasoning before answering—varies by provider. WeKnora abstracts this through the `Thinking()` interface and provides two override mechanisms.

### Provider-Specific Thinking Strategies

Each provider adapter implements `Thinking()` to declare its default strategy:

- **Qwen models** – Use the boolean field `enable_thinking`
- **Other providers** – Use a nested `thinking` object with a `type` field (e.g., `"enabled"`)

The `parseThinkingOverride` function in [`internal/models/chat/thinking.go`](https://github.com/Tencent/WeKnora/blob/main/internal/models/chat/thinking.go) (lines 120–159) decodes the request's `ExtraConfig` to determine which strategy to apply.

### Request-Level Overrides via ExtraConfig

Callers can override the default thinking behavior by supplying `extra_config.thinking_control` in the API request. Valid values are `"enable_thinking"` or `"thinking_type"`, mapping to the concrete `ThinkingStrategy` implementations (`thinkingTypeField` or `enableThinking`).

```go
// Override to use legacy "enable_thinking" flag for Qwen
req.ExtraConfig = map[string]string{
    "thinking_control": "enable_thinking",
}

```

The `Apply` method of the selected strategy then injects the correct fields into the outbound request struct:

- For Qwen: Sets `EnableThinking *bool` (always sent, forced off for non-streaming requests)
- For others: Adds `Thinking *ThinkingConfig` with `{"type": "enabled"}`

This logic resides in [`internal/models/chat/thinking.go`](https://github.com/Tencent/WeKnora/blob/main/internal/models/chat/thinking.go) around lines 61–106.

### Agent-Level Overrides

Custom agents provide a blunt instrument for thinking control. The `CustomAgent.Config.Thinking` field (`*bool`) can explicitly enable or disable reasoning for the entire request.

In [`internal/types/custom_agent.go`](https://github.com/Tencent/WeKnora/blob/main/internal/types/custom_agent.go) (lines 130–132), `applyAgentOverridesToChatManage` copies this boolean into `ChatManage.PipelineRequest.Thinking`. When the provider adapter builds the final request in [`internal/models/chat/provider.go`](https://github.com/Tencent/WeKnora/blob/main/internal/models/chat/provider.go) (lines 136–144), an explicit `thinkingOverride` takes precedence over the provider's default `Thinking()` implementation.

```yaml

# In a custom agent definition (YAML)

thinking: true   # Forces thinking on for this request

```

## Summary

Per-knowledge-base model selection and thinking-mode override in WeKnora follow a clear, deterministic flow:

- **KB Resolution** – `resolveKnowledgeBases` establishes the search scope.
- **Model Priority** – Remote summary models take precedence over local KB models and global fallbacks.
- **Strategy Pattern** – Thinking-mode serialization uses provider-specific strategies (`enable_thinking` vs. `thinking.type`).
- **Dual Override Paths** – Callers can tune reasoning via `extra_config.thinking_control` or enforce it through `CustomAgent.Config.Thinking`.
- **Adapter Enforcement** – The final `ChatAdapter` respects explicit overrides before falling back to provider defaults.

## Frequently Asked Questions

### How does WeKnora handle model selection when multiple knowledge bases are involved?

The system evaluates the resolved KB list in order. The first KB with a Remote summary model (`ModelSourceRemote`) dictates the model ID. If none exist, the first KB's local summary model is used. Only when no KB specifies a model does the pipeline fall back to a generic `ModelTypeKnowledgeQA` instance.

### What is the difference between `thinking_control` and the CustomAgent `Thinking` field?

The `extra_config.thinking_control` string selects the wire-format strategy (`enable_thinking` boolean vs. `thinking` object) but respects the provider's default value. The CustomAgent `Thinking` boolean (`*bool`) is a hard override that forces reasoning on or off regardless of provider defaults, stored in `PipelineRequest.Thinking`.

### Which file contains the logic for applying thinking-mode overrides to the final request?

The `Apply` method implementations in [`internal/models/chat/thinking.go`](https://github.com/Tencent/WeKnora/blob/main/internal/models/chat/thinking.go) (lines 61–106) encode the strategy into provider-specific structs. However, the ultimate precedence logic—checking for `thinkingOverride` before defaulting to `provider.Thinking()`—lives in [`internal/models/chat/provider.go`](https://github.com/Tencent/WeKnora/blob/main/internal/models/chat/provider.go) (lines 136–144).

### Can thinking mode be disabled for non-streaming requests?

Yes. The Qwen adapter explicitly forces `EnableThinking` to `false` when streaming is disabled, regardless of the override strategy. This behavior is hardcoded in the `Apply` method for the `enableThinking` strategy to comply with API constraints.