How Thinking Mode Affects Token Usage and Response Latency in Qwen-Code

Thinking mode generates a separate "thinking" message that increases response latency by adding an extra generation step, but its tokens are excluded from the user-visible usage statistics in the UI.

The QwenLM/qwen-code repository implements an advanced reasoning feature that allows models to emit intermediate reasoning before generating final answers. While this improves response quality, it introduces distinct behaviors regarding token usage and response latency that differ from standard assistant responses. Understanding these mechanics helps developers optimize their interactions with the Qwen coding assistant.

Token Accounting Excludes Thinking Messages

Despite the model generating additional content during thinking mode, the user-visible token count remains unchanged. When the assistant produces a thinking block, the output payload contains a message with role: "thinking", which JSONLAdapter.getMessageType identifies separately from standard assistant responses.

In packages/webui/src/adapters/JSONLAdapter.ts, the adapter explicitly checks for this role:

// JSONLAdapter – maps the incoming role to a unified type
function getMessageType(msg: JSONLMessage): UnifiedMessageType {
  // ...
  if (msg.message?.role === 'thinking') {
    return 'thinking';
  }
}

The ContextIndicator component in packages/webui/src/components/layout/ContextIndicator.tsx receives a ContextUsage object populated directly from the model's usage field. Thinking messages are filtered out before merging into this object, ensuring the displayed statistics only reflect input_tokens and output_tokens from the final answer. Consequently, enabling thinking mode does not inflate the token counter visible to users, even though the LLM internally processes additional tokens.

Response Latency Increases with Extra Generation

Thinking mode adds a noticeable but modest latency penalty—typically a few hundred milliseconds to one second—because the model must complete two distinct generation steps: first emitting the thinking block, then producing the final response. This sequential processing extends the overall round-trip time compared to single-step generation.

During streaming, the useMessageHandling hook in the VS Code companion manages this flow by inserting a temporary placeholder:

// useMessageHandling – inserts a temporary thinking placeholder while streaming
if (streaming && newMsg.role === 'assistant') {
  next.push({ role: 'thinking', content: '', timestamp: Date.now() });
}

This placeholder creates a visual "Thinking..." indicator while the stream remains open, removed only when the final assistant message arrives. The UI renders a collapsible ThinkingMessage component during this period, providing user feedback while the model completes its internal reasoning.

UI Architecture for Thinking Mode

The implementation spans multiple components to handle the distinct message type without disrupting standard chat flows.

Input Controls and State Management

The InputForm component in packages/webui/src/components/layout/InputForm.tsx maintains the thinkingEnabled boolean flag and provides the onToggleThinking callback for enabling the feature:

// InputForm – props include the toggle for thinking mode
export interface InputFormProps {
  // ...
  /** Whether thinking mode is enabled */
  thinkingEnabled: boolean;
  /** Toggle thinking callback */
  onToggleThinking: () => void;
}

Collapsible Reasoning Display

The ThinkingMessage component in packages/webui/src/components/messages/ThinkingMessage.tsx renders the internal reasoning as a collapsible banner featuring a gray dot and "Thinking" label. It displays when status === 'loading' and expands to show reasoning content upon user interaction:

// ThinkingMessage – collapsible UI component
export const ThinkingMessage: FC<ThinkingMessageProps> = ({
  content,
  defaultExpanded = false,
  status = 'default',
}) => {
  const [isExpanded, setIsExpanded] = useState(defaultExpanded);
  // ...
  return (
    <div className={`qwen-message thinking-message thinking-status-${status}`}>
      <button onClick={handleToggle}> … </button>
      {isExpanded && <MessageContent content={content} />}
    </div>
  );
};

Summary

  • Token visibility: Thinking mode tokens are generated by the backend but excluded from ContextIndicator statistics, keeping the displayed token count identical to non-thinking responses.
  • Latency impact: An additional generation step for the thinking block increases total response time by approximately 0.3–1.0 seconds depending on model size and network conditions.
  • Streaming behavior: The useMessageHandling hook inserts temporary thinking placeholders during streaming, creating a visual indicator without affecting token accounting.
  • UI integration: The ThinkingMessage component provides collapsible access to reasoning content while maintaining separation from standard assistant message flows.

Frequently Asked Questions

Does thinking mode increase my token count?

No. Although the model generates additional tokens for the thinking block, the ContextIndicator component filters out messages with role: "thinking" before calculating usage statistics. The UI displays only the input_tokens and output_tokens from the final assistant response.

Why does thinking mode make responses slower?

Thinking mode requires the model to perform two sequential generation steps: first producing the reasoning content, then generating the final answer. This extra round-trip adds latency typically ranging from a few hundred milliseconds to one second, depending on the model size and network speed.

Can I see the thinking content after the response completes?

Yes. The ThinkingMessage component renders the reasoning as a collapsible section that persists after streaming completes. Users can click the "Thinking" banner to expand and view the internal reasoning that led to the final answer.

Is thinking mode available in the VS Code extension?

Yes. The VS Code companion implements thinking mode through the useMessageHandling hook in packages/vscode-ide-companion/src/webview/hooks/message/useMessageHandling.ts, which manages the temporary thinking placeholder during streaming sessions.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →