DeepSeek TUI Thinking-Mode Streaming Chain-of-Thought: Architecture and Implementation
DeepSeek TUI implements DeepSeek's reasoning mode by streaming Thinking blocks through a dedicated event pipeline, buffering chain-of-thought content in memory, and persisting it alongside tool calls to satisfy API validation requirements.
The Hmbown/DeepSeek-TUI repository provides a terminal interface that handles DeepSeek's thinking-mode (reasoning) capability. In this mode, the model emits a chain-of-thought stream before generating visible responses or tool calls. The application manages this thinking-mode streaming chain-of-thought through a sophisticated buffering system that ensures every assistant message containing tool invocations also includes the prior reasoning content—otherwise the DeepSeek API returns HTTP 400 errors.
Core Architecture of Thinking-Mode Streaming
The implementation centers on three interconnected layers: the data model representing thinking content, the event stream processing deltas in real-time, and the UI state management that renders live reasoning to the user.
ContentBlock Enum and the Thinking Variant
At the foundation, the ContentBlock enum in crates/tui/src/models.rs (lines 71-80) unifies all payload types an assistant message can carry. The Thinking variant specifically holds raw chain-of-thought text:
// From crates/tui/src/models.rs
enum ContentBlock {
Thinking { thinking: String },
Text { text: String },
ToolUse { id: String, name: String, input: Value, caller: Option<String> },
// ...
}
This variant stores the accumulated reasoning that must later accompany any tool calls to maintain API compliance.
EngineEvent Streaming Lifecycle
The streaming pipeline processes thinking content through distinct lifecycle events defined in the EngineEvent enum at lines 13-33 of models.rs:
ThinkingStarted: Resetsapp.reasoning_bufferand records timestampsThinkingDelta: Delivers incremental chunks that get sanitized and appended to the bufferThinkingComplete: Finalizes the stream, computing duration and extracting content toapp.last_reasoning
These events mirror DeepSeek's interleaved thinking contract, ensuring the TUI can handle reasoning streams of arbitrary length without blocking the interface.
UI Buffering and Active Cell Management
The UI layer maintains an in-memory buffer (app.reasoning_buffer) that accumulates streaming text until completion. Three helper functions in crates/tui/src/tui/active_cell.rs manage the visual representation:
ensure_streaming_thinking_active_entry: Creates aHistoryCell::Thinkingplaceholder when streaming beginsappend_streaming_thinking: Updates the active cell's content with each deltafinalize_streaming_thinking_active_entry: Marks the cell complete and calculates duration
The HistoryCell::Thinking type tracks content, a streaming flag, and timestamps. Rendering logic in crates/tui/src/history.rs uses thinking_visual_state to determine whether to display a "live" spinner, "done" state, or idle status.
Message Persistence and API Compliance
DeepSeek's API requires that every assistant message containing a tool call must also include the prior reasoning content. The persistence logic in crates/tui/src/tui/ui.rs (lines 669-704) enforces this through the has_sendable_content check:
let has_sendable_content = blocks.iter().any(|b| matches!(
b,
ContentBlock::Text { .. } | ContentBlock::ToolUse { .. }
));
if has_sendable_content {
api_messages.push(Message {
role: "assistant".to_string(),
content: blocks, // Includes the Thinking block
});
}
Turns consisting solely of reasoning blocks are discarded, as the server rejects assistant messages containing only internal monologue without visible output or tool invocations.
Token Accounting for Reasoning Replay
The system tracks the cost of reasoning content that must be replayed in subsequent turns. In models.rs (lines 94-99), the Usage struct includes reasoning_replay_tokens, calculated at approximately 4 characters per token to estimate the overhead of maintaining conversation context.
Practical Implementation Examples
Creating Messages with Thinking Blocks
When constructing assistant messages that combine reasoning with tool calls, assemble the blocks in order:
use deepseek_tui::models::{Message, ContentBlock};
// Create the reasoning block
let thinking = ContentBlock::Thinking {
thinking: "First, I examine the prompt to understand the search requirements...".to_string(),
};
// Append the tool invocation
let tool_use = ContentBlock::ToolUse {
id: "tool-1".to_string(),
name: "search".to_string(),
input: serde_json::json!({ "query": "rust async streams" }),
caller: None,
};
// Construct the complete message
let msg = Message {
role: "assistant".to_string(),
content: vec![thinking, tool_use],
};
The UI renders a live spinner during the Thinking block stream, then transitions to display the tool card once the ToolUse block arrives.
Handling Streaming UI Updates
During the event loop, update the active thinking entry as deltas arrive:
// Inside the UI event loop processing EngineEvent::ThinkingDelta
if let Some(entry_idx) = app.active_cell().find_thinking_entry() {
let cell = app.history.entry_mut(entry_idx);
// Append sanitized content to the live cell
cell.content.push_str(&sanitized_chunk);
app.mark_transcript_dirty();
}
This pattern ensures the user sees reasoning accumulate in real-time without blocking other interface updates.
Persisting Compliant Assistant Turns
Before pushing to api_messages, validate that the turn contains sendable content:
// blocks contains ContentBlock elements including Thinking
let has_sendable_content = blocks.iter().any(|b| matches!(
b,
ContentBlock::Text { .. } | ContentBlock::ToolUse { .. }
));
if has_sendable_content {
app.api_messages.push(Message {
role: "assistant".to_string(),
content: blocks,
});
} else {
// Discard reasoning-only turns to avoid HTTP 400 errors
tracing::debug!("Skipping persistence: reasoning-only turn");
}
This check prevents API errors while ensuring reasoning context is preserved alongside functional responses.
Key Source Files
crates/tui/src/models.rs– DefinesContentBlock,EngineEventvariants (ThinkingStarted,ThinkingDelta,ThinkingComplete),Messagestructure, andUsage::reasoning_replay_tokens(lines 13-33, 71-80, 94-99)crates/tui/src/tui/ui.rs– Main event loop handling thinking events and persisting logic withhas_sendable_contentvalidation (lines 669-704)crates/tui/src/tui/active_cell.rs– Helper functionsensure_streaming_thinking_active_entry,append_streaming_thinking, andfinalize_streaming_thinking_active_entrycrates/tui/src/history.rs–HistoryCell::Thinkingdefinition andthinking_visual_staterendering logiccrates/tui/tests/integration_mock_llm.rs– Integration tests verifying the Thinking→Tool→Tool chain behavior (line 80 containsContentBlock::Thinkingexamples)
Summary
- DeepSeek TUI implements thinking-mode streaming through a dedicated
EngineEventpipeline that processesThinkingStarted,ThinkingDelta, andThinkingCompleteevents - The
ContentBlock::Thinkingvariant inmodels.rsstores raw chain-of-thought text that must accompany tool calls to satisfy API requirements - UI buffering accumulates reasoning in
app.reasoning_bufferand renders it throughHistoryCell::Thinkingentries with live spinner states managed byactive_cell.rshelpers - Persistence logic discards reasoning-only turns while preserving reasoning content alongside
TextorToolUseblocks to prevent HTTP 400 errors - Token accounting tracks
reasoning_replay_tokensto monitor the cost of maintaining conversation context across multi-turn interactions
Frequently Asked Questions
What happens if an assistant turn contains only a reasoning block?
The application discards the turn. According to the logic in crates/tui/src/tui/ui.rs, the code checks has_sendable_content to verify the presence of Text or ToolUse blocks. If only ContentBlock::Thinking is present, the message is not pushed to api_messages, preventing the DeepSeek API from returning an HTTP 400 error.
How does the UI indicate that thinking is in progress?
The UI creates a HistoryCell::Thinking entry via ensure_streaming_thinking_active_entry when EngineEvent::ThinkingStarted fires. The thinking_visual_state function in history.rs determines the visual state based on the streaming flag, displaying a live spinner during active streaming and transitioning to a "done" state when ThinkingComplete finalizes the entry.
Why must reasoning content be replayed with tool calls?
DeepSeek's API contract requires every assistant message containing tool calls to include the prior reasoning content. If an assistant message invokes tools without the accompanying Thinking block that led to that decision, the API returns HTTP 400. The TUI ensures compliance by always including the accumulated app.last_reasoning content when constructing messages with ToolUse blocks.
How are reasoning tokens tracked for usage statistics?
The Usage struct in models.rs (lines 94-99) includes a reasoning_replay_tokens field that estimates token consumption using a ratio of approximately 4 characters per token. This accounting helps users understand the overhead cost of multi-turn conversations where reasoning content must be replayed with each subsequent tool-calling message.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →