How CloddsBot WebChat Implements Context Compacting for Long Conversations

CloddsBot delegates context compaction to the Session Manager, which automatically summarizes older messages into a bullet-point recap when the conversation buffer exceeds 20 messages, preserving only the 10 most recent exchanges while maintaining conversational continuity through a synthetic summary message.

The alsk1992/CloddsBot repository implements an intelligent context management system that prevents token limit errors during extended dialogues. While the WebChat channel handles real-time message ingestion via WebSocket, the actual compaction logic resides in the session layer, ensuring that large language models (LLMs) receive bounded context windows regardless of conversation length.

The Architecture of Context Compacting

CloddsBot's WebChat does not perform compaction internally. Instead, every inbound message flows through the Session Manager (src/sessions/index.ts), which maintains an in-memory LLM context window. When this buffer grows beyond configurable thresholds, the system triggers an extractive summarization process that archives older dialogue while preserving semantic context.

This architecture separates concerns cleanly: the channel handles transport, while the session manager handles token economics and context window optimization.

How Message Flow Triggers Compaction

The compaction process follows a strict pipeline from message reception to history truncation.

Message Reception via WebChat Channel

The WebChat channel (src/channels/webchat/index.ts) converts raw WebSocket payloads into standardized IncomingMessage objects. No token analysis occurs at this layer; the channel simply forwards the message via the onMessage callback:

await callbacks.onMessage({
  id: randomUUID(),
  platform: 'webchat',
  userId: session.userId,
  chatId: session.id,
  chatType: 'dm',
  text: incoming.text,
  timestamp: new Date(),
});

Buffer Management in SessionManager

When sessionManager.addToHistory() receives the message, it appends the exchange to the session's conversationHistory array. This array serves as the LLM context buffer, storing up to MAX_LLM_CONTEXT messages before triggering compaction.

The Compaction Trigger

After each insertion, addToHistory checks whether the buffer exceeds the limit of 20 messages (MAX_LLM_CONTEXT). When overflow is detected, the system evicts the oldest messages, retaining only the 10 most recent exchanges (COMPACT_KEEP_RECENT):

// Inside addToHistory (simplified logic)
if (session.context.conversationHistory.length > MAX_LLM_CONTEXT) {
  const overflow = session.context.conversationHistory.length - COMPACT_KEEP_RECENT;
  const evicted = session.context.conversationHistory.slice(0, overflow);
  const newSummary = compactMessages(evicted);
  session.context.contextSummary = mergeSummaries(session.context.contextSummary, newSummary);
  session.context.conversationHistory = session.context.conversationHistory.slice(overflow);
}

The evicted messages are passed to compactMessages, while the buffer is trimmed to retain only recent history.

The Summarization Algorithm

The compactMessages function in src/sessions/index.ts performs extractive summarization rather than abstractive generation. For each evicted message, it extracts the first meaningful sentence, capped at approximately 120 characters, and formats these into a bullet-point list.

This approach guarantees deterministic performance and prevents token bloat from verbose summarization models. The resulting summary is stored in session.context.contextSummary and capped at 3000 characters to keep the JSON payload small. If the combined summary exceeds this limit, only the most recent portion is retained.

Reconstructing Context for the LLM

When constructing the final prompt via SessionManager.getHistory(), the system reconstructs the full context by prepending the stored summary as a synthetic system message. This creates a "previously on" narrative that informs the model of earlier dialogue without transmitting the full text:

const prompt = sessionManager.getHistory(session);
// Returns:
// [
//   { role: 'user', content: '[Previous conversation summary]\n- User: …\n- Assistant: …' },
//   { role: 'assistant', content: 'Understood. I have context from our earlier conversation.' },
//   // … up to 20 recent messages
// ]

This technique maintains conversational continuity while ensuring the LLM receives a bounded token payload, typically reducing context size by 50% or more during long interactions.

Configuration Constants and Limits

The compaction behavior is governed by constants defined in src/sessions/index.ts lines 23-29:

  • MAX_LLM_CONTEXT = 20: The maximum number of messages held in the active buffer before compaction triggers.
  • COMPACT_KEEP_RECENT = 10: The number of recent messages retained after compaction, ensuring immediate context remains un-summarized for accuracy.
  • Summary cap: Approximately 3000 characters to prevent JSON bloat.
  • Extract length: ~120 characters per message extracted for the summary.

These defaults balance token efficiency against context fidelity, though they can be modified at the session configuration level.

Summary

  • CloddsBot WebChat delegates all compaction logic to the Session Manager rather than handling it within the channel layer.
  • The system uses extractive summarization that extracts the first meaningful sentence from evicted messages and formats them as bullet points.
  • Configuration constants MAX_LLM_CONTEXT (20) and COMPACT_KEEP_RECENT (10) control when compaction triggers and how much recent history remains unsummarized.
  • The compactMessages function in src/sessions/index.ts handles the summarization, while getHistory reconstructs the prompt by prepending the summary as a synthetic message.
  • A hard cap of ~3000 characters on the summary prevents payload bloat during extremely long conversations.

Frequently Asked Questions

Does the WebChat channel handle compaction itself?

No. According to the source code in src/channels/webchat/index.ts, the WebChat channel is responsible only for receiving WebSocket messages and forwarding them via the onMessage callback. All compaction logic resides in src/sessions/index.ts within the Session Manager class, which maintains the conversationHistory buffer and triggers summarization when thresholds are exceeded.

What happens to messages that get compacted?

Evicted messages are processed by the compactMessages function, which extracts the first meaningful sentence (capped at ~120 characters) from each message. These extracts are compiled into a bullet-point summary and stored in session.context.contextSummary. The original full text of the evicted messages is removed from the active buffer but semantically preserved through the summary.

How does the LLM receive context after compaction?

When SessionManager.getHistory() constructs the prompt, it prepends the stored contextSummary as a synthetic user message with the label "[Previous conversation summary]" followed by the bullet-point list. This is followed by up to 10 recent unsummarized messages (as determined by COMPACT_KEEP_RECENT), giving the LLM awareness of both distant and recent dialogue without exceeding token limits.

Can the compaction thresholds be adjusted for different use cases?

Yes. The constants MAX_LLM_CONTEXT and COMPACT_KEEP_RECENT are defined in src/sessions/index.ts (lines 23-29) and can be modified to suit different LLM context windows or conversation styles. However, the summary size guard of ~3000 characters is hardcoded to prevent JSON serialization issues during session persistence.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →