How the WebChat Context-Compacting System Summarizes Messages for LLM Input

The WebChat component in CloddsBot uses a pure extractive summarization routine that preserves the most recent messages while compressing older dialogue into bullet‑point excerpts, all without requiring external LLM calls.

The CloddsBot repository implements an intelligent context management strategy to prevent token limit overflows during LLM interactions. When conversation history exceeds configured thresholds, the WebChat context-compacting system automatically triggers, replacing verbose historical messages with concise summaries derived directly from the source text.

The Compact-Messages Algorithm

The compaction logic follows a two‑stage process that balances context preservation with token efficiency.

Preserving Recent Context with COMPACT_KEEP_RECENT

When the total number of stored messages exceeds the LLM context limit, the system immediately retains the newest COMPACT_KEEP_RECENT (default 10) messages unchanged. These recent exchanges remain in full text to ensure the model maintains awareness of the immediate conversation flow.

Extractive Summarization of Historic Messages

All older messages undergo a local extractive summarization process:

  1. Sentence Extraction – For each historic ConversationMessage, the algorithm searches for the first meaningful sentence (defined as a fragment longer than five characters) by splitting on punctuation marks ([.!?\n]). If no suitable sentence exists, it falls back to the first 120 characters of the content.
  2. Role Prefixing – Each extracted excerpt is prefixed with the speaker role (User or Assistant) to maintain conversational context.
  3. Length Constraints – The final string is truncated to 120 characters and formatted as - <Role>: <Excerpt>.
  4. Concatenation – Entries are joined with newline separators to create a single summary block.

This approach gives the model enough context to understand the conversation’s gist without exceeding token limits.

Core Implementation in src/sessions/index.ts

The compactMessages function in src/sessions/index.ts implements the extractive logic. It iterates through historic messages, applies the sentence extraction rules, and returns a formatted string suitable for LLM consumption.

// src/sessions/index.ts – compacting old messages
function compactMessages(messages: ConversationMessage[]): string {
  const lines: string[] = [];
  for (const msg of messages) {
    const text = msg.content.trim();
    if (!text) continue;

    const prefix = msg.role === 'user' ? 'User' : 'Assistant';
    // Grab the first meaningful sentence (>= 5 chars) or fallback to 120 chars
    const firstLine =
      text.split(/[.!?\n]/).filter(s => s.trim().length > 5)[0]?.trim()
      || text.slice(0, 120);
    lines.push(`- ${prefix}: ${firstLine.slice(0, 120)}`);
  }
  return lines.join('\n');
}

The function handles edge cases such as empty content and role normalization, ensuring robust output even with malformed input.

Context Pipeline Integration

The compaction workflow resides in src/memory/context.ts, where the ContextManager monitors token utilization. When percentUsed exceeds the configured compactThreshold, the system partitions the message array and invokes the summarization routine.

// src/memory/context.ts – triggered when token usage exceeds the threshold
if (percentUsed >= compactThreshold) {
  const recent = messages.slice(-COMPACT_KEEP_RECENT);
  const historic = messages.slice(0, -COMPACT_KEEP_RECENT);
  const summary = compactMessages(historic);           // ← extractive summary
  const compacted = [...recent, { role: 'system', content: summary }];
  // compacted array is then stored and sent to the LLM
}

The resulting compacted array contains the full recent messages followed by a single system message containing the bullet‑point summary of earlier dialogue.

Extensibility and Event Hooks

The system exposes lifecycle events for custom extensions. According to the source in src/hooks/index.ts, the compacting process emits compaction:before and compaction:after events, allowing developers to intercept or modify the summary before it reaches the LLM.

Optional Deep Summarization

While the default implementation uses pure extractive methods, the repository also provides src/memory/summarizer.ts for optional Claude‑powered abstractive summarization. This alternative path activates when the compact mode is explicitly configured to call an external LLM, offering deeper semantic compression at the cost of additional API latency.

Summary

  • Threshold‑driven activation: The system triggers when percentUsed exceeds the configured compactThreshold in src/memory/context.ts.
  • Hybrid retention strategy: It preserves the last COMPACT_KEEP_RECENT (default 10) messages in full while summarizing older content.
  • Pure extractive logic: The compactMessages function in src/sessions/index.ts performs local summarization without external API calls, extracting the first meaningful sentence or falling back to 120 characters.
  • Structured output: Summaries use a consistent - Role: Excerpt format to maintain conversational context for the LLM.
  • Extensible architecture: Event hooks in src/hooks/index.ts and optional Claude integration in src/memory/summarizer.ts provide flexibility for advanced use cases.

Frequently Asked Questions

What triggers the context-compacting system in CloddsBot?

The system activates when the percentage of tokens used (percentUsed) surpasses the compactThreshold configured in the context manager. At this point, src/memory/context.ts partitions the message history and routes older messages through the compactMessages function.

How does the extractive summarizer choose which text to keep?

The algorithm splits message content on sentence terminators ([.!?\n]) and selects the first segment longer than five characters. If no qualifying sentence exists, it defaults to the first 120 characters of the raw content, ensuring every message contributes at least partial information to the final summary.

Can I configure the number of recent messages preserved during compaction?

Yes. The COMPACT_KEEP_RECENT constant controls how many of the newest messages remain unsummarized. The default value is 10, but this can be adjusted in the configuration to trade off between immediate context depth and historical summary length.

Does CloddsBot support LLM-powered summarization instead of extractive methods?

Yes. While the default implementation in src/sessions/index.ts uses pure extractive summarization, the repository includes src/memory/summarizer.ts, which provides Claude‑powered abstractive summarization. This optional path requires external API calls and activates when the compaction mode is configured to use LLM‑based rather than extractive compression.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →