# How the WebChat Context-Compacting System Summarizes Messages for LLM Input

> Discover how CloddsBot's WebChat system compresses older messages into bullet points and preserves recent ones for LLM input, avoiding external LLM calls.

- Repository: [AL/CloddsBot](https://github.com/alsk1992/CloddsBot)
- Tags: how-to-guide
- Published: 2026-09-14

---

**The WebChat component in CloddsBot uses a pure extractive summarization routine that preserves the most recent messages while compressing older dialogue into bullet‑point excerpts, all without requiring external LLM calls.**

The CloddsBot repository implements an intelligent context management strategy to prevent token limit overflows during LLM interactions. When conversation history exceeds configured thresholds, the WebChat context-compacting system automatically triggers, replacing verbose historical messages with concise summaries derived directly from the source text.

## The Compact-Messages Algorithm

The compaction logic follows a two‑stage process that balances context preservation with token efficiency.

### Preserving Recent Context with COMPACT_KEEP_RECENT

When the total number of stored messages exceeds the LLM context limit, the system immediately retains the newest `COMPACT_KEEP_RECENT` (default **10**) messages unchanged. These recent exchanges remain in full text to ensure the model maintains awareness of the immediate conversation flow.

### Extractive Summarization of Historic Messages

All older messages undergo a local extractive summarization process:

1. **Sentence Extraction** – For each historic `ConversationMessage`, the algorithm searches for the first *meaningful* sentence (defined as a fragment longer than five characters) by splitting on punctuation marks (`[.!?\n]`). If no suitable sentence exists, it falls back to the first 120 characters of the content.
2. **Role Prefixing** – Each extracted excerpt is prefixed with the speaker role (`User` or `Assistant`) to maintain conversational context.
3. **Length Constraints** – The final string is truncated to 120 characters and formatted as `- <Role>: <Excerpt>`.
4. **Concatenation** – Entries are joined with newline separators to create a single summary block.

This approach gives the model enough context to understand the conversation’s gist without exceeding token limits.

## Core Implementation in src/sessions/index.ts

The `compactMessages` function in [`src/sessions/index.ts`](https://github.com/alsk1992/CloddsBot/blob/main/src/sessions/index.ts) implements the extractive logic. It iterates through historic messages, applies the sentence extraction rules, and returns a formatted string suitable for LLM consumption.

```typescript
// src/sessions/index.ts – compacting old messages
function compactMessages(messages: ConversationMessage[]): string {
  const lines: string[] = [];
  for (const msg of messages) {
    const text = msg.content.trim();
    if (!text) continue;

    const prefix = msg.role === 'user' ? 'User' : 'Assistant';
    // Grab the first meaningful sentence (>= 5 chars) or fallback to 120 chars
    const firstLine =
      text.split(/[.!?\n]/).filter(s => s.trim().length > 5)[0]?.trim()
      || text.slice(0, 120);
    lines.push(`- ${prefix}: ${firstLine.slice(0, 120)}`);
  }
  return lines.join('\n');
}

```

The function handles edge cases such as empty content and role normalization, ensuring robust output even with malformed input.

## Context Pipeline Integration

The compaction workflow resides in [`src/memory/context.ts`](https://github.com/alsk1992/CloddsBot/blob/main/src/memory/context.ts), where the `ContextManager` monitors token utilization. When `percentUsed` exceeds the configured `compactThreshold`, the system partitions the message array and invokes the summarization routine.

```typescript
// src/memory/context.ts – triggered when token usage exceeds the threshold
if (percentUsed >= compactThreshold) {
  const recent = messages.slice(-COMPACT_KEEP_RECENT);
  const historic = messages.slice(0, -COMPACT_KEEP_RECENT);
  const summary = compactMessages(historic);           // ← extractive summary
  const compacted = [...recent, { role: 'system', content: summary }];
  // compacted array is then stored and sent to the LLM
}

```

The resulting `compacted` array contains the full recent messages followed by a single system message containing the bullet‑point summary of earlier dialogue.

## Extensibility and Event Hooks

The system exposes lifecycle events for custom extensions. According to the source in [`src/hooks/index.ts`](https://github.com/alsk1992/CloddsBot/blob/main/src/hooks/index.ts), the compacting process emits `compaction:before` and `compaction:after` events, allowing developers to intercept or modify the summary before it reaches the LLM.

## Optional Deep Summarization

While the default implementation uses pure extractive methods, the repository also provides [`src/memory/summarizer.ts`](https://github.com/alsk1992/CloddsBot/blob/main/src/memory/summarizer.ts) for optional Claude‑powered abstractive summarization. This alternative path activates when the `compact` mode is explicitly configured to call an external LLM, offering deeper semantic compression at the cost of additional API latency.

## Summary

- **Threshold‑driven activation**: The system triggers when `percentUsed` exceeds the configured `compactThreshold` in [`src/memory/context.ts`](https://github.com/alsk1992/CloddsBot/blob/main/src/memory/context.ts).
- **Hybrid retention strategy**: It preserves the last `COMPACT_KEEP_RECENT` (default 10) messages in full while summarizing older content.
- **Pure extractive logic**: The `compactMessages` function in [`src/sessions/index.ts`](https://github.com/alsk1992/CloddsBot/blob/main/src/sessions/index.ts) performs local summarization without external API calls, extracting the first meaningful sentence or falling back to 120 characters.
- **Structured output**: Summaries use a consistent `- Role: Excerpt` format to maintain conversational context for the LLM.
- **Extensible architecture**: Event hooks in [`src/hooks/index.ts`](https://github.com/alsk1992/CloddsBot/blob/main/src/hooks/index.ts) and optional Claude integration in [`src/memory/summarizer.ts`](https://github.com/alsk1992/CloddsBot/blob/main/src/memory/summarizer.ts) provide flexibility for advanced use cases.

## Frequently Asked Questions

### What triggers the context-compacting system in CloddsBot?

The system activates when the percentage of tokens used (`percentUsed`) surpasses the `compactThreshold` configured in the context manager. At this point, [`src/memory/context.ts`](https://github.com/alsk1992/CloddsBot/blob/main/src/memory/context.ts) partitions the message history and routes older messages through the `compactMessages` function.

### How does the extractive summarizer choose which text to keep?

The algorithm splits message content on sentence terminators (`[.!?\n]`) and selects the first segment longer than five characters. If no qualifying sentence exists, it defaults to the first 120 characters of the raw content, ensuring every message contributes at least partial information to the final summary.

### Can I configure the number of recent messages preserved during compaction?

Yes. The `COMPACT_KEEP_RECENT` constant controls how many of the newest messages remain unsummarized. The default value is 10, but this can be adjusted in the configuration to trade off between immediate context depth and historical summary length.

### Does CloddsBot support LLM-powered summarization instead of extractive methods?

Yes. While the default implementation in [`src/sessions/index.ts`](https://github.com/alsk1992/CloddsBot/blob/main/src/sessions/index.ts) uses pure extractive summarization, the repository includes [`src/memory/summarizer.ts`](https://github.com/alsk1992/CloddsBot/blob/main/src/memory/summarizer.ts), which provides Claude‑powered abstractive summarization. This optional path requires external API calls and activates when the compaction mode is configured to use LLM‑based rather than extractive compression.