# How CloddsBot WebChat Implements Context Compacting for Long Conversations

> Discover how CloddsBot WebChat uses context compacting to manage long conversations. The Session Manager summarizes older messages, preserving continuity and the 10 most recent exchanges.

- Repository: [AL/CloddsBot](https://github.com/alsk1992/CloddsBot)
- Tags: internals
- Published: 2026-09-11

---

**CloddsBot delegates context compaction to the Session Manager, which automatically summarizes older messages into a bullet-point recap when the conversation buffer exceeds 20 messages, preserving only the 10 most recent exchanges while maintaining conversational continuity through a synthetic summary message.**

The **alsk1992/CloddsBot** repository implements an intelligent context management system that prevents token limit errors during extended dialogues. While the WebChat channel handles real-time message ingestion via WebSocket, the actual compaction logic resides in the session layer, ensuring that large language models (LLMs) receive bounded context windows regardless of conversation length.

## The Architecture of Context Compacting

CloddsBot's WebChat does not perform compaction internally. Instead, every inbound message flows through the **Session Manager** ([`src/sessions/index.ts`](https://github.com/alsk1992/CloddsBot/blob/main/src/sessions/index.ts)), which maintains an in-memory LLM context window. When this buffer grows beyond configurable thresholds, the system triggers an **extractive summarization** process that archives older dialogue while preserving semantic context.

This architecture separates concerns cleanly: the channel handles transport, while the session manager handles token economics and context window optimization.

## How Message Flow Triggers Compaction

The compaction process follows a strict pipeline from message reception to history truncation.

### Message Reception via WebChat Channel

The WebChat channel ([`src/channels/webchat/index.ts`](https://github.com/alsk1992/CloddsBot/blob/main/src/channels/webchat/index.ts)) converts raw WebSocket payloads into standardized `IncomingMessage` objects. No token analysis occurs at this layer; the channel simply forwards the message via the `onMessage` callback:

```typescript
await callbacks.onMessage({
  id: randomUUID(),
  platform: 'webchat',
  userId: session.userId,
  chatId: session.id,
  chatType: 'dm',
  text: incoming.text,
  timestamp: new Date(),
});

```

### Buffer Management in SessionManager

When `sessionManager.addToHistory()` receives the message, it appends the exchange to the session's `conversationHistory` array. This array serves as the LLM context buffer, storing up to `MAX_LLM_CONTEXT` messages before triggering compaction.

### The Compaction Trigger

After each insertion, `addToHistory` checks whether the buffer exceeds the limit of **20 messages** (`MAX_LLM_CONTEXT`). When overflow is detected, the system evicts the oldest messages, retaining only the **10 most recent** exchanges (`COMPACT_KEEP_RECENT`):

```typescript
// Inside addToHistory (simplified logic)
if (session.context.conversationHistory.length > MAX_LLM_CONTEXT) {
  const overflow = session.context.conversationHistory.length - COMPACT_KEEP_RECENT;
  const evicted = session.context.conversationHistory.slice(0, overflow);
  const newSummary = compactMessages(evicted);
  session.context.contextSummary = mergeSummaries(session.context.contextSummary, newSummary);
  session.context.conversationHistory = session.context.conversationHistory.slice(overflow);
}

```

The evicted messages are passed to `compactMessages`, while the buffer is trimmed to retain only recent history.

## The Summarization Algorithm

The `compactMessages` function in [`src/sessions/index.ts`](https://github.com/alsk1992/CloddsBot/blob/main/src/sessions/index.ts) performs **extractive summarization** rather than abstractive generation. For each evicted message, it extracts the first meaningful sentence, capped at approximately **120 characters**, and formats these into a bullet-point list.

This approach guarantees deterministic performance and prevents token bloat from verbose summarization models. The resulting summary is stored in `session.context.contextSummary` and capped at **3000 characters** to keep the JSON payload small. If the combined summary exceeds this limit, only the most recent portion is retained.

## Reconstructing Context for the LLM

When constructing the final prompt via `SessionManager.getHistory()`, the system reconstructs the full context by prepending the stored summary as a synthetic system message. This creates a "previously on" narrative that informs the model of earlier dialogue without transmitting the full text:

```typescript
const prompt = sessionManager.getHistory(session);
// Returns:
// [
//   { role: 'user', content: '[Previous conversation summary]\n- User: …\n- Assistant: …' },
//   { role: 'assistant', content: 'Understood. I have context from our earlier conversation.' },
//   // … up to 20 recent messages
// ]

```

This technique maintains **conversational continuity** while ensuring the LLM receives a bounded token payload, typically reducing context size by 50% or more during long interactions.

## Configuration Constants and Limits

The compaction behavior is governed by constants defined in [`src/sessions/index.ts`](https://github.com/alsk1992/CloddsBot/blob/main/src/sessions/index.ts) lines 23-29:

- **`MAX_LLM_CONTEXT = 20`**: The maximum number of messages held in the active buffer before compaction triggers.
- **`COMPACT_KEEP_RECENT = 10`**: The number of recent messages retained after compaction, ensuring immediate context remains un-summarized for accuracy.
- **Summary cap**: Approximately 3000 characters to prevent JSON bloat.
- **Extract length**: ~120 characters per message extracted for the summary.

These defaults balance token efficiency against context fidelity, though they can be modified at the session configuration level.

## Summary

- **CloddsBot WebChat** delegates all compaction logic to the Session Manager rather than handling it within the channel layer.
- The system uses **extractive summarization** that extracts the first meaningful sentence from evicted messages and formats them as bullet points.
- Configuration constants **`MAX_LLM_CONTEXT`** (20) and **`COMPACT_KEEP_RECENT`** (10) control when compaction triggers and how much recent history remains unsummarized.
- The **`compactMessages`** function in [`src/sessions/index.ts`](https://github.com/alsk1992/CloddsBot/blob/main/src/sessions/index.ts) handles the summarization, while **`getHistory`** reconstructs the prompt by prepending the summary as a synthetic message.
- A **hard cap of ~3000 characters** on the summary prevents payload bloat during extremely long conversations.

## Frequently Asked Questions

### Does the WebChat channel handle compaction itself?

No. According to the source code in [`src/channels/webchat/index.ts`](https://github.com/alsk1992/CloddsBot/blob/main/src/channels/webchat/index.ts), the WebChat channel is responsible only for receiving WebSocket messages and forwarding them via the `onMessage` callback. All compaction logic resides in [`src/sessions/index.ts`](https://github.com/alsk1992/CloddsBot/blob/main/src/sessions/index.ts) within the Session Manager class, which maintains the `conversationHistory` buffer and triggers summarization when thresholds are exceeded.

### What happens to messages that get compacted?

Evicted messages are processed by the `compactMessages` function, which extracts the first meaningful sentence (capped at ~120 characters) from each message. These extracts are compiled into a bullet-point summary and stored in `session.context.contextSummary`. The original full text of the evicted messages is removed from the active buffer but semantically preserved through the summary.

### How does the LLM receive context after compaction?

When `SessionManager.getHistory()` constructs the prompt, it prepends the stored `contextSummary` as a synthetic user message with the label "[Previous conversation summary]" followed by the bullet-point list. This is followed by up to 10 recent unsummarized messages (as determined by `COMPACT_KEEP_RECENT`), giving the LLM awareness of both distant and recent dialogue without exceeding token limits.

### Can the compaction thresholds be adjusted for different use cases?

Yes. The constants `MAX_LLM_CONTEXT` and `COMPACT_KEEP_RECENT` are defined in [`src/sessions/index.ts`](https://github.com/alsk1992/CloddsBot/blob/main/src/sessions/index.ts) (lines 23-29) and can be modified to suit different LLM context windows or conversation styles. However, the summary size guard of ~3000 characters is hardcoded to prevent JSON serialization issues during session persistence.