# How MiniSearch Manages Long Conversations: Conversation Memory, Rolling Summaries, and Token Budgets

> Discover how MiniSearch manages long conversations with rolling summaries and token budgets. Learn how it preserves context up to 800 tokens within its active chat window.

- Repository: [Victor Nogueira/minisearch](https://github.com/felladrin/minisearch)
- Tags: deep-dive
- Published: 2026-03-01

---

**MiniSearch maintains a persistent, rolling conversation summary that compresses historical messages into an 800-token limit, ensuring the active chat window stays within 75% of the LLM's context size while preserving context across browser sessions.**

MiniSearch implements a sophisticated conversation memory and summary system designed to prevent context window overflows without sacrificing historical awareness. By combining a PubSub-based persistence layer with intelligent token budgeting and hybrid summarization strategies, the system enables extended chat sessions that survive page reloads. This architecture continuously folds aging messages into a compressed summary while keeping the live prompt within safe token limits.

## The Conversation Summary Store

MiniSearch persists conversation state using the **create-pubsub** library, storing a lightweight object in the browser's `localStorage`. This store maintains both the current summary text and its associated conversation ID.

In [`client/modules/pubSub.ts`](https://github.com/felladrin/minisearch/blob/main/client/modules/pubSub.ts) (lines 49‑62), the system initializes a dedicated PubSub for summary management:

```typescript
// client/modules/pubSub.ts
const conversationSummaryPubSub = createPubSub({
  conversationId: "",
  summary: "",
});
export const [
  updateConversationSummary,
  ,
  getConversationSummary,
] = conversationSummaryPubSub;

```

The `updateConversationSummary` function writes new `{conversationId, summary}` pairs to storage, while `getConversationSummary` retrieves the current state. Because this store leverages `localStorage`, the rolling summary survives page reloads and remains available across multiple LLM calls within the same browser session.

## Rolling Summary Algorithm

When chat history grows beyond the available token budget, MiniSearch **folds** older messages into the summary rather than discarding them. This process involves three distinct phases: detecting overflow, generating the summary, and handling failure scenarios.

### Detecting Token Overflow

During `generateChatResponse`, the system builds `processedMessages` by iterating through the conversation history in reverse chronological order. The algorithm stops adding messages once the token count would exceed the **available token budget**, calculated as 75% of the model's context size minus reserved tokens for system prompts.

This logic appears in [`client/modules/textGeneration.ts`](https://github.com/felladrin/minisearch/blob/main/client/modules/textGeneration.ts) (lines 297‑312):

```typescript
// client/modules/textGeneration.ts – token-budget calculation
const availableTokenBudget = defaultContextSize * 0.75 - reservedTokens;
...
if (currentTokenCount + messageTokens > availableTokenBudget) {
  break; // remaining older messages become “dropped”
}

```

Messages that do not fit within this budget are classified as *dropped* and queued for summarization.

### LLM-Based Summary Generation

When messages are dropped, MiniSearch invokes `createLlmSummary` (lines 72‑84 of [`client/modules/textGeneration.ts`](https://github.com/felladrin/minisearch/blob/main/client/modules/textGeneration.ts)) to generate an updated summary. The function sends a specialized prompt to the LLM requesting a condensed update under the **SUMMARY_TOKEN_LIMIT** of 800 tokens:

```typescript
// client/modules/textGeneration.ts – createLlmSummary implementation
const instructionLines = [
  "You are the conversation memory manager.",
  `Update the running summary under ${SUMMARY_TOKEN_LIMIT} tokens.`,
  // …additional instructions…
];
...
const chat: ChatMessage[] = [{ role: "user", content: prompt }];
...
return (await generateChatWithOpenAi(chat, () => {})).trim();

```

The resulting summary replaces the previous version in the PubSub store, ensuring the compressed history remains available for subsequent turns.

### Fallback Extractive Summarizer

If the LLM call fails or times out, MiniSearch falls back to `summarizeDroppedMessages` (lines 148‑176 of [`client/modules/textGeneration.ts`](https://github.com/felladrin/minisearch/blob/main/client/modules/textGeneration.ts)). This extractive approach concatenates the dropped messages with any existing summary, then retains as many trailing lines as will fit under the 800-token ceiling:

```typescript
// client/modules/textGeneration.ts – summarizeDroppedMessages excerpt
for (let i = parts.length - 1; i >= 0; i--) {
  const candidate = [parts[i], ...kept].join("\n\n");
  const nextTokens = gptTokenizer.encode(candidate).length;
  if (nextTokens > tokenLimit) break;
  kept.unshift(parts[i]);
}
const summary = kept.join("\n\n");
addLogEntry(`Updated rolling summary (${tokens} tokens)`);

```

This fallback ensures the system remains robust even when LLM services are unavailable, preserving the most recent conversation context within the token constraint.

## Token Budget Management

MiniSearch enforces strict token budgets through two primary constants defined in the source code:

- **`SUMMARY_TOKEN_LIMIT`**: Set to **800** tokens (lines 41‑43 of [`client/modules/textGeneration.ts`](https://github.com/felladrin/minisearch/blob/main/client/modules/textGeneration.ts)), this caps the stored rolling summary size.
- **`defaultContextSize`**: Set to **4096** tokens (lines 16‑17 of [`client/modules/textGenerationUtilities.ts`](https://github.com/felladrin/minisearch/blob/main/client/modules/textGenerationUtilities.ts)), this defines the base context window for various LLM backends.

The **available token budget** for the active chat window uses a conservative 75% utilization rule, calculated as:

```typescript
// client/modules/textGeneration.ts
const availableTokenBudget = defaultContextSize * 0.75 - reservedTokens;

```

The `reservedTokens` variable accounts for the system prompt and initial acknowledgment responses, ensuring sufficient headroom for the model's completion tokens. This prevents API rejections due to context length violations while maximizing the usable conversation history.

## Practical Implementation Examples

### Retrieving the Current Conversation Summary

To access the stored summary programmatically, import the getter function from the PubSub module:

```typescript
import { getConversationSummary } from "./pubSub";

function printCurrentSummary() {
  const { conversationId, summary } = getConversationSummary();
  console.log(`Conversation ${conversationId || "<none>"} summary:`);
  console.log(summary || "(empty)");
}

```

*This implementation references the store defined in* **[`client/modules/pubSub.ts`](https://github.com/felladrin/minisearch/blob/main/client/modules/pubSub.ts)** *(lines 49‑62).*

### Manually Triggering Summary Updates

For custom workflows requiring explicit summary generation, combine the LLM and fallback methods:

```typescript
import { createLlmSummary, summarizeDroppedMessages } from "./textGeneration";
import { updateConversationSummary, getConversationSummary } from "./pubSub";
import { gptTokenizer } from "gpt-tokenizer";

async function foldOldMessages(
  dropped: ChatMessage[],
  conversationId: string,
) {
  const { summary: previous } = getConversationSummary();
  
  // Attempt LLM-based summarization first
  const updated = await createLlmSummary(dropped, previous);
  
  // Persist the new summary
  updateConversationSummary({ conversationId, summary: updated });
}

```

*The LLM-based path is implemented in* **[`client/modules/textGeneration.ts`](https://github.com/felladrin/minisearch/blob/main/client/modules/textGeneration.ts)** *(lines 72‑84), while the extractive fallback resides in* **[`client/modules/textGeneration.ts`](https://github.com/felladrin/minisearch/blob/main/client/modules/textGeneration.ts)** *(lines 148‑176).*

### Calculating Token Budgets for Custom Windows

To implement custom message windowing logic that respects MiniSearch's constraints:

```typescript
import { defaultContextSize } from "./textGenerationUtilities";
import gptTokenizer from "gpt-tokenizer";

function buildActiveWindow(messages: ChatMessage[], reservedTokens: number) {
  const budget = defaultContextSize * 0.75 - reservedTokens;
  const kept: ChatMessage[] = [];

  let used = 0;
  for (const msg of [...messages].reverse()) {
    const msgTokens = gptTokenizer.encode(msg.content).length;
    if (used + msgTokens > budget) break;
    kept.unshift(msg);
    used += msgTokens;
  }
  return kept;
}

```

*This logic mirrors the token-budget calculations in* **[`client/modules/textGeneration.ts`](https://github.com/felladrin/minisearch/blob/main/client/modules/textGeneration.ts)** *(lines 297‑312), using the* **`defaultContextSize`** *constant from* **[`client/modules/textGenerationUtilities.ts`](https://github.com/felladrin/minisearch/blob/main/client/modules/textGenerationUtilities.ts)** *(line 16).*

## Summary

- **Persistent Storage**: MiniSearch uses a PubSub-based store ([`client/modules/pubSub.ts`](https://github.com/felladrin/minisearch/blob/main/client/modules/pubSub.ts)) backed by `localStorage` to maintain conversation summaries across browser sessions.
- **Rolling Compression**: The system detects token overflow during `generateChatResponse` and folds excess messages into an 800-token summary using either LLM generation (`createLlmSummary`) or extractive fallback (`summarizeDroppedMessages`).
- **Conservative Budgeting**: Token limits enforce a 75% utilization cap of the 4096-token context window, reserving space for system prompts and model completions.
- **Dual Summarization**: A hybrid approach prioritizes LLM-generated summaries for coherence but falls back to line-based extraction for reliability.
- **Context Preservation**: By continuously updating the rolling summary, MiniSearch maintains awareness of distant conversation history without exceeding model context limits.

## Frequently Asked Questions

### How does MiniSearch preserve conversation memory across page reloads?

MiniSearch persists the conversation summary using a **create-pubsub** store that serializes data to the browser's `localStorage`. When a user reloads the page, the `getConversationSummary` function retrieves the last saved summary and conversation ID from [`client/modules/pubSub.ts`](https://github.com/felladrin/minisearch/blob/main/client/modules/pubSub.ts), allowing the chat to resume with full historical context intact.

### What happens when the LLM fails to generate a summary?

If the `createLlmSummary` call fails or returns an error, MiniSearch invokes `summarizeDroppedMessages` (lines 148‑176 of [`client/modules/textGeneration.ts`](https://github.com/felladrin/minisearch/blob/main/client/modules/textGeneration.ts)) as a fallback. This extractive method keeps the most recent messages that fit within the 800-token `SUMMARY_TOKEN_LIMIT`, ensuring the conversation can continue even when LLM services are unavailable.

### How is the token budget calculated for each chat response?

The system calculates the **available token budget** as 75% of `defaultContextSize` (4096 tokens) minus reserved tokens for the system prompt and initial responses. During message processing in `generateChatResponse`, MiniSearch counts tokens using `gptTokenizer` and stops adding messages once the budget is reached, marking remaining older messages as "dropped" for summarization.

### Why does MiniSearch limit summaries to 800 tokens?

The **`SUMMARY_TOKEN_LIMIT`** constant (defined at line 41 of [`client/modules/textGeneration.ts`](https://github.com/felladrin/minisearch/blob/main/client/modules/textGeneration.ts)) caps summaries at 800 tokens to ensure the compressed history consumes only 20% of the 4096-token context window. This leaves substantial headroom for the active conversation window, system instructions, and model response generation while still preserving salient historical details.