How MiniSearch Manages Long Conversations: Conversation Memory, Rolling Summaries, and Token Budgets
MiniSearch maintains a persistent, rolling conversation summary that compresses historical messages into an 800-token limit, ensuring the active chat window stays within 75% of the LLM's context size while preserving context across browser sessions.
MiniSearch implements a sophisticated conversation memory and summary system designed to prevent context window overflows without sacrificing historical awareness. By combining a PubSub-based persistence layer with intelligent token budgeting and hybrid summarization strategies, the system enables extended chat sessions that survive page reloads. This architecture continuously folds aging messages into a compressed summary while keeping the live prompt within safe token limits.
The Conversation Summary Store
MiniSearch persists conversation state using the create-pubsub library, storing a lightweight object in the browser's localStorage. This store maintains both the current summary text and its associated conversation ID.
In client/modules/pubSub.ts (lines 49‑62), the system initializes a dedicated PubSub for summary management:
// client/modules/pubSub.ts
const conversationSummaryPubSub = createPubSub({
conversationId: "",
summary: "",
});
export const [
updateConversationSummary,
,
getConversationSummary,
] = conversationSummaryPubSub;
The updateConversationSummary function writes new {conversationId, summary} pairs to storage, while getConversationSummary retrieves the current state. Because this store leverages localStorage, the rolling summary survives page reloads and remains available across multiple LLM calls within the same browser session.
Rolling Summary Algorithm
When chat history grows beyond the available token budget, MiniSearch folds older messages into the summary rather than discarding them. This process involves three distinct phases: detecting overflow, generating the summary, and handling failure scenarios.
Detecting Token Overflow
During generateChatResponse, the system builds processedMessages by iterating through the conversation history in reverse chronological order. The algorithm stops adding messages once the token count would exceed the available token budget, calculated as 75% of the model's context size minus reserved tokens for system prompts.
This logic appears in client/modules/textGeneration.ts (lines 297‑312):
// client/modules/textGeneration.ts – token-budget calculation
const availableTokenBudget = defaultContextSize * 0.75 - reservedTokens;
...
if (currentTokenCount + messageTokens > availableTokenBudget) {
break; // remaining older messages become “dropped”
}
Messages that do not fit within this budget are classified as dropped and queued for summarization.
LLM-Based Summary Generation
When messages are dropped, MiniSearch invokes createLlmSummary (lines 72‑84 of client/modules/textGeneration.ts) to generate an updated summary. The function sends a specialized prompt to the LLM requesting a condensed update under the SUMMARY_TOKEN_LIMIT of 800 tokens:
// client/modules/textGeneration.ts – createLlmSummary implementation
const instructionLines = [
"You are the conversation memory manager.",
`Update the running summary under ${SUMMARY_TOKEN_LIMIT} tokens.`,
// …additional instructions…
];
...
const chat: ChatMessage[] = [{ role: "user", content: prompt }];
...
return (await generateChatWithOpenAi(chat, () => {})).trim();
The resulting summary replaces the previous version in the PubSub store, ensuring the compressed history remains available for subsequent turns.
Fallback Extractive Summarizer
If the LLM call fails or times out, MiniSearch falls back to summarizeDroppedMessages (lines 148‑176 of client/modules/textGeneration.ts). This extractive approach concatenates the dropped messages with any existing summary, then retains as many trailing lines as will fit under the 800-token ceiling:
// client/modules/textGeneration.ts – summarizeDroppedMessages excerpt
for (let i = parts.length - 1; i >= 0; i--) {
const candidate = [parts[i], ...kept].join("\n\n");
const nextTokens = gptTokenizer.encode(candidate).length;
if (nextTokens > tokenLimit) break;
kept.unshift(parts[i]);
}
const summary = kept.join("\n\n");
addLogEntry(`Updated rolling summary (${tokens} tokens)`);
This fallback ensures the system remains robust even when LLM services are unavailable, preserving the most recent conversation context within the token constraint.
Token Budget Management
MiniSearch enforces strict token budgets through two primary constants defined in the source code:
SUMMARY_TOKEN_LIMIT: Set to 800 tokens (lines 41‑43 ofclient/modules/textGeneration.ts), this caps the stored rolling summary size.defaultContextSize: Set to 4096 tokens (lines 16‑17 ofclient/modules/textGenerationUtilities.ts), this defines the base context window for various LLM backends.
The available token budget for the active chat window uses a conservative 75% utilization rule, calculated as:
// client/modules/textGeneration.ts
const availableTokenBudget = defaultContextSize * 0.75 - reservedTokens;
The reservedTokens variable accounts for the system prompt and initial acknowledgment responses, ensuring sufficient headroom for the model's completion tokens. This prevents API rejections due to context length violations while maximizing the usable conversation history.
Practical Implementation Examples
Retrieving the Current Conversation Summary
To access the stored summary programmatically, import the getter function from the PubSub module:
import { getConversationSummary } from "./pubSub";
function printCurrentSummary() {
const { conversationId, summary } = getConversationSummary();
console.log(`Conversation ${conversationId || "<none>"} summary:`);
console.log(summary || "(empty)");
}
This implementation references the store defined in client/modules/pubSub.ts (lines 49‑62).
Manually Triggering Summary Updates
For custom workflows requiring explicit summary generation, combine the LLM and fallback methods:
import { createLlmSummary, summarizeDroppedMessages } from "./textGeneration";
import { updateConversationSummary, getConversationSummary } from "./pubSub";
import { gptTokenizer } from "gpt-tokenizer";
async function foldOldMessages(
dropped: ChatMessage[],
conversationId: string,
) {
const { summary: previous } = getConversationSummary();
// Attempt LLM-based summarization first
const updated = await createLlmSummary(dropped, previous);
// Persist the new summary
updateConversationSummary({ conversationId, summary: updated });
}
The LLM-based path is implemented in client/modules/textGeneration.ts (lines 72‑84), while the extractive fallback resides in client/modules/textGeneration.ts (lines 148‑176).
Calculating Token Budgets for Custom Windows
To implement custom message windowing logic that respects MiniSearch's constraints:
import { defaultContextSize } from "./textGenerationUtilities";
import gptTokenizer from "gpt-tokenizer";
function buildActiveWindow(messages: ChatMessage[], reservedTokens: number) {
const budget = defaultContextSize * 0.75 - reservedTokens;
const kept: ChatMessage[] = [];
let used = 0;
for (const msg of [...messages].reverse()) {
const msgTokens = gptTokenizer.encode(msg.content).length;
if (used + msgTokens > budget) break;
kept.unshift(msg);
used += msgTokens;
}
return kept;
}
This logic mirrors the token-budget calculations in client/modules/textGeneration.ts (lines 297‑312), using the defaultContextSize constant from client/modules/textGenerationUtilities.ts (line 16).
Summary
- Persistent Storage: MiniSearch uses a PubSub-based store (
client/modules/pubSub.ts) backed bylocalStorageto maintain conversation summaries across browser sessions. - Rolling Compression: The system detects token overflow during
generateChatResponseand folds excess messages into an 800-token summary using either LLM generation (createLlmSummary) or extractive fallback (summarizeDroppedMessages). - Conservative Budgeting: Token limits enforce a 75% utilization cap of the 4096-token context window, reserving space for system prompts and model completions.
- Dual Summarization: A hybrid approach prioritizes LLM-generated summaries for coherence but falls back to line-based extraction for reliability.
- Context Preservation: By continuously updating the rolling summary, MiniSearch maintains awareness of distant conversation history without exceeding model context limits.
Frequently Asked Questions
How does MiniSearch preserve conversation memory across page reloads?
MiniSearch persists the conversation summary using a create-pubsub store that serializes data to the browser's localStorage. When a user reloads the page, the getConversationSummary function retrieves the last saved summary and conversation ID from client/modules/pubSub.ts, allowing the chat to resume with full historical context intact.
What happens when the LLM fails to generate a summary?
If the createLlmSummary call fails or returns an error, MiniSearch invokes summarizeDroppedMessages (lines 148‑176 of client/modules/textGeneration.ts) as a fallback. This extractive method keeps the most recent messages that fit within the 800-token SUMMARY_TOKEN_LIMIT, ensuring the conversation can continue even when LLM services are unavailable.
How is the token budget calculated for each chat response?
The system calculates the available token budget as 75% of defaultContextSize (4096 tokens) minus reserved tokens for the system prompt and initial responses. During message processing in generateChatResponse, MiniSearch counts tokens using gptTokenizer and stops adding messages once the budget is reached, marking remaining older messages as "dropped" for summarization.
Why does MiniSearch limit summaries to 800 tokens?
The SUMMARY_TOKEN_LIMIT constant (defined at line 41 of client/modules/textGeneration.ts) caps summaries at 800 tokens to ensure the compressed history consumes only 20% of the 4096-token context window. This leaves substantial headroom for the active conversation window, system instructions, and model response generation while still preserving salient historical details.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →