How to Optimize Token Usage and Manage Context Windows in AnythingLLM

AnythingLLM uses a three-layer architecture—TokenManager, prompt-window limits, and a "cannonball" compression algorithm—to automatically fit any conversation into the model's context window while preserving critical information.

Managing token usage and context windows is essential when building applications with the Mintplex-Labs/anything-llm repository. The platform automatically compresses prompts when they approach the model's maximum token limit, ensuring reliable responses without manual truncation. This guide explains the internal mechanics of this system and how to leverage it for custom implementations.

The Three Pillars of Token Management

AnythingLLM structures its token optimization around three core components that work together to enforce context window limits.

TokenManager: The Counting Engine

The TokenManager class in server/utils/helpers/tiktoken.js provides a singleton wrapper around the js-tiktoken library. It exposes methods to count, encode, and estimate tokens for any string or message array.

const { TokenManager } = require('./tiktoken');
const tm = new TokenManager('gpt-3.5-turbo');

const tokenCount = tm.countFromString(userInput);
const tokens = tm.tokensFromString(userInput);

The implementation (lines 70-76) reuses the same encoder instance per model to avoid initialization overhead. Key methods include countFromString() for raw numbers and statsFrom() for array analysis.

Prompt Window Limits: The Budget Enforcer

Every LLM provider (OpenAI, Anthropic, Ollama, etc.) declares a static promptWindowLimit() method that returns the model-specific token budget. For example, in server/utils/AiProviders/openAi/index.js (lines 56-62):

static promptWindowLimit(modelName) {
  return MODEL_MAP.get('openai', modelName) ?? 4_096;
}
promptWindowLimit() {
  return MODEL_MAP.get('openai', this.model) ?? 4_096;
}

This value drives the token buffer calculation and determines when compression triggers. The compressor works with promptWindowLimit - tokenBuffer to guarantee headroom for the model's reply.

MessageArrayCompressor: The Optimization Engine

The messageArrayCompressor function in server/utils/helpers/chat/index.js (lines 49-90) orchestrates the compression flow. It measures total token counts against the budget and aggressively trims the system prompt, user prompt, and chat history until the payload fits.

The Cannonball Compression Algorithm

When a message exceeds its allocated budget, AnythingLLM applies a middle-out truncation called "cannonball." This technique removes tokens from the center of the text, preserving the beginning and end of the message, then inserts a truncation marker.

function cannonball({ input, targetTokenSize, tiktokenInstance }) {
  const tokenManager = tiktokenInstance || new TokenManager();
  const delta = tokenManager.countFromString(input) - targetTokenSize;
  const tokenChunks = tokenManager.tokensFromString(input);
  const middleIdx = Math.floor(tokenChunks.length / 2);
  const leftChunks  = tokenChunks.slice(0, middleIdx - Math.round(delta / 2));
  const rightChunks = tokenChunks.slice(middleIdx + Math.round(delta / 2));
  return tokenManager.bytesFromTokens(leftChunks) +
         '\n\n--prompt truncated for brevity--\n\n' +
         tokenManager.bytesFromTokens(rightChunks);
}

Located at lines 12-45 of chat/index.js, this function splits the input into token chunks, calculates how many to remove from the middle, and reconstructs the string with a delimiter indicating truncation.

The Compression Pipeline Step-by-Step

The messageArrayCompressor processes messages through a specific priority hierarchy to minimize information loss.

1. Early Exit Validation

The system first checks if statsFrom(messages) + tokenBuffer < promptWindowLimit. If true, it returns the original payload unchanged, avoiding unnecessary processing.

2. User Prompt Compression

If the user prompt alone exceeds its portion of the limit, it undergoes cannonball compression targeting 80% of the prompt window limit (0.8 × promptWindowLimit). This ensures the primary query always fits, even if truncated.

3. System Prompt Partitioning

The system prompt splits into two sections: the "Context:" section and the remaining instructions. The compressor allocates 25% of the system budget to context and 75% to instructions, trimming each independently if necessary.

4. History Trimming

The compressor iterates through chat history from newest to oldest, keeping recent pairs until the history budget is exhausted. If the budget is still exceeded after this, it applies cannonball compression to the last three history pairs individually, capping each at approximately half the history budget.

5. The 600-Token Buffer

A static 600-token buffer is reserved for the model's response generation. This buffer is subtracted from the prompt window limit before any compression calculations, ensuring the LLM always has sufficient tokens to generate an answer, even for models with small context windows (like 4,096 tokens). Developers can modify this value in server/utils/helpers/chat/index.js to trade prompt space for longer outputs.

Managing RAG Source Windows

For Retrieval-Augmented Generation (RAG) scenarios, the fillSourceWindow function (starting at line 82 in chat/index.js) caps source documents at a default of four entries. If fewer sources are available, it backfills from recent chat history to prevent empty citation lists. This step executes after message compression, ensuring the final payload—including citations—remains within the context window.

Practical Implementation Examples

Manual Token Budget Verification

For custom requests outside the standard chat flow, manually validate token counts before sending:

const { TokenManager } = require('../server/utils/helpers/tiktoken');
const tm = new TokenManager('gpt-4');

const userPrompt = "Explain quantum entanglement in simple terms...";
const systemPrompt = "You are a helpful assistant.";
const tokenBudget = 8192;               // GPT-4 context window
const tokenBuffer = 600;

if (tm.countFromString(userPrompt) > tokenBudget * 0.7) {
  // Too big – trim with cannonball
  const trimmed = cannonball({
    input: userPrompt,
    targetTokenSize: tokenBudget * 0.7,
    tiktokenInstance: tm,
  });
  // use `trimmed` as the user message
}

Reference the TokenManager implementation at lines 70-76 of tiktoken.js and the cannonball function at lines 12-45 of chat/index.js.

Using the Built-in Compressor

Leverage the provider's compressMessages method for automatic optimization:

const LLMConnector = require('anything-llm/server/utils/AiProviders/openAi');
const rawHistory = await getChatHistory(); // [{messages…}]
const promptArgs = {
  systemPrompt: "You are a legal advisor.",
  userPrompt: longUserInput,
  contextTexts: docs,
};

const compressed = await LLMConnector.compressMessages(promptArgs, rawHistory);
// `compressed` is now an array ready for `openai.createChatCompletion`
await LLMConnector.getChatCompletion(compressed);

The compressMessages method (lines 92-96 in the OpenAI provider) delegates to messageArrayCompressor (lines 49-90 in chat/index.js).

Adjusting the Token Buffer

To accommodate longer model responses, modify the buffer constant:

// In server/utils/helpers/chat/index.js change:
const tokenBuffer = 600;   // ← raise to 1000 for larger replies

This adjustment reserves 1,000 tokens for output generation, reducing the available prompt space by 400 tokens.

Summary

  • TokenManager in server/utils/helpers/tiktoken.js provides accurate token counting via js-tiktoken with methods like countFromString() and tokensFromString().
  • Prompt window limits are enforced by each provider's promptWindowLimit() method, establishing the token budget for compression calculations.
  • MessageArrayCompressor applies a hierarchy of compression: early exit for small payloads, user prompt truncation to 80% of limit, system prompt partitioning (25%/75%), and newest-first history trimming.
  • The cannonball algorithm implements middle-out truncation with explicit markers to reduce token count while preserving message boundaries.
  • A 600-token buffer is reserved for model responses by default, configurable by editing the source.
  • RAG source windows are managed via fillSourceWindow, which caps documents at four and backfills from history if needed.

Frequently Asked Questions

How does AnythingLLM prevent "context window exceeded" errors?

AnythingLLM prevents context window errors through proactive message array compression. The system calculates the total token count of system prompts, user queries, chat history, and RAG sources, then applies the cannonball truncation algorithm to any component exceeding its allocated budget. This ensures the final payload always satisfies total_tokens + buffer < promptWindowLimit before reaching the LLM API.

Can I adjust how aggressively AnythingLLM compresses prompts?

Yes, you can adjust compression behavior by modifying the token buffer or implementing custom compression logic. The default 600-token buffer in server/utils/helpers/chat/index.js can be increased to prioritize longer model responses, or decreased to allow more input context. For granular control, fork the messageArrayCompressor function to adjust the percentage allocations for system prompts (currently 25%/75%) or history pairs.

What happens to my chat history when the context window fills up?

When the context window approaches capacity, AnythingLLM applies newest-first preservation to chat history. The compressor iterates from the most recent messages backward, keeping pairs until the history budget is exhausted. If the history budget is still exceeded, the system applies middle-out truncation to the last three pairs individually. This prioritizes recent conversation context over older messages while maintaining semantic continuity through the cannonball truncation markers.

Does the token counting work for non-English languages?

Yes, because TokenManager relies on js-tiktoken, which implements the same Byte Pair Encoding (BPE) tokenization used by OpenAI models. This tokenizer accurately counts tokens for Unicode text across all languages, including CJK (Chinese, Japanese, Korean) characters and special symbols. The countFromString() method returns the exact token count the LLM provider will see, regardless of character set.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →