How Apache Maka Implements Context Pruning and LLM Compaction

Apache Maka uses a layered context-budget system that combines active tool-result pruning, stale payload archiving, LLM-driven semantic rewriting, and periodic history compaction to keep token counts within model limits while preserving semantic meaning.

Long-running conversational AI systems face a critical challenge: preventing conversation history from exceeding LLM token limits while retaining enough context for coherent responses. The apache/maka repository solves this through a sophisticated context pruning and LLM compaction pipeline orchestrated by the ContextBudgetPolicy interface in packages/runtime/src/context-budget.ts. This system processes every turn through four distinct stages that progressively reduce token footprint without losing essential information.

The Context Budget Architecture

At the heart of Maka’s compaction strategy lies the ContextBudgetPolicy interface defined in packages/runtime/src/context-budget.ts. This configuration object specifies thresholds, model selections, and enablement flags for each pruning stage.

The policy structure includes:

  • maxHistoryEstimatedTokens: The overall token ceiling for the entire conversation
  • maxHistoryTurns: Maximum number of turns to retain before forced compaction
  • minRecentTurns: Protected window of recent turns exempt from pruning
  • activeToolResultPrune: Configuration for immediate tool-result replacement
  • staleToolResultPrune: Settings for archiving old tool outputs
  • semanticCompact: LLM-driven rewriting parameters
  • historyCompact: Periodic summarization settings

During each turn processing cycle, the runtime executes these policies in a specific sequence: active pruning first, followed by stale pruning, semantic compaction, and finally history compaction.

Stage 1: Active Tool-Result Pruning

The first defense against token bloat occurs in packages/runtime/src/active-tool-result-prune.ts through the rewriteActiveToolResultsInMessages function. This mechanism intercepts tool results during the current turn before they reach the LLM.

When a tool response arrives, the system estimates its token size using estimateTokens. If the payload exceeds maxCurrentResultEstimatedTokens (defaulting to 2048 tokens) and the current SDK step number meets minStepNumber, the original payload is serialized and archived via archiveToolResult. The runtime replaces it with a lightweight placeholder object of kind maka.active_archived_tool_result containing a SHA-256 hash, the original token estimate, and a supersession record explaining why the result was pruned.

This immediate replacement prevents a single oversized tool call from consuming the entire context window while preserving the ability to retrieve the full payload later through the archive hash.

Stage 2: Stale Tool-Result Pruning

Before the runtime compacts entire turns, it processes historical tool results in packages/runtime/src/tool-result-archive.ts via pruneStaleToolResultsBeforeCompact. This function targets tool outputs from previous turns that remain in the active context.

The function applies two critical thresholds: maxResultEstimatedTokens (default 2048) and minRecentTurnsFull (defining a protected recent window). For every event outside the protected window, the system checks if tool results exceed the size limit. Oversized results are archived and replaced with ARCHIVED_TOOL_RESULT_PLACEHOLDER_KIND placeholders.

The function returns diagnostics including prunedToolResults, archiveWriteFailures, and updated token estimates, allowing the runtime to track exactly what was removed from the conversation history.

Stage 3: LLM-Driven Semantic Compaction

After mechanical pruning removes oversized payloads, packages/runtime/src/semantic-compact.ts implements intelligent compression through compactSemanticMessages. This stage invokes a secondary LLM to rewrite verbose assistant messages into concise summaries.

The policy specifies which model to use (e.g., gpt-4o-mini) and provides a compactPrompt instructing the LLM how to condense the text—for example, "Summarise the previous assistant messages in fewer than 150 tokens, preserving all facts." The LLM’s response replaces the original message segment, reducing token count while maintaining semantic equivalence.

This step runs after active pruning to ensure the compaction LLM receives a manageable context size, preventing it from exceeding its own token limits during the summarization task.

Stage 4: History Compaction

For long-running sessions, packages/runtime/src/history-compact.ts provides buildHistoryCompactBlockFromSummary, which collapses older runtime events into a single history-compact block. This block stores an LLM-generated summary of prior conversation segments along with source references for replay.

The system records special historyCompactCheckpoint events that mark compaction boundaries, enabling deterministic replay and diagnostic gating. These blocks are replay-friendly and can be re-expanded (hydrated) on demand via the archive-retrieval subsystem, ensuring that while the active context remains small, the full conversation history remains reconstructible for debugging or auditing.

Configuring the Compaction Pipeline

Developers enable context pruning and LLM compaction by assembling a ContextBudgetPolicy and passing it during session creation:

import { ContextBudgetPolicy } from '@maka/runtime';

const budgetPolicy: ContextBudgetPolicy = {
  maxHistoryEstimatedTokens: 16000,
  maxHistoryTurns: 200,
  minRecentTurns: 5,
  
  activeToolResultPrune: {
    enabled: true,
    maxCurrentResultEstimatedTokens: 2000,
    minSupersededResultEstimatedTokens: 256,
    minStepNumber: 1,
  },
  
  staleToolResultPrune: {
    enabled: true,
    maxResultEstimatedTokens: 2048,
    minRecentTurnsFull: 2,
    archiveRefs: [],
  },
  
  semanticCompact: {
    enabled: true,
    model: 'gpt-4o-mini',
    compactPrompt: `Summarise the previous assistant messages in fewer than 150 tokens, preserving all facts.`,
  },
  
  historyCompact: {
    enabled: true,
    maxTurnsPerBlock: 50,
    summaryModel: 'gpt-4o',
  },
};

const session = await maka.createSession({
  budgetPolicy,
});

This configuration activates all four stages: immediate replacement of large tool results, archival of stale outputs from previous turns, LLM-based message rewriting, and periodic folding of ancient history into compact blocks.

Summary

Apache Maka’s context management system employs a surgical approach to token efficiency:

  • Active pruning in active-tool-result-prune.ts replaces oversized current-turn payloads with hash-based placeholders
  • Stale pruning via tool-result-archive.ts archives historical tool results exceeding size thresholds before whole-turn compaction
  • Semantic compaction uses secondary LLM calls to rewrite verbose messages while preserving facts
  • History compaction collapses aged events into replay-friendly summary blocks tracked by checkpoint events
  • All stages are orchestrated through the ContextBudgetPolicy interface in context-budget.ts

Frequently Asked Questions

What triggers active tool-result pruning in Maka?

Active pruning triggers when a tool result in the current turn exceeds maxCurrentResultEstimatedTokens (default 2048 tokens) and the SDK step number is at least minStepNumber. The rewriteActiveToolResultsInMessages function in packages/runtime/src/active-tool-result-prune.ts performs this check, archiving the original payload and substituting a lightweight placeholder containing a SHA-256 hash for later retrieval.

How does semantic compaction preserve conversation meaning?

Semantic compaction relies on explicit prompting strategies defined in the SemanticCompactPolicy. When enabled in packages/runtime/src/semantic-compact.ts, the runtime sends specified message ranges to a configured LLM with instructions to condense content while preserving key facts. The compacted text replaces the original only after successful LLM response validation, ensuring semantic continuity through controlled summarization rather than truncation.

Can I disable specific pruning stages while keeping others?

Yes, the ContextBudgetPolicy interface allows independent toggling of each stage through boolean enabled flags. You can enable activeToolResultPrune while disabling semanticCompact, or configure staleToolResultPrune thresholds without activating historyCompact. This modular design lets developers tune the compaction pipeline to their specific latency, cost, and context-preservation requirements.

Where are archived tool results stored during compaction?

Archived tool results are persisted through the archiveToolResult utility functions found in packages/runtime/src/tool-result-archive.ts and packages/runtime/src/active-tool-result-prune.ts. The system generates SHA-256 hashes as reference keys and stores the original payloads in the archive-retrieval subsystem. These archived values can be rehydrated on demand to reconstruct original conversation states for replay or debugging, even after placeholders replace them in the active context.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →