How Maka Handles Context Pruning Without Deleting History

Maka employs a two-phase prune-then-archive strategy that removes oversized tool results from the active LLM prompt while persisting them to an immutable archive, ensuring the original execution history remains retrievable via placeholder references.

Maka, the open-source agent runtime under the Apache Software Foundation, implements a sophisticated context management system that keeps LLM prompts within token limits. Unlike naive truncation methods that discard valuable execution data, Maka context pruning without deleting history ensures that every tool result remains accessible even after removal from the active conversation window.

The Prune-Then-Archive Architecture

Maka’s runtime separates the concept of active context (what the LLM sees) from archival history (what the system remembers). According to the docs/blogs/log-is-the-runtime.md source, the system never deletes data; it only relocates it. This approach maintains bounded archive hydration, allowing the model to request previously pruned data on demand without forward-scanning the entire log.

The strategy operates in two distinct phases: active-turn pruning, which manages the immediate token budget during execution, and stale-turn pruning, which cleans up redundant data after ledger checkpoints.

Active-Turn Pruning and Context Budgets

Before each reasoning step, Maka evaluates whether tool results will exceed the configured context budget. The default threshold is 2 048 tokens, though this is configurable via runtime options. Any result that would push the prompt beyond this limit is flagged for immediate pruning.

As implemented in docs/eval/terminal-bench-2.1-maka-vs-kimi-code-v11.md, the runtime logs "active prune diagnostics" when this occurs, indicating which specific payloads were removed from the in-flight message list. This mechanism prevents prompt overflow while preserving conversation continuity through compact placeholder references.

Configuring the Context Budget

To enable active-turn pruning, instantiate the agent with a defined contextBudget and archive storage path:

import { MakaAgent } from '@maka/agent';

const agent = new MakaAgent({
  llm: 'openai-chat',
  contextBudget: 2048,          // activate active‑turn pruning
  archivePath: './archive',    // where pruned results are stored
  enableStalePrune: true,      // prune stale tool results after checkpoints
});

// The agent will automatically prune oversized tool results
// and replace them with a placeholder reference.
await agent.runTask(myTask);

Archive-First, Placeholder-Second Safety

The critical safety mechanism ensuring zero data loss is the archive-first, placeholder-second sequence described in docs/blogs/log-is-the-runtime.md. When Maka identifies a result for pruning, it first writes the full payload to an immutable archive store (a durable append-only log). Only after receiving confirmation of successful persistence does the system replace the original data in the active context with a small placeholder—typically a hash or unique ID.

If the archival step fails for any reason, the full result remains in the active context, preventing the LLM from receiving unresolvable references. This guarantees that placeholders always point to retrievable data.

Bounded Archive Hydration

When the model requires data that was previously pruned, Maka performs hydration by fetching the archived record using the placeholder reference. This operation requires only a single lookup against the immutable log, ensuring constant-time retrieval regardless of total historical data volume. The original content is reconstructed exactly as it existed before pruning.

The hydration mechanics are detailed in docs/architecture/llm-compaction-events-log-projection-draft.md, which specifies that the archive supports bounded lookups and maintains cryptographic consistency with the active ledger.

Hydrating Pruned Results

While hydration typically occurs internally within the agent loop, the archive API exposes direct access for manual recovery:

import { Archive } from '@maka/archive';

async function getPrunedResult(ref: string) {
  // `ref` is the placeholder (e.g., a hash) left in the active context
  const payload = await Archive.read(ref);   // fetches the archived data
  return payload;
}

Stale-Turn Pruning After Checkpoints

After the runtime creates a checkpoint—a compacted snapshot of the runtime ledger—Maka initiates stale-turn pruning. This process removes raw, oversized tool results that are no longer referenced by the current "live" window of execution. Importantly, this does not constitute deletion because the checkpoint already contains a summarized, semantically equivalent representation of those events.

As documented in docs/architecture/llm-compaction-events-log-projection-draft.md, this pruning targets only redundant payloads that would otherwise duplicate information captured in the compacted log. The underlying execution history remains intact within the checkpoint and archival layers, ensuring context pruning without deleting history even during aggressive cleanup phases.

Manual Checkpoint Creation

Force a compacted snapshot to enable stale-prune operations and trigger ledger compaction:

// Manual checkpoint creation – forces a compacted snapshot and enables stale‑prune
await agent.checkpoint();   // creates a ledger checkpoint
// After this call, any tool results older than the checkpoint can be pruned safely.

The checkpointing infrastructure utilized by this process is implemented in the native/runtime-host-windows-task-launcher/ directory, which manages low-level disk I/O for the durable archive and ledger snapshotting.

Summary

  • Maka employs a two-phase prune-then-archive strategy to maintain token budgets while preserving complete execution history.
  • Active-turn pruning monitors the contextBudget (default 2 048 tokens) and removes oversized results from the prompt before each reasoning step.
  • The archive-first, placeholder-second protocol ensures data is durably persisted before being replaced with lightweight references, preventing broken links.
  • Bounded archive hydration enables constant-time retrieval of pruned data via single lookups against the immutable log.
  • Stale-turn pruning cleans up redundant raw payloads after checkpoints while maintaining summarized state in the compacted ledger.
  • All pruning logic is coordinated through components in docs/blogs/log-is-the-runtime.md and docs/architecture/llm-compaction-events-log-projection-draft.md.

Frequently Asked Questions

Does pruning remove data from Maka's permanent history?

No. Pruning only relocates data from the active context (the LLM's prompt window) to the immutable archive. As specified in docs/blogs/log-is-the-runtime.md, the system guarantees durability by confirming archive persistence before replacing content with placeholders. The full execution ledger remains available for hydration or audit.

What happens if the archive fails during active-turn pruning?

If the archival write fails, Maka retains the full result in the active context and does not insert a placeholder. According to the runtime safety logic, the system only substitutes references after receiving confirmation of successful persistence, ensuring the model never encounters broken links to missing data.

How does Maka reconstruct pruned tool results during execution?

When the model requires previously pruned data, Maka performs bounded archive hydration by resolving the placeholder reference—typically a content hash—through the Archive.read() interface. This fetches the exact original payload from the immutable store in a single operation, as detailed in docs/architecture/llm-compaction-events-log-projection-draft.md.

Where does Maka store archived tool results?

Archived data is written to the filesystem path specified by the archivePath configuration property. The underlying storage implementation resides in native/runtime-host-windows-task-launcher/, which handles the durable log format and checkpointing used by the pruning system to ensure data survives process restarts.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →