How Context Pruning and LLM Compaction Alter Provider Input Projections in Apache Maka

Context pruning and LLM compaction in Apache Maka reshape provider input projections through a deterministic pipeline that transforms messages and activeTools while leaving the immutable completedSteps history untouched.

Apache Maka handles large language model (LLM) interactions through a sophisticated request-projection system that optimizes token usage without sacrificing conversation history. Understanding how context pruning and LLM compaction alter provider input projections in Maka ensures developers can maintain reliable replay and resumability guarantees. The framework achieves this through a three-stage pipeline defined in packages/runtime/src/request-projection.ts that creates transient views of conversation state.

The Deterministic Request-Projection Pipeline

Maka builds provider requests using the composeRequestProjection function (lines 45-55 in packages/runtime/src/request-projection.ts). This function assembles a deterministic request-projection pipeline that processes conversation state through ordered transformation stages.

The pipeline accepts an input context containing completedSteps, messages, activeTools, and model configuration. Rather than mutating the original history, each stage returns a new RequestProjection object containing potentially modified messages and activeTools arrays. The underlying completedSteps array remains read-only throughout the entire process.

The Three Stages of Projection Transformation

The pipeline composes three optional stages in a specific execution order:

1. Tool-Availability Filtering

The toolAvailability stage (implemented in packages/runtime/src/tool-availability.ts) removes tools that are not currently usable from the projection. This stage filters entries from activeTools and drops associated tool calls from the messages array.

Crucially, this filtering does not modify the completedSteps history. The provider receives a narrowed tool set without losing the record of which tools were previously invoked.

2. Mid-Turn Capacity Compaction

The midTurnCapacityCompact stage (defined in packages/runtime/src/ai-sdk-compaction.ts) performs LLM compaction by trimming the messages list to respect token-budget limits. This stage rewrites the messages array—potentially removing or shortening earlier assistant messages—to fit within the model's context window.

Despite aggressive message truncation, the stage preserves the history of completed steps (completedSteps) intact. The compaction only affects the transient view sent to the provider, not the persistent execution record.

3. Active-Tool-Result Pruning

The activeToolResultPrune stage (located in packages/runtime/src/active-tool-result-prune.ts) handles oversized tool results that would exceed turn budgets. When tool results exceed token thresholds (typically identified via utilities in packages/runtime/src/context-budget.ts), this stage re-archives large results into a new "tail" of the projection.

The stage updates activeTools and messages with placeholders or summaries, but leaves the original completedSteps unchanged. This prevents massive tool outputs from consuming the entire context window while maintaining accurate historical records.

How History Remains Immutable

The immutability guarantee relies on Maka's functional pipeline architecture. Each RequestProjectionStage receives the context and returns either undefined (no changes) or a partial projection object containing only the modified fields.

The critical flow works as follows:

// Assemble the projection stages
const projection = composeRequestProjection(
  toolAvailability,            // may filter unavailable tools
  midTurnCapacityCompact,      // may shrink the message list
  activeToolResultPrune        // may prune large tool results
);

// Execute the pipeline
const projected = await projection({
  completedSteps,   // immutable history passed through unchanged
  stepNumber,
  model,
  messages,         // full turn history subject to compaction
  activeTools,
});

Because the pipeline returns a new RequestProjection instance, the original completedSteps reference remains untouched. The final capacity decision occurs after the pipeline finishes via buildMidTurnFinalRequestVerdict, ensuring that any pruning or compaction can be undone or adjusted by later processing stages if needed.

Implementation Examples

Composing the Complete Pipeline

To implement all three transformation stages together:

import { composeRequestProjection } from '@/runtime/src/request-projection';
import { toolAvailability } from '@/runtime/src/tool-availability';
import { midTurnCapacityCompact } from '@/runtime/src/ai-sdk-compaction';
import { activeToolResultPrune } from '@/runtime/src/active-tool-result-prune';

const project = composeRequestProjection(
  toolAvailability,
  midTurnCapacityCompact,
  activeToolResultPrune,
);

const projection = await project({
  completedSteps,
  stepNumber: 3,
  model,
  messages: fullTurnMessages,
  activeTools: ['browser', 'search'],
});

// projection.messages may be shortened (compaction) and
// projection.activeTools may have pruned entries.
// completedSteps remains unchanged.

Implementing a Custom Pruning Stage

A pruning stage that replaces oversized tool results:

// src/active-tool-result-prune.ts (simplified)
export const activeToolResultPrune: RequestProjectionStage = (ctx) => {
  const oversized = ctx.messages.filter(
    m => m.toolCalls?.length && estimateTokens(m) > 2048
  );
  
  if (oversized.length === 0) return undefined;

  // Replace the large tool result with a token-budget-friendly placeholder
  const pruned = oversized.map(m => ({
    ...m,
    toolCalls: [{ 
      toolCallId: m.toolCalls[0].toolCallId, 
      content: '[pruned]' 
    }],
  }));
  
  return { messages: pruned };
};

Mid-Turn Token Budget Management

The compaction stage respects model-specific token budgets:

// src/ai-sdk-compaction.ts (simplified)
export const midTurnCapacityCompact: RequestProjectionStage = (ctx) => {
  const budget = ctx.model?.tokenBudget ?? 4096;
  const trimmed = trimMessagesToFit(ctx.messages, budget);
  return { messages: trimmed };
};

Summary

  • Immutable history: The completedSteps array passes through the projection pipeline unchanged, ensuring reliable conversation replay and resumability.
  • Transient projections: Context pruning and LLM compaction modify only the RequestProjection output—specifically messages and activeTools—not the persistent conversation state.
  • Ordered processing: The pipeline executes tool-availability filtering, then mid-turn capacity compaction, then active-tool-result pruning in sequence.
  • Budget compliance: Utilities in context-budget.ts inform token-aware pruning decisions without altering historical records.
  • Undo capability: The buildMidTurnFinalRequestVerdict function evaluates the final projection after all stages complete, allowing reversible optimization decisions.

Frequently Asked Questions

What is the exact execution order of the projection stages in Maka?

The pipeline executes three stages in a fixed sequence defined in composeRequestProjection: first toolAvailability filters unavailable tools, then midTurnCapacityCompact performs LLM compaction on messages, and finally activeToolResultPrune handles oversized tool results. This ordering ensures that tool filtering happens before token budget calculations, and compaction occurs before result pruning.

Does context pruning in Maka affect the ability to replay conversations?

No. Context pruning only modifies the transient RequestProjection sent to the provider. The completedSteps history remains immutable throughout the pipeline, preserving the complete execution record necessary for accurate replay and resumption. The provider sees optimized content while Maka retains the full conversation state.

How does Maka handle token budget violations during the projection pipeline?

Each stage independently checks token budgets using utilities from context-budget.ts. The midTurnCapacityCompact stage trims messages to fit within ctx.model.tokenBudget, while activeToolResultPrune replaces individual oversized tool results with placeholders. If violations persist after all stages, buildMidTurnFinalRequestVerdict makes the final capacity determination without mutating the underlying history.

Which specific files implement the compaction and pruning logic?

The mid-turn capacity compaction resides in packages/runtime/src/ai-sdk-compaction.ts, while active-tool-result pruning lives in packages/runtime/src/active-tool-result-prune.ts. Tool-availability filtering is implemented in packages/runtime/src/tool-availability.ts. The orchestration logic and composeRequestProjection function are defined in packages/runtime/src/request-projection.ts (lines 45-55).

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →