How Craft Agents Handles Large Tool Responses with Claude Haiku Summarization

Craft Agents automatically detects oversized tool results, persists them to disk, and generates concise summaries using Claude Haiku to prevent context window overflow while preserving full data accessibility.

When working with the craft-ai-agents/craft-agents-oss repository, developers frequently encounter tool outputs—from web APIs to MCP calls—that exceed comfortable context window limits. The framework implements a sophisticated large response handling utility that leverages Craft Agents Claude Haiku summarization to process responses up to ~60KB and beyond. This system ensures that massive payloads never poison the LLM's context while remaining fully retrievable through standard tools.

Detecting and Processing Binary Data

The large response utility begins by inspecting the payload structure to determine if it contains unprocessable binary content.

In packages/shared/src/utils/large-response.ts, lines 1,558–1,567 implement binary detection logic that identifies raw bytes, data-URLs, or large base-64 blobs. When binary data is detected, the utility immediately saves the payload to disk and returns a minimal "binary saved" message to the agent, preventing token pollution from unprocessable data.

For structured JSON containing embedded media, the system walks the object tree (lines 2,992–3,066) to extract each base-64 asset. It writes both the original file and a linked JSON reference (where blobs are replaced with file paths), returning a concise description of the extracted assets rather than the raw data.

Token Estimation and Threshold Management

Before invoking expensive summarization operations, the system calculates whether the response truly requires handling.

The estimateTokensDensityAware function (lines 1,098–1,121) approximates token counts with special handling for base-64-dense payloads. When the payload contains long runs of base-64 characters, it applies a stricter density calculation (referencing constants at lines 1,076–1,093) to avoid underestimating token usage.

The threshold comparison logic uses tokenLimitFor to check against model-aware limits. The default TOKEN_LIMIT is set to 12,000 tokens (line 1,042), scaled dynamically based on the active model's context window via PER_RESULT_CONTEXT_FRACTION (line 1,059). When the estimated token count exceeds this limit, the response is flagged as large and processed accordingly.

Claude Haiku Summarization Pipeline

Once a response is deemed large, Craft Agents orchestrates a multi-step summarization flow using Claude Haiku.

First, the saveLargeResponse function (lines 1,610–1,625) writes the full payload to <session>/long_responses/…txt. The system then checks if the payload size is below MAX_SUMMARIZATION_INPUT (100,000 tokens ≈ 400KB, defined at line 1,045). If within this bound, buildSummarizationPrompt (lines 1,416–1,460) constructs a prompt instructing Claude Haiku to "extract the most relevant information and provide a concise but comprehensive summary."

The summarization callback—typically agent.runMiniCompletion.bind(agent)—is invoked (lines 1,554–1,560), and the returned text becomes the summary. Finally, formatLargeResponseMessage (lines 1,680–1,700) assembles the user-facing output:

[Large response (~12500 tokens) summarized]

Full data saved to: /home/user/.craft/sessions/abc123/long_responses/2026-07-06_...txt
- Use Read/Grep to access specific content
- Use transform_data with inputFiles: ["long_responses/...txt"] for analysis

Summary:
- 42 pull requests opened in the last week
- Most changed files are *.js and *.tsx
- Highest-impact PR: #42 (adds new authentication flow)

If the payload exceeds MAX_SUMMARIZATION_INPUT, summarization is skipped and only a preview of the first 2,000 characters is displayed.

Implementing the Guard in Your Tools

The guardLargeResult function (lines 1,444–1,527) serves as the primary entry point for large response handling. Tool developers wrap their outputs to automatically gain summarization capabilities:

import { guardLargeResult } from '../utils/large-response';

// Example: Processing a large API response
const rawResult = await callExternalApi(); // Potentially > 60KB
const processed = await guardLargeResult(rawResult, {
  sessionPath: sessionDir,
  toolName: 'github',
  input: { repo: 'craft-ai-agents/craft-agents-oss' },
  intent: 'List recent commits',
  summarize: agent.runMiniCompletion.bind(agent), // Claude Haiku binding
  contextWindow: 128_000, // Model's context window (e.g., Claude 3)
});

This guard is invoked automatically by every tool that may produce large results, including the MCP client pool (packages/shared/src/mcp/mcp-pool.ts), API tools (packages/shared/src/api-tools.ts), and Claude-Agent SDK wrappers (packages/shared/src/claude-agent.ts).

Summary

  • Automatic Detection: The guardLargeResult wrapper in packages/shared/src/utils/large-response.ts intercepts tool outputs exceeding 12,000 tokens (adjustable via TOKEN_LIMIT).
  • Binary Handling: Raw bytes and base-64 blobs are extracted and saved to disk rather than being passed to the LLM.
  • Claude Haiku Integration: Payloads under 100,000 tokens are summarized via runMiniCompletion using prompts built by buildSummarizationPrompt.
  • Persistent Storage: Large responses are written to long_responses/ directories with full paths provided for later retrieval via Read or transform_data tools.
  • Fallback Behavior: Responses exceeding 400KB receive truncated previews instead of full summarization to prevent timeout issues.

Frequently Asked Questions

What is the token threshold that triggers large response handling in Craft Agents?

The default threshold is 12,000 tokens, defined by the TOKEN_LIMIT constant at line 1,042 of packages/shared/src/utils/large-response.ts. However, this limit scales dynamically based on the active model's context window using the PER_RESULT_CONTEXT_FRACTION calculation (line 1,059), ensuring that larger context windows adjust the threshold proportionally.

How does Craft Agents handle binary data in tool responses?

When the utility detects binary content at lines 1,558–1,567, it immediately saves the raw payload to disk under the session's long_responses/ directory and returns a short description referencing the saved file. For JSON containing embedded base-64 media, the system extracts each asset (lines 2,992–3,066), writes them to separate files, and returns a JSON reference structure instead of the raw data.

What happens if a response exceeds the Claude Haiku summarization limit?

If the estimated token count exceeds MAX_SUMMARIZATION_INPUT (100,000 tokens or approximately 400KB), the system skips the Claude Haiku summarization step entirely. Instead, formatLargeResponseMessage generates a preview of the first 2,000 characters and provides the file path to the full saved response, allowing users to access specific sections via the Read or Grep tools.

Which tools in Craft Agents use the large response guard?

According to the source code, guardLargeResult is invoked by the MCP client pool (packages/shared/src/mcp/mcp-pool.ts), API tools (packages/shared/src/api-tools.ts), and Claude-Agent SDK wrappers (packages/shared/src/claude-agent.ts). This ensures consistent handling of large outputs across web APIs, MCP servers, and internal SDK tool calls.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →