# How Craft Agents Handles Large Tool Responses with Claude Haiku Summarization

> Discover how Craft Agents overcomes large tool responses with Claude Haiku summarization. Learn how it efficiently manages data, keeps context clear, and ensures full accessibility.

- Repository: [Craft Ai Agents/craft-agents-oss](https://github.com/craft-ai-agents/craft-agents-oss)
- Tags: how-to-guide
- Published: 2026-07-06

---

**Craft Agents automatically detects oversized tool results, persists them to disk, and generates concise summaries using Claude Haiku to prevent context window overflow while preserving full data accessibility.**

When working with the `craft-ai-agents/craft-agents-oss` repository, developers frequently encounter tool outputs—from web APIs to MCP calls—that exceed comfortable context window limits. The framework implements a sophisticated **large response handling utility** that leverages **Craft Agents Claude Haiku summarization** to process responses up to ~60KB and beyond. This system ensures that massive payloads never poison the LLM's context while remaining fully retrievable through standard tools.

## Detecting and Processing Binary Data

The large response utility begins by inspecting the payload structure to determine if it contains unprocessable binary content.

In [`packages/shared/src/utils/large-response.ts`](https://github.com/craft-ai-agents/craft-agents-oss/blob/main/packages/shared/src/utils/large-response.ts), lines 1,558–1,567 implement **binary detection** logic that identifies raw bytes, data-URLs, or large base-64 blobs. When binary data is detected, the utility immediately saves the payload to disk and returns a minimal "binary saved" message to the agent, preventing token pollution from unprocessable data.

For structured JSON containing embedded media, the system walks the object tree (lines 2,992–3,066) to extract each base-64 asset. It writes both the original file and a linked JSON reference (where blobs are replaced with file paths), returning a concise description of the extracted assets rather than the raw data.

## Token Estimation and Threshold Management

Before invoking expensive summarization operations, the system calculates whether the response truly requires handling.

The `estimateTokensDensityAware` function (lines 1,098–1,121) approximates token counts with special handling for base-64-dense payloads. When the payload contains long runs of base-64 characters, it applies a stricter density calculation (referencing constants at lines 1,076–1,093) to avoid underestimating token usage.

The **threshold comparison** logic uses `tokenLimitFor` to check against model-aware limits. The default `TOKEN_LIMIT` is set to **12,000 tokens** (line 1,042), scaled dynamically based on the active model's context window via `PER_RESULT_CONTEXT_FRACTION` (line 1,059). When the estimated token count exceeds this limit, the response is flagged as large and processed accordingly.

## Claude Haiku Summarization Pipeline

Once a response is deemed large, Craft Agents orchestrates a multi-step summarization flow using Claude Haiku.

First, the `saveLargeResponse` function (lines 1,610–1,625) writes the full payload to `<session>/long_responses/…txt`. The system then checks if the payload size is below `MAX_SUMMARIZATION_INPUT` (100,000 tokens ≈ 400KB, defined at line 1,045). If within this bound, `buildSummarizationPrompt` (lines 1,416–1,460) constructs a prompt instructing Claude Haiku to "extract the most relevant information and provide a concise but comprehensive summary."

The summarization callback—typically `agent.runMiniCompletion.bind(agent)`—is invoked (lines 1,554–1,560), and the returned text becomes the summary. Finally, `formatLargeResponseMessage` (lines 1,680–1,700) assembles the user-facing output:

```text
[Large response (~12500 tokens) summarized]

Full data saved to: /home/user/.craft/sessions/abc123/long_responses/2026-07-06_...txt
- Use Read/Grep to access specific content
- Use transform_data with inputFiles: ["long_responses/...txt"] for analysis

Summary:
- 42 pull requests opened in the last week
- Most changed files are *.js and *.tsx
- Highest-impact PR: #42 (adds new authentication flow)

```

If the payload exceeds `MAX_SUMMARIZATION_INPUT`, summarization is skipped and only a preview of the first 2,000 characters is displayed.

## Implementing the Guard in Your Tools

The `guardLargeResult` function (lines 1,444–1,527) serves as the primary entry point for large response handling. Tool developers wrap their outputs to automatically gain summarization capabilities:

```typescript
import { guardLargeResult } from '../utils/large-response';

// Example: Processing a large API response
const rawResult = await callExternalApi(); // Potentially > 60KB
const processed = await guardLargeResult(rawResult, {
  sessionPath: sessionDir,
  toolName: 'github',
  input: { repo: 'craft-ai-agents/craft-agents-oss' },
  intent: 'List recent commits',
  summarize: agent.runMiniCompletion.bind(agent), // Claude Haiku binding
  contextWindow: 128_000, // Model's context window (e.g., Claude 3)
});

```

This guard is invoked automatically by every tool that may produce large results, including the MCP client pool ([`packages/shared/src/mcp/mcp-pool.ts`](https://github.com/craft-ai-agents/craft-agents-oss/blob/main/packages/shared/src/mcp/mcp-pool.ts)), API tools ([`packages/shared/src/api-tools.ts`](https://github.com/craft-ai-agents/craft-agents-oss/blob/main/packages/shared/src/api-tools.ts)), and Claude-Agent SDK wrappers ([`packages/shared/src/claude-agent.ts`](https://github.com/craft-ai-agents/craft-agents-oss/blob/main/packages/shared/src/claude-agent.ts)).

## Summary

- **Automatic Detection**: The `guardLargeResult` wrapper in [`packages/shared/src/utils/large-response.ts`](https://github.com/craft-ai-agents/craft-agents-oss/blob/main/packages/shared/src/utils/large-response.ts) intercepts tool outputs exceeding 12,000 tokens (adjustable via `TOKEN_LIMIT`).
- **Binary Handling**: Raw bytes and base-64 blobs are extracted and saved to disk rather than being passed to the LLM.
- **Claude Haiku Integration**: Payloads under 100,000 tokens are summarized via `runMiniCompletion` using prompts built by `buildSummarizationPrompt`.
- **Persistent Storage**: Large responses are written to `long_responses/` directories with full paths provided for later retrieval via `Read` or `transform_data` tools.
- **Fallback Behavior**: Responses exceeding 400KB receive truncated previews instead of full summarization to prevent timeout issues.

## Frequently Asked Questions

### What is the token threshold that triggers large response handling in Craft Agents?

The default threshold is **12,000 tokens**, defined by the `TOKEN_LIMIT` constant at line 1,042 of [`packages/shared/src/utils/large-response.ts`](https://github.com/craft-ai-agents/craft-agents-oss/blob/main/packages/shared/src/utils/large-response.ts). However, this limit scales dynamically based on the active model's context window using the `PER_RESULT_CONTEXT_FRACTION` calculation (line 1,059), ensuring that larger context windows adjust the threshold proportionally.

### How does Craft Agents handle binary data in tool responses?

When the utility detects binary content at lines 1,558–1,567, it immediately saves the raw payload to disk under the session's `long_responses/` directory and returns a short description referencing the saved file. For JSON containing embedded base-64 media, the system extracts each asset (lines 2,992–3,066), writes them to separate files, and returns a JSON reference structure instead of the raw data.

### What happens if a response exceeds the Claude Haiku summarization limit?

If the estimated token count exceeds `MAX_SUMMARIZATION_INPUT` (100,000 tokens or approximately 400KB), the system skips the Claude Haiku summarization step entirely. Instead, `formatLargeResponseMessage` generates a preview of the first 2,000 characters and provides the file path to the full saved response, allowing users to access specific sections via the `Read` or `Grep` tools.

### Which tools in Craft Agents use the large response guard?

According to the source code, `guardLargeResult` is invoked by the **MCP client pool** ([`packages/shared/src/mcp/mcp-pool.ts`](https://github.com/craft-ai-agents/craft-agents-oss/blob/main/packages/shared/src/mcp/mcp-pool.ts)), **API tools** ([`packages/shared/src/api-tools.ts`](https://github.com/craft-ai-agents/craft-agents-oss/blob/main/packages/shared/src/api-tools.ts)), and **Claude-Agent SDK wrappers** ([`packages/shared/src/claude-agent.ts`](https://github.com/craft-ai-agents/craft-agents-oss/blob/main/packages/shared/src/claude-agent.ts)). This ensures consistent handling of large outputs across web APIs, MCP servers, and internal SDK tool calls.