How Large Response Handling and Automatic Summarization Works in Craft Agents
Craft Agents automatically intercepts oversized tool outputs, extracts embedded binary assets, estimates token density using a density-aware heuristic, and optionally summarizes large responses using a mini-completion LLM call to prevent context window overflow while preserving data accessibility.
The craft-ai-agents/craft-agents-oss repository implements a sophisticated pipeline to manage arbitrarily large tool results—such as big PDFs, long JSON payloads, or binary blobs—without exhausting the LLM's context window. The core logic resides in packages/shared/src/utils/large-response.ts and operates through three distinct phases: detection, size validation with optional summarization, and formatted persistence.
The Three-Phase Pipeline
Phase 1: Binary Detection and Asset Extraction
When a tool returns a result, the pipeline first determines if the content requires special handling. The guardLargeResult function (lines 443‑475 in packages/shared/src/utils/large-response.ts) orchestrates this initial screening:
- Binary buffer detection: The
looksLikeBinaryutility identifies raw binary data and immediately saves it as a downloadable file viasaveBinaryResponse, returning only a short description to the LLM. - Structured JSON extraction: If the result is parseable JSON containing base64 blobs,
extractAssetsFromStructuredJson(lines 292‑368) extracts each asset, replaces the blob with a reference object, and stores both the original data and the linked JSON file. - Base64 decoding: Data-URL or raw base64 strings are decoded to binary files using
extractBase64Binary(lines 194‑266 inpackages/shared/src/utils/binary-detection.ts).
Phase 2: Token Estimation and Summarization Triggers
After extraction, the pipeline validates the text size against the model's token limit. The default heuristic estimates 4 characters per token, but for base64-dense payloads, the system uses estimateTokensDensityAware (lines 98‑121) to avoid "token poisoning" where inflated character counts would otherwise bypass limits.
If the estimated token count exceeds TOKEN_LIMIT (approximately 12,000 tokens, defined in lines 31‑38), the pipeline checks for a summarization callback (agent.runMiniCompletion). When provided and the response fits within MAX_SUMMARIZATION_INPUT (~400 KB), the system builds a summarization prompt via buildSummarizationPrompt and sends it to the LLM.
Phase 3: File Persistence and Message Formatting
The final phase persists the full response to disk and formats the user-facing message:
- Save operation:
saveLargeResponse(lines 61‑80) writes the content tosession/long_responses/...with a timestamped filename. - Message construction:
formatLargeResponseMessage(lines 84‑101) combines a "large response summarized" header, the file reference path, and either the LLM-generated summary or a 2,000-character preview if summarization failed. - Return: The formatted string replaces the original bulky output in the conversation context.
Key Implementation Details
Density-Aware Token Estimation
Standard character-count heuristics fail for base64-encoded data, which inflates token counts disproportionately. The estimateTokensDensityAware function analyzes character density to provide accurate token estimates, preventing oversized payloads from silently filling the context window.
The Summarization Guard
The handleLargeResponse function (lines 120‑176) manages the conditional summarization flow. It only invokes the summarizer when:
- A valid
summarizecallback is provided - The payload size is under
MAX_SUMMARIZATION_INPUT - The token count exceeds
TOKEN_LIMIT
If conditions aren't met, the pipeline gracefully degrades to showing a preview snippet while still saving the full file for later retrieval.
Practical Implementation Examples
Using the Guard in a Tool Implementation
Wrap tool outputs with guardLargeResult to enable automatic handling:
import { guardLargeResult } from '@/utils/large-response';
async function runTool(params: unknown) {
const rawResult = await callExternalApi(params); // string | Buffer
const summarized = await guardLargeResult(rawResult, {
sessionPath: this.sessionPath,
toolName: 'myTool',
input: params,
intent: this.currentIntent,
summarize: this.agent.runMiniCompletion.bind(this.agent), // optional
contextWindow: this.model.contextWindow,
});
// Returns null if small enough, otherwise returns formatted summary string
return summarized ?? rawResult;
}
Manually Invoking the Pipeline
For custom workflows, use handleLargeResponse directly:
import { handleLargeResponse } from '@/utils/large-response';
async function processLargeText(text: string, sessionPath: string) {
const result = await handleLargeResponse({
text,
sessionPath,
context: { toolName: 'search', path: '/v1/docs' },
summarize: async (prompt) => `Summary: ${prompt.slice(0, 100)}...`,
});
console.log(result?.message); // Contains file reference + summary
}
Accessing Saved Large Files Later
Assistant tools can retrieve the full content using standard file operations:
import { readFileSync } from 'fs';
import { join } from 'path';
const filePath = join(
sessionPath,
'long_responses',
'2026-07-03_myTool_example.txt'
);
const fullContent = readFileSync(filePath, 'utf-8');
// Process content with Read or Grep tools
Summary
- Craft Agents prevents context overflow by intercepting large tool outputs before they reach the LLM.
- Binary detection in
packages/shared/src/utils/binary-detection.tshandles raw buffers and base64-encoded assets by extracting them to disk. - Token estimation uses a density-aware algorithm to avoid underestimating base64-heavy payloads.
- Automatic summarization triggers when outputs exceed ~12,000 tokens, using a mini-completion call to condense the content while preserving a file reference.
- Graceful degradation ensures users always receive a preview or summary even when full summarization fails.
Frequently Asked Questions
What happens if a tool returns binary data instead of text?
The looksLikeBinary function detects raw binary buffers and routes them through saveBinaryResponse. The LLM receives only a short description with a file path, while the actual binary is stored in the session's long_responses directory for later access.
How does the system prevent base64 data from consuming the entire context window?
The estimateTokensDensityAware function (lines 98‑121) analyzes character density patterns. Unlike the standard 4-characters-per-token heuristic, this density-aware approach recognizes that base64 strings generate significantly more tokens per character, ensuring accurate size calculations before the data hits the TOKEN_LIMIT threshold.
Can I disable automatic summarization for specific tools?
Yes. The guardLargeResult function accepts an optional summarize parameter. If omitted or set to null, the pipeline skips the LLM summarization step and returns a preview snippet (first 2,000 characters) alongside the file reference, effectively disabling the mini-completion call while still preventing context overflow.
Where are the large response files stored?
The saveLargeResponse function (lines 61‑80) persists files to session/long_responses/ with timestamped filenames (e.g., 2026-07-03_toolname_inputhash.txt). These paths are referenced in the formatted message returned to the LLM, allowing tools like Read or Grep to access the full content on demand.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →