How the MCP Server Exposes `caveman_retrieve` for Decompressing Payloads Inside an Agent
The MCP server exposes caveman_retrieve as a specialized interception tool that resolves decompression requests locally by querying the Caveman Compression Runtime (CCR) store, returning original payload bytes directly to the agent without involving the upstream model.
The caveman_retrieve tool serves as the critical bridge between compressed model responses and agent execution within the JuliusBrussee/caveman repository. When the proxy's compression engine shrinks payloads using the S4 algorithm, it inserts retrievable markers that agents can later resolve through this MCP-exposed function to access full, uncompressed context. Understanding how the server registers, guards, and resolves this tool is essential for implementing custom agents that handle compressed data streams.
Tool Registration and Naming Conventions
The MCP server registers the retrieval tool using multiple canonical names to ensure compatibility across different agent implementations. In proxy/providers/adapter.go, the system defines three distinct identifiers for the same recovery mechanism:
// proxy/providers/adapter.go
const (
recoveryToolName = "caveman_retrieve"
mcpRecoveryToolName = "mcp__caveman__caveman_retrieve"
kiloRecoveryToolName = "caveman_caveman_retrieve"
)
Source: adapter.go lines 364–366
These constants allow the gateway to recognize tool invocation requests regardless of whether the agent uses the short form, fully-qualified MCP namespace, or legacy kilo prefix. The primary identifier caveman_retrieve remains the standard interface exposed to agent code, while the prefixed variants ensure backward compatibility with different MCP client implementations.
Server-Side Resolution Architecture
Unlike standard MCP tools that proxy requests to the underlying model, caveman_retrieve undergoes server-side resolution within the gateway layer. When processing responses in proxy/internal/gateway/server.go, the system examines outgoing tool calls for the recovery identifier and intercepts them before they reach the model provider:
// proxy/internal/gateway/server.go
// resolve caveman_retrieve calls server-side after S4 compression.
// recoveryViaMCP records that the wrapped agent fulfills caveman_retrieve itself.
// RecoveryViaMCP routes S4 recovery through the agent's own caveman_retrieve MCP.
Source: server.go lines 134, 304, 357, 416
This interception pattern ensures that decompression operations remain transparent to the model while remaining accessible to the agent. The server maintains the mapping between compression markers and original payload locations, eliminating the need to transmit large uncompressed data through the model context window.
The Decompression Workflow
When an agent encounters compressed content marked with <<ccr:handle>> tokens, the retrieval process follows a strict server-side sequence:
- Marker Detection: The agent emits a
caveman_retrievetool call containing the handle identifier embedded in the compressed stream. - Store Lookup: The gateway extracts the handle and queries the CCR store for the original bytes associated with that compression marker.
- Direct Response: The server returns the uncompressed payload as the tool result, bypassing the model entirely to prevent token waste.
The tool schema itself carries no arguments because the proxy maintains session state mapping requests to their respective compression contexts. As defined in proxy/internal/gateway/retrieve_tool.go, the tool declaration uses an empty parameter object:
// proxy/internal/gateway/retrieve_tool.go
const retrieveToolName = "caveman_retrieve"
Source: retrieve_tool.go line 17
Agent Integration and Duplicate Prevention
The Pi extension (packages/pi-extension/src/index.ts) automatically registers caveman_retrieve as a model-visible function to ensure agents can always request decompression when needed:
// packages/pi-extension/src/index.ts
const RECOVERY_TOOL = "caveman_retrieve";
Source: index.ts line 21
To prevent conflicts, the gateway implements a strict injection guard that detects when an agent already provides its own caveman_retrieve implementation. The test suite in proxy/internal/gateway/retrieve_tool_test.go validates this behavior:
// proxy/internal/gateway/retrieve_tool_test.go
// MCP recovery must NOT inject a tool (the agent owns caveman_retrieve)
Source: retrieve_tool_test.go lines 65–81
This guard ensures that custom agent implementations retain control over their retrieval logic while preventing duplicate tool registrations that could confuse the MCP client.
Code Implementation Examples
The following JSON structure represents how the tool appears in an agent's available function set:
{
"type": "function",
"function": {
"name": "caveman_retrieve",
"description": "Recover original uncompressed content from CCR store",
"parameters": {
"type": "object",
"properties": {
"handle": {
"type": "string",
"description": "The CCR handle marker from compressed content"
}
}
}
}
}
When the gateway processes a retrieval request, the internal handling logic follows this pattern:
func handleRetrieve(ctx context.Context, req *RetrieveRequest) (*ToolResult, error) {
// Extract the CCR handle from the tool call parameters
handle := req.Parameters["handle"].(string)
// Load original bytes from the compression runtime store
original, err := ccrStore.Load(handle)
if err != nil {
return nil, fmt.Errorf("failed to retrieve payload: %w", err)
}
// Return uncompressed data as the tool output
return &ToolResult{
Name: "caveman_retrieve",
Output: string(original),
}, nil
}
Summary
- The MCP server exposes
caveman_retrievethrough multiple canonical names defined inproxy/providers/adapter.goto support various agent implementations. - Server-side resolution in
proxy/internal/gateway/server.gointercepts retrieval calls to prevent model proxying and reduce token usage. - The tool queries the CCR store using compression markers (
<<ccr:handle>>) to restore original payloads without transmitting uncompressed data through the model context. - The Pi extension automatically registers the tool for agent visibility, while the gateway prevents duplicate injections to avoid MCP client conflicts.
- An empty parameter schema indicates that the proxy maintains compression context internally, requiring only the handle identifier from the agent.
Frequently Asked Questions
How does caveman_retrieve differ from standard MCP tools?
Standard MCP tools forward requests to the underlying model provider for processing, whereas caveman_retrieve is resolved entirely within the gateway. The server intercepts these calls according to the logic in proxy/internal/gateway/server.go to query the CCR store directly, ensuring that large uncompressed payloads never consume model context window tokens.
What prevents duplicate caveman_retrieve registrations when agents define their own?
The gateway implements an injection guard that checks for existing tool definitions before registering the built-in retrieval handler. As validated in proxy/internal/gateway/retrieve_tool_test.go (lines 65–81), the system skips tool injection when it detects that the agent already exposes caveman_retrieve, allowing custom implementations to override the default behavior.
Why does the caveman_retrieve tool use an empty parameter schema?
The tool schema omits required parameters because the proxy maintains session state mapping compression markers to their original payloads. When an agent invokes the tool, the gateway correlates the request with the current conversation context to identify which compressed block to retrieve, eliminating the need for the agent to manually track and transmit complex identifiers beyond the embedded handle marker.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →