# How the MCP Server Exposes `caveman_retrieve` for Decompressing Payloads Inside an Agent

> Discover how the MCP server exposes caveman_retrieve to decompress agent payloads locally via the CCR store. Get original bytes without upstream model involvement.

- Repository: [Julius Brussee/caveman](https://github.com/JuliusBrussee/caveman)
- Tags: how-to-guide
- Published: 2026-09-04

---

**The MCP server exposes `caveman_retrieve` as a specialized interception tool that resolves decompression requests locally by querying the Caveman Compression Runtime (CCR) store, returning original payload bytes directly to the agent without involving the upstream model.**

The `caveman_retrieve` tool serves as the critical bridge between compressed model responses and agent execution within the JuliusBrussee/caveman repository. When the proxy's compression engine shrinks payloads using the S4 algorithm, it inserts retrievable markers that agents can later resolve through this MCP-exposed function to access full, uncompressed context. Understanding how the server registers, guards, and resolves this tool is essential for implementing custom agents that handle compressed data streams.

## Tool Registration and Naming Conventions

The MCP server registers the retrieval tool using multiple canonical names to ensure compatibility across different agent implementations. In [`proxy/providers/adapter.go`](https://github.com/JuliusBrussee/caveman/blob/main/proxy/providers/adapter.go), the system defines three distinct identifiers for the same recovery mechanism:

```go
// proxy/providers/adapter.go
const (
    recoveryToolName     = "caveman_retrieve"
    mcpRecoveryToolName  = "mcp__caveman__caveman_retrieve"
    kiloRecoveryToolName = "caveman_caveman_retrieve"
)

```

*Source: [adapter.go lines 364–366](https://github.com/JuliusBrussee/caveman/blob/main/proxy/providers/adapter.go#L364-L366)*

These constants allow the gateway to recognize tool invocation requests regardless of whether the agent uses the short form, fully-qualified MCP namespace, or legacy kilo prefix. The primary identifier `caveman_retrieve` remains the standard interface exposed to agent code, while the prefixed variants ensure backward compatibility with different MCP client implementations.

## Server-Side Resolution Architecture

Unlike standard MCP tools that proxy requests to the underlying model, `caveman_retrieve` undergoes **server-side resolution** within the gateway layer. When processing responses in [`proxy/internal/gateway/server.go`](https://github.com/JuliusBrussee/caveman/blob/main/proxy/internal/gateway/server.go), the system examines outgoing tool calls for the recovery identifier and intercepts them before they reach the model provider:

```go
// proxy/internal/gateway/server.go
// resolve caveman_retrieve calls server-side after S4 compression.
// recoveryViaMCP records that the wrapped agent fulfills caveman_retrieve itself.
// RecoveryViaMCP routes S4 recovery through the agent's own caveman_retrieve MCP.

```

*Source: [server.go lines 134, 304, 357, 416](https://github.com/JuliusBrussee/caveman/blob/main/proxy/internal/gateway/server.go#L134-L416)*

This interception pattern ensures that decompression operations remain transparent to the model while remaining accessible to the agent. The server maintains the mapping between compression markers and original payload locations, eliminating the need to transmit large uncompressed data through the model context window.

## The Decompression Workflow

When an agent encounters compressed content marked with `<<ccr:handle>>` tokens, the retrieval process follows a strict server-side sequence:

1.  **Marker Detection**: The agent emits a `caveman_retrieve` tool call containing the handle identifier embedded in the compressed stream.
2.  **Store Lookup**: The gateway extracts the handle and queries the **CCR store** for the original bytes associated with that compression marker.
3.  **Direct Response**: The server returns the uncompressed payload as the tool result, bypassing the model entirely to prevent token waste.

The tool schema itself carries no arguments because the proxy maintains session state mapping requests to their respective compression contexts. As defined in [`proxy/internal/gateway/retrieve_tool.go`](https://github.com/JuliusBrussee/caveman/blob/main/proxy/internal/gateway/retrieve_tool.go), the tool declaration uses an empty parameter object:

```go
// proxy/internal/gateway/retrieve_tool.go
const retrieveToolName = "caveman_retrieve"

```

*Source: [retrieve_tool.go line 17](https://github.com/JuliusBrussee/caveman/blob/main/proxy/internal/gateway/retrieve_tool.go#L17)*

## Agent Integration and Duplicate Prevention

The **Pi extension** ([`packages/pi-extension/src/index.ts`](https://github.com/JuliusBrussee/caveman/blob/main/packages/pi-extension/src/index.ts)) automatically registers `caveman_retrieve` as a model-visible function to ensure agents can always request decompression when needed:

```typescript
// packages/pi-extension/src/index.ts
const RECOVERY_TOOL = "caveman_retrieve";

```

*Source: [index.ts line 21](https://github.com/JuliusBrussee/caveman/blob/main/packages/pi-extension/src/index.ts#L21)*

To prevent conflicts, the gateway implements a strict **injection guard** that detects when an agent already provides its own `caveman_retrieve` implementation. The test suite in [`proxy/internal/gateway/retrieve_tool_test.go`](https://github.com/JuliusBrussee/caveman/blob/main/proxy/internal/gateway/retrieve_tool_test.go) validates this behavior:

```go
// proxy/internal/gateway/retrieve_tool_test.go
// MCP recovery must NOT inject a tool (the agent owns caveman_retrieve)

```

*Source: [retrieve_tool_test.go lines 65–81](https://github.com/JuliusBrussee/caveman/blob/main/proxy/internal/gateway/retrieve_tool_test.go#L65-L81)*

This guard ensures that custom agent implementations retain control over their retrieval logic while preventing duplicate tool registrations that could confuse the MCP client.

## Code Implementation Examples

The following JSON structure represents how the tool appears in an agent's available function set:

```json
{
  "type": "function",
  "function": {
    "name": "caveman_retrieve",
    "description": "Recover original uncompressed content from CCR store",
    "parameters": {
      "type": "object",
      "properties": {
        "handle": {
          "type": "string",
          "description": "The CCR handle marker from compressed content"
        }
      }
    }
  }
}

```

When the gateway processes a retrieval request, the internal handling logic follows this pattern:

```go
func handleRetrieve(ctx context.Context, req *RetrieveRequest) (*ToolResult, error) {
    // Extract the CCR handle from the tool call parameters
    handle := req.Parameters["handle"].(string)
    
    // Load original bytes from the compression runtime store
    original, err := ccrStore.Load(handle)
    if err != nil {
        return nil, fmt.Errorf("failed to retrieve payload: %w", err)
    }
    
    // Return uncompressed data as the tool output
    return &ToolResult{
        Name:   "caveman_retrieve",
        Output: string(original),
    }, nil
}

```

## Summary

- The MCP server exposes `caveman_retrieve` through multiple canonical names defined in [`proxy/providers/adapter.go`](https://github.com/JuliusBrussee/caveman/blob/main/proxy/providers/adapter.go) to support various agent implementations.
- Server-side resolution in [`proxy/internal/gateway/server.go`](https://github.com/JuliusBrussee/caveman/blob/main/proxy/internal/gateway/server.go) intercepts retrieval calls to prevent model proxying and reduce token usage.
- The tool queries the **CCR store** using compression markers (`<<ccr:handle>>`) to restore original payloads without transmitting uncompressed data through the model context.
- The **Pi extension** automatically registers the tool for agent visibility, while the gateway prevents duplicate injections to avoid MCP client conflicts.
- An empty parameter schema indicates that the proxy maintains compression context internally, requiring only the handle identifier from the agent.

## Frequently Asked Questions

### How does `caveman_retrieve` differ from standard MCP tools?

Standard MCP tools forward requests to the underlying model provider for processing, whereas `caveman_retrieve` is resolved **entirely within the gateway**. The server intercepts these calls according to the logic in [`proxy/internal/gateway/server.go`](https://github.com/JuliusBrussee/caveman/blob/main/proxy/internal/gateway/server.go) to query the CCR store directly, ensuring that large uncompressed payloads never consume model context window tokens.

### What prevents duplicate `caveman_retrieve` registrations when agents define their own?

The gateway implements an injection guard that checks for existing tool definitions before registering the built-in retrieval handler. As validated in [`proxy/internal/gateway/retrieve_tool_test.go`](https://github.com/JuliusBrussee/caveman/blob/main/proxy/internal/gateway/retrieve_tool_test.go) (lines 65–81), the system skips tool injection when it detects that the agent already exposes `caveman_retrieve`, allowing custom implementations to override the default behavior.

### Why does the `caveman_retrieve` tool use an empty parameter schema?

The tool schema omits required parameters because the **proxy maintains session state** mapping compression markers to their original payloads. When an agent invokes the tool, the gateway correlates the request with the current conversation context to identify which compressed block to retrieve, eliminating the need for the agent to manually track and transmit complex identifiers beyond the embedded handle marker.