How Agent Wrapping Works for Claude Code, Codex, and Other LLM Agents

Caveman’s agent wrapping feature creates a transparent local proxy that intercepts LLM traffic from Claude Code, Codex, and other supported agents, enabling compression, metering, and tracing without modifying the agent’s source code.

Agent wrapping is the core mechanism in JuliusBrussee/caveman that lets you run mainstream AI coding agents through a controlled local pipeline. When you execute caveman wrap <agent>, the CLI dynamically configures environment variables and network routes to redirect all provider traffic through a temporary local proxy.

The Agent Wrapping Execution Flow

The wrapping process follows a strict eight-step pipeline implemented in packages/cli/src/index.ts. Each stage handles authentication, proxy bootstrapping, and process isolation to ensure the wrapped agent operates unaware of the interception layer.

Step 1: Argument Parsing and Profile Resolution

First, parseWrapArgs extracts command-line flags including --off (disable compression) and --pixel ( pixel-based metering), captures an optional workflow identifier, and isolates the target agent name. Immediately after, findAgent(requested) queries the agent registry in packages/cli/src/agents.generated.ts to retrieve the AgentProfile matching the short name (e.g., claude maps to Claude Code, codex maps to Codex). This profile contains the binary name, default CLI arguments, and supported runtime lanes.

Step 2: Binary Resolution and Authentication

The wrapper validates that the agent binary exists on the system PATH using which(bin). If the executable is missing, Caveman renders a "not-found" UI and exits gracefully. For authentication, the behavior splits by provider:

  • Claude Code: Uses native Anthropic authentication (API key or Claude Pro/Max OAuth token) passed unchanged through the proxy
  • Codex: The helper detectCodexWrapAuthMode() in packages/cli/src/index.ts distinguishes between api-key and subscription modes, determining whether to expect a standard OpenAI key or an Azure-backed OAuth session

Step 3: Local Proxy Bootstrapping

The function bootstrapLocalWrapRuntime launches a temporary Caveman proxy instance. By default, it operates in compress mode, applying local token-level metering and byte-stream compression before forwarding requests to upstream providers. When --off is supplied, the proxy switches to record mode, capturing traffic without compression. The proxy listens on a local port and prepares URL rewrite rules for each supported provider.

Step 4: Environment Injection and Process Execution

The runWrapped function (defined around line 4600 in packages/cli/src/index.ts) prepares the execution environment by injecting critical variables:

  • CAVE_WORKFLOW: Propagates the optional workflow identifier for distributed tracing
  • ANTHROPIC_BASE_URL: Redirects Claude Code traffic to http://127.0.0.1:<port>/w/claude
  • GOOGLE_AI_API_URL: Overrides the endpoint for Gemini agents
  • Custom headers (x-cave-api-key, x-cave-upstream-key) when a Caveman Cloud gateway is configured

The wrapper then spawns the agent process with modified arguments. The proxy transparently tracks request/response sizes, applies compression algorithms, and stores a session summary accessible via the caveman tools UI. After the agent exits, Caveman emits telemetry events and can display retroactive scans showing token savings achieved through compression.

Supported Agents and Configuration Matrix

Caveman’s agent wrapping supports multiple LLM providers through standardized profile configurations. The following table details how the proxy handles each major agent:

Agent Default Base URL Rewritten Proxy Path Authentication Mode Compression Default
Claude Code https://api.anthropic.com/v1 http://127.0.0.1:<port>/w/claude API key or OAuth token Enabled (compress)
Codex https://api.openai.com/v1 http://127.0.0.1:<port>/w/codex API key or subscription Enabled (compress)
Gemini https://generativelanguage.googleapis.com http://127.0.0.1:<port>/w/gemini Google AI API key Enabled (compress)

The proxy preserves raw authentication tokens, modifying only the request destination path. For agents supporting MCP (Model Context Protocol), such as those utilizing caveman-mcp, the wrapper optionally injects recovery hooks and tool configurations managed by packages/pi-extension/src/lifecycle.ts.

Pixel Mode and Subscription Restrictions

When using --pixel for visual token metering, the wrapper enforces compatibility checks early in the flow. Pixel mode is explicitly disabled for subscription-based sessions (both Claude Pro/Max and Codex subscription tiers) to prevent authentication conflicts with visual encoding layers.

Practical Agent Wrapping Examples

Use the following commands to wrap agents in different operating modes:


# Wrap Claude Code with default compression enabled

caveman wrap claude --prompt "Explain quantum tunneling"

# Wrap Codex in record mode (no compression) for debugging traffic

caveman wrap codex --off "Implement a binary search tree"

# Attach a workflow identifier for distributed tracing

caveman wrap claude --workflow research-run-42 "Summarize the attached paper"

# Forward all arguments dynamically in a shell script

#!/usr/bin/env bash
caveman wrap claude "$@"

The --off flag disables the default compression layer, useful when debugging raw API traffic or when working with streaming responses that conflict with byte-level compression.

Summary

  • Agent wrapping in Caveman creates an intermediary proxy that intercepts LLM traffic without code modifications to the target agent.
  • The process relies on parseWrapArgs and findAgent to resolve configurations from packages/cli/src/agents.generated.ts.
  • Authentication tokens pass through transparently while base URLs get rewritten to local proxy endpoints like /w/claude or /w/codex.
  • Compress mode (default) enables token-level metering and bandwidth reduction, while record mode (--off) captures raw traffic.
  • Environment variables including ANTHROPIC_BASE_URL and CAVE_WORKFLOW ensure seamless integration with Claude Code, Codex, and Gemini.

Frequently Asked Questions

What is agent wrapping in Caveman?

Agent wrapping is a CLI feature that executes LLM coding agents (Claude Code, Codex, Gemini) through a local transparent proxy. This interception layer enables request logging, compression, and workflow tracing without requiring changes to the agent’s own codebase or configuration files.

How does Caveman handle authentication for wrapped agents?

Caveman forwards authentication tokens unchanged to upstream providers. For Claude Code, it accepts standard API keys or OAuth tokens from Claude Pro/Max subscriptions. For Codex, the detectCodexWrapAuthMode() helper determines whether to expect a standard OpenAI API key or a subscription-based Azure token, applying the appropriate headers accordingly.

Can I use agent wrapping with agents other than Claude Code and Codex?

Yes. Caveman supports a configurable agent registry defined in packages/cli/src/agents.generated.ts. You can wrap any agent that relies on standard HTTP-based LLM APIs by adding its profile, binary name, and base URL to the registry. The proxy system generically handles URL rewriting for any provider following the /w/{agent} path pattern.

What is the difference between compress mode and record mode?

Compress mode is the default behavior where bootstrapLocalWrapRuntime applies local token-level compression and metering before forwarding requests to the provider, reducing bandwidth and enabling cost tracking. Record mode (activated with --off) disables compression to capture raw request/response streams, useful for debugging or when working with non-compressible binary data.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →