# How Agent Wrapping Works for Claude Code, Codex, and Other LLM Agents

> Discover how Caveman's agent wrapping creates a local proxy for LLM traffic, enabling compression, metering, and tracing for Claude Code, Codex, and more without code changes.

- Repository: [Julius Brussee/caveman](https://github.com/JuliusBrussee/caveman)
- Tags: how-to-guide
- Published: 2026-09-04

---

**Caveman’s agent wrapping feature creates a transparent local proxy that intercepts LLM traffic from Claude Code, Codex, and other supported agents, enabling compression, metering, and tracing without modifying the agent’s source code.**

Agent wrapping is the core mechanism in [JuliusBrussee/caveman](https://github.com/JuliusBrussee/caveman) that lets you run mainstream AI coding agents through a controlled local pipeline. When you execute `caveman wrap <agent>`, the CLI dynamically configures environment variables and network routes to redirect all provider traffic through a temporary local proxy.

## The Agent Wrapping Execution Flow

The wrapping process follows a strict eight-step pipeline implemented in [`packages/cli/src/index.ts`](https://github.com/JuliusBrussee/caveman/blob/main/packages/cli/src/index.ts). Each stage handles authentication, proxy bootstrapping, and process isolation to ensure the wrapped agent operates unaware of the interception layer.

### Step 1: Argument Parsing and Profile Resolution

First, `parseWrapArgs` extracts command-line flags including `--off` (disable compression) and `--pixel` ( pixel-based metering), captures an optional workflow identifier, and isolates the target agent name. Immediately after, `findAgent(requested)` queries the agent registry in [`packages/cli/src/agents.generated.ts`](https://github.com/JuliusBrussee/caveman/blob/main/packages/cli/src/agents.generated.ts) to retrieve the **AgentProfile** matching the short name (e.g., `claude` maps to Claude Code, `codex` maps to Codex). This profile contains the binary name, default CLI arguments, and supported runtime lanes.

### Step 2: Binary Resolution and Authentication

The wrapper validates that the agent binary exists on the system PATH using `which(bin)`. If the executable is missing, Caveman renders a "not-found" UI and exits gracefully. For authentication, the behavior splits by provider:

- **Claude Code**: Uses native Anthropic authentication (API key or Claude Pro/Max OAuth token) passed unchanged through the proxy
- **Codex**: The helper `detectCodexWrapAuthMode()` in [`packages/cli/src/index.ts`](https://github.com/JuliusBrussee/caveman/blob/main/packages/cli/src/index.ts) distinguishes between **api-key** and **subscription** modes, determining whether to expect a standard OpenAI key or an Azure-backed OAuth session

### Step 3: Local Proxy Bootstrapping

The function `bootstrapLocalWrapRuntime` launches a temporary Caveman proxy instance. By default, it operates in **compress** mode, applying local token-level metering and byte-stream compression before forwarding requests to upstream providers. When `--off` is supplied, the proxy switches to **record** mode, capturing traffic without compression. The proxy listens on a local port and prepares URL rewrite rules for each supported provider.

### Step 4: Environment Injection and Process Execution

The `runWrapped` function (defined around line 4600 in [`packages/cli/src/index.ts`](https://github.com/JuliusBrussee/caveman/blob/main/packages/cli/src/index.ts)) prepares the execution environment by injecting critical variables:

- `CAVE_WORKFLOW`: Propagates the optional workflow identifier for distributed tracing
- `ANTHROPIC_BASE_URL`: Redirects Claude Code traffic to `http://127.0.0.1:<port>/w/claude`
- `GOOGLE_AI_API_URL`: Overrides the endpoint for Gemini agents
- Custom headers (`x-cave-api-key`, `x-cave-upstream-key`) when a Caveman Cloud gateway is configured

The wrapper then spawns the agent process with modified arguments. The proxy transparently tracks request/response sizes, applies compression algorithms, and stores a session summary accessible via the `caveman tools` UI. After the agent exits, Caveman emits telemetry events and can display retroactive scans showing token savings achieved through compression.

## Supported Agents and Configuration Matrix

Caveman’s agent wrapping supports multiple LLM providers through standardized profile configurations. The following table details how the proxy handles each major agent:

| Agent | Default Base URL | Rewritten Proxy Path | Authentication Mode | Compression Default |
|-------|-----------------|---------------------|---------------------|-------------------|
| **Claude Code** | `https://api.anthropic.com/v1` | `http://127.0.0.1:<port>/w/claude` | API key or OAuth token | Enabled (compress) |
| **Codex** | `https://api.openai.com/v1` | `http://127.0.0.1:<port>/w/codex` | API key or subscription | Enabled (compress) |
| **Gemini** | `https://generativelanguage.googleapis.com` | `http://127.0.0.1:<port>/w/gemini` | Google AI API key | Enabled (compress) |

The proxy preserves raw authentication tokens, modifying only the request destination path. For agents supporting **MCP** (Model Context Protocol), such as those utilizing `caveman-mcp`, the wrapper optionally injects recovery hooks and tool configurations managed by [`packages/pi-extension/src/lifecycle.ts`](https://github.com/JuliusBrussee/caveman/blob/main/packages/pi-extension/src/lifecycle.ts).

### Pixel Mode and Subscription Restrictions

When using `--pixel` for visual token metering, the wrapper enforces compatibility checks early in the flow. **Pixel mode is explicitly disabled for subscription-based sessions** (both Claude Pro/Max and Codex subscription tiers) to prevent authentication conflicts with visual encoding layers.

## Practical Agent Wrapping Examples

Use the following commands to wrap agents in different operating modes:

```bash

# Wrap Claude Code with default compression enabled

caveman wrap claude --prompt "Explain quantum tunneling"

# Wrap Codex in record mode (no compression) for debugging traffic

caveman wrap codex --off "Implement a binary search tree"

# Attach a workflow identifier for distributed tracing

caveman wrap claude --workflow research-run-42 "Summarize the attached paper"

# Forward all arguments dynamically in a shell script

#!/usr/bin/env bash
caveman wrap claude "$@"

```

The `--off` flag disables the default compression layer, useful when debugging raw API traffic or when working with streaming responses that conflict with byte-level compression.

## Summary

- **Agent wrapping** in Caveman creates an intermediary proxy that intercepts LLM traffic without code modifications to the target agent.
- The process relies on `parseWrapArgs` and `findAgent` to resolve configurations from [`packages/cli/src/agents.generated.ts`](https://github.com/JuliusBrussee/caveman/blob/main/packages/cli/src/agents.generated.ts).
- Authentication tokens pass through transparently while base URLs get rewritten to local proxy endpoints like `/w/claude` or `/w/codex`.
- **Compress mode** (default) enables token-level metering and bandwidth reduction, while **record mode** (`--off`) captures raw traffic.
- Environment variables including `ANTHROPIC_BASE_URL` and `CAVE_WORKFLOW` ensure seamless integration with Claude Code, Codex, and Gemini.

## Frequently Asked Questions

### What is agent wrapping in Caveman?

Agent wrapping is a CLI feature that executes LLM coding agents (Claude Code, Codex, Gemini) through a local transparent proxy. This interception layer enables request logging, compression, and workflow tracing without requiring changes to the agent’s own codebase or configuration files.

### How does Caveman handle authentication for wrapped agents?

Caveman forwards authentication tokens unchanged to upstream providers. For Claude Code, it accepts standard API keys or OAuth tokens from Claude Pro/Max subscriptions. For Codex, the `detectCodexWrapAuthMode()` helper determines whether to expect a standard OpenAI API key or a subscription-based Azure token, applying the appropriate headers accordingly.

### Can I use agent wrapping with agents other than Claude Code and Codex?

Yes. Caveman supports a configurable agent registry defined in [`packages/cli/src/agents.generated.ts`](https://github.com/JuliusBrussee/caveman/blob/main/packages/cli/src/agents.generated.ts). You can wrap any agent that relies on standard HTTP-based LLM APIs by adding its profile, binary name, and base URL to the registry. The proxy system generically handles URL rewriting for any provider following the `/w/{agent}` path pattern.

### What is the difference between compress mode and record mode?

**Compress mode** is the default behavior where `bootstrapLocalWrapRuntime` applies local token-level compression and metering before forwarding requests to the provider, reducing bandwidth and enabling cost tracking. **Record mode** (activated with `--off`) disables compression to capture raw request/response streams, useful for debugging or when working with non-compressible binary data.