Internal Request Flow Through caveman-proxy: From Agent to Provider Explained

The internal request flow through caveman-proxy routes agent requests through HTTP entry, authentication, credential resolution, compression, intelligent routing, and provider-specific adapters before returning metered, byte-preserved responses.

The caveman repository by JuliusBrussee implements a byte-preserving, request-metering gateway that mediates between AI agents and large language model providers. Understanding the internal request flow through caveman-proxy reveals how raw HTTP requests from agents transform into provider-compatible payloads while enforcing authentication, tracking spend, and preserving critical code structure.

HTTP Entry Point and RequestContext Initialization

The proxy binary initializes in cmd/caveman-proxy/main.go, binding to the port specified by the CAVEMAN_PROXY_PORT environment variable (default 8787).

Incoming HTTP requests hit the Server struct defined in proxy/internal/gateway/server.go at line 283, which implements the http.Handler interface. For every connection, the server constructs a RequestContext (defined at line 33 in the same file) that aggregates:

  • The raw HTTP request object
  • A unique request identifier
  • The target provider selection
  • Authentication metadata
  • Telemetry hooks for the request lifecycle

Authentication and Credential Resolution

Before any upstream network traffic occurs, the proxy enforces security boundaries. The Authenticator interface (line 55 in server.go) validates Authorization headers or proxy-specific bearer tokens. Upon successful authentication, the CredentialResolver (line 62) translates the agent's credential hint into concrete provider API keys or secrets.

If authentication fails at this stage, the proxy returns 401/403 errors immediately without contacting the upstream provider, preventing unnecessary network egress and credential exposure.

Routing and Cost Estimation

With valid credentials, the proxy consults the routing package to determine provider selection and escalation policies (e.g., routing from cheaper to more expensive models based on context). The Estimator (line 123 in server.go) predicts the anticipated cost of the request, while the PrefixCache (line 162) may short-circuit repeated identical prompts, returning cached responses without upstream calls.

Request Transformation and Compression

Tool Schema Stripping

For requests containing function-calling or tool definitions, the ToolSchemaStripper (line 113) removes these schemas from the payload sent to the provider. This ensures the upstream model receives a clean prompt while the proxy retains the tooling metadata for post-processing.

Compression Pipeline

The Compressor hierarchy (lines 86-102) applies caveman compression rules to the request payload. This process preserves code blocks, URLs, and identifiers while stripping filler text and unnecessary whitespace. In proxy/internal/gateway/proxy.go, the system tracks requestEvidence at line 573 and later records compressionOutcome at line 907 to manage streaming response states.

Upstream Dispatch and Provider Adaptation

After pre-processing, the proxy constructs a new HTTP request targeting the selected provider's API endpoint. The actual network call executes through a roundTripFunc (referenced in server_test.go at line 57) that wraps http.DefaultTransport for telemetry injection.

Provider-specific adapters reside under proxy/providers/*. For instance, proxy/providers/openai/openai.go translates the generic internal request into OpenAI's exact JSON schema and parses the provider's response back into the standardized Caveman format.

Response Handling and Telemetry

The provider's raw response flows through the reverse compression pipeline to expand compressed markers and re-insert tool results. Throughout this lifecycle, the TelemetrySink (line 68) and PayloadSink (line 74) record structured metrics including token usage, latency, and spend. This data persists to the local CCR database specified by CAVEMAN_CCR_DB for cost analysis via caveman-proxy stats.

Practical Code Examples

Starting the proxy locally:


# Start the proxy on default port 8787

export CAVEMAN_PROXY_PORT=8787
caveman-proxy serve

Configuring a TypeScript agent to use the proxy:

import { createCavemanAgent } from '@caveman/sdk';

const agent = await createCavemanAgent({
  apiBaseUrl: 'http://127.0.0.1:8787',  // Proxy endpoint
  provider: 'openai',                   // Upstream target
});

Manual request flow via curl:


# Agent sends request to proxy

curl -X POST http://127.0.0.1:8787/v1/chat/completions \
  -H "Authorization: Bearer <agent-token>" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "gpt-4o-mini",
    "messages": [{"role": "user", "content": "Explain quantum tunneling"}],
    "max_tokens": 200
  }'

Retrieving accumulated statistics:

caveman-proxy stats --json

# Output: {"total_input_tokens":12345,"total_output_tokens":6789,"spend_usd":12.34}

Key Implementation Files

The following source files define the complete internal request flow through caveman-proxy:

Summary

The internal request flow through caveman-proxy follows a rigorous pipeline:

  • Entry: HTTP requests hit the Server struct in server.go, creating a RequestContext
  • Security: Authenticator and CredentialResolver validate identity before upstream contact
  • Optimization: Estimator and PrefixCache handle routing decisions and caching
  • Transformation: ToolSchemaStripper and Compressor modify payloads while preserving bytes
  • Dispatch: Provider adapters in proxy/providers/ translate requests to vendor-specific formats
  • Telemetry: TelemetrySink records metrics to the CCR database for spend analysis

Frequently Asked Questions

How does caveman-proxy handle authentication failures?

The Authenticator interface (line 55 in server.go) validates incoming Authorization headers. If validation fails, the proxy returns HTTP 401 or 403 errors immediately during the credential resolution phase, preventing any network traffic to upstream providers and protecting API keys from exposure.

What compression techniques does caveman-proxy apply to requests?

The Compressor hierarchy (lines 86-102) applies domain-specific rules that preserve code blocks, URLs, and identifiers while removing non-essential prose. For streaming responses, proxy.go tracks requestEvidence (line 573) and compressionOutcome (line 907) to manage marker expansion and ensure byte-accurate reconstruction.

Which components determine which LLM provider receives the request?

The routing package consults the Estimator (line 123) for cost prediction and the PrefixCache (line 162) for deduplication. The CredentialResolver (line 62) then selects the appropriate provider adapter from proxy/providers/* (such as openai.go) to handle protocol translation.

Where does caveman-proxy store request telemetry and spend data?

The TelemetrySink (line 68) and PayloadSink (line 74) capture metrics including token usage and latency. This data persists to the local CCR database path specified by the CAVEMAN_CCR_DB environment variable, accessible via the caveman-proxy stats command for aggregated spend analysis.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →