# Internal Request Flow Through caveman-proxy: From Agent to Provider Explained

> Understand the internal request flow through caveman-proxy. See how agent requests are routed from entry to provider, including authentication, routing, and adapters.

- Repository: [Julius Brussee/caveman](https://github.com/JuliusBrussee/caveman)
- Tags: internals
- Published: 2026-09-04

---

**The internal request flow through caveman-proxy routes agent requests through HTTP entry, authentication, credential resolution, compression, intelligent routing, and provider-specific adapters before returning metered, byte-preserved responses.**

The `caveman` repository by **JuliusBrussee** implements a byte-preserving, request-metering gateway that mediates between AI agents and large language model providers. Understanding the internal request flow through caveman-proxy reveals how raw HTTP requests from agents transform into provider-compatible payloads while enforcing authentication, tracking spend, and preserving critical code structure.

## HTTP Entry Point and RequestContext Initialization

The proxy binary initializes in [`cmd/caveman-proxy/main.go`](https://github.com/JuliusBrussee/caveman/blob/main/cmd/caveman-proxy/main.go), binding to the port specified by the `CAVEMAN_PROXY_PORT` environment variable (default **8787**).

Incoming HTTP requests hit the `Server` struct defined in [`proxy/internal/gateway/server.go`](https://github.com/JuliusBrussee/caveman/blob/main/proxy/internal/gateway/server.go) at **line 283**, which implements the `http.Handler` interface. For every connection, the server constructs a **`RequestContext`** (defined at **line 33** in the same file) that aggregates:

- The raw HTTP request object
- A unique request identifier
- The target provider selection
- Authentication metadata
- Telemetry hooks for the request lifecycle

## Authentication and Credential Resolution

Before any upstream network traffic occurs, the proxy enforces security boundaries. The **`Authenticator`** interface (line **55** in [`server.go`](https://github.com/JuliusBrussee/caveman/blob/main/server.go)) validates `Authorization` headers or proxy-specific bearer tokens. Upon successful authentication, the **`CredentialResolver`** (line **62**) translates the agent's credential hint into concrete provider API keys or secrets.

If authentication fails at this stage, the proxy returns **401/403** errors immediately without contacting the upstream provider, preventing unnecessary network egress and credential exposure.

## Routing and Cost Estimation

With valid credentials, the proxy consults the **`routing`** package to determine provider selection and escalation policies (e.g., routing from cheaper to more expensive models based on context). The **`Estimator`** (line **123** in [`server.go`](https://github.com/JuliusBrussee/caveman/blob/main/server.go)) predicts the anticipated cost of the request, while the **`PrefixCache`** (line **162**) may short-circuit repeated identical prompts, returning cached responses without upstream calls.

## Request Transformation and Compression

### Tool Schema Stripping

For requests containing function-calling or tool definitions, the **`ToolSchemaStripper`** (line **113**) removes these schemas from the payload sent to the provider. This ensures the upstream model receives a clean prompt while the proxy retains the tooling metadata for post-processing.

### Compression Pipeline

The **`Compressor`** hierarchy (lines **86-102**) applies **caveman compression rules** to the request payload. This process preserves code blocks, URLs, and identifiers while stripping filler text and unnecessary whitespace. In [`proxy/internal/gateway/proxy.go`](https://github.com/JuliusBrussee/caveman/blob/main/proxy/internal/gateway/proxy.go), the system tracks `requestEvidence` at **line 573** and later records `compressionOutcome` at **line 907** to manage streaming response states.

## Upstream Dispatch and Provider Adaptation

After pre-processing, the proxy constructs a new HTTP request targeting the selected provider's API endpoint. The actual network call executes through a **`roundTripFunc`** (referenced in [`server_test.go`](https://github.com/JuliusBrussee/caveman/blob/main/server_test.go) at **line 57**) that wraps `http.DefaultTransport` for telemetry injection.

Provider-specific adapters reside under `proxy/providers/*`. For instance, [`proxy/providers/openai/openai.go`](https://github.com/JuliusBrussee/caveman/blob/main/proxy/providers/openai/openai.go) translates the generic internal request into OpenAI's exact JSON schema and parses the provider's response back into the standardized Caveman format.

## Response Handling and Telemetry

The provider's raw response flows through the reverse compression pipeline to expand compressed markers and re-insert tool results. Throughout this lifecycle, the **`TelemetrySink`** (line **68**) and **`PayloadSink`** (line **74**) record structured metrics including token usage, latency, and spend. This data persists to the local CCR database specified by `CAVEMAN_CCR_DB` for cost analysis via `caveman-proxy stats`.

## Practical Code Examples

Starting the proxy locally:

```bash

# Start the proxy on default port 8787

export CAVEMAN_PROXY_PORT=8787
caveman-proxy serve

```

Configuring a TypeScript agent to use the proxy:

```typescript
import { createCavemanAgent } from '@caveman/sdk';

const agent = await createCavemanAgent({
  apiBaseUrl: 'http://127.0.0.1:8787',  // Proxy endpoint
  provider: 'openai',                   // Upstream target
});

```

Manual request flow via curl:

```bash

# Agent sends request to proxy

curl -X POST http://127.0.0.1:8787/v1/chat/completions \
  -H "Authorization: Bearer <agent-token>" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "gpt-4o-mini",
    "messages": [{"role": "user", "content": "Explain quantum tunneling"}],
    "max_tokens": 200
  }'

```

Retrieving accumulated statistics:

```bash
caveman-proxy stats --json

# Output: {"total_input_tokens":12345,"total_output_tokens":6789,"spend_usd":12.34}

```

## Key Implementation Files

The following source files define the complete internal request flow through caveman-proxy:

- **[`proxy/internal/gateway/server.go`](https://github.com/JuliusBrussee/caveman/blob/main/proxy/internal/gateway/server.go)** – Contains the `Server` struct (line 283), `RequestContext` (line 33), and core authentication interfaces
- **[`proxy/internal/gateway/proxy.go`](https://github.com/JuliusBrussee/caveman/blob/main/proxy/internal/gateway/proxy.go)** – Implements request evidence tracking (line 573) and compression outcome handling (line 907)
- **[`proxy/providers/openai/openai.go`](https://github.com/JuliusBrussee/caveman/blob/main/proxy/providers/openai/openai.go)** – Provider-specific adapter for OpenAI API translation
- **[`proxy/internal/gateway/server_test.go`](https://github.com/JuliusBrussee/caveman/blob/main/proxy/internal/gateway/server_test.go)** – Defines `roundTripFunc` wrapper (line 57) for HTTP transport testing
- **[`cmd/caveman-proxy/main.go`](https://github.com/JuliusBrussee/caveman/blob/main/cmd/caveman-proxy/main.go)** – CLI entry point that initializes the HTTP listener

## Summary

The internal request flow through caveman-proxy follows a rigorous pipeline:

- **Entry**: HTTP requests hit the `Server` struct in [`server.go`](https://github.com/JuliusBrussee/caveman/blob/main/server.go), creating a `RequestContext`
- **Security**: `Authenticator` and `CredentialResolver` validate identity before upstream contact
- **Optimization**: `Estimator` and `PrefixCache` handle routing decisions and caching
- **Transformation**: `ToolSchemaStripper` and `Compressor` modify payloads while preserving bytes
- **Dispatch**: Provider adapters in `proxy/providers/` translate requests to vendor-specific formats
- **Telemetry**: `TelemetrySink` records metrics to the CCR database for spend analysis

## Frequently Asked Questions

### How does caveman-proxy handle authentication failures?

The `Authenticator` interface (line 55 in [`server.go`](https://github.com/JuliusBrussee/caveman/blob/main/server.go)) validates incoming `Authorization` headers. If validation fails, the proxy returns HTTP 401 or 403 errors immediately during the credential resolution phase, preventing any network traffic to upstream providers and protecting API keys from exposure.

### What compression techniques does caveman-proxy apply to requests?

The `Compressor` hierarchy (lines 86-102) applies domain-specific rules that preserve code blocks, URLs, and identifiers while removing non-essential prose. For streaming responses, [`proxy.go`](https://github.com/JuliusBrussee/caveman/blob/main/proxy.go) tracks `requestEvidence` (line 573) and `compressionOutcome` (line 907) to manage marker expansion and ensure byte-accurate reconstruction.

### Which components determine which LLM provider receives the request?

The **routing** package consults the `Estimator` (line 123) for cost prediction and the `PrefixCache` (line 162) for deduplication. The `CredentialResolver` (line 62) then selects the appropriate provider adapter from `proxy/providers/*` (such as [`openai.go`](https://github.com/JuliusBrussee/caveman/blob/main/openai.go)) to handle protocol translation.

### Where does caveman-proxy store request telemetry and spend data?

The `TelemetrySink` (line 68) and `PayloadSink` (line 74) capture metrics including token usage and latency. This data persists to the local CCR database path specified by the `CAVEMAN_CCR_DB` environment variable, accessible via the `caveman-proxy stats` command for aggregated spend analysis.