# How Caveman Telemetry Reporting Works and What Data Is Never Sent

> Discover how Caveman telemetry reporting captures operational metadata and securely filters out sensitive data like prompts and credentials before persistence. Learn what data is never sent.

- Repository: [Julius Brussee/caveman](https://github.com/JuliusBrussee/caveman)
- Tags: deep-dive
- Published: 2026-09-04

---

**Caveman telemetry reporting captures operational metadata in immutable `Span` structs while aggressively filtering out sensitive data—including raw prompts, credentials, and content-bearing attributes—before any persistence occurs.**

The JuliusBrussee/caveman repository implements a privacy-first telemetry pipeline designed to record system performance without compromising user confidentiality. Each telemetry event is constructed as a metadata-only record in [`shared/platform/telemetry/span.go`](https://github.com/JuliusBrussee/caveman/blob/main/shared/platform/telemetry/span.go), ensuring that raw user inputs and authentication secrets never leave the secure execution boundary. Understanding exactly how this sanitization pipeline functions is essential for operators auditing their data exhaust.

## How Telemetry Spans Capture Operational Data

The foundation of Caveman telemetry lies in the `Span` struct defined in [`shared/platform/telemetry/span.go`](https://github.com/JuliusBrussee/caveman/blob/main/shared/platform/telemetry/span.go). This immutable data structure stores high-level operational metadata such as timestamps, trace identifiers, provider names, model types, token counts, and cost metrics. The `ApplyProvenance` function (lines 46-96) enriches each span with source context—including the source kind, system identifier, ingestion ID, and normalization version—establishing a complete audit trail without exposing payload contents.

## The Sanitization Pipeline: What Gets Filtered

Before any span reaches persistent storage, the `SanitizeAttributes` function processes its attribute map through a multi-layered filtering system. This pipeline enforces strict exclusion rules based on key naming patterns, content classifications, and size constraints defined in the same source file.

### AttributeAllowlist Logic

The `AttributeAllowed` function (lines 20-45) serves as the primary gatekeeper, returning `false` for any key that matches sensitive patterns. This function implements a default-deny posture for high-risk identifiers, ensuring that only explicitly permitted metadata keys survive the filtering process.

### Content Prefixes That Are Always Stripped

Attributes bearing content-related prefixes are unconditionally removed. According to the source code (lines 98-119), any key matching patterns such as `gen_ai.input.messages`, `gen_ai.output.messages`, or `tool.parameters` is flagged as content-bearing and immediately discarded. This ensures that raw prompt text and response bodies never enter the telemetry stream, as emphasized in the repository's privacy documentation at [`docs/technical/security-and-privacy.md`](https://github.com/JuliusBrussee/caveman/blob/main/docs/technical/security-and-privacy.md).

### Credential and Secret Detection

The filtering logic explicitly scans for credential-related substrings. Keys containing `authorization`, `apikey`, `api`, `secret`, `password`, `passwd`, `api_key`, `access_token`, or `refresh_token` are rejected (lines 33-41). Additionally, any attribute starting with the reserved namespace prefix `cave.reserved.` is automatically excluded from telemetry to prevent leakage of platform-internal state.

### Size Limits and Validation

The system imposes hard boundaries on attribute dimensions to prevent data exfiltration or storage abuse. Constants defined at lines 15-44 enforce maximum limits on the number of attributes per span, individual key length, value length, and total byte size. Any attribute violating these constraints is truncated or removed during the sanitization pass.

## What Data Is Never Sent to Telemetry

Based on the `AttributeAllowed` implementation and content filtering rules in [`shared/platform/telemetry/span.go`](https://github.com/JuliusBrussee/caveman/blob/main/shared/platform/telemetry/span.go), the following categories are categorically prohibited from telemetry transmission:

- **Raw message content**: Any attribute with content prefixes like `gen_ai.input.messages` or `gen_ai.output.messages` (lines 98-119)
- **Authentication credentials**: Keys containing `authorization`, `apikey`, `secret`, or token identifiers (lines 33-41)
- **Reserved internal state**: Keys prefixed with `cave.reserved.`
- **Oversized payloads**: Attributes exceeding the `MaxAttributeKeyLength`, `MaxAttributeValueLength`, or `MaxAttributesPerSpan` limits defined in lines 15-44

## Implementation Example: Building a Sanitized Span

The following example demonstrates constructing a telemetry span, applying provenance metadata, and running the sanitization routine:

```go
package main

import (
    "fmt"
    "github.com/JuliusBrussee/caveman/shared/platform/telemetry"
)

func RecordCompletion() {
    // Initialize span with operational metadata only
    span := telemetry.Span{
        Timestamp:    "2026-01-15T10:30:00Z",
        TraceID:      "trace-abc-123",
        SpanID:       "span-def-456",
        SpanName:     "llm.completion",
        SpanType:     "completion",
        Provider:     "openai",
        Model:        "gpt-4",
        InputTokens:  150,
        OutputTokens: 400,
        TotalCostUSD: 0.015,
        Attributes: map[string]string{
            "user_id":               "user-789",      // Safe: retained
            "authorization":         "Bearer sk-xxx", // Unsafe: stripped by AttributeAllowed
            "gen_ai.input.messages": "[{\"role\":\"user\",\"content\":\"secret\"}]", // Unsafe: content prefix stripped
        },
    }

    // Enrich with provenance (lines 46-96 in span.go)
    telemetry.ApplyProvenance(&span, telemetry.Provenance{
        SourceKind:           telemetry.SourceKindGateway,
        SourceSystem:         telemetry.SourceSystemCaveman,
        NormalizationVersion: telemetry.NormalizationVersion,
        ContentMode:          telemetry.ContentModeMetadataOnly,
    })

    // Sanitize attributes (lines 20-45, 98-119 in span.go)
    cleanAttrs, err := telemetry.SanitizeAttributes(span.Attributes)
    if err != nil {
        panic(err)
    }
    span.Attributes = cleanAttrs

    // Result: span.Attributes contains only {"user_id": "user-789"}
    fmt.Printf("Sanitized attributes: %v\n", span.Attributes)
}

```

## Summary

- **Span-centric architecture**: Caveman uses immutable `Span` structs in [`shared/platform/telemetry/span.go`](https://github.com/JuliusBrussee/caveman/blob/main/shared/platform/telemetry/span.go) to record operational telemetry without raw payloads.
- **Aggressive content filtering**: The `AttributeAllowed` function and content prefix checks ensure raw prompts and responses never persist (lines 98-119).
- **Credential protection**: Automatic detection of `authorization`, `apikey`, and similar substrings prevents secret leakage (lines 33-41).
- **Hard size limits**: Constants enforcing maximum attribute counts and byte sizes mitigate storage risks (lines 15-44).
- **Provenance tracking**: `ApplyProvenance` adds source metadata without exposing sensitive execution context (lines 46-96).

## Frequently Asked Questions

### Does Caveman telemetry store raw LLM prompts?

No. The telemetry system explicitly strips any attribute matching content-bearing prefixes such as `gen_ai.input.messages` or `gen_ai.output.messages` (lines 98-119 in [`shared/platform/telemetry/span.go`](https://github.com/JuliusBrussee/caveman/blob/main/shared/platform/telemetry/span.go)). Only token counts, model identifiers, and other metadata are retained, as documented in the repository's security guidelines.

### How does Caveman detect and remove API keys from telemetry?

The `AttributeAllowed` function scans attribute keys for credential-related substrings including `apikey`, `api_key`, `authorization`, `secret`, `password`, `access_token`, and `refresh_token` (lines 33-41). Any matching key is discarded before the span is serialized, ensuring that authentication materials never reach the telemetry backend.

### What happens if a span exceeds attribute size limits?

The `SanitizeAttributes` function enforces bounds defined by constants at lines 15-44, including `MaxAttributesPerSpan`, `MaxAttributeKeyLength`, and `MaxAttributeValueLength`. Attributes violating these limits are truncated or removed entirely to prevent storage abuse and potential denial-of-service vectors.

### Is the telemetry data structure immutable once created?

While the `Span` struct itself is designed as an immutable record, the `Attributes` map is explicitly sanitized via `SanitizeAttributes` before persistence. The provenance metadata is applied through `ApplyProvenance` (lines 46-96), after which the span is treated as read-only for downstream consumers and telemetry sinks.