How Caveman Telemetry Reporting Works and What Data Is Never Sent

Caveman telemetry reporting captures operational metadata in immutable Span structs while aggressively filtering out sensitive data—including raw prompts, credentials, and content-bearing attributes—before any persistence occurs.

The JuliusBrussee/caveman repository implements a privacy-first telemetry pipeline designed to record system performance without compromising user confidentiality. Each telemetry event is constructed as a metadata-only record in shared/platform/telemetry/span.go, ensuring that raw user inputs and authentication secrets never leave the secure execution boundary. Understanding exactly how this sanitization pipeline functions is essential for operators auditing their data exhaust.

How Telemetry Spans Capture Operational Data

The foundation of Caveman telemetry lies in the Span struct defined in shared/platform/telemetry/span.go. This immutable data structure stores high-level operational metadata such as timestamps, trace identifiers, provider names, model types, token counts, and cost metrics. The ApplyProvenance function (lines 46-96) enriches each span with source context—including the source kind, system identifier, ingestion ID, and normalization version—establishing a complete audit trail without exposing payload contents.

The Sanitization Pipeline: What Gets Filtered

Before any span reaches persistent storage, the SanitizeAttributes function processes its attribute map through a multi-layered filtering system. This pipeline enforces strict exclusion rules based on key naming patterns, content classifications, and size constraints defined in the same source file.

AttributeAllowlist Logic

The AttributeAllowed function (lines 20-45) serves as the primary gatekeeper, returning false for any key that matches sensitive patterns. This function implements a default-deny posture for high-risk identifiers, ensuring that only explicitly permitted metadata keys survive the filtering process.

Content Prefixes That Are Always Stripped

Attributes bearing content-related prefixes are unconditionally removed. According to the source code (lines 98-119), any key matching patterns such as gen_ai.input.messages, gen_ai.output.messages, or tool.parameters is flagged as content-bearing and immediately discarded. This ensures that raw prompt text and response bodies never enter the telemetry stream, as emphasized in the repository's privacy documentation at docs/technical/security-and-privacy.md.

Credential and Secret Detection

The filtering logic explicitly scans for credential-related substrings. Keys containing authorization, apikey, api, secret, password, passwd, api_key, access_token, or refresh_token are rejected (lines 33-41). Additionally, any attribute starting with the reserved namespace prefix cave.reserved. is automatically excluded from telemetry to prevent leakage of platform-internal state.

Size Limits and Validation

The system imposes hard boundaries on attribute dimensions to prevent data exfiltration or storage abuse. Constants defined at lines 15-44 enforce maximum limits on the number of attributes per span, individual key length, value length, and total byte size. Any attribute violating these constraints is truncated or removed during the sanitization pass.

What Data Is Never Sent to Telemetry

Based on the AttributeAllowed implementation and content filtering rules in shared/platform/telemetry/span.go, the following categories are categorically prohibited from telemetry transmission:

  • Raw message content: Any attribute with content prefixes like gen_ai.input.messages or gen_ai.output.messages (lines 98-119)
  • Authentication credentials: Keys containing authorization, apikey, secret, or token identifiers (lines 33-41)
  • Reserved internal state: Keys prefixed with cave.reserved.
  • Oversized payloads: Attributes exceeding the MaxAttributeKeyLength, MaxAttributeValueLength, or MaxAttributesPerSpan limits defined in lines 15-44

Implementation Example: Building a Sanitized Span

The following example demonstrates constructing a telemetry span, applying provenance metadata, and running the sanitization routine:

package main

import (
    "fmt"
    "github.com/JuliusBrussee/caveman/shared/platform/telemetry"
)

func RecordCompletion() {
    // Initialize span with operational metadata only
    span := telemetry.Span{
        Timestamp:    "2026-01-15T10:30:00Z",
        TraceID:      "trace-abc-123",
        SpanID:       "span-def-456",
        SpanName:     "llm.completion",
        SpanType:     "completion",
        Provider:     "openai",
        Model:        "gpt-4",
        InputTokens:  150,
        OutputTokens: 400,
        TotalCostUSD: 0.015,
        Attributes: map[string]string{
            "user_id":               "user-789",      // Safe: retained
            "authorization":         "Bearer sk-xxx", // Unsafe: stripped by AttributeAllowed
            "gen_ai.input.messages": "[{\"role\":\"user\",\"content\":\"secret\"}]", // Unsafe: content prefix stripped
        },
    }

    // Enrich with provenance (lines 46-96 in span.go)
    telemetry.ApplyProvenance(&span, telemetry.Provenance{
        SourceKind:           telemetry.SourceKindGateway,
        SourceSystem:         telemetry.SourceSystemCaveman,
        NormalizationVersion: telemetry.NormalizationVersion,
        ContentMode:          telemetry.ContentModeMetadataOnly,
    })

    // Sanitize attributes (lines 20-45, 98-119 in span.go)
    cleanAttrs, err := telemetry.SanitizeAttributes(span.Attributes)
    if err != nil {
        panic(err)
    }
    span.Attributes = cleanAttrs

    // Result: span.Attributes contains only {"user_id": "user-789"}
    fmt.Printf("Sanitized attributes: %v\n", span.Attributes)
}

Summary

  • Span-centric architecture: Caveman uses immutable Span structs in shared/platform/telemetry/span.go to record operational telemetry without raw payloads.
  • Aggressive content filtering: The AttributeAllowed function and content prefix checks ensure raw prompts and responses never persist (lines 98-119).
  • Credential protection: Automatic detection of authorization, apikey, and similar substrings prevents secret leakage (lines 33-41).
  • Hard size limits: Constants enforcing maximum attribute counts and byte sizes mitigate storage risks (lines 15-44).
  • Provenance tracking: ApplyProvenance adds source metadata without exposing sensitive execution context (lines 46-96).

Frequently Asked Questions

Does Caveman telemetry store raw LLM prompts?

No. The telemetry system explicitly strips any attribute matching content-bearing prefixes such as gen_ai.input.messages or gen_ai.output.messages (lines 98-119 in shared/platform/telemetry/span.go). Only token counts, model identifiers, and other metadata are retained, as documented in the repository's security guidelines.

How does Caveman detect and remove API keys from telemetry?

The AttributeAllowed function scans attribute keys for credential-related substrings including apikey, api_key, authorization, secret, password, access_token, and refresh_token (lines 33-41). Any matching key is discarded before the span is serialized, ensuring that authentication materials never reach the telemetry backend.

What happens if a span exceeds attribute size limits?

The SanitizeAttributes function enforces bounds defined by constants at lines 15-44, including MaxAttributesPerSpan, MaxAttributeKeyLength, and MaxAttributeValueLength. Attributes violating these limits are truncated or removed entirely to prevent storage abuse and potential denial-of-service vectors.

Is the telemetry data structure immutable once created?

While the Span struct itself is designed as an immutable record, the Attributes map is explicitly sanitized via SanitizeAttributes before persistence. The provenance metadata is applied through ApplyProvenance (lines 46-96), after which the span is treated as read-only for downstream consumers and telemetry sinks.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →