OpenHuman Observability: Run Journals, Cost Accounting, and Sentry Integration

OpenHuman implements a layered observability subsystem that records every agent turn in run journals, classifies errors to filter transient failures, and integrates with Sentry for error tracking while maintaining accurate cost accounting per user.

OpenHuman is an open-source AI agent framework built with observability as a first-class concern. The OpenHuman observability architecture centralizes error classification, turn-based journaling, and credit tracking within the core crate, ensuring consistent debugging and billing logic across voice, inference, and chat domains.

Core Observability Architecture

The observability engine lives in src/core/observability.rs and serves as the single source of truth for error handling decisions. This module exposes functions such as report_error (line 2011) and report_error_or_expected (line 2032) that determine whether an exception should be forwarded to Sentry, logged as a budget event, or suppressed entirely.

The core provides predicate functions to classify failures semantically:

  • is_transient_provider_http_failure (line 2741) – Detects temporary network glitches that should not trigger alerts.
  • is_budget_event (line 3245) – Identifies quota-related messages that affect user billing.
  • is_quota_exhausted_message (line 3268) and is_quota_exhausted_event (line 3353) – Signal hard credit limits that terminate sessions.
  • backend_error_code_skips_sentry (line 212) and managed_error_skips_sentry (line 135) – Filter provider-specific noise.

Each predicate is exercised by the test suite (see lines 2988–3004 in src/core/observability.rs), ensuring that classification logic remains stable across releases.

Configuration and Sentry Setup

Observability behavior is controlled by the ObservabilityConfig struct defined in src/openhuman/config/schema/observability.rs. This schema supports TOML or environment-based configuration of the Sentry DSN, analytics toggles, and budget tracking flags.

During startup, the core initializes Sentry only when a valid DSN is present; otherwise, it installs a no-op transport for CI environments (see src/core/sentry_transport.rs).


# config.toml

[observability]
sentry_dsn = "https://public@sentry.example.com/1"
analytics_enabled = true
// src/openhuman/platform/app_state/ops.rs (line 704)
let cfg = Config::load_or_init()?;
if let Some(dsn) = &cfg.observability.sentry_dsn {
    let opts = sentry::ClientOptions {
        dsn: Some(dsn.parse().unwrap()),
        ..Default::default()
    };
    let _guard = sentry::init(opts);
}

The minimal transport implementation in src/core/sentry_transport.rs ensures that binaries compiled without the full Sentry SDK can still capture and forward events during testing.

Run Journals and Turn Tracking

Every agent turn is wrapped in a run journal that records execution metadata. The AgentObservation struct defined in src/openhuman/agent/tinyagents/observability.rs tracks start and end timestamps, token consumption, and error classifications. These journals are emitted to the UI, stored in the core log, and optionally forwarded to Sentry with PII stripped by src/openhuman/security/credentials/sentry_scope.rs.

use tinyagents::harness::observability::AgentObservation;

fn run_turn(input: &str) -> Result<String, anyhow::Error> {
    let mut journal = AgentObservation::new("my_turn");
    
    // ... perform inference, tool calls, etc.
    journal.record_token_usage(123);
    
    // On error:
    if let Err(e) = risky_operation() {
        journal.error("provider timeout");
        crate::core::observability::report_error_or_expected(
            &e, "inference", "run_turn", &[]
        )?;
    }
    
    Ok(journal.finalize())
}

Domain modules such as src/openhuman/web_chat/ops.rs and src/openhuman/platform/socket/ws_loop.rs invoke report_error or report_error_or_expected to close journals consistently, ensuring that turn-level telemetry captures both successful completions and classified failures.

Cost Accounting and Quota Management

The cost tracker in src/openhuman/platform/cost/tracker.rs aggregates token usage, API-key limits, and credit consumption per user. It consults the same observability predicates used by Sentry to decide whether a failed request should consume credits.

The CostTracker exposes methods like track_usage (line 144) and record_quota_exhaustion that distinguish between billable and non-billable events. For example, is_quota_exhausted_message (line 3336) signals a hard quota breach that is billed, whereas is_transient_provider_transport_failure (line 2797) ensures temporary network glitches do not drain user credits.

use openhuman::platform::cost::tracker::CostTracker;
use uuid::Uuid;

fn charge_usage(user_id: Uuid, tokens: u64) {
    let tracker = CostTracker::global();
    // Only bills if the observability layer confirms non-transient status
    tracker.record_usage(user_id, tokens);
}

This integration prevents users from paying for provider downtime or transient HTTP timeouts while ensuring that legitimate quota violations are tracked accurately.

Error Classification in Practice

Domain code uses the core observability API to avoid polluting Sentry with expected failures. When handling HTTP errors, developers first check is_transient_provider_http_failure before invoking report_error.

use crate::core::observability::{
    report_error, is_transient_provider_http_failure
};

fn handle_http_error(
    err: reqwest::Error, 
    domain: &str, 
    operation: &str
) {
    // Skip Sentry for temporary network glitches
    if is_transient_provider_http_failure(&err.to_string()) {
        tracing::debug!("Transient error ignored for observability");
        return;
    }
    
    // Forward actionable errors to Sentry with domain tags
    report_error(&err, domain, operation, &[]);
}

The event forwarded to Sentry includes tags such as domain, operation, and optional budget flags, enabling precise filtering in the Sentry UI. The is_updater_transient_http_status predicate (line 3000) provides similar protection for the auto-updater component, ensuring that version-check timeouts do not generate alert fatigue.

Summary

  • Centralized classification: The src/core/observability.rs module provides a single, tested source of truth for error semantics across all domains.
  • Selective Sentry reporting: Transient failures and quota events are filtered by predicates like is_transient_provider_http_failure and is_budget_event to reduce noise.
  • Run journals: Every turn generates an AgentObservation that captures timing, tokens, and errors for UI rendering and debugging.
  • Accurate billing: The cost tracker in src/openhuman/platform/cost/tracker.rs uses the same classification logic to avoid charging users for provider-side failures.
  • Flexible configuration: The ObservabilityConfig schema supports runtime toggling of Sentry, analytics, and budget tracking without recompilation.

Frequently Asked Questions

How does OpenHuman prevent transient errors from flooding Sentry?

OpenHuman uses predicate functions in src/core/observability.rs such as is_transient_provider_http_failure (line 2741) and is_updater_transient_http_status (line 3000) to classify errors before reporting. When an error matches these patterns, the report_error function returns early without constructing a Sentry event, ensuring that temporary network glitches do not create alert noise.

How does cost accounting distinguish between billable and non-billable failures?

The cost tracker in src/openhuman/platform/cost/tracker.rs consults observability predicates to determine billing status. Failures classified by is_transient_provider_transport_failure (line 2797) are excluded from user charges, while is_quota_exhausted_message (line 3336) events are recorded as billable quota breaches. This ensures users only pay for successful inference or legitimate limit violations.

What data structure does OpenHuman use for run journals?

OpenHuman uses the AgentObservation struct defined in src/openhuman/agent/tinyagents/observability.rs to implement run journals. This structure records turn identifiers, start/end timestamps, token usage metrics, and error classifications. Domain code interacts with journals through methods like record_token_usage and finalize, while the core observability layer handles Sentry forwarding and log persistence.

How do you configure Sentry integration in OpenHuman?

Sentry is configured via the ObservabilityConfig struct in src/openhuman/config/schema/observability.rs. Set the sentry_dsn field to a valid DSN string in your TOML configuration or environment variables. If the DSN is absent, OpenHuman automatically installs a minimal transport from src/core/sentry_transport.rs to maintain compatibility in CI environments without sending data to external servers.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →