How AgentsView Handles Model Pricing with LiteLLM: Architecture and Implementation

AgentsView implements a resilient three-tier pricing system that combines static fallback rates, live LiteLLM API fetching, and SQLite persistence with cooldown-protected refresh logic to ensure accurate cost calculations even when offline.

AgentsView is an LLM observability tool that tracks token usage and costs across multiple providers. The system stores per-model token prices in a local SQLite model_pricing table and integrates with LiteLLM's public pricing catalog to maintain up-to-date rates while providing reliable offline functionality.

Static Fallback Catalog in internal/pricing/fallback.go

AgentsView maintains a hard-coded catalog of token prices for essential models including Claude, OpenAI, and Mistral families. This static data serves as the authoritative baseline when network connectivity is unavailable.

The fallback catalog is versioned using the FallbackVersion constant. This versioning prevents unnecessary database writes: the seeder compares the stored meta key "_fallback_version" against the current code version before upserting rows. Only when versions differ does the system rewrite the fallback pricing data, minimizing disk I/O during routine restarts.

Fetching Live Prices from LiteLLM

When network connectivity is available, AgentsView pulls the public LiteLLM price file model_prices_and_context_window.json directly from GitHub. This logic resides in internal/pricing/litellm.go.

The Fetch Operation

The FetchLiteLLMPricing function creates a short-lived HTTP client to GET the JSON payload:

// Pull the latest LiteLLM catalog (used by the background refresh)
prices, err := pricing.FetchLiteLLMPricing()
if err != nil {
    log.Printf("pricing refresh: litellm fetch failed: %v", err)
    return
}
if err := upsertPricing(database, prices); err != nil {
    log.Printf("pricing refresh: upsert failed: %v", err)
}

Parsing and Normalization

The ParseLiteLLMPricing function (lines 60-94) handles JSON unmarshaling and unit conversion:

  • Converts per-token costs from LiteLLM into per-million-token costs (InputPerMTok, OutputPerMTok)
  • Filters out entries lacking both input and output rates to prevent incomplete pricing records
  • Returns a slice of pricing.ModelPricing structs ready for database insertion
// Parsing LiteLLM JSON – converts per-token to per-million-token costs
func ParseLiteLLMPricing(data []byte) ([]ModelPricing, error) {
    var raw map[string]litellmEntry
    if err := json.Unmarshal(data, &raw); err != nil {
        return nil, fmt.Errorf("parsing litellm JSON: %w", err)
    }
    // … conversion loop …
}

Database Seeding and Background Refresh

The pricing lifecycle begins in cmd/agentsview/usage.go through the seedPricing function (lines 73-99).

Server Startup Sequence

  1. Version Check: Reads the stored meta key "_fallback_version" and compares it against pricing.FallbackVersion
  2. Fallback Upsert: If versions differ, calls upsertPricing to populate the database with static rates
  3. Background Goroutine: Launches refreshPricingFromLiteLLM to fetch and update LiteLLM rates asynchronously without blocking startup

This pattern ensures the application remains functional immediately while attempting to enrich data in the background.

CLI Refresh Logic with Cooldown Protection

For CLI-driven usage queries, the ensurePricing function implements intelligent offline/online detection:

// CLI fallback path – used when the user runs “agentsview usage” offline
func ensurePricing(database *db.DB, offline bool) {
    var prices []pricing.ModelPricing
    if offline {
        prices = pricing.FallbackPricing()
    } else {
        var err error
        prices, err = pricing.FetchLiteLLMPricing()
        if err != nil {
            fmt.Fprintf(os.Stderr,
                "warning: pricing fetch failed: %v; using fallback\n", err)
            prices = pricing.FallbackPricing()
        }
    }
    _ = upsertPricing(database, prices) // ignore errors for brevity
}

Cooldown Mechanism

The refreshPricingIfStale function (lines 31-71) prevents aggressive network polling using:

  • Meta key tracking: Stores "_litellm_last_attempt" timestamp in SQLite
  • 1-hour cooldown: Defined by pricingRefreshCooldown constant
  • Conditional refresh: Only fetches new data when the cooldown period has elapsed

For data-directory initialization, insertMissingPricing performs non-destructive inserts, ensuring fallback rows populate only when a model entry is absent.

Cost Calculation and Display

When generating usage reports via printDailyTable in cmd/agentsview/usage.go, the system:

  1. Retrieves token counts from db.DailyUsageResult
  2. Multiplies usage by the per-million-token rates stored in model_pricing
  3. Displays the calculated cost in the COST column of CLI output

This calculation uses the normalized InputPerMTok and OutputPerMTok values, ensuring consistent cost reporting regardless of whether the pricing data originated from LiteLLM or the static fallback catalog.

Summary

  • Three-tier architecture: Static fallback in internal/pricing/fallback.go, live LiteLLM fetching in internal/pricing/litellm.go, and SQLite persistence via cmd/agentsview/usage.go
  • Version-aware seeding: The FallbackVersion constant prevents redundant database writes during application restarts
  • Resilient fetching: FetchLiteLLMPricing converts LiteLLM's per-token JSON into per-million-token database records with automatic filtering of incomplete entries
  • Cooldown protection: The _litellm_last_attempt meta key enforces a 1-hour refresh window to minimize unnecessary network requests
  • Graceful degradation: CLI commands automatically fall back to static pricing when offline mode is enabled or GitHub requests fail

Frequently Asked Questions

How does AgentsView calculate costs when the LiteLLM API is unreachable?

When FetchLiteLLMPricing returns an error, the system logs a warning to stderr and immediately falls back to the static catalog via pricing.FallbackPricing(). This ensures cost calculations remain available using the hard-coded rates in internal/pricing/fallback.go even during complete network outages.

What prevents AgentsView from constantly refreshing pricing data?

The refreshPricingIfStale function implements a cooldown mechanism using the SQLite meta key "_litellm_last_attempt". The system compares the current time against this stored timestamp and only permits fresh fetches when more than pricingRefreshCooldown (1 hour) has elapsed since the last attempt.

How does the pricing data format differ between LiteLLM and AgentsView's database?

LiteLLM provides per-token costs in its JSON API, but AgentsView stores per-million-token rates (InputPerMTok, OutputPerMTok) in the SQLite model_pricing table. The ParseLiteLLMPricing function handles this conversion during ingestion to simplify runtime cost calculations.

Where does AgentsView store its fallback pricing catalog?

The static fallback catalog is hard-coded in the Go source file internal/pricing/fallback.go and embedded directly into the binary. This ensures pricing data is always available without external dependencies, and the FallbackVersion constant allows the application to detect when this embedded catalog has been updated between releases.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →