How AgentsView Handles Model Pricing with LiteLLM: Architecture and Implementation
AgentsView implements a resilient three-tier pricing system that combines static fallback rates, live LiteLLM API fetching, and SQLite persistence with cooldown-protected refresh logic to ensure accurate cost calculations even when offline.
AgentsView is an LLM observability tool that tracks token usage and costs across multiple providers. The system stores per-model token prices in a local SQLite model_pricing table and integrates with LiteLLM's public pricing catalog to maintain up-to-date rates while providing reliable offline functionality.
Static Fallback Catalog in internal/pricing/fallback.go
AgentsView maintains a hard-coded catalog of token prices for essential models including Claude, OpenAI, and Mistral families. This static data serves as the authoritative baseline when network connectivity is unavailable.
The fallback catalog is versioned using the FallbackVersion constant. This versioning prevents unnecessary database writes: the seeder compares the stored meta key "_fallback_version" against the current code version before upserting rows. Only when versions differ does the system rewrite the fallback pricing data, minimizing disk I/O during routine restarts.
Fetching Live Prices from LiteLLM
When network connectivity is available, AgentsView pulls the public LiteLLM price file model_prices_and_context_window.json directly from GitHub. This logic resides in internal/pricing/litellm.go.
The Fetch Operation
The FetchLiteLLMPricing function creates a short-lived HTTP client to GET the JSON payload:
// Pull the latest LiteLLM catalog (used by the background refresh)
prices, err := pricing.FetchLiteLLMPricing()
if err != nil {
log.Printf("pricing refresh: litellm fetch failed: %v", err)
return
}
if err := upsertPricing(database, prices); err != nil {
log.Printf("pricing refresh: upsert failed: %v", err)
}
Parsing and Normalization
The ParseLiteLLMPricing function (lines 60-94) handles JSON unmarshaling and unit conversion:
- Converts per-token costs from LiteLLM into per-million-token costs (
InputPerMTok,OutputPerMTok) - Filters out entries lacking both input and output rates to prevent incomplete pricing records
- Returns a slice of
pricing.ModelPricingstructs ready for database insertion
// Parsing LiteLLM JSON – converts per-token to per-million-token costs
func ParseLiteLLMPricing(data []byte) ([]ModelPricing, error) {
var raw map[string]litellmEntry
if err := json.Unmarshal(data, &raw); err != nil {
return nil, fmt.Errorf("parsing litellm JSON: %w", err)
}
// … conversion loop …
}
Database Seeding and Background Refresh
The pricing lifecycle begins in cmd/agentsview/usage.go through the seedPricing function (lines 73-99).
Server Startup Sequence
- Version Check: Reads the stored meta key
"_fallback_version"and compares it againstpricing.FallbackVersion - Fallback Upsert: If versions differ, calls
upsertPricingto populate the database with static rates - Background Goroutine: Launches
refreshPricingFromLiteLLMto fetch and update LiteLLM rates asynchronously without blocking startup
This pattern ensures the application remains functional immediately while attempting to enrich data in the background.
CLI Refresh Logic with Cooldown Protection
For CLI-driven usage queries, the ensurePricing function implements intelligent offline/online detection:
// CLI fallback path – used when the user runs “agentsview usage” offline
func ensurePricing(database *db.DB, offline bool) {
var prices []pricing.ModelPricing
if offline {
prices = pricing.FallbackPricing()
} else {
var err error
prices, err = pricing.FetchLiteLLMPricing()
if err != nil {
fmt.Fprintf(os.Stderr,
"warning: pricing fetch failed: %v; using fallback\n", err)
prices = pricing.FallbackPricing()
}
}
_ = upsertPricing(database, prices) // ignore errors for brevity
}
Cooldown Mechanism
The refreshPricingIfStale function (lines 31-71) prevents aggressive network polling using:
- Meta key tracking: Stores
"_litellm_last_attempt"timestamp in SQLite - 1-hour cooldown: Defined by
pricingRefreshCooldownconstant - Conditional refresh: Only fetches new data when the cooldown period has elapsed
For data-directory initialization, insertMissingPricing performs non-destructive inserts, ensuring fallback rows populate only when a model entry is absent.
Cost Calculation and Display
When generating usage reports via printDailyTable in cmd/agentsview/usage.go, the system:
- Retrieves token counts from
db.DailyUsageResult - Multiplies usage by the per-million-token rates stored in
model_pricing - Displays the calculated cost in the COST column of CLI output
This calculation uses the normalized InputPerMTok and OutputPerMTok values, ensuring consistent cost reporting regardless of whether the pricing data originated from LiteLLM or the static fallback catalog.
Summary
- Three-tier architecture: Static fallback in
internal/pricing/fallback.go, live LiteLLM fetching ininternal/pricing/litellm.go, and SQLite persistence viacmd/agentsview/usage.go - Version-aware seeding: The
FallbackVersionconstant prevents redundant database writes during application restarts - Resilient fetching:
FetchLiteLLMPricingconverts LiteLLM's per-token JSON into per-million-token database records with automatic filtering of incomplete entries - Cooldown protection: The
_litellm_last_attemptmeta key enforces a 1-hour refresh window to minimize unnecessary network requests - Graceful degradation: CLI commands automatically fall back to static pricing when
offlinemode is enabled or GitHub requests fail
Frequently Asked Questions
How does AgentsView calculate costs when the LiteLLM API is unreachable?
When FetchLiteLLMPricing returns an error, the system logs a warning to stderr and immediately falls back to the static catalog via pricing.FallbackPricing(). This ensures cost calculations remain available using the hard-coded rates in internal/pricing/fallback.go even during complete network outages.
What prevents AgentsView from constantly refreshing pricing data?
The refreshPricingIfStale function implements a cooldown mechanism using the SQLite meta key "_litellm_last_attempt". The system compares the current time against this stored timestamp and only permits fresh fetches when more than pricingRefreshCooldown (1 hour) has elapsed since the last attempt.
How does the pricing data format differ between LiteLLM and AgentsView's database?
LiteLLM provides per-token costs in its JSON API, but AgentsView stores per-million-token rates (InputPerMTok, OutputPerMTok) in the SQLite model_pricing table. The ParseLiteLLMPricing function handles this conversion during ingestion to simplify runtime cost calculations.
Where does AgentsView store its fallback pricing catalog?
The static fallback catalog is hard-coded in the Go source file internal/pricing/fallback.go and embedded directly into the binary. This ensures pricing data is always available without external dependencies, and the FallbackVersion constant allows the application to detect when this embedded catalog has been updated between releases.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →