How Agentsview Tracks Token Usage and Calculates LLM Costs: A Deep Dive
Agentsview records every LLM request in SQLite or PostgreSQL, normalizes token counts from JSON blobs or explicit events, and multiplies them against LiteLLM pricing rates to compute per-row costs that aggregate into project, model, and agent totals.
The agentsview repository provides a comprehensive mechanism to track token usage and calculate LLM costs across AI agent sessions. By intercepting chat messages and explicit billing events, the system builds a unified usage view that supports both real-time CLI reporting and HTTP API consumption. Understanding this pipeline requires examining how raw token counts are persisted, queried, and monetized using external pricing data.
Where Agentsview Stores Token Usage Data
The system captures token consumption in two database tables, each serving different data sources.
The messages Table
When a session syncs chat completions, the messages table stores a token_usage JSON blob containing the raw counts. According to the schema analysis in internal/db/usage.go, this field typically includes:
{
"input_tokens": 123,
"output_tokens": 456,
"cache_creation_input_tokens": 12,
"cache_read_input_tokens": 34
}
Go code extracts these values using the gjson library rather than SQLite JSON functions, parsing the blob in the dailyUsageAmounts function around lines 998-1010.
The usage_events Table
For agents that report explicit billing events, the usage_events table stores denormalized columns: input_tokens, output_tokens, cache_creation_input_tokens, cache_read_input_tokens, and optionally cost_usd. This dual-storage approach ensures that both inferred usage from chat logs and explicit cost reports from external providers are tracked consistently.
Unified Usage Query Implementation
The heart of the aggregation logic lives in internal/db/usage.go. Here, a unified query merges both storage sources using the usageMessageEligibility and usageEventEligibility predicates (lines 26-31).
The query normalizes rows into a common shape containing:
session_idmodelinput_tokens,output_tokens,cache_creation_input_tokens,cache_read_input_tokenscost_usd(when pre-calculated)
This normalization allows the cost calculation layer to treat messages and events identically, regardless of whether they originated from chat logs or direct billing reports.
Cost Calculation Methodology
Agentsview computes costs in three stages: loading pricing maps, calculating per-row costs, and aggregating results.
Loading Pricing Maps from LiteLLM
At the start of every usage operation, the database loads per-model rates via loadPricingMap. The internal/pricing/litellm.go file (lines 36-94) fetches the JSON price list from https://raw.githubusercontent.com/BerriAI/litellm/main/model_prices_and_context_window.json and converts per-token costs to per-million-token rates.
For models not covered by LiteLLM, internal/pricing/fallback.go provides a static map of hard-coded rates, ensuring the system can estimate costs even for proprietary or newly released models.
Per-Row Cost Computation
Inside dailyUsageAmounts (lines 1018-1024 of internal/db/usage.go), the logic follows this priority:
- Explicit cost: If
r.costUSD.Validis true (from a usage_event), use that value directly. - Calculated cost: Otherwise, multiply each token count by the appropriate per-token rate from the pricing map and divide by 1,000,000.
The calculation accounts for all four token categories:
- Standard input/output tokens
- Cache creation tokens (typically cheaper)
- Cache read tokens (often discounted)
Aggregation and Sorting
After scanning rows, the system aggregates costs using helpers in internal/server/usage.go:
foldProjectTotals(groups by project)foldModelTotals(groups by model)foldAgentTotals(groups by agent)
These functions (lines 72-168) sort results by total cost descending, ensuring the UI surfaces the most expensive projects and models first. The aggregated data populates fields like TotalCost, CacheSavings, and CopilotAICredits in the UsageTotals struct.
API Endpoints for Usage Analytics
The HTTP layer exposes aggregated usage data through REST endpoints defined in the routes configuration (via huma_routes_usage.go).
GET /api/v1/usage/summary returns a UsageSummaryResponse containing:
- Date range totals
- Breakdowns by project, model, and agent
- Comparison metrics against prior periods
GET /api/v1/usage/top calls GetTopSessionsByCost (lines 701-820 of internal/server/usage.go) to identify high-cost sessions, using identical cost-calculation logic to maintain consistency across views.
For PostgreSQL deployments, internal/postgres/usage.go mirrors the SQLite logic, ensuring cost calculations remain identical regardless of storage backend.
Working with Usage Data in Practice
Querying Daily Usage via CLI
ctx := context.Background()
filter := db.UsageFilter{
From: "2024-06-01",
To: "2024-06-30",
Breakdowns: true, // include per-project/agent/model breakdowns
}
daily, err := dbInstance.GetDailyUsage(ctx, filter)
if err != nil {
log.Fatalf("usage error: %v", err)
}
fmt.Printf("Total cost for June: $%.2f\n", daily.Totals.TotalCost)
HTTP Response Structure
{
"from": "2024-06-01",
"to": "2024-06-30",
"totals": {
"inputTokens": 123456,
"outputTokens": 234567,
"cacheCreationTokens": 12000,
"cacheReadTokens": 34000,
"totalCost": 45.67,
"cacheSavings": 3.12
},
"projectTotals": [
{"project": "my-app", "inputTokens": 50000, "cost": 22.1}
],
"modelTotals": [
{"model": "gpt-4o", "inputTokens": 100000, "cost": 30.5}
],
"agentTotals": [
{"agent": "claude", "inputTokens": 80000, "cost": 25.0}
],
"comparison": {
"priorFrom": "2024-05-01",
"priorTo": "2024-05-31",
"priorTotalCost": 40.2,
"deltaPct": 14.2
}
}
Manual Cost Calculation (Fallback)
When operating without database access, you can replicate the pricing logic:
// Assume token counts from a message:
inputTok, outputTok, cacheCr, cacheRd := 200, 150, 0, 0
pricing, _ := pricingpkg.LoadDefault() // loads LiteLLM map
rates, _ := pricingpkg.Resolve(pricing, "gpt-4o") // rates for the model
cost := (float64(inputTok)*rates.input +
float64(outputTok)*rates.output +
float64(cacheCr)*rates.cacheCreation +
float64(cacheRd)*rates.cacheRead) / 1_000_000
fmt.Printf("Estimated cost: $%.5f\n", cost)
Summary
- Dual storage: Agentsview writes token counts to the
messagestable (as JSON) and theusage_eventstable (as explicit columns), unified by the query ininternal/db/usage.go. - External pricing: The
internal/pricing/litellm.gomodule downloads and caches per-model rates, withinternal/pricing/fallback.goproviding hard-coded backups. - Calculation logic: The
dailyUsageAmountsfunction divides token counts by 1,000,000 and multiplies by loaded rates, preferring explicitcost_usdvalues when available. - Consistent aggregation: Helpers in
internal/server/usage.gofold rows into project, model, and agent totals, sorting by cost for UI presentation. - API exposure: The
/api/v1/usage/summaryand/api/v1/usage/topendpoints serve these calculations to the web interface and CLI commands.
Frequently Asked Questions
How does Agentsview handle token counts from different LLM providers?
The system normalizes all token counts into four standard categories—input_tokens, output_tokens, cache_creation_input_tokens, and cache_read_input_tokens—regardless of provider. Whether the source is a JSON blob in the messages table or explicit columns in usage_events, the unified query in internal/db/usage.go extracts these values using gjson for messages or direct column access for events, ensuring consistent handling across OpenAI, Anthropic, and other providers.
What happens if a model is not listed in the LiteLLM pricing database?
Agentsview falls back to a static pricing map defined in internal/pricing/fallback.go. When loadPricingMap executes, it attempts to resolve rates from the LiteLLM JSON feed first; if a model is absent, the system checks the fallback map before returning a "pricing not found" error. This ensures cost estimates remain available for newly released or proprietary models not yet indexed by LiteLLM.
Can Agentsview use costs reported directly by an agent instead of calculating them?
Yes. The cost calculation logic in dailyUsageAmounts checks if r.costUSD.Valid before applying mathematical operations. If the usage_events table contains a pre-calculated cost_usd value—such as when an agent reports its own billing data from a cloud provider—that value is used directly without recalculating from token counts and pricing maps, preserving the external provider's exact billing amount.
How does the system maintain consistency between SQLite and PostgreSQL backends?
The repository maintains parallel implementations: internal/db/usage.go handles SQLite while internal/postgres/usage.go provides equivalent logic for PostgreSQL. Both implementations use identical SQL predicates (usageMessageEligibility, usageEventEligibility) and the same Go-based token extraction and cost calculation algorithms, ensuring that switching storage backends does not alter reported costs or usage statistics.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →