How the Apache Maka Model Call Ledger Tracks Token Usage and Attributes Costs

The Apache Maka Model Call Ledger is a SQLite-based projection that records every LLM invocation's token counts and monetary costs by consuming AgentRun events, applying pricing rules from usage-pricing.ts, and storing structured ModelCallAttempt records for queryable usage analytics.

The Model Call Ledger is a critical component of Apache Maka's telemetry and billing infrastructure. It provides a durable, auditable record of all model interactions without duplicating authoritative state. This guide explains how the ledger captures token metrics, computes costs using the pricing engine, and exposes that data for downstream consumption.

What the Model Call Ledger Stores

At its core, the ledger is a read-only projection of the AgentRun event stream. Rather than acting as a primary data store, it materializes relevant slices of the event log into an efficient queryable format.

Each model call produces a ModelCallAttempt object containing:

Field Description
inputTokens Tokens sent to the model (prompt)
outputTokens Tokens received from the model (completion)
totalTokens Combined token count
cacheReadTokens, cacheWriteTokens Optional cache utilization metrics
costBasis Either 'unpriced' or a priced basis identifier
costUsd Computed dollar cost (undefined if unpriced)
attemptId, sessionId, runId Execution context identifiers
completedAt Unix timestamp of call completion

The ledger persists these records in the usage_model_call_attempts SQLite table, with the full JSON representation stored in the record_json column.

Recording Model Calls via Event Projection

The ledger consumes events through the catchUpProjection() method implemented in packages/storage/src/model-call-ledger.ts. This mechanism ensures the ledger stays synchronized with the authoritative event stream.

Projection Workflow

  1. Read checkpoint – The ledger queries usage_model_call_projection_checkpoints to find the last processed sequence number
  2. Fetch new events – It requests AgentRun events with sequence numbers greater than the checkpoint
  3. Deserialize attempts – For each MODEL_CALL_ATTEMPT_EVENT_TYPE event, it extracts the ModelCallAttempt from the payload
  4. Validate context – Verifies that sessionId and runId exist and match expected formats
  5. Upsert record – Writes validated attempts using attemptId as the idempotent key
  6. Update checkpoint – Persists the new high-water mark and cumulative unreadable event count

This design guarantees exactly-once processing semantics: duplicate events with the same attemptId overwrite previous entries without creating duplicates.

// From packages/storage/src/model-call-ledger.ts
// Core projection loop (conceptual)

async catchUpProjection(): Promise<void> {
  const checkpoint = await this.getCheckpoint();
  
  const events = await this.eventStore.readRange({
    fromSeq: checkpoint.lastSequence,
    eventTypes: [MODEL_CALL_ATTEMPT_EVENT_TYPE]
  });

  let unreadableCount = checkpoint.unreadableCount;

  for (const event of events) {
    try {
      const attempt = ModelCallAttemptSchema.parse(event.payload);
      await this.upsertAttempt(attempt);
    } catch (parseError) {
      // Corrupt events are counted but not re-thrown
      unreadableCount++;
    }
  }

  await this.writeCheckpoint({
    lastSequence: events.at(-1)?.sequence ?? checkpoint.lastSequence,
    unreadableCount
  });
}

Computing Token Usage Costs

Cost attribution happens in two stages: token counting during the call itself, and price application via the runtime telemetry system.

Stage 1: Token Collection

The runtime's AI SDK backend (packages/runtime/src/ai-sdk-backend.ts) interfaces with model providers and extracts token usage from their responses. This occurs in functions that wrap provider SDKs and normalize responses into the Maka-internal format.

Stage 2: Cost Calculation

The computeTokenUsageCostUsd() function applies pricing formulas using rates defined in packages/runtime-host/src/protocol/usage-pricing.ts.

// From packages/runtime/src/ai-sdk-backend.ts
// Simplified cost computation flow

import { PricingBucket, computeTokenUsageCostUsd } from '@maka/runtime-host/protocol';

function finalizeModelCallAttempt(
  rawAttempt: Omit<ModelCallAttempt, 'costUsd'>,
  pricingBucket: PricingBucket | undefined
): ModelCallAttempt {
  
  if (!pricingBucket) {
    return {
      ...rawAttempt,
      costBasis: 'unpriced',
      costUsd: undefined
    };
  }

  const costUsd = computeTokenUsageCostUsd({
    inputTokens: rawAttempt.inputTokens,
    outputTokens: rawAttempt.outputTokens,
    cacheReadTokens: rawAttempt.cacheReadTokens,
    cacheWriteTokens: rawAttempt.cacheWriteTokens
  }, pricingBucket.rates);

  return {
    ...rawAttempt,
    costBasis: pricingBucket.basisId,
    costUsd
  };
}

The usage-pricing.ts module defines pricing buckets with per-token rates that may vary by:

  • Model provider (OpenAI, Anthropic, etc.)
  • Model tier (GPT-4, Claude-3, etc.)
  • Token type (input vs. output, cache hit vs. miss)
  • Time period (promotional rates, volume discounts)

When Pricing Is Unavailable

If no pricing bucket matches a call's context, the ledger still records the attempt with costBasis: 'unpriced' and costUsd: undefined. These records contribute to the unreadable count returned by ledger queries, flagging them for manual review or later backfill.

Querying the Ledger for Usage Analytics

The ledger exposes a read() method that returns paginated, time-bounded results with full cost aggregation support.

// From packages/storage/src/model-call-ledger.ts
// Public read API

interface ModelCallLedgerPage {
  attempts: ModelCallAttempt[];
  unreadableCount: number;  // Events that failed parsing in this range
  hasMore: boolean;
  nextCursor?: string;
}

class SqliteModelCallLedger {
  async read(
    range: { from: number; to: number },  // Unix timestamps
    sessionId?: string                     // Optional filter
  ): Promise<ModelCallLedgerPage>;
}

Practical Aggregation Example

import { createSqliteModelCallLedger } from '@maka/storage';

const ledger = createSqliteModelCallLedger(workspaceRoot);

// Usage report for a billing period
const page = await ledger.read(
  { from: Date.parse('2024-01-01'), to: Date.parse('2024-02-01') },
  'session-customer-abc-123'
);

const totals = page.attempts.reduce((acc, attempt) => ({
  inputTokens: acc.inputTokens + (attempt.inputTokens || 0),
  outputTokens: acc.outputTokens + (attempt.outputTokens || 0),
  costUsd: acc.costUsd + (attempt.costUsd || 0),
  unpricedCalls: acc.unpricedCalls + (attempt.costUsd === undefined ? 1 : 0)
}), { inputTokens: 0, outputTokens: 0, costUsd: 0, unpricedCalls: 0 });

console.log(`Total cost: $${totals.costUsd.toFixed(4)}`);
console.log(`Unpriced calls requiring review: ${totals.unpricedCalls}`);

Key Design Guarantees

The Model Call Ledger architecture provides several operational benefits:

  • Idempotent writes – attemptId upserts prevent duplicates during re-projection
  • Crash recovery – Checkpoint-based resume ensures no events are lost or double-counted
  • Auditability – Raw JSON preservation allows reconstruction of exact runtime state
  • Extensibility – New token types or cost dimensions can be added to ModelCallAttempt without schema migrations

Summary

  • The Model Call Ledger in packages/storage/src/model-call-ledger.ts projects AgentRun events into SQLite for efficient usage queries
  • Token usage is captured in ModelCallAttempt records with fields for input, output, and cache tokens
  • Costs are computed by computeTokenUsageCostUsd() in the runtime, using pricing buckets from packages/runtime-host/src/protocol/usage-pricing.ts
  • The ledger handles unpriced calls gracefully, tracking them via costBasis: 'unpriced' and unreadable counters
  • Checkpoint-based projection ensures exactly-once processing with automatic crash recovery

Frequently Asked Questions

What happens if pricing information changes after a call is recorded?

The ledger stores the costBasis identifier, not the rates themselves. If pricing is corrected retroactively, re-projecting the event stream with updated buckets will recalculate costUsd for affected calls. Historical costBasis values remain stable, enabling audit trails.

Can the ledger be used without SQLite?

The current implementation in packages/storage/src/model-call-ledger.ts is SQLite-specific, but the projection interface is abstracted. Alternative backends would need to implement the checkpoint, upsert, and read operations while maintaining the same idempotency guarantees.

How does the ledger handle event stream corruption?

Corrupt events increment the unreadableCount but do not halt projection. The checkpoint advances past them, and the read() API returns the per-query unreadable count alongside valid attempts. This transparency allows operators to detect data quality issues without losing throughput.

Is the Model Call Ledger suitable for real-time billing?

The ledger's eventual consistency model suits billing reconciliation better than real-time authorization. For sub-second latency requirements, the runtime's in-memory telemetry (packages/runtime/src/telemetry/) provides faster signals, with the ledger serving as the authoritative audit source.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →