How the Apache Maka Model Call Ledger Tracks Token Usage and Attributes Costs
The Apache Maka Model Call Ledger is a SQLite-based projection that records every LLM invocation's token counts and monetary costs by consuming AgentRun events, applying pricing rules from usage-pricing.ts, and storing structured ModelCallAttempt records for queryable usage analytics.
The Model Call Ledger is a critical component of Apache Maka's telemetry and billing infrastructure. It provides a durable, auditable record of all model interactions without duplicating authoritative state. This guide explains how the ledger captures token metrics, computes costs using the pricing engine, and exposes that data for downstream consumption.
What the Model Call Ledger Stores
At its core, the ledger is a read-only projection of the AgentRun event stream. Rather than acting as a primary data store, it materializes relevant slices of the event log into an efficient queryable format.
Each model call produces a ModelCallAttempt object containing:
| Field | Description |
|---|---|
inputTokens |
Tokens sent to the model (prompt) |
outputTokens |
Tokens received from the model (completion) |
totalTokens |
Combined token count |
cacheReadTokens, cacheWriteTokens |
Optional cache utilization metrics |
costBasis |
Either 'unpriced' or a priced basis identifier |
costUsd |
Computed dollar cost (undefined if unpriced) |
attemptId, sessionId, runId |
Execution context identifiers |
completedAt |
Unix timestamp of call completion |
The ledger persists these records in the usage_model_call_attempts SQLite table, with the full JSON representation stored in the record_json column.
Recording Model Calls via Event Projection
The ledger consumes events through the catchUpProjection() method implemented in packages/storage/src/model-call-ledger.ts. This mechanism ensures the ledger stays synchronized with the authoritative event stream.
Projection Workflow
- Read checkpoint – The ledger queries
usage_model_call_projection_checkpointsto find the last processed sequence number - Fetch new events – It requests AgentRun events with sequence numbers greater than the checkpoint
- Deserialize attempts – For each
MODEL_CALL_ATTEMPT_EVENT_TYPEevent, it extracts theModelCallAttemptfrom the payload - Validate context – Verifies that
sessionIdandrunIdexist and match expected formats - Upsert record – Writes validated attempts using
attemptIdas the idempotent key - Update checkpoint – Persists the new high-water mark and cumulative unreadable event count
This design guarantees exactly-once processing semantics: duplicate events with the same attemptId overwrite previous entries without creating duplicates.
// From packages/storage/src/model-call-ledger.ts
// Core projection loop (conceptual)
async catchUpProjection(): Promise<void> {
const checkpoint = await this.getCheckpoint();
const events = await this.eventStore.readRange({
fromSeq: checkpoint.lastSequence,
eventTypes: [MODEL_CALL_ATTEMPT_EVENT_TYPE]
});
let unreadableCount = checkpoint.unreadableCount;
for (const event of events) {
try {
const attempt = ModelCallAttemptSchema.parse(event.payload);
await this.upsertAttempt(attempt);
} catch (parseError) {
// Corrupt events are counted but not re-thrown
unreadableCount++;
}
}
await this.writeCheckpoint({
lastSequence: events.at(-1)?.sequence ?? checkpoint.lastSequence,
unreadableCount
});
}
Computing Token Usage Costs
Cost attribution happens in two stages: token counting during the call itself, and price application via the runtime telemetry system.
Stage 1: Token Collection
The runtime's AI SDK backend (packages/runtime/src/ai-sdk-backend.ts) interfaces with model providers and extracts token usage from their responses. This occurs in functions that wrap provider SDKs and normalize responses into the Maka-internal format.
Stage 2: Cost Calculation
The computeTokenUsageCostUsd() function applies pricing formulas using rates defined in packages/runtime-host/src/protocol/usage-pricing.ts.
// From packages/runtime/src/ai-sdk-backend.ts
// Simplified cost computation flow
import { PricingBucket, computeTokenUsageCostUsd } from '@maka/runtime-host/protocol';
function finalizeModelCallAttempt(
rawAttempt: Omit<ModelCallAttempt, 'costUsd'>,
pricingBucket: PricingBucket | undefined
): ModelCallAttempt {
if (!pricingBucket) {
return {
...rawAttempt,
costBasis: 'unpriced',
costUsd: undefined
};
}
const costUsd = computeTokenUsageCostUsd({
inputTokens: rawAttempt.inputTokens,
outputTokens: rawAttempt.outputTokens,
cacheReadTokens: rawAttempt.cacheReadTokens,
cacheWriteTokens: rawAttempt.cacheWriteTokens
}, pricingBucket.rates);
return {
...rawAttempt,
costBasis: pricingBucket.basisId,
costUsd
};
}
The usage-pricing.ts module defines pricing buckets with per-token rates that may vary by:
- Model provider (OpenAI, Anthropic, etc.)
- Model tier (GPT-4, Claude-3, etc.)
- Token type (input vs. output, cache hit vs. miss)
- Time period (promotional rates, volume discounts)
When Pricing Is Unavailable
If no pricing bucket matches a call's context, the ledger still records the attempt with costBasis: 'unpriced' and costUsd: undefined. These records contribute to the unreadable count returned by ledger queries, flagging them for manual review or later backfill.
Querying the Ledger for Usage Analytics
The ledger exposes a read() method that returns paginated, time-bounded results with full cost aggregation support.
// From packages/storage/src/model-call-ledger.ts
// Public read API
interface ModelCallLedgerPage {
attempts: ModelCallAttempt[];
unreadableCount: number; // Events that failed parsing in this range
hasMore: boolean;
nextCursor?: string;
}
class SqliteModelCallLedger {
async read(
range: { from: number; to: number }, // Unix timestamps
sessionId?: string // Optional filter
): Promise<ModelCallLedgerPage>;
}
Practical Aggregation Example
import { createSqliteModelCallLedger } from '@maka/storage';
const ledger = createSqliteModelCallLedger(workspaceRoot);
// Usage report for a billing period
const page = await ledger.read(
{ from: Date.parse('2024-01-01'), to: Date.parse('2024-02-01') },
'session-customer-abc-123'
);
const totals = page.attempts.reduce((acc, attempt) => ({
inputTokens: acc.inputTokens + (attempt.inputTokens || 0),
outputTokens: acc.outputTokens + (attempt.outputTokens || 0),
costUsd: acc.costUsd + (attempt.costUsd || 0),
unpricedCalls: acc.unpricedCalls + (attempt.costUsd === undefined ? 1 : 0)
}), { inputTokens: 0, outputTokens: 0, costUsd: 0, unpricedCalls: 0 });
console.log(`Total cost: $${totals.costUsd.toFixed(4)}`);
console.log(`Unpriced calls requiring review: ${totals.unpricedCalls}`);
Key Design Guarantees
The Model Call Ledger architecture provides several operational benefits:
- Idempotent writes –
attemptIdupserts prevent duplicates during re-projection - Crash recovery – Checkpoint-based resume ensures no events are lost or double-counted
- Auditability – Raw JSON preservation allows reconstruction of exact runtime state
- Extensibility – New token types or cost dimensions can be added to
ModelCallAttemptwithout schema migrations
Summary
- The Model Call Ledger in
packages/storage/src/model-call-ledger.tsprojects AgentRun events into SQLite for efficient usage queries - Token usage is captured in
ModelCallAttemptrecords with fields for input, output, and cache tokens - Costs are computed by
computeTokenUsageCostUsd()in the runtime, using pricing buckets frompackages/runtime-host/src/protocol/usage-pricing.ts - The ledger handles unpriced calls gracefully, tracking them via
costBasis: 'unpriced'and unreadable counters - Checkpoint-based projection ensures exactly-once processing with automatic crash recovery
Frequently Asked Questions
What happens if pricing information changes after a call is recorded?
The ledger stores the costBasis identifier, not the rates themselves. If pricing is corrected retroactively, re-projecting the event stream with updated buckets will recalculate costUsd for affected calls. Historical costBasis values remain stable, enabling audit trails.
Can the ledger be used without SQLite?
The current implementation in packages/storage/src/model-call-ledger.ts is SQLite-specific, but the projection interface is abstracted. Alternative backends would need to implement the checkpoint, upsert, and read operations while maintaining the same idempotency guarantees.
How does the ledger handle event stream corruption?
Corrupt events increment the unreadableCount but do not halt projection. The checkpoint advances past them, and the read() API returns the per-query unreadable count alongside valid attempts. This transparency allows operators to detect data quality issues without losing throughput.
Is the Model Call Ledger suitable for real-time billing?
The ledger's eventual consistency model suits billing reconciliation better than real-time authorization. For sub-second latency requirements, the runtime's in-memory telemetry (packages/runtime/src/telemetry/) provides faster signals, with the ledger serving as the authoritative audit source.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →