Real-Time Cost Tracking and Token Budget Attribution in Munder-Difflin: A Technical Deep Dive
Munder-Difflin implements a two-layer cost control system that tracks LLM token usage in real-time and enforces configurable token budgets at both global and per-agent levels.
The chaitanyagiri/munder-difflin repository contains a sophisticated cost management infrastructure designed for autonomous agent swarms. This article examines how the real-time cost tracking and token budget attribution systems work together to prevent runaway spending while providing full transparency into where tokens are consumed.
Architecture Overview: Three Core Components
The cost control system rests on three pillars:
- Live cost aggregation – captures every token as it flows through the system
- Budget enforcement – compares live totals against caps and triggers protective actions
- Attribution and telemetry – maps costs to specific agents, tasks, and time periods
Understanding how these components interact is essential for configuring effective spend controls in production deployments.
Real-Time Cost Tracking: The Cost Store
At the heart of the system sits a pure, dependency-free cost store that maintains running totals of token consumption and estimated USD spend.
Data Structure and Core Functions
The cost store in [src/renderer/src/realtime/costStore.ts](https://github.com/chaitanyagiri/munder-difflin/blob/main/src/renderer/src/realtime/costStore.ts) exposes three primary operations:
// Initialize or reset tracking for an agent
initRealtimeCost(agentId: string): void
// Record new usage from a completed LLM response
recordRealtimeUsage(agentId: string, delta: TokenDelta): void
// Retrieve current snapshot for enforcement decisions
getRealtimeCostSnapshot(agentId: string): CostSnapshot
The TokenDelta interface captures the essential metrics:
interface TokenDelta {
inputTokens: number; // Prompt tokens sent to the model
outputTokens: number; // Completion tokens generated
cachedTokens?: number; // Tokens served from cache (discounted pricing)
}
The CostSnapshot returned by getRealtimeCostSnapshot contains the enforcement-critical fields:
interface CostSnapshot {
tokens: number; // Cumulative total tokens this session
usd: number; // Estimated dollar cost using model-specific rates
overCap: boolean; // Pre-computed flag: true if budget exceeded
}
The USD estimation uses a static pricing table defined in src/shared/realtimePricing.ts, applying per-model rates that distinguish between input, output, and cached token costs.
Integration with Voice Sessions
Token tracking activates automatically when an agent enters a voice session. The session controller in [src/renderer/src/realtime/session.ts](https://github.com/chaitanyagiri/munder-difflin/blob/main/src/renderer/src/realtime/session.ts) wires the cost store to the WebRTC data channel:
// Inside session.ts: cost tracking initialization
this.dataChannel.on('response.done', (event) => {
const delta = {
inputTokens: event.response.usage.input_tokens,
outputTokens: event.response.usage.output_tokens,
cachedTokens: event.response.usage.cache_read_input_tokens,
};
recordRealtimeUsage(this.agentId, delta);
});
This event-driven approach ensures zero-latency cost updates—every LLM response immediately contributes to the running total without batching delays.
Cost Guard: Automated Session Protection
Beyond passive tracking, the session layer implements an active cost guard that prevents overspend in real-time:
// Cost-guard tick runs every 5 seconds during active sessions
this.costGuardInterval = setInterval(() => {
const snapshot = getRealtimeCostSnapshot(this.agentId);
if (snapshot.overCap) {
this.disconnect('cost-cap-exceeded');
this.emit('budget:violated', snapshot);
}
}, 5000);
This protective mechanism ensures that even during high-throughput streaming, the system can terminate expensive sessions before significant budget overrun occurs.
Token Budget Attribution: Configuration and Enforcement
While the cost store tracks usage, the token budget attribution system defines what "too expensive" means and who pays the price when limits are breached.
Configuration Hierarchy
Budget settings are loaded from the hive configuration in [src/main/config.ts](https://github.com/chaitanyagiri/munder-difflin/blob/main/src/main/config.ts):
interface CostConfig {
// Global floor-wide token budget (undefined = unlimited)
costCapTokens?: number;
// Per-agent overrides: specific limits for individual agents
agentTokenCaps?: Record<string, number>;
}
This two-tier structure enables fine-grained spend control:
- Floor cap (
costCapTokens): The default maximum for all agents in the hive - Agent caps (
agentTokenCaps): Optional overrides for specific high-value or restricted agents
Circuit Breaker Enforcement
The enforcement logic resides in [src/main/breaker.ts](https://github.com/chaitanyagiri/munder-difflin/blob/main/src/main/breaker.ts), which evaluates whether the current token consumption justifies tripping the circuit breaker.
Global cap evaluation:
// breaker.ts: checking floor-wide token budget
if (typeof cfg.costCapTokens === 'number' && cfg.costCapTokens > 0) {
const total = getFloorTokenTotal(allAgents);
if (total > cfg.costCapTokens) {
const topSpender = identifyTopSpender(allAgents);
return {
tripping: true,
reason: `token cap: floor total ${total} exceeds ${cfg.costCapTokens}`,
responsibleAgent: topSpender,
};
}
}
Per-agent cap evaluation takes precedence when configured:
// Per-agent override check
const agentCap = cfg.agentTokenCaps?.[agentId];
if (typeof agentCap === 'number' && agentCap > 0) {
const agentTotal = getAgentTokenTotal(agentId);
if (agentTotal > agentCap) {
return {
tripping: true,
reason: `token cap: agent ${agentId} ${agentTotal} exceeds ${agentCap}`,
responsibleAgent: agentId,
};
}
}
The breaker runs on a 30-second evaluation cycle, balancing responsiveness against computational overhead.
Attribution to Specific Agents
When a cap violation occurs, the system identifies the top token spender using cumulative usage data from the cost store:
function identifyTopSpender(agents: Agent[]): string | null {
return agents
.map(a => ({ id: a.id, tokens: getRealtimeCostSnapshot(a.id).tokens }))
.sort((a, b) => b.tokens - a.tokens)[0]?.id ?? null;
}
This attribution enables targeted intervention—only the overspending agent faces breaker consequences while others continue operating.
Querying Costs: The get_cost Tool
Agents and users can inspect current costs at any time through the get_cost tool endpoint:
// src/renderer/src/realtime/tools.ts
export const get_cost = {
name: 'get_cost',
description: 'Get current token usage and estimated cost for this agent',
handler: async (context: ToolContext) => {
const snapshot = getRealtimeCostSnapshot(context.agentId);
return {
tokens: snapshot.tokens,
usd: snapshot.usd,
overCap: snapshot.overCap,
remaining: context.config.costCapTokens
? context.config.costCapTokens - snapshot.tokens
: null,
};
},
};
Usage example from an agent:
const costStatus = await runTool('get_cost');
if (costStatus.remaining !== null && costStatus.remaining < 5000) {
await say(`Warning: Only ${costStatus.remaining} tokens remain in budget.`);
await requestBudgetIncrease();
}
This self-aware cost querying enables agents to implement sophisticated budget-aware behaviors, such as switching to cheaper models or requesting human approval before expensive operations.
Historical Attribution: The Cost Ledger
While real-time tracking prevents immediate overspend, historical attribution provides accountability. The main process writes every token delta to a durable SQLite ledger in src/main/usage.ts:
// Persist usage for audit and reporting
await costLedger.insert({
agent_id: agentId,
task_id: currentTaskId,
model: response.model,
input_tokens: delta.inputTokens,
output_tokens: delta.outputTokens,
cached_tokens: delta.cachedTokens,
estimated_usd: computedCost,
timestamp: Date.now(),
});
This ledger enables per-task cost attribution—when a multi-agent task completes, the system aggregates all related ledger entries and attaches the total to the task record in tasks.json. The "Angela-budget" hire manifest uses this data to generate spending reports ranked by agent, model, and task type.
UI Integration: Live Cost HUD
The React-based UI renders real-time cost information through components in src/renderer/src/components/:
- Cost HUD overlay: Floating meter showing
current / captokens with color-coded warnings - Settings modal ([
SettingsModal.tsx](https://github.com/chaitanyagiri/munder-difflin/blob/main/src/renderer/src/components/SettingsModal.tsx)): ConfigurecostCapTokensand per-agent caps - Fleet dashboard: Reads
fleet.jsonto display breaker status and per-agent spend
Configuring the token budget via SettingsModal:
// SettingsModal.tsx excerpt
const [tokenCap, setTokenCap] = useState(
config.costCapTokens?.toString() ?? ''
);
const handleSave = () => {
const parsed = parseInt(tokenCap, 10);
setConfig({
...config,
costCapTokens: Number.isFinite(parsed) ? parsed : undefined,
});
};
Changes take effect immediately—the breaker re-evaluates on its next cycle with the new limits.
Performance Characteristics
| Aspect | Implementation Detail | Impact |
|---|---|---|
| Update latency | Event-driven, no batching | Sub-100ms cost visibility |
| Memory footprint | In-memory maps per agent | ~200 bytes per active agent |
| Persistence overhead | Async SQLite writes | Non-blocking, 5-10ms per write |
| Breaker evaluation | 30-second intervals | Amortizes check cost across sessions |
| Cost guard frequency | 5-second polling during sessions | Rapid overspend prevention |
The design prioritizes low-latency visibility for enforcement decisions while accepting slightly delayed persistence for historical records.
Summary
- Real-time cost tracking centers on the cost store in
src/renderer/src/realtime/costStore.ts, which aggregates token deltas and computes USD estimates using model-specific pricing - Token budget attribution combines global caps (
costCapTokens) and per-agent overrides (agentTokenCaps) defined insrc/main/config.ts - Enforcement occurs through the circuit breaker (
src/main/breaker.ts) for floor-wide violations and the session cost guard (src/renderer/src/realtime/session.ts) for immediate session protection - Transparency is provided via the
get_costtool, live UI components, and the durable cost ledger for historical analysis - Attribution precision maps every token to its consuming agent, enabling targeted interventions and detailed spend reporting
Frequently Asked Questions
How quickly does the system detect a budget violation?
The session-level cost guard checks every 5 seconds during active voice sessions, while the global circuit breaker evaluates every 30 seconds. This two-speed design ensures rapid protection for expensive streaming sessions without overwhelming the system with enforcement checks during idle periods.
Can different agents have different token budgets?
Yes. The agentTokenCaps configuration object allows assigning specific limits to individual agent IDs. When both a global cap and per-agent cap exist for the same agent, the more restrictive limit applies. Per-agent caps are evaluated before global caps in the breaker logic.
What happens when a token cap is exceeded?
For session-level violations, the cost guard immediately disconnects the voice session and emits a budget:violated event. For floor-wide violations, the circuit breaker trips and identifies the top-spending agent, potentially isolating that agent while allowing others to continue. Both events are logged to the cost ledger and surfaced in the UI.
How accurate are the USD cost estimates?
Estimates use static pricing tables that match current provider rates for Claude, GPT-4, and other supported models. The system distinguishes input, output, and cached token pricing. While estimates are typically within 5% of actual invoices, they exclude factors like rate limits, overages, or enterprise discounts—treat them as planning estimates rather than billing guarantees.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →