# Real-Time Cost Tracking and Token Budget Attribution in Munder-Difflin: A Technical Deep Dive

> Explore Munder-Difflin's real-time cost tracking and token budget attribution system. Discover how it enforces global and per-agent token budgets for effective LLM cost control.

- Repository: [Chaitanya Giri/munder-difflin](https://github.com/chaitanyagiri/munder-difflin)
- Tags: deep-dive
- Published: 2026-08-20

---

**Munder-Difflin implements a two-layer cost control system that tracks LLM token usage in real-time and enforces configurable token budgets at both global and per-agent levels.**

The [`chaitanyagiri/munder-difflin`](https://github.com/chaitanyagiri/munder-difflin) repository contains a sophisticated cost management infrastructure designed for autonomous agent swarms. This article examines how the **real-time cost tracking** and **token budget attribution** systems work together to prevent runaway spending while providing full transparency into where tokens are consumed.

## Architecture Overview: Three Core Components

The cost control system rests on three pillars:

1. **Live cost aggregation** – captures every token as it flows through the system
2. **Budget enforcement** – compares live totals against caps and triggers protective actions
3. **Attribution and telemetry** – maps costs to specific agents, tasks, and time periods

Understanding how these components interact is essential for configuring effective spend controls in production deployments.

## Real-Time Cost Tracking: The Cost Store

At the heart of the system sits a pure, dependency-free **cost store** that maintains running totals of token consumption and estimated USD spend.

### Data Structure and Core Functions

The cost store in [[`src/renderer/src/realtime/costStore.ts`](https://github.com/chaitanyagiri/munder-difflin/blob/main/src/renderer/src/realtime/costStore.ts)](https://github.com/chaitanyagiri/munder-difflin/blob/main/src/renderer/src/realtime/costStore.ts) exposes three primary operations:

```typescript
// Initialize or reset tracking for an agent
initRealtimeCost(agentId: string): void

// Record new usage from a completed LLM response
recordRealtimeUsage(agentId: string, delta: TokenDelta): void

// Retrieve current snapshot for enforcement decisions
getRealtimeCostSnapshot(agentId: string): CostSnapshot

```

The **TokenDelta** interface captures the essential metrics:

```typescript
interface TokenDelta {
  inputTokens: number;      // Prompt tokens sent to the model
  outputTokens: number;     // Completion tokens generated
  cachedTokens?: number;    // Tokens served from cache (discounted pricing)
}

```

The **CostSnapshot** returned by `getRealtimeCostSnapshot` contains the enforcement-critical fields:

```typescript
interface CostSnapshot {
  tokens: number;           // Cumulative total tokens this session
  usd: number;              // Estimated dollar cost using model-specific rates
  overCap: boolean;         // Pre-computed flag: true if budget exceeded
}

```

The USD estimation uses a static pricing table defined in [`src/shared/realtimePricing.ts`](https://github.com/chaitanyagiri/munder-difflin/blob/main/src/shared/realtimePricing.ts), applying per-model rates that distinguish between input, output, and cached token costs.

### Integration with Voice Sessions

Token tracking activates automatically when an agent enters a voice session. The session controller in [[`src/renderer/src/realtime/session.ts`](https://github.com/chaitanyagiri/munder-difflin/blob/main/src/renderer/src/realtime/session.ts)](https://github.com/chaitanyagiri/munder-difflin/blob/main/src/renderer/src/realtime/session.ts) wires the cost store to the WebRTC data channel:

```typescript
// Inside session.ts: cost tracking initialization
this.dataChannel.on('response.done', (event) => {
  const delta = {
    inputTokens: event.response.usage.input_tokens,
    outputTokens: event.response.usage.output_tokens,
    cachedTokens: event.response.usage.cache_read_input_tokens,
  };
  recordRealtimeUsage(this.agentId, delta);
});

```

This event-driven approach ensures **zero-latency cost updates**—every LLM response immediately contributes to the running total without batching delays.

### Cost Guard: Automated Session Protection

Beyond passive tracking, the session layer implements an active **cost guard** that prevents overspend in real-time:

```typescript
// Cost-guard tick runs every 5 seconds during active sessions
this.costGuardInterval = setInterval(() => {
  const snapshot = getRealtimeCostSnapshot(this.agentId);
  if (snapshot.overCap) {
    this.disconnect('cost-cap-exceeded');
    this.emit('budget:violated', snapshot);
  }
}, 5000);

```

This protective mechanism ensures that even during high-throughput streaming, the system can terminate expensive sessions before significant budget overrun occurs.

## Token Budget Attribution: Configuration and Enforcement

While the cost store tracks usage, the **token budget attribution system** defines what "too expensive" means and who pays the price when limits are breached.

### Configuration Hierarchy

Budget settings are loaded from the hive configuration in [[`src/main/config.ts`](https://github.com/chaitanyagiri/munder-difflin/blob/main/src/main/config.ts)](https://github.com/chaitanyagiri/munder-difflin/blob/main/src/main/config.ts):

```typescript
interface CostConfig {
  // Global floor-wide token budget (undefined = unlimited)
  costCapTokens?: number;
  
  // Per-agent overrides: specific limits for individual agents
  agentTokenCaps?: Record<string, number>;
}

```

This two-tier structure enables **fine-grained spend control**:

- **Floor cap** (`costCapTokens`): The default maximum for all agents in the hive
- **Agent caps** (`agentTokenCaps`): Optional overrides for specific high-value or restricted agents

### Circuit Breaker Enforcement

The enforcement logic resides in [[`src/main/breaker.ts`](https://github.com/chaitanyagiri/munder-difflin/blob/main/src/main/breaker.ts)](https://github.com/chaitanyagiri/munder-difflin/blob/main/src/main/breaker.ts), which evaluates whether the current token consumption justifies tripping the circuit breaker.

**Global cap evaluation:**

```typescript
// breaker.ts: checking floor-wide token budget
if (typeof cfg.costCapTokens === 'number' && cfg.costCapTokens > 0) {
  const total = getFloorTokenTotal(allAgents);
  if (total > cfg.costCapTokens) {
    const topSpender = identifyTopSpender(allAgents);
    return {
      tripping: true,
      reason: `token cap: floor total ${total} exceeds ${cfg.costCapTokens}`,
      responsibleAgent: topSpender,
    };
  }
}

```

**Per-agent cap evaluation** takes precedence when configured:

```typescript
// Per-agent override check
const agentCap = cfg.agentTokenCaps?.[agentId];
if (typeof agentCap === 'number' && agentCap > 0) {
  const agentTotal = getAgentTokenTotal(agentId);
  if (agentTotal > agentCap) {
    return {
      tripping: true,
      reason: `token cap: agent ${agentId} ${agentTotal} exceeds ${agentCap}`,
      responsibleAgent: agentId,
    };
  }
}

```

The breaker runs on a **30-second evaluation cycle**, balancing responsiveness against computational overhead.

### Attribution to Specific Agents

When a cap violation occurs, the system identifies the **top token spender** using cumulative usage data from the cost store:

```typescript
function identifyTopSpender(agents: Agent[]): string | null {
  return agents
    .map(a => ({ id: a.id, tokens: getRealtimeCostSnapshot(a.id).tokens }))
    .sort((a, b) => b.tokens - a.tokens)[0]?.id ?? null;
}

```

This attribution enables **targeted intervention**—only the overspending agent faces breaker consequences while others continue operating.

## Querying Costs: The `get_cost` Tool

Agents and users can inspect current costs at any time through the [`get_cost`](https://github.com/chaitanyagiri/munder-difflin/blob/main/src/renderer/src/realtime/tools.ts) tool endpoint:

```typescript
// src/renderer/src/realtime/tools.ts
export const get_cost = {
  name: 'get_cost',
  description: 'Get current token usage and estimated cost for this agent',
  handler: async (context: ToolContext) => {
    const snapshot = getRealtimeCostSnapshot(context.agentId);
    return {
      tokens: snapshot.tokens,
      usd: snapshot.usd,
      overCap: snapshot.overCap,
      remaining: context.config.costCapTokens 
        ? context.config.costCapTokens - snapshot.tokens 
        : null,
    };
  },
};

```

**Usage example from an agent:**

```typescript
const costStatus = await runTool('get_cost');
if (costStatus.remaining !== null && costStatus.remaining < 5000) {
  await say(`Warning: Only ${costStatus.remaining} tokens remain in budget.`);
  await requestBudgetIncrease();
}

```

This **self-aware cost querying** enables agents to implement sophisticated budget-aware behaviors, such as switching to cheaper models or requesting human approval before expensive operations.

## Historical Attribution: The Cost Ledger

While real-time tracking prevents immediate overspend, **historical attribution** provides accountability. The main process writes every token delta to a durable SQLite ledger in [`src/main/usage.ts`](https://github.com/chaitanyagiri/munder-difflin/blob/main/src/main/usage.ts):

```typescript
// Persist usage for audit and reporting
await costLedger.insert({
  agent_id: agentId,
  task_id: currentTaskId,
  model: response.model,
  input_tokens: delta.inputTokens,
  output_tokens: delta.outputTokens,
  cached_tokens: delta.cachedTokens,
  estimated_usd: computedCost,
  timestamp: Date.now(),
});

```

This ledger enables **per-task cost attribution**—when a multi-agent task completes, the system aggregates all related ledger entries and attaches the total to the task record in [`tasks.json`](https://github.com/chaitanyagiri/munder-difflin/blob/main/tasks.json). The "Angela-budget" hire manifest uses this data to generate spending reports ranked by agent, model, and task type.

## UI Integration: Live Cost HUD

The React-based UI renders real-time cost information through components in `src/renderer/src/components/`:

- **Cost HUD overlay**: Floating meter showing `current / cap` tokens with color-coded warnings
- **Settings modal** ([[`SettingsModal.tsx`](https://github.com/chaitanyagiri/munder-difflin/blob/main/SettingsModal.tsx)](https://github.com/chaitanyagiri/munder-difflin/blob/main/src/renderer/src/components/SettingsModal.tsx)): Configure `costCapTokens` and per-agent caps
- **Fleet dashboard**: Reads [`fleet.json`](https://github.com/chaitanyagiri/munder-difflin/blob/main/fleet.json) to display breaker status and per-agent spend

**Configuring the token budget via SettingsModal:**

```tsx
// SettingsModal.tsx excerpt
const [tokenCap, setTokenCap] = useState(
  config.costCapTokens?.toString() ?? ''
);

const handleSave = () => {
  const parsed = parseInt(tokenCap, 10);
  setConfig({
    ...config,
    costCapTokens: Number.isFinite(parsed) ? parsed : undefined,
  });
};

```

Changes take effect immediately—the breaker re-evaluates on its next cycle with the new limits.

## Performance Characteristics

| Aspect | Implementation Detail | Impact |
|--------|----------------------|--------|
| **Update latency** | Event-driven, no batching | Sub-100ms cost visibility |
| **Memory footprint** | In-memory maps per agent | ~200 bytes per active agent |
| **Persistence overhead** | Async SQLite writes | Non-blocking, 5-10ms per write |
| **Breaker evaluation** | 30-second intervals | Amortizes check cost across sessions |
| **Cost guard frequency** | 5-second polling during sessions | Rapid overspend prevention |

The design prioritizes **low-latency visibility** for enforcement decisions while accepting slightly delayed persistence for historical records.

## Summary

- **Real-time cost tracking** centers on the cost store in [`src/renderer/src/realtime/costStore.ts`](https://github.com/chaitanyagiri/munder-difflin/blob/main/src/renderer/src/realtime/costStore.ts), which aggregates token deltas and computes USD estimates using model-specific pricing
- **Token budget attribution** combines global caps (`costCapTokens`) and per-agent overrides (`agentTokenCaps`) defined in [`src/main/config.ts`](https://github.com/chaitanyagiri/munder-difflin/blob/main/src/main/config.ts)
- **Enforcement** occurs through the circuit breaker ([`src/main/breaker.ts`](https://github.com/chaitanyagiri/munder-difflin/blob/main/src/main/breaker.ts)) for floor-wide violations and the session cost guard ([`src/renderer/src/realtime/session.ts`](https://github.com/chaitanyagiri/munder-difflin/blob/main/src/renderer/src/realtime/session.ts)) for immediate session protection
- **Transparency** is provided via the `get_cost` tool, live UI components, and the durable cost ledger for historical analysis
- **Attribution precision** maps every token to its consuming agent, enabling targeted interventions and detailed spend reporting

## Frequently Asked Questions

### How quickly does the system detect a budget violation?

The session-level **cost guard** checks every 5 seconds during active voice sessions, while the global **circuit breaker** evaluates every 30 seconds. This two-speed design ensures rapid protection for expensive streaming sessions without overwhelming the system with enforcement checks during idle periods.

### Can different agents have different token budgets?

Yes. The `agentTokenCaps` configuration object allows assigning specific limits to individual agent IDs. When both a global cap and per-agent cap exist for the same agent, the **more restrictive limit applies**. Per-agent caps are evaluated before global caps in the breaker logic.

### What happens when a token cap is exceeded?

For **session-level violations**, the cost guard immediately disconnects the voice session and emits a `budget:violated` event. For **floor-wide violations**, the circuit breaker trips and identifies the top-spending agent, potentially isolating that agent while allowing others to continue. Both events are logged to the cost ledger and surfaced in the UI.

### How accurate are the USD cost estimates?

Estimates use static pricing tables that match current provider rates for Claude, GPT-4, and other supported models. The system distinguishes **input**, **output**, and **cached** token pricing. While estimates are typically within 5% of actual invoices, they exclude factors like rate limits, overages, or enterprise discounts—treat them as **planning estimates** rather than billing guarantees.