# How Apache Maka Handles Streaming Output and Token Usage Tracking with Model Providers

> Apache Maka efficiently streams LLM outputs and tracks token usage across providers with delta functions and unified storage. Discover how Maka optimizes your model interactions.

- Repository: [The Apache Software Foundation/maka](https://github.com/apache/maka)
- Tags: internals
- Published: 2026-08-26

---

**Apache Maka processes LLM responses as discrete streaming events, using delta functions to update live turn state while aggregating token usage metrics across providers through a unified storage layer.**

Apache Maka is an open-source AI runtime that normalizes interactions across diverse model providers. The repository implements a sophisticated event-driven architecture to handle streaming output and token usage tracking with model providers, ensuring real-time UI updates and accurate cost accounting regardless of whether you are calling OpenAI, Anthropic, or Groq.

## Streaming Output Architecture

Maka treats every LLM call as a sequence of events streamed back from the provider. The UI-side stream modules consume these events incrementally to maintain responsive interfaces.

### Processing Incremental Delta Events

When a model provider sends incremental chunks, Maka routes them through specialized stream handlers. The [`assistant-stream.ts`](https://github.com/apache/maka/blob/main/assistant-stream.ts), [`thinking-stream.ts`](https://github.com/apache/maka/blob/main/thinking-stream.ts), and [`tool-output-stream.ts`](https://github.com/apache/maka/blob/main/tool-output-stream.ts) modules each implement an **`apply…Delta`** function that consumes a single delta and mutates a **LiveTurn** state. These handlers process `text_delta` and `tool_output` events as they arrive, inserting **redaction** and **output-cap** logic via [`streaming-display-redaction.ts`](https://github.com/apache/maka/blob/main/streaming-display-redaction.ts).

The delta application follows this pattern:

```typescript
import { applyAssistantDelta } from '@maka/ui/assistant-stream';

function onDelta(event: { turnId: string; delta: string; seq: number }) {
  liveTurn = applyAssistantDelta(liveTurn, {
    turnId: event.turnId,
    delta: event.delta,
    seq: event.seq,
    createdAt: Date.now(),
  });
}

```

### Managing Live Turn State

The **LiveTurnProjection** in [`live-turn-projection.ts`](https://github.com/apache/maka/blob/main/live-turn-projection.ts) merges incoming deltas into a coherent state object. It maintains a **monotonic list of chunks** where each entry tracks its `seq` and `createdAt` timestamp. As deltas arrive, the projection updates the relevant stream—whether `stdout`, `stderr`, or the assistant’s primary stream—and tracks the turn’s *phase* transitioning from `waiting` to `streamed` and finally to *terminal*.

Once a turn becomes terminal, the stream freezes to prevent further mutations, ensuring data integrity for downstream storage operations.

### Progressive Rendering and Safety Mechanisms

The [`streaming-presentation.ts`](https://github.com/apache/maka/blob/main/streaming-presentation.ts) module contains the gate that enables or disables **progressive re-rendering**. When enabled, the React renderer updates on every token; when disabled, the UI only displays the final assembled text.

To prevent runaway memory consumption, Maka enforces a safety cap of approximately **256KB** per assistant turn. The stream modules insert **truncation markers** such as `assistantTailTruncated` and `assistantChunkTruncated` when content exceeds limits. As implemented in [`assistant-stream.ts`](https://github.com/apache/maka/blob/main/assistant-stream.ts), these caps use `copy.assistantTailTruncated` to signal truncation boundaries to the UI.

## Token Usage Tracking Implementation

Maka aggregates token metrics through a pipeline that spans the UI layer and persistent storage, enabling accurate cost reporting across all supported providers.

### Capturing Provider Metrics

Each LLM response concludes with a `token_usage` message containing `inputTokens` and `outputTokens`. The **[`materialize.ts`](https://github.com/apache/maka/blob/main/materialize.ts)** module walks the message list of a turn; when it encounters a `token_usage` entry, it aggregates the numbers into `turn.tokens`.

When handling turn completion, the system materializes the live state into a stored message:

```typescript
import { applyAssistantComplete } from '@maka/ui/assistant-stream';
import { materializeTurn } from '@maka/ui/materialize';

function onComplete(event) {
  liveTurn = applyAssistantComplete(liveTurn, {
    turnId: event.turnId,
    text: event.text,
    createdAt: Date.now(),
  });

  const stored = materializeTurn(liveTurn);
  // Access aggregated totals via stored.tokens?.input and stored.tokens?.output
}

```

### Persistent Aggregation and Cost Calculation

The storage layer collects all `token_usage` messages belonging to a session. The **[`usage-stats-store.ts`](https://github.com/apache/maka/blob/main/usage-stats-store.ts)** filters the message log for `token_usage` events and updates an in-memory map; on persistence, it writes to the SQLite `usage_stats` table.

Pricing data for every supported model lives in **[`model-pricing.generated.ts`](https://github.com/apache/maka/blob/main/model-pricing.generated.ts)**, which contains per-model `inputUsdPer1M` and `outputUsdPer1M` rates. The daily-review components multiply summed token counts by these rates to produce USD costs:

```typescript
import { getModelPricing } from '@maka/runtime/telemetry/model-pricing.generated';

function costForTurn(modelKey: string, inputTokens: number, outputTokens: number) {
  const pricing = getModelPricing(modelKey);
  const inputCost = (inputTokens / 1_000_000) * pricing.inputUsdPer1M;
  const outputCost = (outputTokens / 1_000_000) * pricing.outputUsdPer1M;
  return inputCost + outputCost;
}

```

### Provider-Agnostic Normalization

Because every provider ultimately emits the same `token_usage` shape, Maka presents a single token-summary UI regardless of the backend. The **[`chat-model-helpers.ts`](https://github.com/apache/maka/blob/main/chat-model-helpers.ts)** module normalizes provider-specific connection objects to a common event schema, ensuring that OpenAI, Anthropic, DeepInfra, and Groq responses are processed identically.

## Recording Usage to Persistent Storage

The storage layer provides a dedicated interface for updating the token-usage ledger:

```typescript
import { UsageStatsStore } from '@maka/storage/usage-stats-store';

async function recordUsage(sessionId, connectionId, usage) {
  const store = new UsageStatsStore();
  await store.addMessage({
    sessionId,
    connectionId,
    type: 'token_usage',
    inputTokens: usage.input,
    outputTokens: usage.output,
    createdAt: Date.now(),
  });
}

```

This implementation in [`usage-stats-store.ts`](https://github.com/apache/maka/blob/main/usage-stats-store.ts) ensures that token consumption data survives application restarts and remains available for historical analysis and billing reports.

## Summary

- Maka processes LLM responses as discrete delta events through specialized stream modules ([`assistant-stream.ts`](https://github.com/apache/maka/blob/main/assistant-stream.ts), [`thinking-stream.ts`](https://github.com/apache/maka/blob/main/thinking-stream.ts), [`tool-output-stream.ts`](https://github.com/apache/maka/blob/main/tool-output-stream.ts)).
- The **LiveTurnProjection** maintains mutable state that transitions from `waiting` to `streamed` to terminal phases, ensuring monotonic chunk ordering.
- **Token usage aggregates** persist across sessions via `UsageStatsStore`, which writes to an SQLite `usage_stats` table for durable accounting.
- **Model pricing data** in [`model-pricing.generated.ts`](https://github.com/apache/maka/blob/main/model-pricing.generated.ts) enables real-time cost calculation regardless of which provider serves the request.
- **Safety caps** and truncation markers prevent runaway memory consumption, while [`chat-model-helpers.ts`](https://github.com/apache/maka/blob/main/chat-model-helpers.ts) ensures cross-provider uniformity.

## Frequently Asked Questions

### How does Maka normalize streaming formats between different providers like OpenAI and Anthropic?

Maka uses the **[`chat-model-helpers.ts`](https://github.com/apache/maka/blob/main/chat-model-helpers.ts)** module to normalize provider-specific response shapes into a common event schema. Whether the backend sends OpenAI-style deltas or Anthropic extended-thinking chunks, the helper transforms them into uniform `text_delta` and `thinking_delta` events that the stream modules can process identically.

### What happens when a streaming response exceeds the size limit?

Maka enforces an output cap of approximately **256KB** per assistant turn. When content approaches this limit, the system inserts **truncation markers** such as `assistantTailTruncated` via the stream modules in [`assistant-stream.ts`](https://github.com/apache/maka/blob/main/assistant-stream.ts). This prevents memory exhaustion while signaling to users that content has been cut off.

### How is token usage data persisted across application restarts?

The **[`usage-stats-store.ts`](https://github.com/apache/maka/blob/main/usage-stats-store.ts)** module maintains an in-memory aggregation map that flushes to the SQLite `usage_stats` table. By calling `addMessage` with `type: 'token_usage'`, the system records per-connection totals that survive process termination, enabling historical cost analysis and usage reporting.

### Can progressive streaming be disabled to improve rendering performance?

Yes. The **[`streaming-presentation.ts`](https://github.com/apache/maka/blob/main/streaming-presentation.ts)** module contains a configuration gate for `isProgressiveStreamingEnabled`. When disabled, the React renderer skips intermediate updates and only displays the final assembled text after the stream completes, reducing CPU load for applications handling high-frequency token generation.