How Apache Maka Handles Streaming Output and Token Usage Tracking with Model Providers
Apache Maka processes LLM responses as discrete streaming events, using delta functions to update live turn state while aggregating token usage metrics across providers through a unified storage layer.
Apache Maka is an open-source AI runtime that normalizes interactions across diverse model providers. The repository implements a sophisticated event-driven architecture to handle streaming output and token usage tracking with model providers, ensuring real-time UI updates and accurate cost accounting regardless of whether you are calling OpenAI, Anthropic, or Groq.
Streaming Output Architecture
Maka treats every LLM call as a sequence of events streamed back from the provider. The UI-side stream modules consume these events incrementally to maintain responsive interfaces.
Processing Incremental Delta Events
When a model provider sends incremental chunks, Maka routes them through specialized stream handlers. The assistant-stream.ts, thinking-stream.ts, and tool-output-stream.ts modules each implement an apply…Delta function that consumes a single delta and mutates a LiveTurn state. These handlers process text_delta and tool_output events as they arrive, inserting redaction and output-cap logic via streaming-display-redaction.ts.
The delta application follows this pattern:
import { applyAssistantDelta } from '@maka/ui/assistant-stream';
function onDelta(event: { turnId: string; delta: string; seq: number }) {
liveTurn = applyAssistantDelta(liveTurn, {
turnId: event.turnId,
delta: event.delta,
seq: event.seq,
createdAt: Date.now(),
});
}
Managing Live Turn State
The LiveTurnProjection in live-turn-projection.ts merges incoming deltas into a coherent state object. It maintains a monotonic list of chunks where each entry tracks its seq and createdAt timestamp. As deltas arrive, the projection updates the relevant stream—whether stdout, stderr, or the assistant’s primary stream—and tracks the turn’s phase transitioning from waiting to streamed and finally to terminal.
Once a turn becomes terminal, the stream freezes to prevent further mutations, ensuring data integrity for downstream storage operations.
Progressive Rendering and Safety Mechanisms
The streaming-presentation.ts module contains the gate that enables or disables progressive re-rendering. When enabled, the React renderer updates on every token; when disabled, the UI only displays the final assembled text.
To prevent runaway memory consumption, Maka enforces a safety cap of approximately 256KB per assistant turn. The stream modules insert truncation markers such as assistantTailTruncated and assistantChunkTruncated when content exceeds limits. As implemented in assistant-stream.ts, these caps use copy.assistantTailTruncated to signal truncation boundaries to the UI.
Token Usage Tracking Implementation
Maka aggregates token metrics through a pipeline that spans the UI layer and persistent storage, enabling accurate cost reporting across all supported providers.
Capturing Provider Metrics
Each LLM response concludes with a token_usage message containing inputTokens and outputTokens. The materialize.ts module walks the message list of a turn; when it encounters a token_usage entry, it aggregates the numbers into turn.tokens.
When handling turn completion, the system materializes the live state into a stored message:
import { applyAssistantComplete } from '@maka/ui/assistant-stream';
import { materializeTurn } from '@maka/ui/materialize';
function onComplete(event) {
liveTurn = applyAssistantComplete(liveTurn, {
turnId: event.turnId,
text: event.text,
createdAt: Date.now(),
});
const stored = materializeTurn(liveTurn);
// Access aggregated totals via stored.tokens?.input and stored.tokens?.output
}
Persistent Aggregation and Cost Calculation
The storage layer collects all token_usage messages belonging to a session. The usage-stats-store.ts filters the message log for token_usage events and updates an in-memory map; on persistence, it writes to the SQLite usage_stats table.
Pricing data for every supported model lives in model-pricing.generated.ts, which contains per-model inputUsdPer1M and outputUsdPer1M rates. The daily-review components multiply summed token counts by these rates to produce USD costs:
import { getModelPricing } from '@maka/runtime/telemetry/model-pricing.generated';
function costForTurn(modelKey: string, inputTokens: number, outputTokens: number) {
const pricing = getModelPricing(modelKey);
const inputCost = (inputTokens / 1_000_000) * pricing.inputUsdPer1M;
const outputCost = (outputTokens / 1_000_000) * pricing.outputUsdPer1M;
return inputCost + outputCost;
}
Provider-Agnostic Normalization
Because every provider ultimately emits the same token_usage shape, Maka presents a single token-summary UI regardless of the backend. The chat-model-helpers.ts module normalizes provider-specific connection objects to a common event schema, ensuring that OpenAI, Anthropic, DeepInfra, and Groq responses are processed identically.
Recording Usage to Persistent Storage
The storage layer provides a dedicated interface for updating the token-usage ledger:
import { UsageStatsStore } from '@maka/storage/usage-stats-store';
async function recordUsage(sessionId, connectionId, usage) {
const store = new UsageStatsStore();
await store.addMessage({
sessionId,
connectionId,
type: 'token_usage',
inputTokens: usage.input,
outputTokens: usage.output,
createdAt: Date.now(),
});
}
This implementation in usage-stats-store.ts ensures that token consumption data survives application restarts and remains available for historical analysis and billing reports.
Summary
- Maka processes LLM responses as discrete delta events through specialized stream modules (
assistant-stream.ts,thinking-stream.ts,tool-output-stream.ts). - The LiveTurnProjection maintains mutable state that transitions from
waitingtostreamedto terminal phases, ensuring monotonic chunk ordering. - Token usage aggregates persist across sessions via
UsageStatsStore, which writes to an SQLiteusage_statstable for durable accounting. - Model pricing data in
model-pricing.generated.tsenables real-time cost calculation regardless of which provider serves the request. - Safety caps and truncation markers prevent runaway memory consumption, while
chat-model-helpers.tsensures cross-provider uniformity.
Frequently Asked Questions
How does Maka normalize streaming formats between different providers like OpenAI and Anthropic?
Maka uses the chat-model-helpers.ts module to normalize provider-specific response shapes into a common event schema. Whether the backend sends OpenAI-style deltas or Anthropic extended-thinking chunks, the helper transforms them into uniform text_delta and thinking_delta events that the stream modules can process identically.
What happens when a streaming response exceeds the size limit?
Maka enforces an output cap of approximately 256KB per assistant turn. When content approaches this limit, the system inserts truncation markers such as assistantTailTruncated via the stream modules in assistant-stream.ts. This prevents memory exhaustion while signaling to users that content has been cut off.
How is token usage data persisted across application restarts?
The usage-stats-store.ts module maintains an in-memory aggregation map that flushes to the SQLite usage_stats table. By calling addMessage with type: 'token_usage', the system records per-connection totals that survive process termination, enabling historical cost analysis and usage reporting.
Can progressive streaming be disabled to improve rendering performance?
Yes. The streaming-presentation.ts module contains a configuration gate for isProgressiveStreamingEnabled. When disabled, the React renderer skips intermediate updates and only displays the final assembled text after the stream completes, reducing CPU load for applications handling high-frequency token generation.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →