# What Is Ergodic Streaming in Moonshine’s Real-Time Transcription?

> Discover Moonshine's ergodic streaming for real-time transcription. Achieve sub-200ms latency by processing audio frames once and reusing cached data for efficient transcription.

- Repository: [Moonshine AI/moonshine](https://github.com/moonshine-ai/moonshine)
- Tags: deep-dive
- Published: 2026-02-16

---

**Ergodic streaming is Moonshine’s architectural technique that achieves sub-200 ms end-to-end latency by processing each audio frame exactly once and reusing cached encoder outputs and cross-attention KV states across overlapping audio windows.**

Ergodic streaming powers the real-time transcription capabilities of Moonshine, an open-source automatic speech recognition (ASR) engine developed by moonshine-ai. Unlike conventional streaming ASR systems that recompute encoder features for every new audio chunk, Moonshine’s ergodic streaming encoder treats the audio stream as an ergodic process—reusing computations and maintaining persistent state to minimize latency on edge devices.

## How Ergodic Streaming Works

### The Problem with Traditional Streaming ASR

In conventional streaming ASR pipelines, the encoder re-processes every audio frame each time a new chunk arrives. When overlapping windows are used to maintain acoustic context, this redundancy wastes compute and adds unnecessary latency. Every inference step repeats matrix operations on audio segments that have already been analyzed.

### Frame-Level Reuse and KV Cache Persistence

Moonshine’s **ergodic encoder**, implemented in [`core/moonshine-streaming-model.cpp`](https://github.com/moonshine-ai/moonshine/blob/main/core/moonshine-streaming-model.cpp), eliminates this redundancy by processing each audio frame exactly once. It stores intermediate representations in `accumulated_features` within the `MoonshineStreamingState` struct (defined in [`core/moonshine-streaming-model.h`](https://github.com/moonshine-ai/moonshine/blob/main/core/moonshine-streaming-model.h)).

The system stitches together results from overlapping windows so the decoder consumes a continuously updated "memory" of the utterance. The decoder leverages a pre-computed cross-attention KV cache (the `cross_kv.onnx` model) via `compute_cross_kv()` and `run_decoder_with_cross_kv()`. After the initial setup, heavy matrix multiplications are already cached; only a minimal auto-regressive step via `decode_step()` is needed for each new token.

## Implementation in Moonshine’s C++ Core

The ergodic streaming architecture centers on two primary structures in [`core/moonshine-streaming-model.h`](https://github.com/moonshine-ai/moonshine/blob/main/core/moonshine-streaming-model.h):

- **`MoonshineStreamingConfig`**: Holds model hyper-parameters including encoder dimensions and look-ahead settings.
- **`MoonshineStreamingState`**: Maintains per-stream buffers including raw samples, convolution state, accumulated encoder features, decoder memory, and self- and cross-attention KV caches (with a `cross_kv_valid` flag).

The streaming loop in [`core/moonshine-streaming-model.cpp`](https://github.com/moonshine-ai/moonshine/blob/main/core/moonshine-streaming-model.cpp) orchestrates the process:

1. **`process_audio_chunk()`**: Accepts new audio data, updates frontend buffers, and triggers the ergodic encoder on only the new frames.
2. **`encode()`**: Updates encoder memory with new features and appends to `accumulated_features`.
3. **`compute_cross_kv()`**: Pre-computes cross-attention keys and values from the accumulated memory.
4. **`decode_step()`**: Performs minimal auto-regressive decoding using the cached cross-attention states.

This design enables the **tiny-streaming**, **small-streaming**, and **medium-streaming** model variants to deliver **sub-200 ms end-to-end latency** while the speaker is still talking.

## Code Examples

### Python Streaming API

The Python wrapper in [`python/src/moonshine_voice/transcriber.py`](https://github.com/moonshine-ai/moonshine/blob/main/python/src/moonshine_voice/transcriber.py) exposes the ergodic streaming engine via `Transcriber.add_audio()`, which maps to the C++ `moonshine_transcribe` function.

```python
from moonshine_voice import Transcriber, TranscriptEventListener

class PrintListener(TranscriptEventListener):
    def on_line_started(self, event):
        print("[START] ", event.line.text)

    def on_line_text_changed(self, event):
        print("[UPDATE] ", event.line.text)

    def on_line_completed(self, event):
        print("[DONE]   ", event.line.text)

# Load a streaming model (e.g. tiny-streaming)

transcriber = Transcriber(
    model_path="path/to/tiny-streaming-en",
    model_arch="tiny-streaming",
)

transcriber.add_listener(PrintListener())
transcriber.start()

# Simulate live audio by feeding 100 ms chunks from a wav file

audio, sr = load_wav_file("examples/python/two_cities.wav")
chunk_ms = 0.1
chunk_samples = int(chunk_ms * sr)

for i in range(0, len(audio), chunk_samples):
    transcriber.add_audio(audio[i:i + chunk_samples], sr)

transcriber.stop()

```

Each `add_audio` call triggers `process_audio_chunk` in the C++ core, running the ergodic encoder on only the new audio data.

### Swift Real-Time Transcription

The Swift bindings in [`swift/Sources/MoonshineVoice/Transcriber.swift`](https://github.com/moonshine-ai/moonshine/blob/main/swift/Sources/MoonshineVoice/Transcriber.swift) and [`MicTranscriber.swift`](https://github.com/moonshine-ai/moonshine/blob/main/MicTranscriber.swift) forward audio chunks to the same C++ streaming engine.

```swift
import MoonshineVoice

let transcriber = try Transcriber(
    modelPath: "/path/to/tiny-streaming-en",
    modelArch: .tinyStreaming
)

let mic = MicTranscriber(transcriber: transcriber)

class Listener: TranscriptEventListener {
    func onLineStarted(_ event: TranscriptEvent) {
        print("▶︎", event.line.text)
    }
    func onLineTextChanged(_ event: TranscriptEvent) {
        print("…", event.line.text)
    }
    func onLineCompleted(_ event: TranscriptEvent) {
        print("✔︎", event.line.text)
    }
}
transcriber.addListener(Listener())

try mic.start()          // captures microphone and streams chunks internally
// … later …
try mic.stop()

```

The `MicTranscriber` class manages the audio session and calls the underlying `MoonshineStreamingModel` to process chunks via the ergodic encoder.

### Direct C++ Usage

For applications requiring direct integration, the C++ API in [`core/moonshine-streaming-model.h`](https://github.com/moonshine-ai/moonshine/blob/main/core/moonshine-streaming-model.h) provides fine-grained control over the ergodic streaming state.

```cpp
MoonshineStreamingModel model;
model.load("model_dir", "tokenizer.bin", /*model_type=*/1);
auto *state = model.create_state();

std::vector<float> chunk = get_next_audio_chunk();
model.process_audio_chunk(state, chunk.data(), chunk.size(), nullptr);
model.encode(state, false, nullptr);   // updates encoder memory
model.compute_cross_kv(state);         // pre-compute cross-attention KV
int token = 1; // BOS
float logits[config.vocab_size];
model.decode_step(state, token, logits); // tiny auto-regressive step

```

Key methods include:
- `process_audio_chunk`: Handles new audio and frontend processing.
- `encode`: Runs the ergodic encoder on new frames only.
- `compute_cross_kv`: Generates the reusable cross-attention cache.
- `decode_step`: Performs minimal decoding using cached states.

## Summary

Ergodic streaming is the architectural innovation that enables Moonshine’s real-time ASR models to achieve sub-200 ms latency on edge devices. Key takeaways include:

- **Frame-level computation reuse**: The ergodic encoder in [`core/moonshine-streaming-model.cpp`](https://github.com/moonshine-ai/moonshine/blob/main/core/moonshine-streaming-model.cpp) processes each audio frame exactly once, storing representations in `MoonshineStreamingState.accumulated_features`.
- **Persistent KV caching**: The decoder reuses pre-computed cross-attention keys and values via `compute_cross_kv()` and `run_decoder_with_cross_kv()`, minimizing per-token matrix operations.
- **State management**: The `MoonshineStreamingState` struct in [`core/moonshine-streaming-model.h`](https://github.com/moonshine-ai/moonshine/blob/main/core/moonshine-streaming-model.h) maintains convolution state, encoder memory, and attention caches across chunks.
- **Multi-language support**: Python ([`transcriber.py`](https://github.com/moonshine-ai/moonshine/blob/main/transcriber.py)), Swift ([`Transcriber.swift`](https://github.com/moonshine-ai/moonshine/blob/main/Transcriber.swift)), and direct C++ APIs all expose the same ergodic streaming engine.

## Frequently Asked Questions

### What does "ergodic" mean in this context?

In Moonshine’s architecture, "ergodic" refers to the statistical mechanics principle that the time average of a process equals its ensemble average. Practically, this means the encoder can safely reuse past computations—treating the accumulated history of processed frames as representative of the full acoustic context—without reprocessing audio segments. This ergodic property allows the system to maintain transcription accuracy while computing each frame only once.

### How does ergodic streaming achieve sub-200 ms latency?

Ergodic streaming achieves ultra-low latency through two primary mechanisms implemented in [`core/moonshine-streaming-model.cpp`](https://github.com/moonshine-ai/moonshine/blob/main/core/moonshine-streaming-model.cpp). First, the **ergodic encoder** eliminates redundant computation by processing each audio frame exactly once and caching results in `accumulated_features`. Second, the decoder uses **pre-computed cross-attention KV caches** (via `compute_cross_kv()`) so that after initial setup, each new token requires only a minimal auto-regressive step (`decode_step()`) rather than full attention recomputation.

### What is the difference between Moonshine’s streaming and non-streaming models?

Moonshine’s streaming models (tiny-streaming, small-streaming, medium-streaming) utilize the ergodic streaming encoder and maintain persistent state via `MoonshineStreamingState` to process audio chunks in real-time with sub-200 ms latency. Non-streaming models process entire utterances in a single forward pass without maintaining cross-chunk state, resulting in higher latency for real-time applications but potentially higher accuracy for offline transcription. The streaming models specifically leverage the `cross_kv.onnx` cache and the ergodic encoder’s frame-reuse strategy defined in [`core/moonshine-streaming-model.h`](https://github.com/moonshine-ai/moonshine/blob/main/core/moonshine-streaming-model.h).

### Where can I find the research paper describing ergodic streaming?

The technical details of the ergodic streaming architecture are documented in the research paper **"Moonshine v2: Ergodic Streaming Encoder ASR for Latency-Critical Speech"** (arXiv:2602.12241). The paper explains the statistical foundations of the ergodic approach and provides benchmarks comparing latency and accuracy against traditional streaming ASR systems. The implementation in the moonshine-ai/moonshine repository directly reflects the architecture described in this paper.