What Is Ergodic Streaming in Moonshine’s Real-Time Transcription?

Ergodic streaming is Moonshine’s architectural technique that achieves sub-200 ms end-to-end latency by processing each audio frame exactly once and reusing cached encoder outputs and cross-attention KV states across overlapping audio windows.

Ergodic streaming powers the real-time transcription capabilities of Moonshine, an open-source automatic speech recognition (ASR) engine developed by moonshine-ai. Unlike conventional streaming ASR systems that recompute encoder features for every new audio chunk, Moonshine’s ergodic streaming encoder treats the audio stream as an ergodic process—reusing computations and maintaining persistent state to minimize latency on edge devices.

How Ergodic Streaming Works

The Problem with Traditional Streaming ASR

In conventional streaming ASR pipelines, the encoder re-processes every audio frame each time a new chunk arrives. When overlapping windows are used to maintain acoustic context, this redundancy wastes compute and adds unnecessary latency. Every inference step repeats matrix operations on audio segments that have already been analyzed.

Frame-Level Reuse and KV Cache Persistence

Moonshine’s ergodic encoder, implemented in core/moonshine-streaming-model.cpp, eliminates this redundancy by processing each audio frame exactly once. It stores intermediate representations in accumulated_features within the MoonshineStreamingState struct (defined in core/moonshine-streaming-model.h).

The system stitches together results from overlapping windows so the decoder consumes a continuously updated "memory" of the utterance. The decoder leverages a pre-computed cross-attention KV cache (the cross_kv.onnx model) via compute_cross_kv() and run_decoder_with_cross_kv(). After the initial setup, heavy matrix multiplications are already cached; only a minimal auto-regressive step via decode_step() is needed for each new token.

Implementation in Moonshine’s C++ Core

The ergodic streaming architecture centers on two primary structures in core/moonshine-streaming-model.h:

  • MoonshineStreamingConfig: Holds model hyper-parameters including encoder dimensions and look-ahead settings.
  • MoonshineStreamingState: Maintains per-stream buffers including raw samples, convolution state, accumulated encoder features, decoder memory, and self- and cross-attention KV caches (with a cross_kv_valid flag).

The streaming loop in core/moonshine-streaming-model.cpp orchestrates the process:

  1. process_audio_chunk(): Accepts new audio data, updates frontend buffers, and triggers the ergodic encoder on only the new frames.
  2. encode(): Updates encoder memory with new features and appends to accumulated_features.
  3. compute_cross_kv(): Pre-computes cross-attention keys and values from the accumulated memory.
  4. decode_step(): Performs minimal auto-regressive decoding using the cached cross-attention states.

This design enables the tiny-streaming, small-streaming, and medium-streaming model variants to deliver sub-200 ms end-to-end latency while the speaker is still talking.

Code Examples

Python Streaming API

The Python wrapper in python/src/moonshine_voice/transcriber.py exposes the ergodic streaming engine via Transcriber.add_audio(), which maps to the C++ moonshine_transcribe function.

from moonshine_voice import Transcriber, TranscriptEventListener

class PrintListener(TranscriptEventListener):
    def on_line_started(self, event):
        print("[START] ", event.line.text)

    def on_line_text_changed(self, event):
        print("[UPDATE] ", event.line.text)

    def on_line_completed(self, event):
        print("[DONE]   ", event.line.text)

# Load a streaming model (e.g. tiny-streaming)

transcriber = Transcriber(
    model_path="path/to/tiny-streaming-en",
    model_arch="tiny-streaming",
)

transcriber.add_listener(PrintListener())
transcriber.start()

# Simulate live audio by feeding 100 ms chunks from a wav file

audio, sr = load_wav_file("examples/python/two_cities.wav")
chunk_ms = 0.1
chunk_samples = int(chunk_ms * sr)

for i in range(0, len(audio), chunk_samples):
    transcriber.add_audio(audio[i:i + chunk_samples], sr)

transcriber.stop()

Each add_audio call triggers process_audio_chunk in the C++ core, running the ergodic encoder on only the new audio data.

Swift Real-Time Transcription

The Swift bindings in swift/Sources/MoonshineVoice/Transcriber.swift and MicTranscriber.swift forward audio chunks to the same C++ streaming engine.

import MoonshineVoice

let transcriber = try Transcriber(
    modelPath: "/path/to/tiny-streaming-en",
    modelArch: .tinyStreaming
)

let mic = MicTranscriber(transcriber: transcriber)

class Listener: TranscriptEventListener {
    func onLineStarted(_ event: TranscriptEvent) {
        print("▶︎", event.line.text)
    }
    func onLineTextChanged(_ event: TranscriptEvent) {
        print("…", event.line.text)
    }
    func onLineCompleted(_ event: TranscriptEvent) {
        print("✔︎", event.line.text)
    }
}
transcriber.addListener(Listener())

try mic.start()          // captures microphone and streams chunks internally
// … later …
try mic.stop()

The MicTranscriber class manages the audio session and calls the underlying MoonshineStreamingModel to process chunks via the ergodic encoder.

Direct C++ Usage

For applications requiring direct integration, the C++ API in core/moonshine-streaming-model.h provides fine-grained control over the ergodic streaming state.

MoonshineStreamingModel model;
model.load("model_dir", "tokenizer.bin", /*model_type=*/1);
auto *state = model.create_state();

std::vector<float> chunk = get_next_audio_chunk();
model.process_audio_chunk(state, chunk.data(), chunk.size(), nullptr);
model.encode(state, false, nullptr);   // updates encoder memory
model.compute_cross_kv(state);         // pre-compute cross-attention KV
int token = 1; // BOS
float logits[config.vocab_size];
model.decode_step(state, token, logits); // tiny auto-regressive step

Key methods include:

  • process_audio_chunk: Handles new audio and frontend processing.
  • encode: Runs the ergodic encoder on new frames only.
  • compute_cross_kv: Generates the reusable cross-attention cache.
  • decode_step: Performs minimal decoding using cached states.

Summary

Ergodic streaming is the architectural innovation that enables Moonshine’s real-time ASR models to achieve sub-200 ms latency on edge devices. Key takeaways include:

  • Frame-level computation reuse: The ergodic encoder in core/moonshine-streaming-model.cpp processes each audio frame exactly once, storing representations in MoonshineStreamingState.accumulated_features.
  • Persistent KV caching: The decoder reuses pre-computed cross-attention keys and values via compute_cross_kv() and run_decoder_with_cross_kv(), minimizing per-token matrix operations.
  • State management: The MoonshineStreamingState struct in core/moonshine-streaming-model.h maintains convolution state, encoder memory, and attention caches across chunks.
  • Multi-language support: Python (transcriber.py), Swift (Transcriber.swift), and direct C++ APIs all expose the same ergodic streaming engine.

Frequently Asked Questions

What does "ergodic" mean in this context?

In Moonshine’s architecture, "ergodic" refers to the statistical mechanics principle that the time average of a process equals its ensemble average. Practically, this means the encoder can safely reuse past computations—treating the accumulated history of processed frames as representative of the full acoustic context—without reprocessing audio segments. This ergodic property allows the system to maintain transcription accuracy while computing each frame only once.

How does ergodic streaming achieve sub-200 ms latency?

Ergodic streaming achieves ultra-low latency through two primary mechanisms implemented in core/moonshine-streaming-model.cpp. First, the ergodic encoder eliminates redundant computation by processing each audio frame exactly once and caching results in accumulated_features. Second, the decoder uses pre-computed cross-attention KV caches (via compute_cross_kv()) so that after initial setup, each new token requires only a minimal auto-regressive step (decode_step()) rather than full attention recomputation.

What is the difference between Moonshine’s streaming and non-streaming models?

Moonshine’s streaming models (tiny-streaming, small-streaming, medium-streaming) utilize the ergodic streaming encoder and maintain persistent state via MoonshineStreamingState to process audio chunks in real-time with sub-200 ms latency. Non-streaming models process entire utterances in a single forward pass without maintaining cross-chunk state, resulting in higher latency for real-time applications but potentially higher accuracy for offline transcription. The streaming models specifically leverage the cross_kv.onnx cache and the ergodic encoder’s frame-reuse strategy defined in core/moonshine-streaming-model.h.

Where can I find the research paper describing ergodic streaming?

The technical details of the ergodic streaming architecture are documented in the research paper "Moonshine v2: Ergodic Streaming Encoder ASR for Latency-Critical Speech" (arXiv:2602.12241). The paper explains the statistical foundations of the ergodic approach and provides benchmarks comparing latency and accuracy against traditional streaming ASR systems. The implementation in the moonshine-ai/moonshine repository directly reflects the architecture described in this paper.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →