Moonshine Streaming State Management for Incremental Decoding: Architecture and Implementation
Moonshine's streaming state management for incremental decoding relies on two tightly-coupled Python classes—Transcriber and Stream—which maintain decoding state, interface with native C-library calls, and dispatch real-time transcript events.
The moonshine-ai/moonshine repository implements a sophisticated event-driven architecture for live audio transcription. This system enables incremental decoding by buffering audio chunks, tracking stream timing, and emitting structured events whenever the underlying speech recognition model produces new or updated hypotheses.
Core Components of Moonshine Streaming Architecture
The Transcriber Class
The Transcriber class serves as the high-level façade for streaming operations. Located in python/src/moonshine_voice/transcriber.py, it loads the native Moonshine library via _MoonshineLib, creates a transcriber handle, and manages a default Stream instance in self._default_stream.
Key responsibilities include:
- Model initialization: Loads weights and creates the native transcriber context.
- Stream lifecycle: Provides
create_stream()to instantiate new streaming sessions and stores the default stream for convenience. - Convenience methods: Shortcuts like
add_audio(),start(), andstop()forward calls to the default stream, simplifying single-stream use cases.
The Stream Class
The Stream class represents an individual live transcription session. It maintains the mutable state required for incremental decoding and wraps all native C API interactions. Each stream instance holds a native handle in self._handle and tracks temporal state through three critical fields:
_stream_time: Accumulated duration of audio processed._last_update_time: Timestamp of the last decoding update._update_interval: Configurable threshold (in seconds) that triggers automatic transcription updates when elapsed audio exceeds this value.
The Stream class buffers incoming audio and delegates decoding to native functions: moonshine_add_audio_to_stream, moonshine_transcribe_stream, and moonshine_free_stream.
Native C API Bindings
The Python layer communicates with the optimized decoder through ctypes bindings defined in python/src/moonshine_voice/moonshine_api.py. This module declares low-level functions including:
moonshine_create_streammoonshine_start_streammoonshine_add_audio_to_streammoonshine_transcribe_streammoonshine_free_stream
Error handling is centralized in the check_error function, which validates return codes from the native library and raises Python exceptions accordingly.
Event System and Listeners
Moonshine implements an event-driven architecture for delivering transcription results. The system defines structured event types in python/src/moonshine_voice/transcriber.py:
LineStarted: Emitted when a new transcript line begins.LineUpdated: Emitted when an existing line's hypothesis changes.LineTextChanged: Emitted specifically when the text content of a line changes.LineCompleted: Emitted when a line is finalized.Error: Emitted when an exception occurs during processing.
All events inherit from TranscriptEvent and carry the stream_handle and associated TranscriptLine. User code can implement the TranscriptEventListener interface or provide callable functions via add_listener() to receive these events.
State Management Flow for Incremental Decoding
Moonshine's streaming state management follows a precise lifecycle that enables real-time incremental decoding:
-
Stream Creation:
Transcriber.create_stream()invokesmoonshine_create_streamvia the native API, storing the returned handle in a newStreaminstance. The stream initializes_stream_timeto0.0and sets the update interval (defaulting to 0.2 seconds). -
Audio Ingestion: The
add_audio()method converts Python float arrays to C-compatible buffers and forwards them tomoonshine_add_audio_to_stream. It increments_stream_timeby the audio duration. If the elapsed time since_last_update_timeexceeds_update_interval, it automatically triggersupdate_transcription(). -
Incremental Decoding:
update_transcription()callsmoonshine_transcribe_stream, receiving aTranscriptCstruct pointer. The method parses this native structure into PythonTranscriptandTranscriptLineobjects, then invokes_notify_from_transcript()to process state changes. -
Event Emission:
_notify_from_transcript()examines boolean flags on each line (is_new,is_updated,has_text_changed,is_complete) to determine which events to emit. It iterates overself._listenersand dispatches the appropriate event objects. Errors during listener execution are caught and re-emitted asErrorevents without terminating the stream. -
Stream Termination:
stop()finalizes the stream by optionally flushing remaining audio through a finalupdate_transcription()call, then invokesmoonshine_free_streamto release native resources. The method returns the finalTranscriptcontaining all completed lines.
Implementation Examples
Basic Real-Time Transcription
The following example demonstrates initializing a transcriber, registering a listener, and processing audio chunks incrementally:
from moonshine_voice.transcriber import Transcriber
# Load a streaming model (e.g., tiny-streaming)
trans = Transcriber(model_path="models/tiny-streaming-en")
# Start the default stream
trans.start()
# Register a simple listener
def print_line(event):
print(f"[{event.line.start_time:.2f}s] {event.line.text}")
trans.add_listener(print_line)
# Simulate microphone chunks (replace with real audio)
audio, sr = load_wav_file("examples/audio.wav")
chunk_sz = int(0.1 * sr) # 100 ms chunks
for i in range(0, len(audio), chunk_sz):
trans.add_audio(audio[i:i+chunk_sz], sr)
# Finish and retrieve the final transcript
final = trans.stop()
for line in final.lines:
print(line.text)
trans.close()
Custom Event Listeners
For type-safe event handling, implement the TranscriptEventListener interface:
from moonshine_voice.transcriber import TranscriptEventListener, LineStarted, LineCompleted
class MyListener(TranscriptEventListener):
def on_line_started(self, event: LineStarted):
print(">>> Start:", event.line.text)
def on_line_completed(self, event: LineCompleted):
print("<<< Done:", event.line.text)
trans = Transcriber(...)
trans.add_listener(MyListener())
trans.start()
# … feed audio …
trans.stop()
Configuring Update Intervals
Control the frequency of incremental decoding by specifying the update_interval parameter when creating a stream:
# Create a stream that emits updates every 0.2 seconds
stream = trans.create_stream(update_interval=0.2)
stream.start()
# feed audio via stream.add_audio(...)
transcript = stream.stop()
Key Source Files
| File | Purpose |
|---|---|
python/src/moonshine_voice/transcriber.py |
Implements Transcriber, Stream, event types (LineStarted, LineCompleted, etc.), and listener dispatch logic. |
python/src/moonshine_voice/moonshine_api.py |
ctypes bindings to the native Moonshine library, declaring stream creation, audio addition, and transcription functions. |
python/src/moonshine_voice/utils.py |
Helpers for locating model assets and loading WAV files. |
examples/python/basic_transcription.py |
Minimal example demonstrating streaming transcription usage. |
python/src/moonshine_voice/mic_transcriber.py |
Convenience wrapper that binds a microphone source to a Transcriber stream. |
Summary
- Moonshine's streaming state management for incremental decoding centers on the
TranscriberandStreamclasses inpython/src/moonshine_voice/transcriber.py. - The
Streamclass maintains temporal state (_stream_time,_last_update_time,_update_interval) and manages a native handle for C-library interaction. - Audio ingestion triggers automatic transcription updates when buffered audio exceeds the configured
_update_interval. - Event-driven architecture emits typed events (
LineStarted,LineUpdated,LineCompleted) through theTranscriptEventListenerinterface, enabling real-time UI updates. - Resource lifecycle is strictly managed:
moonshine_create_streaminitializes native resources, whilemoonshine_free_streamreleases them onstop().
Frequently Asked Questions
How does Moonshine decide when to perform incremental decoding?
Moonshine uses a time-based threshold mechanism. The Stream class tracks _stream_time (total audio duration added) and _last_update_time (timestamp of last decode). When the difference exceeds _update_interval (defaulting to 0.2 seconds), add_audio() automatically triggers update_transcription(), which calls the native moonshine_transcribe_stream function.
What is the difference between the Transcriber and Stream classes?
The Transcriber acts as a high-level factory and façade. It loads the model, manages the native library via _MoonshineLib, and provides convenience methods that delegate to a default Stream instance. The Stream class represents a single transcription session, maintaining the native handle, temporal state, and listener registry. While Transcriber manages global resources, Stream manages the lifecycle of a specific audio input.
How can I handle transcription events in real-time?
Implement the TranscriptEventListener interface or pass a callable to add_listener(). The stream emits typed events—LineStarted, LineUpdated, LineTextChanged, and LineCompleted—based on flags in the native transcript (is_new, is_updated, has_text_changed, is_complete). These events contain the TranscriptLine and stream_handle, allowing you to update UIs or trigger downstream processing as audio is processed incrementally.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →