How VoiceStudio Implements Silent-Model Fallback in Dictation

VoiceStudio detects silent Sherpa ONNX models during WebSocket transcription and automatically recovers by re-running audio through a non-silent fallback model using the _recover_silent_sherpa coroutine, while ensuring no redundant network downloads occur through strict installed-model checks.

VoiceStudio's dictation pipeline relies on the Sherpa ONNX backend (services/sherpa_dictation.py) to convert speech to text. When a user selects a "silent" model—one that transcribes without producing audible output—the system must seamlessly switch to an audible alternative without interrupting the streaming capture session. This fallback mechanism is orchestrated within the WebSocket capture handler in backend/api/routers/capture_ws.py.

Silent-Model Detection in the Capture Pipeline

Identifying Silent Output

The detection logic resides in the CaptureWebSocket handler. When the Sherpa recognizer finishes processing an audio buffer, it returns a result dictionary containing a model_silent boolean flag. This flag is populated based on the model descriptor in services/sherpa_dictation.py, where specific entries in the _MODELS list are marked with "silent": true.

The Detection Check

At approximately line 897 in backend/api/routers/capture_ws.py, the handler inspects the transcription result:

final = await recognizer.transcribe(audio_chunk)
if final.get("model_silent"):
    # Trigger recovery to audible fallback

    recovered, segments = await self._recover_silent_sherpa(audio_chunk)

Fallback Recovery Implementation

The _recover_silent_sherpa Coroutine

Defined at line 644 in backend/api/routers/capture_ws.py, this private asynchronous method handles the transition from silent to audible output. It fetches the fallback model ID—typically the default non-silent dictation variant—and retrieves a warm backend instance.

Backend Caching and Reuse

The fallback leverages the singleton factory get_sherpa_dictation_backend() implemented in services/sherpa_dictation.py. This function caches initialized backends per model_id, ensuring that the fallback instantiation reuses existing memory structures rather than reloading weights from disk.

async def _recover_silent_sherpa(self, audio_chunk: bytes):
    """
    Recover from silent model by re-transcribing with fallback.
    Source: backend/api/routers/capture_ws.py#L644
    """
    fallback_id = dictation_fallback_model_id()
    fallback_backend = get_sherpa_dictation_backend(model_id=fallback_id)
    
    # Re-run transcription with non-silent model

    recovered = await fallback_backend.transcribe(audio_chunk)
    return recovered, recovered.segments

Preventing Redundant Model Downloads

Installed Model Verification

The system avoids network activity during fallback by enforcing the skip_sherpa and require_installed flags. Before attempting to load any model, the capture handler verifies that the fallback model exists in the local cache. If the model is missing, the recovery aborts rather than triggering an automatic download.

Test Coverage for Offline Operation

The test suite in tests/test_sherpa_ws_streaming.py validates this behavior around line 234. The test patches sherpa_dictation.download_model to raise an assertion error if invoked, confirming that silent-model fallback operations remain strictly offline when the fallback model is already installed.

def test_offline_silent_model_does_not_download_a_fallback(monkeypatch):
    """Source: tests/test_sherpa_ws_streaming.py#L234"""
    monkeypatch.setattr(
        sherpa_dictation, 
        "download_model", 
        lambda *_: pytest.fail("Download should not trigger")
    )
    # Streaming execution proceeds without network calls

Key Source Files and Line References

Code Implementation Examples

Complete Detection and Recovery Flow


# Inside CaptureWebSocket._transcribe_buffer_full

final = await recognizer.transcribe(audio_chunk)

if final.get("model_silent"):
    # Line ~897: Switch to audible fallback model

    result, segments = await self._recover_silent_sherpa(audio_chunk)
    
    return {
        "text": result.text,
        "model_silent": False,
        "segments": segments,
        "fallback_applied": True
    }
return final

Model Configuration Structure


# services/sherpa_dictation.py

_MODELS = [
    {
        "id": "sherpa-whisper-tiny-silent",
        "name": "Whisper Tiny Silent",
        "silent": True,
        "fallback_id": "sherpa-whisper-tiny"
    },
    {
        "id": "sherpa-whisper-tiny",
        "name": "Whisper Tiny",
        "silent": False
    }
]

def dictation_fallback_model_id() -> str:
    """Returns the default non-silent model identifier."""
    return "sherpa-whisper-tiny"

Summary

  • Detection: The system checks final["model_silent"] in the WebSocket capture handler to identify when a silent model produces transcription output
  • Recovery: The _recover_silent_sherpa coroutine at line 644 switches to a non-silent fallback model and re-transcribes the audio buffer without client interruption
  • Optimization: Backend caching via get_sherpa_dictation_backend() prevents duplicate model loading overhead and memory duplication
  • Offline-First: Flags like require_installed ensure fallback operations never trigger network downloads when models are cached locally
  • Verification: Comprehensive tests in test_sherpa_ws_streaming.py confirm the fallback mechanism respects offline constraints and completes within streaming latency requirements

Frequently Asked Questions

What triggers the silent-model fallback in VoiceStudio?

The fallback triggers when the Sherpa recognizer returns a transcription result with model_silent set to true. This occurs when the active model's descriptor in services/sherpa_dictation.py contains "silent": true, indicating the model transcribes speech without producing audible feedback. The CaptureWebSocket handler detects this flag at approximately line 897 and immediately invokes _recover_silent_sherpa to re-process the audio through an audible model.

Does the fallback mechanism download new models automatically?

No. The implementation explicitly prevents automatic downloads during fallback by checking require_installed flags before invoking the recovery. The test suite verifies this by patching the download_model function to fail if called, ensuring the fallback only proceeds when the target model already exists in the local filesystem cache.

How does VoiceStudio avoid performance penalties when switching models?

VoiceStudio uses a backend caching pattern through get_sherpa_dictation_backend() in services/sherpa_dictation.py. This function maintains a dictionary of initialized model instances, so when _recover_silent_sherpa requests the fallback model, it retrieves a warm backend rather than loading weights from storage. This ensures the switch adds only the latency of a second transcription pass, not model initialization time.

Where is the silent-model fallback logic implemented in the source code?

The primary implementation resides in backend/api/routers/capture_ws.py at line 644, where the _recover_silent_sherpa coroutine is defined. The invocation occurs around line 897 in the same file, while model definitions and backend factories are located in services/sherpa_dictation.py. Integration tests confirming offline-only behavior exist in tests/test_sherpa_ws_streaming.py between lines 169 and 236.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →