Family-Specific Extensions in transcribe.cpp: Parakeet, Whisper, Canary, and Voxtral

transcribe.cpp implements a modular, plugin-style architecture where each model family—such as Parakeet, Whisper, Canary, and Voxtral—provides its own dedicated extension comprising specialized C++ classes for hyperparameters, weight layouts, and runtime behavior that the core dispatcher loads dynamically based on GGUF metadata tags.

The handy-computer/transcribe.cpp repository unifies diverse speech-to-text and text-generation architectures under a single inference framework. By isolating family-specific extensions into separate translation units under src/arch/, the library maintains a clean separation between generic orchestration logic and model-specific optimizations.

How Family Extensions Work

At load time, the GGUF parser in src/transcribe-bin-loader.cpp extracts the family key-value pair from model metadata. The dispatcher in src/transcribe.cpp (around line 1086) then maps this string to a family-specific factory function—such as parakeet_view, whisper_view, or voxtral_view—which instantiates the appropriate model and session objects.

Each extension exposes a public C ABI through arch/<family>/public.cpp, allowing the core library to cast opaque pointers (e.g., transcribe_session) to family-specific types (e.g., WhisperSession) when advanced features are required. This design guarantees that new families can be added without modifying src/transcribe.cpp, provided they implement the standard interface of model.cpp, weights.cpp, and optional encoder.cpp or decoder.cpp components.

Supported Model Families

Parakeet

Parakeet extensions support NeMo-derived automatic speech recognition (ASR) models with optimized streaming capabilities. The implementation resides in src/arch/parakeet/model.cpp, weights.cpp, and hparams.cpp.

Key characteristics include:

  • Chunked streaming modes: Supports regular chunked inference, ChunkedLimited, and ChunkedLimitedWithRc for controlled latency.
  • Batch-norm fusion: Optional fusion and float-32 promotion of depth-wise convolutions.
  • Fixed frontend: Uses an 80-bin mel spectrogram frontend defined in src/transcribe-mel.cpp (line 813).

Whisper

The Whisper extension implements the classic OpenAI Whisper architecture in src/arch/whisper/model.cpp, weights.cpp, and public.cpp.

Distinctive features include:

  • Tokenization: GPT-2 style byte-level BPE tokenizer.
  • Audio processing: Fixed 16 kHz sample rate with 80-bin mel filters and n_fft = 400.
  • Weight naming: Encoder and decoder share the attn.out.weight naming convention for attention projections.

Canary

Canary extensions handle multitask audio event detection (AED) supporting translation, language identification, and voice activity detection. Core files include src/arch/canary/model.cpp, weights.cpp, encoder.cpp, and decoder.cpp.

Notable capabilities:

  • Flexible normalization: Selectable layer-norm or batch-norm via CanaryHParams::norm_type.
  • Task tokens: Built-in support for <| transcribe |> and <| translate |> style task prefixes.

Canary-Qwen

A variant of Canary that substitutes the transformer backbone with Qwen-style attention layers. Implemented in src/arch/canary_qwen/model.cpp and weights.cpp, this extension shares most logic with the base Canary implementation but uses Qwen-specific projection layers for improved multilingual performance.

Voxtral

Voxtral provides Llama-style decoder architectures for text generation with optional multimodal inputs. The implementation spans src/arch/voxtral/model.cpp, weights.cpp, encoder.cpp, decoder.cpp, and capabilities.cpp.

Key limitations and features:

  • No timestamp output: Explicitly disabled in src/arch/voxtral/capabilities.cpp (unlike Whisper or Parakeet).
  • Instruction following: Exposes a "prompt-only" API for chat-style interactions.

Voxtral-Realtime

This streaming variant of Voxtral enables real-time mel-frame processing. Files are located in src/arch/voxtral_realtime/model.cpp, encoder.cpp, and decoder.cpp.

Technical specifics:

Qwen3-ASR (Experimental)

An experimental extension combining Whisper-style ASR with Qwen attention mechanisms. While not fully integrated into the main build, test references indicate the presence of src/arch/qwen3_asr/model.cpp. This family shares the standard Whisper frontend but replaces transformer blocks with Qwen3 architectures.

Runtime Dispatch Implementation

The dispatch mechanism centers on a family switch statement in src/transcribe.cpp at line 1086. When transcribe_load_model() processes a GGUF file, it extracts the family identifier and routes to the corresponding extension:

  1. Metadata extraction: src/transcribe-bin-loader.cpp parses the family string.
  2. Factory selection: The switch maps identifiers like "parakeet", "whisper", or "canary" to their respective view factories.
  3. Object initialization: The factory constructs family-specific Model and Session objects, initializing weights, hyperparameters, and streaming caches according to the extension's hparams.cpp.

This architecture ensures that memory layouts and computational paths remain optimal for each family while presenting a unified transcribe_model/transcribe_session interface to consumers.

Working with Family-Specific APIs

While the core API is family-agnostic, extensions expose specialized functionality through casting helpers defined in their public.cpp files.

Generic Model Loading

// Family-agnostic loading works for any supported GGUF
transcribe_model *model = nullptr;
transcribe_status s = transcribe_load_model("model.gguf", &model);
if (s != TRANSCRIBE_OK) {
    // Handle loading error
}

Accessing Whisper-Specific Features

To access Whisper-only functionality like token-level timestamps, cast the session using the helper from src/arch/whisper/public.cpp (line 30):

#include "arch/whisper/public.cpp"

const transcribe::whisper::WhisperSession *whisper_ctx =
    maybe_whisper_context(session);

if (whisper_ctx) {
    // Access Whisper-specific methods for timestamp retrieval
}

Configuring Voxtral-Realtime Streaming

For real-time inference with Voxtral, set the specific model path before initialization:

// Reference: tests/voxtral_realtime_real_smoke.cpp
transcribe_set_option(session, "TRANSCRIBE_VOXTRAL_REALTIME_GGUF", 
                      "path/to/voxtral-realtime.gguf");

Summary

  • transcribe.cpp uses a plugin architecture where each model family (Parakeet, Whisper, Canary, Voxtral, etc.) resides in src/arch/<family>/.
  • The dispatcher in src/transcribe.cpp (line 1086) routes to family-specific factories based on GGUF metadata.
  • Parakeet extensions support chunked streaming ASR with batch-norm fusion.
  • Whisper extensions implement OpenAI’s original encoder-decoder with GPT-2 BPE tokenization.
  • Canary and Canary-Qwen handle multitask AED with configurable normalization backends.
  • Voxtral provides text-generation capabilities without timestamp support, while Voxtral-Realtime adds fixed-buffer streaming for low-latency scenarios.
  • Family-specific APIs are accessed through casting helpers defined in each extension’s public.cpp.

Frequently Asked Questions

How does transcribe.cpp determine which family extension to load?

The library reads the family string from the GGUF metadata header via src/transcribe-bin-loader.cpp. It then executes a switch statement in src/transcribe.cpp (around line 1086) that instantiates the appropriate extension-specific Model and Session classes (e.g., ParakeetModel or WhisperSession).

Can I use Whisper-specific features like token timestamps with the generic API?

No, timestamp extraction requires casting the generic transcribe_session to a WhisperSession pointer using maybe_whisper_context(), defined in src/arch/whisper/public.cpp (line 30). This guard ensures that family-specific features are only accessed when the loaded model actually supports them.

What is the difference between Canary and Canary-Qwen extensions?

The base Canary extension in src/arch/canary/ uses a traditional transformer encoder-decoder optimized for audio event detection. Canary-Qwen (src/arch/canary_qwen/) replaces the transformer backbone with Qwen-style attention and projection layers while retaining the same AED task heads, offering enhanced multilingual capabilities through the Qwen architecture.

Is Voxtral-Realtime suitable for low-latency streaming applications?

Yes. Voxtral-Realtime implements a fixed-size global mel buffer (see src/transcribe-mel.cpp line 774) that processes audio chunks as they arrive, minimizing latency compared to full-context models. It is activated by setting the TRANSCRIBE_VOXTRAL_REALTIME_GGUF environment variable before session creation, as shown in the test suite.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →