# Family-Specific Extensions in transcribe.cpp: Parakeet, Whisper, Canary, and Voxtral

> Explore family-specific extensions in transcribe.cpp like Parakeet, Whisper, Canary, and Voxtral. Learn about specialized classes for hyperparameters, weights, and runtime behavior.

- Repository: [handy-computer/transcribe.cpp](https://github.com/handy-computer/transcribe.cpp)
- Tags: deep-dive
- Published: 2026-07-21

---

**transcribe.cpp** implements a modular, plugin-style architecture where each model family—such as Parakeet, Whisper, Canary, and Voxtral—provides its own dedicated extension comprising specialized C++ classes for hyperparameters, weight layouts, and runtime behavior that the core dispatcher loads dynamically based on GGUF metadata tags.

The [`handy-computer/transcribe.cpp`](https://github.com/handy-computer/transcribe.cpp/blob/main/handy-computer/transcribe.cpp) repository unifies diverse speech-to-text and text-generation architectures under a single inference framework. By isolating **family-specific extensions** into separate translation units under `src/arch/`, the library maintains a clean separation between generic orchestration logic and model-specific optimizations.

## How Family Extensions Work

At load time, the GGUF parser in [`src/transcribe-bin-loader.cpp`](https://github.com/handy-computer/transcribe.cpp/blob/main/src/transcribe-bin-loader.cpp) extracts the `family` key-value pair from model metadata. The dispatcher in [`src/transcribe.cpp`](https://github.com/handy-computer/transcribe.cpp/blob/main/src/transcribe.cpp) (around line 1086) then maps this string to a family-specific factory function—such as `parakeet_view`, `whisper_view`, or `voxtral_view`—which instantiates the appropriate model and session objects.

Each extension exposes a **public C ABI** through `arch/<family>/public.cpp`, allowing the core library to cast opaque pointers (e.g., `transcribe_session`) to family-specific types (e.g., `WhisperSession`) when advanced features are required. This design guarantees that new families can be added without modifying [`src/transcribe.cpp`](https://github.com/handy-computer/transcribe.cpp/blob/main/src/transcribe.cpp), provided they implement the standard interface of [`model.cpp`](https://github.com/handy-computer/transcribe.cpp/blob/main/model.cpp), [`weights.cpp`](https://github.com/handy-computer/transcribe.cpp/blob/main/weights.cpp), and optional [`encoder.cpp`](https://github.com/handy-computer/transcribe.cpp/blob/main/encoder.cpp) or [`decoder.cpp`](https://github.com/handy-computer/transcribe.cpp/blob/main/decoder.cpp) components.

## Supported Model Families

### Parakeet

Parakeet extensions support NeMo-derived automatic speech recognition (ASR) models with optimized streaming capabilities. The implementation resides in [`src/arch/parakeet/model.cpp`](https://github.com/handy-computer/transcribe.cpp/blob/main/src/arch/parakeet/model.cpp), [`weights.cpp`](https://github.com/handy-computer/transcribe.cpp/blob/main/weights.cpp), and [`hparams.cpp`](https://github.com/handy-computer/transcribe.cpp/blob/main/hparams.cpp).

Key characteristics include:
- **Chunked streaming modes**: Supports regular chunked inference, `ChunkedLimited`, and `ChunkedLimitedWithRc` for controlled latency.
- **Batch-norm fusion**: Optional fusion and float-32 promotion of depth-wise convolutions.
- **Fixed frontend**: Uses an 80-bin mel spectrogram frontend defined in [`src/transcribe-mel.cpp`](https://github.com/handy-computer/transcribe.cpp/blob/main/src/transcribe-mel.cpp) (line 813).

### Whisper

The Whisper extension implements the classic OpenAI Whisper architecture in [`src/arch/whisper/model.cpp`](https://github.com/handy-computer/transcribe.cpp/blob/main/src/arch/whisper/model.cpp), [`weights.cpp`](https://github.com/handy-computer/transcribe.cpp/blob/main/weights.cpp), and [`public.cpp`](https://github.com/handy-computer/transcribe.cpp/blob/main/public.cpp).

Distinctive features include:
- **Tokenization**: GPT-2 style byte-level BPE tokenizer.
- **Audio processing**: Fixed 16 kHz sample rate with 80-bin mel filters and `n_fft = 400`.
- **Weight naming**: Encoder and decoder share the `attn.out.weight` naming convention for attention projections.

### Canary

Canary extensions handle multitask audio event detection (AED) supporting translation, language identification, and voice activity detection. Core files include [`src/arch/canary/model.cpp`](https://github.com/handy-computer/transcribe.cpp/blob/main/src/arch/canary/model.cpp), [`weights.cpp`](https://github.com/handy-computer/transcribe.cpp/blob/main/weights.cpp), [`encoder.cpp`](https://github.com/handy-computer/transcribe.cpp/blob/main/encoder.cpp), and [`decoder.cpp`](https://github.com/handy-computer/transcribe.cpp/blob/main/decoder.cpp).

Notable capabilities:
- **Flexible normalization**: Selectable layer-norm or batch-norm via `CanaryHParams::norm_type`.
- **Task tokens**: Built-in support for `<| transcribe |>` and `<| translate |>` style task prefixes.

### Canary-Qwen

A variant of Canary that substitutes the transformer backbone with Qwen-style attention layers. Implemented in [`src/arch/canary_qwen/model.cpp`](https://github.com/handy-computer/transcribe.cpp/blob/main/src/arch/canary_qwen/model.cpp) and [`weights.cpp`](https://github.com/handy-computer/transcribe.cpp/blob/main/weights.cpp), this extension shares most logic with the base Canary implementation but uses Qwen-specific projection layers for improved multilingual performance.

### Voxtral

Voxtral provides Llama-style decoder architectures for text generation with optional multimodal inputs. The implementation spans [`src/arch/voxtral/model.cpp`](https://github.com/handy-computer/transcribe.cpp/blob/main/src/arch/voxtral/model.cpp), [`weights.cpp`](https://github.com/handy-computer/transcribe.cpp/blob/main/weights.cpp), [`encoder.cpp`](https://github.com/handy-computer/transcribe.cpp/blob/main/encoder.cpp), [`decoder.cpp`](https://github.com/handy-computer/transcribe.cpp/blob/main/decoder.cpp), and [`capabilities.cpp`](https://github.com/handy-computer/transcribe.cpp/blob/main/capabilities.cpp).

Key limitations and features:
- **No timestamp output**: Explicitly disabled in [`src/arch/voxtral/capabilities.cpp`](https://github.com/handy-computer/transcribe.cpp/blob/main/src/arch/voxtral/capabilities.cpp) (unlike Whisper or Parakeet).
- **Instruction following**: Exposes a "prompt-only" API for chat-style interactions.

### Voxtral-Realtime

This streaming variant of Voxtral enables real-time mel-frame processing. Files are located in [`src/arch/voxtral_realtime/model.cpp`](https://github.com/handy-computer/transcribe.cpp/blob/main/src/arch/voxtral_realtime/model.cpp), [`encoder.cpp`](https://github.com/handy-computer/transcribe.cpp/blob/main/encoder.cpp), and [`decoder.cpp`](https://github.com/handy-computer/transcribe.cpp/blob/main/decoder.cpp).

Technical specifics:
- **Fixed-size mel buffer**: Implements a global mel buffer managed in [`src/transcribe-mel.cpp`](https://github.com/handy-computer/transcribe.cpp/blob/main/src/transcribe-mel.cpp) (line 774) to maintain streaming state.
- **Environment configuration**: Activated via the `TRANSCRIBE_VOXTRAL_REALTIME_GGUF` environment variable, as demonstrated in [`tests/voxtral_realtime_real_smoke.cpp`](https://github.com/handy-computer/transcribe.cpp/blob/main/tests/voxtral_realtime_real_smoke.cpp).

### Qwen3-ASR (Experimental)

An experimental extension combining Whisper-style ASR with Qwen attention mechanisms. While not fully integrated into the main build, test references indicate the presence of [`src/arch/qwen3_asr/model.cpp`](https://github.com/handy-computer/transcribe.cpp/blob/main/src/arch/qwen3_asr/model.cpp). This family shares the standard Whisper frontend but replaces transformer blocks with Qwen3 architectures.

## Runtime Dispatch Implementation

The dispatch mechanism centers on a family switch statement in [`src/transcribe.cpp`](https://github.com/handy-computer/transcribe.cpp/blob/main/src/transcribe.cpp) at line 1086. When `transcribe_load_model()` processes a GGUF file, it extracts the family identifier and routes to the corresponding extension:

1. **Metadata extraction**: [`src/transcribe-bin-loader.cpp`](https://github.com/handy-computer/transcribe.cpp/blob/main/src/transcribe-bin-loader.cpp) parses the `family` string.
2. **Factory selection**: The switch maps identifiers like `"parakeet"`, `"whisper"`, or `"canary"` to their respective view factories.
3. **Object initialization**: The factory constructs family-specific `Model` and `Session` objects, initializing weights, hyperparameters, and streaming caches according to the extension's [`hparams.cpp`](https://github.com/handy-computer/transcribe.cpp/blob/main/hparams.cpp).

This architecture ensures that memory layouts and computational paths remain optimal for each family while presenting a unified `transcribe_model`/`transcribe_session` interface to consumers.

## Working with Family-Specific APIs

While the core API is family-agnostic, extensions expose specialized functionality through casting helpers defined in their [`public.cpp`](https://github.com/handy-computer/transcribe.cpp/blob/main/public.cpp) files.

### Generic Model Loading

```cpp
// Family-agnostic loading works for any supported GGUF
transcribe_model *model = nullptr;
transcribe_status s = transcribe_load_model("model.gguf", &model);
if (s != TRANSCRIBE_OK) {
    // Handle loading error
}

```

### Accessing Whisper-Specific Features

To access Whisper-only functionality like token-level timestamps, cast the session using the helper from [`src/arch/whisper/public.cpp`](https://github.com/handy-computer/transcribe.cpp/blob/main/src/arch/whisper/public.cpp) (line 30):

```cpp
#include "arch/whisper/public.cpp"

const transcribe::whisper::WhisperSession *whisper_ctx =
    maybe_whisper_context(session);

if (whisper_ctx) {
    // Access Whisper-specific methods for timestamp retrieval
}

```

### Configuring Voxtral-Realtime Streaming

For real-time inference with Voxtral, set the specific model path before initialization:

```cpp
// Reference: tests/voxtral_realtime_real_smoke.cpp
transcribe_set_option(session, "TRANSCRIBE_VOXTRAL_REALTIME_GGUF", 
                      "path/to/voxtral-realtime.gguf");

```

## Summary

- **transcribe.cpp** uses a plugin architecture where each model family (Parakeet, Whisper, Canary, Voxtral, etc.) resides in `src/arch/<family>/`.
- The dispatcher in [`src/transcribe.cpp`](https://github.com/handy-computer/transcribe.cpp/blob/main/src/transcribe.cpp) (line 1086) routes to family-specific factories based on GGUF metadata.
- Parakeet extensions support chunked streaming ASR with batch-norm fusion.
- Whisper extensions implement OpenAI’s original encoder-decoder with GPT-2 BPE tokenization.
- Canary and Canary-Qwen handle multitask AED with configurable normalization backends.
- Voxtral provides text-generation capabilities without timestamp support, while Voxtral-Realtime adds fixed-buffer streaming for low-latency scenarios.
- Family-specific APIs are accessed through casting helpers defined in each extension’s [`public.cpp`](https://github.com/handy-computer/transcribe.cpp/blob/main/public.cpp).

## Frequently Asked Questions

### How does transcribe.cpp determine which family extension to load?

The library reads the `family` string from the GGUF metadata header via [`src/transcribe-bin-loader.cpp`](https://github.com/handy-computer/transcribe.cpp/blob/main/src/transcribe-bin-loader.cpp). It then executes a switch statement in [`src/transcribe.cpp`](https://github.com/handy-computer/transcribe.cpp/blob/main/src/transcribe.cpp) (around line 1086) that instantiates the appropriate extension-specific `Model` and `Session` classes (e.g., `ParakeetModel` or `WhisperSession`).

### Can I use Whisper-specific features like token timestamps with the generic API?

No, timestamp extraction requires casting the generic `transcribe_session` to a `WhisperSession` pointer using `maybe_whisper_context()`, defined in [`src/arch/whisper/public.cpp`](https://github.com/handy-computer/transcribe.cpp/blob/main/src/arch/whisper/public.cpp) (line 30). This guard ensures that family-specific features are only accessed when the loaded model actually supports them.

### What is the difference between Canary and Canary-Qwen extensions?

The base **Canary** extension in `src/arch/canary/` uses a traditional transformer encoder-decoder optimized for audio event detection. **Canary-Qwen** (`src/arch/canary_qwen/`) replaces the transformer backbone with Qwen-style attention and projection layers while retaining the same AED task heads, offering enhanced multilingual capabilities through the Qwen architecture.

### Is Voxtral-Realtime suitable for low-latency streaming applications?

Yes. **Voxtral-Realtime** implements a fixed-size global mel buffer (see [`src/transcribe-mel.cpp`](https://github.com/handy-computer/transcribe.cpp/blob/main/src/transcribe-mel.cpp) line 774) that processes audio chunks as they arrive, minimizing latency compared to full-context models. It is activated by setting the `TRANSCRIBE_VOXTRAL_REALTIME_GGUF` environment variable before session creation, as shown in the test suite.