# How the Transcribe Command Uses Whisper Providers in Agent-Reach

> Learn how the Agent-Reach transcribe command uses Whisper providers like Groq and OpenAI for efficient audio to text conversion. Understand its chunked pipeline and fallback handling.

- Repository: [Pnant/Agent-Reach](https://github.com/Panniantong/Agent-Reach)
- Tags: internals
- Published: 2026-06-29

---

**The `transcribe` command converts audio to text by routing requests to Groq or OpenAI Whisper APIs with automatic fallback handling, processing both remote URLs and local files through a chunked pipeline implemented in [`agent_reach/transcribe.py`](https://github.com/Panniantong/Agent-Reach/blob/main/agent_reach/transcribe.py).**

The Agent-Reach CLI provides a robust transcription interface that leverages Whisper-compatible speech-to-text APIs. This functionality supports multiple providers and implements intelligent fallback mechanisms to maximize transcription reliability across varying network conditions and API availability.

## Supported Whisper Providers and Configuration

Agent-Reach integrates with two distinct Whisper providers, each configured through specific API keys loaded from environment variables or [`.agent-reach.yaml`](https://github.com/Panniantong/Agent-Reach/blob/main/.agent-reach.yaml) via [`agent_reach/config.py`](https://github.com/Panniantong/Agent-Reach/blob/main/agent_reach/config.py).

### Groq Provider Configuration

The **Groq** provider utilizes the `whisper-large-v3` model and sends HTTP POST requests to `https://api.groq.com/openai/v1/audio/transcriptions`. This provider requires the `groq_api_key` configuration value and serves as the default in `auto` mode due to its generous free tier.

### OpenAI Provider Configuration

The **OpenAI** provider uses the `whisper-1` model with the endpoint `https://api.openai.com/v1/audio/transcriptions`. Configuration requires the `openai_api_key` setting, and this provider acts as the secondary fallback when Groq requests fail.

## Provider Selection Logic and Fallback Strategy

The provider selection mechanism is implemented in [`agent_reach/transcribe.py`](https://github.com/Panniantong/Agent-Reach/blob/main/agent_reach/transcribe.py) through the `_provider_order()` function, which determines the execution sequence based on CLI arguments parsed in [`agent_reach/cli.py`](https://github.com/Panniantong/Agent-Reach/blob/main/agent_reach/cli.py).

### Auto Mode Resolution

When the `--provider auto` flag is specified (or defaulted), `_provider_order()` returns `["groq", "openai"]`, prioritizing Groq while maintaining OpenAI as a backup. Before any audio processing begins, `_provider_key()` validates that at least one API key exists for the chosen providers, raising `NoProviderConfigured` if validation fails.

### Explicit Provider Selection

Users can force a specific provider by passing `--provider groq` or `--provider openai`, which returns a single-element list (`["groq"]` or `["openai"]` respectively), bypassing the fallback mechanism entirely.

## The Transcription Pipeline (Implementation Details)

The complete transcription workflow in [`agent_reach/transcribe.py`](https://github.com/Panniantong/Agent-Reach/blob/main/agent_reach/transcribe.py) follows a multi-stage pipeline designed to handle large audio files efficiently while maintaining provider flexibility.

### Audio Acquisition and Preparation

For remote URLs, the `download_audio()` function utilizes `yt-dlp` to fetch the source content. The `compress_audio()` function then re-encodes the audio to optimize file size, followed by `chunk_audio()` which splits the content into segments of ≤10 minutes to comply with API constraints.

### Chunk Processing with Fallback

The `_transcribe_with_fallback()` function iterates through the provider order for each audio chunk. It calls `transcribe_chunk()`, which POSTs the binary audio data to the provider's endpoint with a **Bearer** token derived from the corresponding API key. If a request fails due to network errors or quota limits, the error is stored and the loop continues to the next provider in the sequence.

### Result Aggregation

After successful transcription of all chunks, the main `transcribe()` function concatenates the individual text segments into a single continuous string, preserving the temporal order of the original audio source.

## Code Examples

CLI usage demonstrating provider selection and output directory specification:

```bash

# Auto-select provider (Groq first, OpenAI fallback)

agent-reach transcribe "https://www.youtube.com/watch?v=VIDEO_ID"

# Force specific provider

agent-reach transcribe "https://example.com/podcast.mp3" --provider openai

# Specify output directory for intermediate files

agent-reach transcribe local_audio.wav -o /tmp/transcription/

```

Python API usage for programmatic integration:

```python
from agent_reach.transcribe import transcribe, TranscribeError

try:
    # Auto provider selection with fallback

    text = transcribe("https://youtu.be/abcdef", provider="auto")
    print(text)
except TranscribeError as exc:
    print(f"Transcription failed: {exc}")

```

## Summary

- **Agent-Reach supports two Whisper providers**: Groq (`whisper-large-v3`) and OpenAI (`whisper-1`), configured via `groq_api_key` and `openai_api_key` respectively.
- **Automatic fallback prioritizes Groq**: The `auto` mode attempts Groq first, falling back to OpenAI only if the primary request fails, as implemented in `_provider_order()`.
- **Chunked processing handles large files**: Audio is automatically compressed and split into 10-minute segments by `chunk_audio()` before API submission.
- **Resilient error handling**: The `_transcribe_with_fallback()` function ensures transcription attempts continue across providers until success or exhaustion of options.

## Frequently Asked Questions

### What Whisper models does Agent-Reach support?

Agent-Reach specifically utilizes `whisper-large-v3` for Groq API requests and `whisper-1` for OpenAI requests. These model selections are hardcoded in [`agent_reach/transcribe.py`](https://github.com/Panniantong/Agent-Reach/blob/main/agent_reach/transcribe.py) to optimize for accuracy and API availability across both providers.

### How does the transcribe command handle network errors?

The `_transcribe_with_fallback()` function implements resilient error handling by iterating through the provider order established by `_provider_order()`. If a request to the primary provider fails due to network issues, rate limits, or quota exhaustion, the system automatically retries with the secondary provider before returning an error to the user.

### Can I use the transcribe functionality without the CLI?

Yes. The `transcribe()` function in [`agent_reach/transcribe.py`](https://github.com/Panniantong/Agent-Reach/blob/main/agent_reach/transcribe.py) is fully importable and can be called directly from Python scripts or channel implementations such as [`agent_reach/channels/youtube.py`](https://github.com/Panniantong/Agent-Reach/blob/main/agent_reach/channels/youtube.py). It accepts the same parameters as the CLI, including `provider` selection and output directory options.

### What is the maximum audio file size supported?

While the APIs themselves have varying limits, Agent-Reach preprocesses all audio by compressing and chunking it into segments of no more than 10 minutes using `chunk_audio()`. This preprocessing ensures compatibility with provider constraints and allows processing of arbitrarily long audio files through sequential chunk transcription and concatenation.