How the Transcribe Command Uses Whisper Providers in Agent-Reach

The transcribe command converts audio to text by routing requests to Groq or OpenAI Whisper APIs with automatic fallback handling, processing both remote URLs and local files through a chunked pipeline implemented in agent_reach/transcribe.py.

The Agent-Reach CLI provides a robust transcription interface that leverages Whisper-compatible speech-to-text APIs. This functionality supports multiple providers and implements intelligent fallback mechanisms to maximize transcription reliability across varying network conditions and API availability.

Supported Whisper Providers and Configuration

Agent-Reach integrates with two distinct Whisper providers, each configured through specific API keys loaded from environment variables or .agent-reach.yaml via agent_reach/config.py.

Groq Provider Configuration

The Groq provider utilizes the whisper-large-v3 model and sends HTTP POST requests to https://api.groq.com/openai/v1/audio/transcriptions. This provider requires the groq_api_key configuration value and serves as the default in auto mode due to its generous free tier.

OpenAI Provider Configuration

The OpenAI provider uses the whisper-1 model with the endpoint https://api.openai.com/v1/audio/transcriptions. Configuration requires the openai_api_key setting, and this provider acts as the secondary fallback when Groq requests fail.

Provider Selection Logic and Fallback Strategy

The provider selection mechanism is implemented in agent_reach/transcribe.py through the _provider_order() function, which determines the execution sequence based on CLI arguments parsed in agent_reach/cli.py.

Auto Mode Resolution

When the --provider auto flag is specified (or defaulted), _provider_order() returns ["groq", "openai"], prioritizing Groq while maintaining OpenAI as a backup. Before any audio processing begins, _provider_key() validates that at least one API key exists for the chosen providers, raising NoProviderConfigured if validation fails.

Explicit Provider Selection

Users can force a specific provider by passing --provider groq or --provider openai, which returns a single-element list (["groq"] or ["openai"] respectively), bypassing the fallback mechanism entirely.

The Transcription Pipeline (Implementation Details)

The complete transcription workflow in agent_reach/transcribe.py follows a multi-stage pipeline designed to handle large audio files efficiently while maintaining provider flexibility.

Audio Acquisition and Preparation

For remote URLs, the download_audio() function utilizes yt-dlp to fetch the source content. The compress_audio() function then re-encodes the audio to optimize file size, followed by chunk_audio() which splits the content into segments of ≤10 minutes to comply with API constraints.

Chunk Processing with Fallback

The _transcribe_with_fallback() function iterates through the provider order for each audio chunk. It calls transcribe_chunk(), which POSTs the binary audio data to the provider's endpoint with a Bearer token derived from the corresponding API key. If a request fails due to network errors or quota limits, the error is stored and the loop continues to the next provider in the sequence.

Result Aggregation

After successful transcription of all chunks, the main transcribe() function concatenates the individual text segments into a single continuous string, preserving the temporal order of the original audio source.

Code Examples

CLI usage demonstrating provider selection and output directory specification:


# Auto-select provider (Groq first, OpenAI fallback)

agent-reach transcribe "https://www.youtube.com/watch?v=VIDEO_ID"

# Force specific provider

agent-reach transcribe "https://example.com/podcast.mp3" --provider openai

# Specify output directory for intermediate files

agent-reach transcribe local_audio.wav -o /tmp/transcription/

Python API usage for programmatic integration:

from agent_reach.transcribe import transcribe, TranscribeError

try:
    # Auto provider selection with fallback

    text = transcribe("https://youtu.be/abcdef", provider="auto")
    print(text)
except TranscribeError as exc:
    print(f"Transcription failed: {exc}")

Summary

  • Agent-Reach supports two Whisper providers: Groq (whisper-large-v3) and OpenAI (whisper-1), configured via groq_api_key and openai_api_key respectively.
  • Automatic fallback prioritizes Groq: The auto mode attempts Groq first, falling back to OpenAI only if the primary request fails, as implemented in _provider_order().
  • Chunked processing handles large files: Audio is automatically compressed and split into 10-minute segments by chunk_audio() before API submission.
  • Resilient error handling: The _transcribe_with_fallback() function ensures transcription attempts continue across providers until success or exhaustion of options.

Frequently Asked Questions

What Whisper models does Agent-Reach support?

Agent-Reach specifically utilizes whisper-large-v3 for Groq API requests and whisper-1 for OpenAI requests. These model selections are hardcoded in agent_reach/transcribe.py to optimize for accuracy and API availability across both providers.

How does the transcribe command handle network errors?

The _transcribe_with_fallback() function implements resilient error handling by iterating through the provider order established by _provider_order(). If a request to the primary provider fails due to network issues, rate limits, or quota exhaustion, the system automatically retries with the secondary provider before returning an error to the user.

Can I use the transcribe functionality without the CLI?

Yes. The transcribe() function in agent_reach/transcribe.py is fully importable and can be called directly from Python scripts or channel implementations such as agent_reach/channels/youtube.py. It accepts the same parameters as the CLI, including provider selection and output directory options.

What is the maximum audio file size supported?

While the APIs themselves have varying limits, Agent-Reach preprocesses all audio by compressing and chunking it into segments of no more than 10 minutes using chunk_audio(). This preprocessing ensures compatibility with provider constraints and allows processing of arbitrarily long audio files through sequential chunk transcription and concatenation.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →