How the Transcribe Command Uses Groq Whisper API with OpenAI Fallback in Agent Reach

The transcribe command prioritizes Groq's free Whisper API and automatically falls back to OpenAI's Whisper API if Groq fails, using a provider-order loop in agent_reach/transcribe.py to ensure reliable speech-to-text conversion without manual intervention.

Agent Reach's transcription system, implemented in the agent_reach/transcribe.py module, provides robust audio-to-text conversion by orchestrating multiple AI providers. The system is designed to maximize availability by first attempting transcription through Groq's Whisper API, then seamlessly degrading to OpenAI's Whisper API when necessary. This article examines the complete fallback mechanism, from provider configuration to error handling, based on the actual source code implementation in the Panniantong/Agent-Reach repository.

Provider Configuration and Routing

The Provider Definition Table

The foundation of the fallback system rests in the PROVIDERS dictionary, which maps logical provider names to their concrete HTTP endpoints and model configurations. Located in agent_reach/transcribe.py, this table defines the endpoint URLs, model names (whisper-large-v3 for Groq and whisper-1 for OpenAI), and the configuration keys required to retrieve API tokens from the user's settings in agent_reach/config.py.

Provider Selection Logic

The _provider_order function determines which providers to attempt and in what sequence. When provider="auto" is specified (the default), the system returns ["groq", "openai"], ensuring Groq is always attempted first. If a specific provider name is passed, only that provider is used, bypassing the fallback mechanism entirely.

Fallback Orchestration Mechanism

The core resilience logic resides in _transcribe_with_fallback, which implements a retry loop that attempts transcription with each provider in sequence. The function iterates over the provider order, skips any provider lacking a configured API key, and calls transcribe_chunk for the active provider.

If transcribe_chunk raises a TranscribeError (triggered by any non-2xx HTTP response), the loop captures the exception and proceeds to the next provider. Only when all providers fail does the function raise an aggregated error containing all captured exceptions, ensuring the user receives a transcription rather than a partial failure whenever possible.

Audio Processing Pipeline

Input Handling and Validation

Before initiating transcription, the transcribe function validates that at least one provider has a configured API key by checking any(_provider_key(p, cfg) for p in order). The function then handles input sources through download_audio (using yt-dlp for URLs) or accepts local file paths directly.

Audio Compression and Chunking

To ensure compatibility with Whisper's requirements, compress_audio re-encodes input to mono, 16 kHz, 32 kbps format. If the resulting file exceeds Whisper's size limit (~24 MiB), chunk_audio splits the audio into segments of ≤10 minutes each, processing each chunk independently through the fallback pipeline.

Chunk Transcription and Aggregation

For each chunk, transcribe_chunk constructs a multipart POST request containing the audio file, model parameter, and response_format=text setting. The request includes an Authorization: Bearer <api_key> header and targets the provider-specific endpoint (e.g., https://api.groq.com/openai/v1/audio/transcriptions for Groq).

Upon successful completion (HTTP 200), the raw text response is returned. The transcribe function then aggregates all chunk transcripts, stripping whitespace and joining them with newlines to produce the final result.

Practical Implementation Examples

Using the Python API

from agent_reach import transcribe
from agent_reach.config import Config

# Assumes keys are stored via `agent-reach configure`

cfg = Config()  # Loads keys from default config file

text = transcribe(
    "https://youtu.be/example",  # Can also be a local .m4a path

    provider="auto",  # Default: Groq → OpenAI fallback

    config=cfg,
)
print(text)

Using the CLI

agent-reach transcribe "https://youtu.be/example" --provider auto

Testing the Fallback Behavior

The unit tests in tests/test_transcribe.py validate the fallback mechanism through specific scenarios:

  • test_groq_succeeds_no_openai_call: Verifies OpenAI is never contacted when Groq returns HTTP 200
  • test_groq_429_falls_back_to_openai: Confirms fallback occurs when Groq returns HTTP 429 (rate limit)
  • test_skip_unconfigured_provider: Ensures providers without API keys are skipped automatically

Summary

  • The transcribe command in Agent Reach implements a robust fallback system that prioritizes Groq's Whisper API (whisper-large-v3) over OpenAI's (whisper-1)
  • Provider selection and routing are controlled by the PROVIDERS dictionary and _provider_order function in agent_reach/transcribe.py, with auto mode defaulting to the Groq-first sequence
  • The _transcribe_with_fallback function implements the resilience logic, attempting each provider in order and only failing after all options are exhausted
  • Audio preprocessing includes format conversion to mono/16kHz/32kbps and automatic chunking for files exceeding ~24 MiB or 10 minutes
  • The system is validated by unit tests in tests/test_transcribe.py that verify correct behavior during network failures, rate limiting, and missing API keys

Frequently Asked Questions

What happens if both Groq and OpenAI API keys are missing?

The transcribe function validates API key availability before attempting transcription. If neither provider has a configured key (checked via _provider_key), the function raises a configuration error indicating that at least one provider must be properly authenticated in agent_reach/config.py.

Does the fallback mechanism work for local audio files as well as URLs?

Yes. The transcription pipeline handles both inputs identically after the initial download step. Whether the source is a YouTube URL processed by download_audio or a local Path object, the audio undergoes the same compression, chunking, and fallback orchestration through _transcribe_with_fallback.

How does the system handle rate limiting from Groq?

When Groq returns an HTTP 429 status code (rate limit exceeded) or any other non-2xx response, the transcribe_chunk function raises a TranscribeError. This triggers the fallback loop in _transcribe_with_fallback, which immediately proceeds to attempt transcription with OpenAI without exposing the error to the user unless both providers fail.

Can I force the system to use only OpenAI and skip Groq?

Yes. By specifying provider="openai" in the function call or CLI command, you bypass the automatic fallback ordering. The _provider_order function returns a single-element list containing only the specified provider, ensuring Groq is never attempted.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →