How the Agent-Reach Transcribe Command Uses Groq vs OpenAI Whisper APIs
The transcribe command routes audio transcription requests to Groq first, then falls back to OpenAI Whisper, using a provider-agnostic orchestration layer that handles chunking, authentication, and error handling automatically.
The Agent-Reach repository provides a robust CLI and Python library for AI-powered automation tasks. Its transcribe command demonstrates a clean implementation of provider abstraction, allowing users to convert audio from local files or URLs into text using either Groq's or OpenAI's Whisper-compatible APIs with intelligent fallback handling.
Provider Catalog and Configuration
The transcription system defines its supported backends in a static PROVIDERS dictionary located in agent_reach/transcribe.py (lines 30-40). This catalog maps each vendor to its endpoint URL, model name, and the configuration key that stores its API token.
Before processing any audio, the system validates that at least one selected provider has its API key configured in the Config class (sourced from agent_reach/config.py). If no keys are present for the chosen providers, the function raises NoProviderConfigured immediately (lines 22-26), failing fast before any network requests or audio processing begins.
Dynamic Provider Selection and Fallback Order
Provider selection logic resides in the _provider_order helper function (lines 99-104). This function converts the user-supplied provider argument into an ordered list of vendors to attempt.
auto(default): Returns["groq", "openai"], prioritizing Groq's free tiergroqoropenai: Returns a single-element list, forcing that specific provider and skipping fallback
This design allows users to either leverage the automatic failover strategy or enforce a specific backend for cost or compliance reasons.
Audio Processing and Chunking Strategy
The command handles audio ingestion through a robust preprocessing pipeline (lines 77-91). For remote URLs, it utilizes yt-dlp to fetch content, then processes all audio through ffmpeg to compress it into a Whisper-friendly format.
To comply with the 24 MiB Whisper API limit, the system splits large files into chunks of ≤ 10 minutes each (lines 101-112). This chunking ensures compatibility with both Groq and OpenAI's file size restrictions while maintaining transcription quality across long-form content.
The Fallback Mechanism and API Execution
The core resilience logic lives in _transcribe_with_fallback (lines 49-61). This helper iterates over the provider list from _provider_order, skipping any provider whose API key is missing from the configuration. It calls transcribe_chunk for each provider until it receives a successful response, or re-raises the last encountered exception if all providers fail.
Individual API requests are constructed in transcribe_chunk (lines 63-70, 71-89). This function:
- Builds a multipart POST request containing the audio file and selected Whisper model
- Injects the
Authorization: Bearer <API-key>header using the configured token - Applies a per-chunk timeout of 120 seconds by default
- Wraps network errors and non-200 HTTP status codes in
TranscribeError
Result Assembly and Output Handling
Once all chunks are processed successfully, the system strips whitespace, filters empty segments, and concatenates the transcripts with newline separators (lines 42-46) to produce the final output.
The CLI entry point in agent_reach/cli.py (lines 1113-1120) acts as a thin wrapper, forwarding arguments to the library function and handling stdout printing or file writing based on user preferences.
Usage Examples
CLI Auto-Selection with Fallback
Use the default provider order (Groq first, OpenAI as backup) to transcribe a YouTube video:
agent-reach transcribe "https://www.youtube.com/watch?v=example"
Force a Specific Provider
Bypass the fallback mechanism and use only Groq:
agent-reach transcribe "https://example.com/podcast.mp3" --provider groq
Programmatic Auto-Selection
Call the transcription function from Python with automatic provider routing:
from agent_reach.transcribe import transcribe, TranscribeError
try:
text = transcribe("https://youtu.be/example")
print(text)
except TranscribeError as exc:
print(f"Transcription failed: {exc}")
Custom Output Directory
Specify where temporary and final files should be stored:
from pathlib import Path
from agent_reach.transcribe import transcribe
out_dir = Path("/tmp/my_transcribe")
result = transcribe("local_file.m4a", out_dir=out_dir)
print(result)
Direct Provider API Call
Bypass the fallback logic and call a specific provider directly:
from agent_reach.transcribe import transcribe_chunk, Config
cfg = Config() # Loads API keys from config file or environment
chunk_path = Path("chunk_001.m4a")
groq_text = transcribe_chunk(chunk_path, "groq", config=cfg)
print(groq_text)
Summary
- Dual-provider support: The
transcribecommand supports both Groq and OpenAI Whisper APIs through a unified interface defined inagent_reach/transcribe.py. - Intelligent fallback: The default
automode prioritizes Groq (["groq", "openai"]) and automatically fails over to OpenAI if Groq returns an error or is unconfigured. - Robust preprocessing: Audio is fetched via yt-dlp, compressed with ffmpeg, and automatically chunked to stay within the 24 MiB API limit.
- Fail-fast validation: The system checks for API keys before downloading or processing audio, raising
NoProviderConfiguredif no backends are available. - Flexible usage: Available both as a CLI command (
agent-reach transcribe) and as a Python library function with granular control over output directories and provider selection.
Frequently Asked Questions
What is the default provider order when using auto mode?
When the provider parameter is set to auto (the default), the system attempts Groq first, then falls back to OpenAI if Groq fails or is not configured. This ordering is hardcoded in the _provider_order function within agent_reach/transcribe.py (lines 99-104).
How does the transcribe command handle large audio files?
The command automatically splits audio exceeding the 24 MiB Whisper limit into chunks of 10 minutes or less using ffmpeg. Each chunk is transcribed separately through the provider fallback chain, and the results are concatenated with newline separators to form the complete transcript (lines 42-46 and 101-112).
What happens if neither Groq nor OpenAI API keys are configured?
Before initiating any downloads or processing, the function validates that at least one selected provider has a valid API key in the configuration. If no keys are found, it raises a NoProviderConfigured error immediately (lines 22-26), preventing wasted compute on audio fetching.
Can I use the transcription functionality programmatically without the CLI?
Yes. The core logic is exposed through the transcribe function in agent_reach/transcribe.py, which can be imported directly into Python applications. For advanced use cases, you can also call transcribe_chunk to process individual audio segments or bypass the fallback mechanism entirely.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →