How the Agent-Reach Transcribe Command Uses Groq and OpenAI Whisper Providers
The transcribe command routes audio transcription through Groq and OpenAI Whisper APIs using a provider-agnostic orchestration layer that automatically falls back from Groq to OpenAI when errors occur.
The transcribe command in the Agent-Reach repository converts audio files or URLs into text by leveraging Whisper-compatible APIs from Groq and OpenAI. This implementation follows a robust provider-routing strategy that prioritizes cost-effective transcription while ensuring reliability through automatic fallback mechanisms.
Provider Architecture and Routing Logic
The transcription system maintains a static registry of supported providers and implements dynamic selection logic to determine which API to call.
The Provider Catalog
In agent_reach/transcribe.py, the PROVIDERS dictionary (lines 30-40) defines the endpoint URLs, model names, and configuration keys for each vendor. This catalog enables the system to treat Groq and OpenAI as interchangeable transcription backends while maintaining their specific API requirements.
Dynamic Provider Selection
The _provider_order helper function translates user input into an ordered list of providers. When the CLI receives --provider auto (the default), the function returns ["groq", "openai"], prioritizing Groq's typically faster and cheaper Whisper implementation (lines 99-104). Users can override this by specifying --provider groq or --provider openai to force a single vendor.
Transcription Workflow and Fallback Strategy
The command processes audio through a multi-stage pipeline that handles large files via chunking and implements resilient error handling across providers.
Input Processing and Chunking
Before sending data to any API, the system downloads remote audio using yt-dlp, compresses it to Whisper-friendly formats with ffmpeg, and splits files exceeding the 24 MiB Whisper limit into segments of 10 minutes or less (lines 77-91). This preprocessing ensures compatibility with both Groq and OpenAI's file size constraints.
The Fallback Mechanism
The _transcribe_with_fallback function implements the core resilience logic (lines 49-70). It iterates through the provider list established by _provider_order, skipping any provider lacking a configured API key. The function attempts transcription with each available provider in sequence, returning the first successful response or re-raising the final exception if all providers fail.
Individual Chunk Transcription
For each audio chunk, transcribe_chunk constructs a multipart POST request containing the audio file, the selected Whisper model, and an Authorization: Bearer <API-key> header (lines 71-89). The function enforces a default 120-second timeout per chunk and wraps HTTP errors or non-200 status codes in a TranscribeError for consistent error handling.
Configuration and API Key Validation
The system validates API key availability before processing audio. In agent_reach/transcribe.py (lines 22-26), the function checks Config for at least one valid provider key, raising NoProviderConfigured if neither Groq nor OpenAI credentials are present. This early validation prevents wasted computation on audio preprocessing when transcription is impossible.
CLI Usage and Programmatic Integration
The CLI entry point in agent_reach/cli.py (lines 1113-1120) forwards arguments to the library function, handling output formatting and file writing.
Auto-select provider (Groq preferred):
agent-reach transcribe "https://www.youtube.com/watch?v=example"
Force specific provider:
agent-reach transcribe "https://example.com/podcast.mp3" --provider groq
Programmatic usage with error handling:
from agent_reach.transcribe import transcribe, TranscribeError
try:
text = transcribe("https://youtu.be/example")
print(text)
except TranscribeError as exc:
print(f"Transcription failed: {exc}")
Direct provider bypass (advanced):
from agent_reach.transcribe import transcribe_chunk, Config
from pathlib import Path
cfg = Config()
chunk_path = Path("chunk_001.m4a")
groq_text = transcribe_chunk(chunk_path, "groq", config=cfg)
print(groq_text)
Summary
- The
transcribecommand uses a provider-agnostic architecture defined inagent_reach/transcribe.pyto abstract Groq and OpenAI Whisper implementations. - Provider selection follows a priority order (Groq first, then OpenAI) when using
--provider auto, with fallback logic in_transcribe_with_fallback. - Audio preprocessing includes yt-dlp downloading, ffmpeg compression, and intelligent chunking to respect the 24 MiB API limits.
- API keys are validated early against the
Configclass to prevent unnecessary processing, raisingNoProviderConfiguredwhen credentials are missing. - Each chunk is transmitted via multipart POST with Bearer authentication and 120-second timeouts, with errors normalized to
TranscribeError.
Frequently Asked Questions
What happens if both Groq and OpenAI API keys are configured?
When both keys are present and --provider auto is used (the default), the system attempts transcription with Groq first. Only if Groq returns an error or timeout does it fall back to OpenAI. If Groq succeeds, OpenAI is never called for that specific chunk.
Can I use the transcribe function without the CLI?
Yes. Import transcribe from agent_reach.transcribe to use the function programmatically. The function accepts file paths, URLs, and optional parameters for output directory and provider selection, returning the complete transcript as a string.
Why does the command split audio into chunks?
Both Groq and OpenAI impose a 24 MiB file size limit on Whisper API requests. The transcribe command automatically splits larger files into 10-minute segments using ffmpeg, processes each chunk independently through the provider fallback chain, and concatenates the results with newline separators.
How does the system handle network timeouts?
The transcribe_chunk function implements a default 120-second timeout per chunk. If a provider fails to respond within this window or returns a non-200 HTTP status, the error is caught and wrapped in TranscribeError, triggering the fallback to the next configured provider if available.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →