How the Agent-Reach `transcribe` Command Utilizes Whisper Through Groq or OpenAI for Audio Transcription
The transcribe command routes audio files to Whisper-compatible APIs from Groq and OpenAI, automatically falling back between providers while chunking large files into 10-minute segments to handle the 24 MiB size limit.
The transcribe command in the Agent-Reach repository provides a robust audio-to-text pipeline that leverages Whisper models hosted by Groq and OpenAI. This functionality, implemented primarily in agent_reach/transcribe.py, abstracts provider-specific API differences behind a unified interface with intelligent fallback logic.
Provider Configuration and Routing Strategy
The system defines a static PROVIDERS dictionary at lines 30-40 of agent_reach/transcribe.py that catalogs each vendor’s endpoint, model name, and the configuration key storing its API token. This central registry enables the command to treat Groq and OpenAI as interchangeable backends despite their differing authentication schemes.
Dynamic provider selection occurs through the _provider_order helper function (lines 99-104). When the user specifies provider="auto", the function returns ["groq", "openai"], prioritizing Groq’s free tier before falling back to OpenAI. Explicit provider selection bypasses this ordering, creating a single-item list containing only the requested vendor.
API Key Validation
Before initiating any network operations or audio processing, the command validates that at least one selected provider has its API key configured in the Config object. If neither GROQ_API_KEY nor OPENAI_API_KEY is present, the function raises NoProviderConfigured (lines 22-26), failing fast to avoid wasting resources on downloads and compression.
Audio Ingestion and Chunking Pipeline
The command handles both local files and remote URLs through yt-dlp integration. Once fetched, audio undergoes compression via ffmpeg to ensure Whisper compatibility.
To respect the 24 MiB Whisper API limit, the implementation splits files exceeding this threshold into ≤10-minute chunks (lines 77-91 and 101-112). This chunking strategy ensures that even lengthy podcasts or interviews can be processed without hitting provider size restrictions.
Fallback Mechanism and API Communication
The core transcription logic resides in _transcribe_with_fallback (lines 49-61 and 63-70). This helper iterates over the provider list established by _provider_order, skipping any provider lacking a configured API key. For each chunk, it calls transcribe_chunk until receiving a successful response, ensuring that a Groq outage or rate limit automatically triggers an OpenAI retry.
Request Construction and Error Handling
The transcribe_chunk function (lines 71-89) constructs a multipart POST request containing:
- The audio file binary
- The specified Whisper model identifier
- An
Authorization: Bearer <API-key>header
Each request uses a default 120-second timeout per chunk. Network failures and non-200 HTTP status codes are captured and wrapped in TranscribeError, providing consistent error semantics regardless of which provider generated the failure.
Result Assembly
After successfully transcribing all chunks, the command strips whitespace from individual transcripts, filters out empty strings, and concatenates the results with newline separators (lines 42-46). This produces a clean, contiguous text output even when the source audio required segmentation.
CLI Integration
The command-line interface entry point in agent_reach/cli.py (lines 1113-1120) forwards user arguments directly to the library function, handling output formatting and file writing. This thin wrapper ensures that CLI behavior remains synchronized with the underlying Python API.
Usage Examples
Auto-Select Provider via CLI
Groq is attempted first; OpenAI serves as backup if Groq fails or lacks configuration:
agent-reach transcribe "https://www.youtube.com/watch?v=example"
Force Specific Provider
Bypass the fallback chain to use OpenAI exclusively:
agent-reach transcribe "https://example.com/podcast.mp3" --provider openai
Programmatic Usage with Error Handling
from agent_reach.transcribe import transcribe, TranscribeError
try:
# Uses Groq → OpenAI fallback order
text = transcribe("https://youtu.be/example")
print(text)
except TranscribeError as exc:
print(f"Transcription failed: {exc}")
Process Local File with Custom Output
from pathlib import Path
from agent_reach.transcribe import transcribe
out_dir = Path("/tmp/transcriptions")
result = transcribe("interview.m4a", out_dir=out_dir)
print(result)
Direct Provider Access
Bypass the fallback logic to call a specific provider directly:
from agent_reach.transcribe import transcribe_chunk, Config
cfg = Config()
chunk_path = Path("segment_001.m4a")
text = transcribe_chunk(chunk_path, "groq", config=cfg)
Summary
- Provider Agnostic: The
PROVIDERSdictionary and_provider_orderfunction abstract differences between Groq and OpenAI Whisper implementations. - Intelligent Fallback: The system automatically retries with secondary providers when primary endpoints fail, configured via the
provider="auto"parameter. - Size Management: Audio exceeding 24 MiB is automatically split into 10-minute chunks to comply with API constraints.
- Fast Failure: API key validation occurs before any network or disk-heavy operations, raising
NoProviderConfiguredimmediately if credentials are missing. - Unified Interface: Both CLI (
agent_reach/cli.py) and Python API (agent_reach/transcribe.py) share the same chunking, fallback, and error-handling logic.
Frequently Asked Questions
How does the transcribe command choose between Groq and OpenAI?
When using the default provider="auto" setting, the _provider_order function returns ["groq", "openai"], attempting Groq first due to its free tier availability, then falling back to OpenAI if Groq returns an error or lacks an API key.
What happens if both API keys are missing?
The command raises NoProviderConfigured during initialization (lines 22-26 of agent_reach/transcribe.py) before downloading or processing any audio, ensuring immediate feedback rather than mid-process failures.
How does the command handle large audio files?
Files exceeding the 24 MiB Whisper limit are compressed with ffmpeg and split into chunks of no more than 10 minutes each. Each chunk is transcribed sequentially and concatenated into the final transcript.
Can I force a specific provider instead of using auto?
Yes. Pass --provider groq or --provider openai via the CLI, or specify the provider parameter in Python: transcribe(url, provider="openai"). This creates a single-provider list in _provider_order, disabling the fallback mechanism.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →