How the Agent-Reach Transcribe Command Uses Whisper with Groq or OpenAI
The transcribe command routes audio to Groq or OpenAI Whisper APIs with automatic fallback, chunking files over 24 MiB and validating API keys before processing.
The Agent-Reach repository provides a robust CLI tool for converting audio to text using OpenAI-compatible Whisper endpoints. The implementation in agent_reach/transcribe.py orchestrates provider selection, audio preprocessing, and resilient API communication to deliver transcripts from either Groq or OpenAI.
Provider Catalog and Configuration
The transcription system relies on a static provider catalogue defined at the top of agent_reach/transcribe.py. The PROVIDERS dictionary (lines 30-40) maps each vendor to its endpoint URL, model identifier, and the configuration key that stores its API token.
Before any audio processing begins, the system validates that at least one provider has a configured API key in the Config object. If neither the Groq nor the OpenAI key is present, the function raises a NoProviderConfigured error immediately (lines 22-26), preventing unnecessary network traffic and file compression overhead.
Dynamic Provider Selection and Fallback Strategy
The _provider_order function (lines 99-104) translates the user-supplied provider argument into an ordered list of vendors to attempt. When the user specifies auto (the default), the function returns ["groq", "openai"], prioritizing Groq's free tier.
The core fallback logic resides in _transcribe_with_fallback (lines 49-61 and 63-70). This helper iterates over the provider list, skips any for which the API key is missing, and calls transcribe_chunk for each. The first successful response is returned immediately; if all providers fail, the last exception is re-raised as a TranscribeError.
Audio Processing and Chunking Pipeline
For remote sources, the command uses yt-dlp to fetch audio, then compresses it to a Whisper-friendly format using ffmpeg. Files exceeding the 24 MiB Whisper limit are automatically split into chunks of approximately 10 minutes each (lines 77-91 and 101-112).
Each chunk is processed sequentially. The transcribe_chunk function (lines 71-89) constructs a multipart POST request containing:
- The audio file blob
- The selected Whisper model name (e.g.,
whisper-large-v3) - An
Authorization: Bearer <API-key>header
Requests use a default timeout of 120 seconds per chunk. Network failures and non-200 HTTP status codes are wrapped in TranscribeError to provide uniform error handling.
After transcription, chunk results are stripped of whitespace, filtered for empty strings, and concatenated with newline separators to form the final transcript (lines 42-46).
CLI and Programmatic Usage
Command Line Interface
The CLI entry point in agent_reach/cli.py (lines 1113-1120) forwards arguments to the library function.
Transcribe a YouTube video with automatic provider selection (Groq first, then OpenAI):
agent-reach transcribe "https://www.youtube.com/watch?v=example"
Force a specific provider to fail fast if its key is unavailable:
agent-reach transcribe "https://example.com/podcast.mp3" --provider groq
Python API
Basic usage with automatic fallback:
from agent_reach.transcribe import transcribe, TranscribeError
try:
text = transcribe("https://youtu.be/example")
print(text)
except TranscribeError as exc:
print(f"Transcription failed: {exc}")
Process a local file with a custom output directory:
from pathlib import Path
from agent_reach.transcribe import transcribe
out_dir = Path("/tmp/my_transcribe")
result = transcribe("local_file.m4a", out_dir=out_dir)
print(result)
Bypass the fallback logic and call a specific provider directly:
from agent_reach.transcribe import transcribe_chunk, Config
cfg = Config() # Loads API keys from config file or env vars
chunk_path = Path("chunk_001.m4a")
groq_text = transcribe_chunk(chunk_path, "groq", config=cfg)
print(groq_text)
Summary
- Provider routing: The
PROVIDERSdictionary inagent_reach/transcribe.pydefines Groq and OpenAI endpoints, withautodefaulting to Groq-first ordering. - Pre-validation: API keys are checked before audio download to avoid wasted bandwidth; missing keys trigger
NoProviderConfigured. - Chunking strategy: Audio exceeding 24 MiB is split into ≤10-minute segments via ffmpeg to comply with Whisper API limits.
- Resilient execution:
_transcribe_with_fallbackiterates through providers, returning the first successful transcript or the final exception. - CLI integration: The
agent-reach transcribecommand exposes all functionality, handling file I/O and result printing.
Frequently Asked Questions
Which provider does the transcribe command use by default?
When the provider argument is set to auto (the default), the _provider_order function returns ["groq", "openai"], meaning Groq is attempted first. If Groq returns an error or its API key is not configured, the system automatically falls back to OpenAI.
How does the transcribe command handle large audio files?
Files larger than 24 MiB are automatically split into chunks of approximately 10 minutes using ffmpeg. The _transcribe_with_fallback function processes each chunk sequentially through the selected provider(s), then concatenates the results with newline separators to produce the final transcript.
What happens if both Groq and OpenAI API keys are missing?
The function validates available API keys immediately after determining the provider order. If no valid key is found for any candidate provider, it raises a NoProviderConfigured error before downloading or processing any audio data.
Can I use the transcribe function programmatically without the CLI?
Yes. Import transcribe from agent_reach.transcribe to access the full orchestration layer, or import transcribe_chunk to target a specific provider directly. Both functions accept a Config object containing your API keys and support local file paths as well as remote URLs.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →