How the Agent-Reach Transcribe Command Works with Groq and OpenAI Whisper
The Agent-Reach transcribe command leverages a provider-agnostic orchestrator that automatically routes audio transcription requests to Groq Whisper by default, with OpenAI Whisper serving as a seamless fallback when the primary provider fails or lacks configuration.
The transcribe functionality in Panniantong/Agent-Reach offers a unified interface for converting audio content—including YouTube videos and local files—into text using Whisper-compatible APIs. Implemented in agent_reach/transcribe.py, the system abstracts provider-specific details behind a robust configuration layer and intelligent fallback mechanism.
Provider Architecture and Configuration
The transcription system defines supported backends in a centralized provider map that specifies endpoints, model identifiers, and required API keys.
Groq vs OpenAI Whisper Provider Definitions
In agent_reach/transcribe.py, the PROVIDERS dictionary encapsulates the two Whisper-compatible backends:
PROVIDERS = {
"groq": {
"endpoint": "https://api.groq.com/openai/v1/audio/transcriptions",
"model": "whisper-large-v3",
"key_field": "groq_api_key",
},
"openai": {
"endpoint": "https://api.openai.com/v1/audio/transcriptions",
"model": "whisper-1",
"key_field": "openai_api_key",
},
}
Each entry stores the HTTP endpoint for audio transcriptions, the specific model identifier, and the configuration key name that holds the required authentication token.
Configuration Requirements and API Keys
The configuration system in agent_reach/config.py manages user secrets through a YAML file located at ~/.agent-reach/config.yaml. The transcribe command looks for groq_api_key and openai_api_key within the feature requirements mapping. Before attempting transcription, the Config.is_configured helper validates that at least one provider has its corresponding API key set, raising NoProviderConfigured early if neither credential is present.
Transcription Pipeline and Chunking Strategy
The transcribe() function orchestrates a complete pipeline from source acquisition to text output, handling large files through automatic segmentation.
Audio Download and Compression
When the source is a URL, download_audio invokes yt-dlp to fetch the audio stream as an M4A file. The compress_audio function then normalizes the audio to mono, 16 kHz, and 32 kbps, ensuring the output remains under the Whisper API’s 25 MiB limit while maintaining transcription quality.
Automatic Chunking for Large Files
If compression still yields a file exceeding the size threshold, chunk_audio splits the audio into segments of CHUNK_SECONDS = 600 (10 minutes) maximum duration. The transcribe_chunk function processes each segment individually through HTTP POST requests to the selected provider, concatenating results with newline characters to form the complete transcript.
Provider Selection and Fallback Logic
The transcribe command supports explicit provider selection or automatic routing with built-in resilience.
The Auto Provider and Priority Order
The public function signature transcribe(source, *, provider="auto", …) accepts a provider argument where "auto" triggers the internal _provider_order list: ["groq", "openai"]. This prioritizes Groq’s fast inference while maintaining OpenAI as a reliable backup.
Error Handling and TranscribeError
The _transcribe_with_fallback helper implements silent failover logic. If the first provider returns a 429 error or other failure, the system automatically retries with the next configured provider until one succeeds or all options are exhausted. Specific exception classes—including TranscribeError for API failures and NoProviderConfigured for missing credentials—enable precise error handling for both CLI and programmatic callers.
Usage Examples
The transcribe command is accessible via both the CLI entry point and direct Python API calls.
CLI Usage with Explicit Providers
Install the optional transcription dependencies:
pip install "agent-reach[transcribe]"
Configure your preferred provider:
agent-reach configure groq-key YOUR_GROQ_KEY
# or
agent-reach configure openai-key YOUR_OPENAI_KEY
Transcribe a YouTube video with explicit provider selection:
agent-reach transcribe https://youtu.be/abc123 --provider groq
Use automatic fallback routing:
agent-reach transcribe https://youtu.be/abc123
Python API Integration
Configure credentials and transcribe programmatically:
from agent_reach import transcribe as tr
from agent_reach.config import Config
cfg = Config()
cfg.set("groq_api_key", "YOUR_GROQ_KEY")
text = tr.transcribe(
"https://youtu.be/abc123",
provider="groq",
config=cfg,
)
print(text)
For automatic fallback when Groq is unavailable:
cfg.set("openai_api_key", "YOUR_OPENAI_KEY")
text = tr.transcribe("my_podcast.m4a", provider="auto", config=cfg)
Integration with YouTube Channel
The YouTube channel implementation in agent_reach/channels/youtube.py delegates directly to the transcribe orchestrator. Its check method advertises transcription capability only when at least one Whisper provider is configured and ffmpeg is present on the system, allowing the doctor command to surface readiness diagnostics.
Summary
- Provider abstraction: The
PROVIDERSdictionary inagent_reach/transcribe.pydefines Groq and OpenAI endpoints with their respective API key requirements (groq_api_keyandopenai_api_key). - Automatic fallback: The
provider="auto"setting routes requests to Groq first, then OpenAI if errors occur via_transcribe_with_fallback. - Robust preprocessing: Audio undergoes compression to mono 16 kHz/32 kbps and automatic chunking into 10-minute segments (
CHUNK_SECONDS = 600) to satisfy API limits. - Configuration management: API keys stored in
~/.agent-reach/config.yamlenable provider selection and early validation throughConfig.is_configured. - Clear error handling: Distinct exceptions (
NoProviderConfigured,TranscribeError) separate configuration issues from runtime API failures.
Frequently Asked Questions
What is the default provider order when using agent-reach transcribe with provider="auto"?
When provider is set to "auto", the system follows the _provider_order list defined in agent_reach/transcribe.py, which prioritizes Groq ("groq") first and OpenAI ("openai") second. This ensures fast, free-tier transcription attempts before falling back to alternative providers.
How does Agent-Reach handle audio files larger than 25 MiB?
The transcribe command automatically compresses audio to mono, 16 kHz, and 32 kbps using compress_audio. If the file remains above the 25 MiB limit, chunk_audio splits it into segments of maximum 600 seconds (10 minutes) each, processing them sequentially via transcribe_chunk and concatenating the results with newlines.
Where does Agent-Reach store API keys for Groq and OpenAI Whisper?
Credentials are stored in the YAML configuration file at ~/.agent-reach/config.yaml, managed by agent_reach/config.py. The system looks for groq_api_key and openai_api_key entries, which can be set via the CLI using agent-reach configure groq-key or agent-reach configure openai-key.
Can I use the transcribe functionality without installing YouTube-specific dependencies?
Yes, the transcribe command works with local audio files without requiring yt-dlp. However, downloading from YouTube URLs requires the optional [transcribe] extras package which includes yt-dlp and ffmpeg bindings for audio extraction.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →