How the Agent-Reach Transcribe Command Uses Groq and OpenAI Whisper Providers

The transcribe command routes audio transcription through Groq and OpenAI Whisper APIs using a provider-agnostic orchestration layer that automatically falls back from Groq to OpenAI when errors occur.

The transcribe command in the Agent-Reach repository converts audio files or URLs into text by leveraging Whisper-compatible APIs from Groq and OpenAI. This implementation follows a robust provider-routing strategy that prioritizes cost-effective transcription while ensuring reliability through automatic fallback mechanisms.

Provider Architecture and Routing Logic

The transcription system maintains a static registry of supported providers and implements dynamic selection logic to determine which API to call.

The Provider Catalog

In agent_reach/transcribe.py, the PROVIDERS dictionary (lines 30-40) defines the endpoint URLs, model names, and configuration keys for each vendor. This catalog enables the system to treat Groq and OpenAI as interchangeable transcription backends while maintaining their specific API requirements.

Dynamic Provider Selection

The _provider_order helper function translates user input into an ordered list of providers. When the CLI receives --provider auto (the default), the function returns ["groq", "openai"], prioritizing Groq's typically faster and cheaper Whisper implementation (lines 99-104). Users can override this by specifying --provider groq or --provider openai to force a single vendor.

Transcription Workflow and Fallback Strategy

The command processes audio through a multi-stage pipeline that handles large files via chunking and implements resilient error handling across providers.

Input Processing and Chunking

Before sending data to any API, the system downloads remote audio using yt-dlp, compresses it to Whisper-friendly formats with ffmpeg, and splits files exceeding the 24 MiB Whisper limit into segments of 10 minutes or less (lines 77-91). This preprocessing ensures compatibility with both Groq and OpenAI's file size constraints.

The Fallback Mechanism

The _transcribe_with_fallback function implements the core resilience logic (lines 49-70). It iterates through the provider list established by _provider_order, skipping any provider lacking a configured API key. The function attempts transcription with each available provider in sequence, returning the first successful response or re-raising the final exception if all providers fail.

Individual Chunk Transcription

For each audio chunk, transcribe_chunk constructs a multipart POST request containing the audio file, the selected Whisper model, and an Authorization: Bearer <API-key> header (lines 71-89). The function enforces a default 120-second timeout per chunk and wraps HTTP errors or non-200 status codes in a TranscribeError for consistent error handling.

Configuration and API Key Validation

The system validates API key availability before processing audio. In agent_reach/transcribe.py (lines 22-26), the function checks Config for at least one valid provider key, raising NoProviderConfigured if neither Groq nor OpenAI credentials are present. This early validation prevents wasted computation on audio preprocessing when transcription is impossible.

CLI Usage and Programmatic Integration

The CLI entry point in agent_reach/cli.py (lines 1113-1120) forwards arguments to the library function, handling output formatting and file writing.

Auto-select provider (Groq preferred):

agent-reach transcribe "https://www.youtube.com/watch?v=example"

Force specific provider:

agent-reach transcribe "https://example.com/podcast.mp3" --provider groq

Programmatic usage with error handling:

from agent_reach.transcribe import transcribe, TranscribeError

try:
    text = transcribe("https://youtu.be/example")
    print(text)
except TranscribeError as exc:
    print(f"Transcription failed: {exc}")

Direct provider bypass (advanced):

from agent_reach.transcribe import transcribe_chunk, Config
from pathlib import Path

cfg = Config()
chunk_path = Path("chunk_001.m4a")
groq_text = transcribe_chunk(chunk_path, "groq", config=cfg)
print(groq_text)

Summary

  • The transcribe command uses a provider-agnostic architecture defined in agent_reach/transcribe.py to abstract Groq and OpenAI Whisper implementations.
  • Provider selection follows a priority order (Groq first, then OpenAI) when using --provider auto, with fallback logic in _transcribe_with_fallback.
  • Audio preprocessing includes yt-dlp downloading, ffmpeg compression, and intelligent chunking to respect the 24 MiB API limits.
  • API keys are validated early against the Config class to prevent unnecessary processing, raising NoProviderConfigured when credentials are missing.
  • Each chunk is transmitted via multipart POST with Bearer authentication and 120-second timeouts, with errors normalized to TranscribeError.

Frequently Asked Questions

What happens if both Groq and OpenAI API keys are configured?

When both keys are present and --provider auto is used (the default), the system attempts transcription with Groq first. Only if Groq returns an error or timeout does it fall back to OpenAI. If Groq succeeds, OpenAI is never called for that specific chunk.

Can I use the transcribe function without the CLI?

Yes. Import transcribe from agent_reach.transcribe to use the function programmatically. The function accepts file paths, URLs, and optional parameters for output directory and provider selection, returning the complete transcript as a string.

Why does the command split audio into chunks?

Both Groq and OpenAI impose a 24 MiB file size limit on Whisper API requests. The transcribe command automatically splits larger files into 10-minute segments using ffmpeg, processes each chunk independently through the provider fallback chain, and concatenates the results with newline separators.

How does the system handle network timeouts?

The transcribe_chunk function implements a default 120-second timeout per chunk. If a provider fails to respond within this window or returns a non-200 HTTP status, the error is caught and wrapped in TranscribeError, triggering the fallback to the next configured provider if available.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →