How Agent Reach's Transcribe Command Uses Groq Whisper for Podcast Transcription

Agent Reach leverages Groq's Whisper model (whisper-large-v3) as the primary transcription engine for its transcribe command, automatically processing podcast audio through the Groq API with OpenAI Whisper as a fallback.

Agent Reach provides a high-level transcription pipeline that converts podcast audio into plain text using Groq's Whisper API. The implementation in agent_reach/transcribe.py orchestrates provider selection, audio preprocessing, and API communication to deliver fast, cost-effective transcription with minimal configuration.

Provider Configuration and Routing

The transcription system is provider-agnostic, supporting both Groq and OpenAI Whisper APIs through a unified interface.

Provider Definitions and Model Mapping

In agent_reach/transcribe.py, a PROVIDERS dictionary (lines 32-43) maps logical provider names to their respective API endpoints and model identifiers:

  • Groq: Uses endpoint https://api.groq.com/openai/v1/audio/transcriptions with model whisper-large-v3
  • OpenAI: Uses the standard OpenAI audio transcriptions endpoint with model whisper-1

Automatic Provider Selection

Users specify the --provider flag with values groq, openai, or auto. When auto is selected, the _provider_order function (lines 49-54) returns ["groq", "openai"], ensuring Groq is attempted first for faster, cost-free transcription. This prioritization makes Groq Whisper the default engine for podcast transcription in Agent Reach.

API Key Management

The _provider_key function (lines 66-70) retrieves credentials from the Agent Reach configuration. It expects either groq_api_key or openai_api_key in your settings. If neither is configured, the system raises a NoProviderConfigured exception before attempting any API calls.

Audio Preprocessing Pipeline

Before sending audio to Groq Whisper, Agent Reach prepares the source material to meet API requirements.

Downloading and Compression

The download_audio function fetches the source URL using yt-dlp, supporting both direct audio files and podcatching scenarios. The compress_audio function then converts the audio to a mono, 16 kHz, 32 kbps M4A format. This compression minimizes upload size while preserving speech clarity for transcription accuracy.

Chunking Strategy

Groq's API imposes file size limits, so chunk_audio splits large audio files into segments of ≤25MB. If the compressed audio fits within this limit, it remains a single chunk. Each chunk is processed independently and concatenated in the final output.

Groq Whisper API Integration

When Groq is selected as the provider, Agent Reach communicates directly with Groq's Whisper implementation.

Endpoint and Request Format

The transcribe_chunk function POSTs audio data to https://api.groq.com/openai/v1/audio/transcriptions with the following parameters:

  • model: whisper-large-v3
  • response_format: text
  • Authorization: Bearer token from groq_api_key

The request includes the audio file as multipart/form-data, following the OpenAI-compatible API specification that Groq provides.

Chunk Processing

For each audio chunk, the system streams the binary data to Groq's endpoint. The API returns plain text which is collected and concatenated with other chunks to form the complete transcript.

Fallback Mechanism and Error Handling

Agent Reach implements resilient error handling through the _transcribe_with_fallback function (lines 106-118).

Provider Failover

If Groq returns a non-2xx status code or encounters a network error, the system immediately attempts transcription with the next provider in the order (OpenAI). This fallback happens seamlessly without user intervention.

Error States

If all configured providers fail, the system raises a TranscribeError with details from the last attempted provider. This ensures users receive clear feedback when transcription is impossible due to API outages or authentication issues.

CLI Integration and Usage

The transcription functionality is exposed through the agent-reach transcribe subcommand defined in agent_reach/cli.py (lines 1135-1144).

Command-Line Interface

The CLI parses arguments and calls the transcribe function with user-specified options:


# Transcribe with auto provider selection (Groq first)

agent-reach transcribe "https://example.com/podcast.mp3"

# Specify output file

agent-reach transcribe "https://example.com/podcast.mp3" -o transcript.txt

# Force specific provider

agent-reach transcribe "https://example.com/podcast.mp3" --provider groq

Library Usage

You can also use the transcription engine programmatically:

from agent_reach.transcribe import transcribe, TranscribeError

try:
    transcript = transcribe(
        "https://xiaoyuzhoufm.com/podcast/12345",
        provider="auto",  # Uses Groq first, OpenAI fallback

    )
    print(transcript)
except TranscribeError as exc:
    print(f"Transcription failed: {exc}")

Channel Integration

Podcast-specific channels like XiaoyuzhouChannel delegate to the core transcription function. In agent_reach/channels/xiaoyuzhou.py (lines 10-15), the transcribe method simply forwards the request to the generic transcribe function, allowing Xiaoyuzhou podcast URLs to benefit from Groq Whisper processing automatically.


# Example from channel implementation

from agent_reach.transcribe import transcribe as _transcribe

def transcribe(self, url: str, *, provider: str = "auto", config=None) -> str:
    return _transcribe(url, provider=provider, config=config)

Summary

  • Agent Reach uses Groq Whisper (whisper-large-v3) as the primary transcription provider, configured in the PROVIDERS dictionary within agent_reach/transcribe.py.
  • The audio pipeline downloads via yt-dlp, compresses to mono 16kHz M4A, and chunks files to ≤25MB to meet API constraints.
  • Automatic fallback to OpenAI Whisper occurs if Groq fails, managed by _transcribe_with_fallback in the transcription module.
  • API keys are read from Agent Reach configuration fields groq_api_key and openai_api_key, raising NoProviderConfigured if unavailable.
  • The CLI command agent-reach transcribe and podcast channels like XiaoyuzhouChannel provide direct access to Groq-powered transcription.

Frequently Asked Questions

What Whisper model does Agent Reach use for transcription?

Agent Reach uses Groq's whisper-large-v3 model when the Groq provider is selected. If Groq fails and OpenAI is configured as a fallback, it uses OpenAI's whisper-1 model. The model selection is hardcoded in the PROVIDERS dictionary in agent_reach/transcribe.py.

How does Agent Reach handle large podcast files?

The system automatically compresses audio to mono 16kHz, 32kbps M4A format and splits files into chunks of 25MB or smaller using the chunk_audio function. Each chunk is transcribed independently and concatenated, allowing arbitrarily long podcasts to be processed within API size limits.

What happens if the Groq API is unavailable?

Agent Reach implements automatic failover through _transcribe_with_fallback. If Groq returns a non-2xx status or network error, the system immediately retries the transcription with OpenAI Whisper. If both providers fail, a TranscribeError is raised with diagnostic details.

How do I configure API keys for the transcribe command?

Set groq_api_key or openai_api_key in your Agent Reach configuration file. The _provider_key function checks these fields when initializing the transcription client. If using provider auto, only the Groq key is required unless a fallback occurs, though configuring both ensures maximum reliability.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →