How to Transcribe Podcasts from Xiaoyuzhou Using Groq Whisper with Agent Reach

Agent Reach automates the entire workflow—detecting Xiaoyuzhou URLs, downloading audio, compressing it to meet Groq Whisper's limits, and returning a complete transcript via the transcribe() function.

The Panniantong/Agent-Reach repository provides a modular Python framework for extracting and transcribing content from Chinese podcast platforms. By leveraging Groq's free whisper-large-v3 model, you can convert Xiaoyuzhou (小宇宙) episodes into searchable text without managing complex audio pipelines manually. This guide explains the internal architecture and exact implementation steps based on the source code.

Prerequisites and Environment Setup

Before processing any Xiaoyuzhou URLs, you must satisfy three hard dependencies enforced by the XiaoyuzhouChannel.check() method in agent_reach/channels/xiaoyuzhou.py.

Installing Required Dependencies

The transcription pipeline requires two external binaries and one helper script:

  • ffmpeg: Required by compress_audio() and chunk_audio() in agent_reach/transcribe.py to re-encode audio to mono, 16 kHz, 32 kbps.
  • yt-dlp: Used by download_audio() to fetch the original podcast audio as an M4A file.
  • transcribe.sh: A helper script installed at ~/.agent-reach/tools/xiaoyuzhou/transcribe.sh via the CLI command agent-reach install --env=auto.

If any binary is missing from your PATH, the code raises a MissingDependency exception with a clear error message.

Configuring the Groq API Key

The system checks for credentials in two locations, in order of priority:

  1. Environment variable: GROQ_API_KEY
  2. Agent Reach configuration: groq_api_key (managed via agent_reach/config.py)

You can set this permanently using the CLI configure command or export it per session:

export GROQ_API_KEY="gsk_XXXXXXXXXXXXXXXX"

The Transcription Pipeline Architecture

Understanding the data flow helps debug failures and extend the system. The pipeline consists of three coordinated layers: channel detection, audio preparation, and API orchestration.

Channel Detection and URL Routing

When you submit a URL, the routing logic in agent_reach/channels/xiaoyuzhou.py triggers:

  • can_handle(url): Returns True for any string containing xiaoyuzhoufm.com.
  • check(): Validates that ffmpeg is executable, the helper script exists, and the Groq API key is configured. Returns a status tuple (off, warn, error, or ok) with a diagnostic message.

If the channel reports ok, the CLI forwards the request to the transcription engine.

Audio Acquisition and Preparation

The heavy lifting occurs in agent_reach/transcribe.py:

  1. download_audio(): Invokes yt-dlp to retrieve the podcast audio.
  2. compress_audio(): Runs ffmpeg to convert the file to mono, 16 kHz, 32 kbps, ensuring the output stays under Groq's 25 MiB limit.
  3. chunk_audio(): If compression alone does not meet the size requirement, splits the audio into ≤10-minute segments.

Each helper function raises MissingDependency if the underlying binary is unavailable.

Groq Whisper API Integration

The transcribe() function selects providers based on the provider parameter:

  • provider="auto" (default): Attempts Groq first, then falls back to OpenAI if the request fails.
  • provider="groq": Forces Groq usage only; raises an error if the API call fails.

For Groq-specific execution, the code:

  • Retrieves the API key from agent_reach/config.py or the environment.
  • Sends a POST request to https://api.groq.com/openai/v1/audio/transcriptions.
  • Uses the model identifier whisper-large-v3.
  • Concatenates all chunk texts with newlines into the final transcript.

Step-by-Step Implementation

Command-Line Usage

The simplest method uses the CLI entry point in agent_reach/cli.py:


# 1. Install the helper script

agent-reach install --env=auto

# 2. Set your API key

export GROQ_API_KEY="gsk_XXXXXXXXXXXXXXXX"

# 3. Transcribe and save to file

python -m agent_reach.cli transcribe https://xiaoyuzhoufm.com/podcast/12345 > transcript.txt

The CLI automatically discovers the Xiaoyuzhou channel, validates prerequisites, and streams the final text to stdout.

Programmatic Python Usage

For integration into larger applications, import the transcription functions directly:

from agent_reach.transcribe import transcribe
from agent_reach.config import Config

# Configure credentials via the Config object

cfg = Config()
cfg.set("groq_api_key", "gsk_XXXXXXXXXXXXXXXX")

# Execute transcription with explicit provider selection

text = transcribe(
    "https://xiaoyuzhoufm.com/podcast/12345",
    provider="groq",      # Forces Groq only; remove to enable OpenAI fallback

    out_dir=None,         # Uses temporary directory; set path to preserve chunks

    config=cfg,
)

print(text)

Advanced Channel Interaction

For custom workflows requiring fine-grained control over channel state:

from agent_reach.channels.xiaoyuzhou import XiaoyuzhouChannel
from agent_reach.transcribe import transcribe

channel = XiaoyuzhouChannel()
url = "https://xiaoyuzhoufm.com/podcast/12345"

if channel.can_handle(url):
    status, msg = channel.check()
    if status == "ok":
        result = transcribe(url, provider="groq")
        print(result)
    else:
        print(f"Channel unavailable: {msg}")

Summary

  • Agent Reach provides a channel-based architecture where XiaoyuzhouChannel handles URL detection and prerequisite validation via can_handle() and check().
  • Audio processing relies on ffmpeg and yt-dlp, orchestrated through download_audio(), compress_audio(), and chunk_audio() in agent_reach/transcribe.py.
  • Groq Whisper is the default provider for transcribe(), using the whisper-large-v3 model with automatic fallback to OpenAI when provider="auto".
  • Size limits are enforced by compressing audio to mono, 16 kHz, 32 kbps and optionally chunking into ≤10-minute segments to stay under the 25 MiB API limit.
  • Configuration supports both environment variables (GROQ_API_KEY) and the internal Config class in agent_reach/config.py.

Frequently Asked Questions

What audio format does Groq Whisper require?

Groq Whisper accepts common audio formats (MP3, WAV, M4A), but Agent Reach specifically compresses Xiaoyuzhou podcasts to mono, 16 kHz, 32 kbps using ffmpeg. This ensures file sizes remain under the 25 MiB limit imposed by the Groq API while maintaining transcription accuracy.

Why does the transcription fail with a MissingDependency error?

The MissingDependency exception indicates that ffmpeg, yt-dlp, or the Xiaoyuzhou helper script is not found in your system PATH. Run agent-reach install --env=auto to install the script, and verify that ffmpeg and yt-dlp are installed and accessible from your shell before invoking the CLI or Python API.

Can I use Agent Reach without a Groq API key?

No. While Agent Reach supports fallback to OpenAI Whisper when provider="auto", the XiaoyuzhouChannel.check() method specifically requires a Groq API key to be set in either the GROQ_API_KEY environment variable or the groq_api_key configuration setting. Without this, the channel will report an error status and refuse to process URLs.

How does Agent Reach handle long podcasts that exceed the 25 MiB limit?

The chunk_audio() function in agent_reach/transcribe.py automatically splits compressed audio into segments of approximately 10 minutes each. The transcribe() function then processes each chunk sequentially through the Groq API and concatenates the results with newline separators, producing a single coherent transcript regardless of the original episode length.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →