How to Perform Audio Transcription Using Agent Reach: A Complete Guide

Agent Reach provides a built-in transcription pipeline that automatically downloads audio from URLs, compresses it to Whisper-compatible format, and transcribes it via Groq or OpenAI Whisper APIs with intelligent fallback handling.

Agent Reach is an open-source automation framework that includes a robust audio transcription module for converting speech to text. The transcription engine, implemented in agent_reach/transcribe.py, wraps Whisper-compatible APIs to handle everything from YouTube downloads to large file chunking. Whether you need to transcribe podcasts, lectures, or meeting recordings, this guide covers every method to perform audio transcription using Agent Reach.

How the Agent Reach Transcription Pipeline Works

The transcription logic in agent_reach/transcribe.py orchestrates a four-stage pipeline that handles edge cases like large files and provider failures automatically.

Stage 1: Source Acquisition with yt-dlp

When provided with a URL, the download_audio function (lines 77-95) uses yt-dlp to extract audio streams into M4A format. For local files, this step is skipped.

Stage 2: Compression and Format Standardization

The compress_audio function (lines 101-112) invokes ffmpeg to re-encode audio to mono 16 kHz at 32 kbps, ensuring compatibility with Whisper API requirements.

Stage 3: Intelligent Chunking for Large Files

Files exceeding Whisper's 25 MiB limit are processed by chunk_audio (lines 126-150), which splits audio into 10-minute segments without losing continuity.

Stage 4: Provider API Integration with Fallback

The transcribe_chunk function (lines 163-190) sends each segment to the selected provider's /v1/audio/transcriptions endpoint. The _transcribe_with_fallback method (lines 249-261) automatically retries with OpenAI if Groq fails.

Configuration and Prerequisites

Before running transcription tasks, you must install system dependencies and configure API credentials.

First, ensure ffmpeg and yt-dlp are available:

agent-reach install --channels=youtube

Then configure your provider keys in ~/.agent-reach/config.yaml or via environment variables:

agent-reach configure groq-key gsk_XXXXXXXXXXXXXXXX
agent-reach configure openai-key sk-XXXXXXXXXXXXXXXX

The Config class in agent_reach/config.py (lines 21-27) manages these credentials, supporting both Groq and OpenAI API keys.

Methods to Transcribe Audio Using Agent Reach

Agent Reach offers two primary interfaces for transcription: the CLI for quick tasks and the Python SDK for embedded workflows.

CLI Method (Fastest Implementation)

The transcribe sub-command in agent_reach/cli.py (lines 1113-1130) provides the simplest entry point:


# Transcribe a YouTube video automatically

agent-reach transcribe "https://www.youtube.com/watch?v=VIDEO_ID"

# Transcribe local file with specific provider

agent-reach transcribe ./interview.wav --provider openai -o ./output.txt

Python Library Integration

For programmatic use, import the transcribe function from agent_reach/transcribe.py:

from agent_reach.transcribe import transcribe, TranscribeError
from agent_reach.config import Config

cfg = Config()

try:
    transcript = transcribe(
        "https://example.com/podcast.mp3",
        provider="auto",  # Uses Groq first, falls back to OpenAI

        config=cfg
    )
    print(transcript)
except TranscribeError as exc:
    print(f"Transcription failed: {exc}")

Provider-Specific Transcription

To bypass automatic fallback and force a specific provider:

from agent_reach.transcribe import transcribe

# Force Groq whisper-large-v3

text = transcribe("./meeting.m4a", provider="groq")

# Force OpenAI whisper-1

text = transcribe("./lecture.mp3", provider="openai")

Core Architecture and Source Code References

Understanding the internal implementation helps with debugging and customization:

  • agent_reach/transcribe.py: Contains the complete pipeline implementation including download_audio, compress_audio, chunk_audio, and the public transcribe function (lines 207-246) that orchestrates the workflow.
  • agent_reach/config.py: Defines the configuration schema for API keys and provider settings.
  • agent_reach/cli.py: Implements argument parsing and output handling for the CLI interface.

The default provider priority places Groq (groq) first for its speed and cost efficiency, with OpenAI (openai) serving as the fallback when the provider="auto" parameter is used.

Summary

  • Agent Reach bundles a complete transcription pipeline in agent_reach/transcribe.py that handles downloading, compression, chunking, and API communication.
  • The workflow automatically splits files larger than 25 MiB into 10-minute chunks to comply with Whisper API limits.
  • Groq is the default provider (using whisper-large-v3), with automatic fallback to OpenAI (whisper-1) if the primary provider fails.
  • Configuration is stored in ~/.agent-reach/config.yaml and managed through the Config class.
  • Both CLI (agent-reach transcribe) and Python library interfaces support local files and URLs, with yt-dlp handling video platform downloads.

Frequently Asked Questions

What audio formats does Agent Reach support for transcription?

Agent Reach accepts any format that ffmpeg can decode, including MP3, WAV, M4A, and OGG. When processing URLs, yt-dlp extracts the best available audio stream automatically. The internal compress_audio function standardizes all inputs to mono 16 kHz 32 kbps M4A before sending to the Whisper API.

How does Agent Reach handle audio files larger than 25 MB?

The chunk_audio function in agent_reach/transcribe.py automatically segments audio into 10-minute chunks when files exceed the Whisper API's 25 MiB limit. Each chunk is transcribed separately and concatenated into a single coherent transcript without requiring manual intervention.

Can I use Agent Reach transcription without Groq or OpenAI API keys?

No, Agent Reach requires valid API credentials for at least one supported provider. The system checks for keys in ~/.agent-reach/config.yaml or corresponding environment variables (GROQ_API_KEY, OPENAI_API_KEY). Attempting to transcribe without configured credentials raises a configuration error before any audio processing begins.

Is it possible to transcribe YouTube videos directly using Agent Reach?

Yes, Agent Reach integrates yt-dlp to extract audio from YouTube URLs and other video platforms automatically. Simply pass the video URL to the transcribe function or CLI command, and the pipeline handles download, extraction, and transcription in one operation.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →