# How to Perform Audio Transcription Using Agent Reach: A Complete Guide

> Learn how to perform audio transcription with Agent Reach. This guide covers automatic audio download, format compression, and transcription using Groq or OpenAI Whisper APIs.

- Repository: [Pnant/Agent-Reach](https://github.com/Panniantong/Agent-Reach)
- Tags: how-to-guide
- Published: 2026-06-25

---

**Agent Reach provides a built-in transcription pipeline that automatically downloads audio from URLs, compresses it to Whisper-compatible format, and transcribes it via Groq or OpenAI Whisper APIs with intelligent fallback handling.**

Agent Reach is an open-source automation framework that includes a robust audio transcription module for converting speech to text. The transcription engine, implemented in [`agent_reach/transcribe.py`](https://github.com/Panniantong/Agent-Reach/blob/main/agent_reach/transcribe.py), wraps Whisper-compatible APIs to handle everything from YouTube downloads to large file chunking. Whether you need to transcribe podcasts, lectures, or meeting recordings, this guide covers every method to perform audio transcription using Agent Reach.

## How the Agent Reach Transcription Pipeline Works

The transcription logic in [`agent_reach/transcribe.py`](https://github.com/Panniantong/Agent-Reach/blob/main/agent_reach/transcribe.py) orchestrates a four-stage pipeline that handles edge cases like large files and provider failures automatically.

### Stage 1: Source Acquisition with yt-dlp

When provided with a URL, the `download_audio` function (lines 77-95) uses **yt-dlp** to extract audio streams into M4A format. For local files, this step is skipped.

### Stage 2: Compression and Format Standardization

The `compress_audio` function (lines 101-112) invokes **ffmpeg** to re-encode audio to mono 16 kHz at 32 kbps, ensuring compatibility with Whisper API requirements.

### Stage 3: Intelligent Chunking for Large Files

Files exceeding Whisper's 25 MiB limit are processed by `chunk_audio` (lines 126-150), which splits audio into 10-minute segments without losing continuity.

### Stage 4: Provider API Integration with Fallback

The `transcribe_chunk` function (lines 163-190) sends each segment to the selected provider's `/v1/audio/transcriptions` endpoint. The `_transcribe_with_fallback` method (lines 249-261) automatically retries with OpenAI if Groq fails.

## Configuration and Prerequisites

Before running transcription tasks, you must install system dependencies and configure API credentials.

First, ensure **ffmpeg** and **yt-dlp** are available:

```bash
agent-reach install --channels=youtube

```

Then configure your provider keys in `~/.agent-reach/config.yaml` or via environment variables:

```bash
agent-reach configure groq-key gsk_XXXXXXXXXXXXXXXX
agent-reach configure openai-key sk-XXXXXXXXXXXXXXXX

```

The `Config` class in [`agent_reach/config.py`](https://github.com/Panniantong/Agent-Reach/blob/main/agent_reach/config.py) (lines 21-27) manages these credentials, supporting both Groq and OpenAI API keys.

## Methods to Transcribe Audio Using Agent Reach

Agent Reach offers two primary interfaces for transcription: the CLI for quick tasks and the Python SDK for embedded workflows.

### CLI Method (Fastest Implementation)

The `transcribe` sub-command in [`agent_reach/cli.py`](https://github.com/Panniantong/Agent-Reach/blob/main/agent_reach/cli.py) (lines 1113-1130) provides the simplest entry point:

```bash

# Transcribe a YouTube video automatically

agent-reach transcribe "https://www.youtube.com/watch?v=VIDEO_ID"

# Transcribe local file with specific provider

agent-reach transcribe ./interview.wav --provider openai -o ./output.txt

```

### Python Library Integration

For programmatic use, import the `transcribe` function from [`agent_reach/transcribe.py`](https://github.com/Panniantong/Agent-Reach/blob/main/agent_reach/transcribe.py):

```python
from agent_reach.transcribe import transcribe, TranscribeError
from agent_reach.config import Config

cfg = Config()

try:
    transcript = transcribe(
        "https://example.com/podcast.mp3",
        provider="auto",  # Uses Groq first, falls back to OpenAI

        config=cfg
    )
    print(transcript)
except TranscribeError as exc:
    print(f"Transcription failed: {exc}")

```

### Provider-Specific Transcription

To bypass automatic fallback and force a specific provider:

```python
from agent_reach.transcribe import transcribe

# Force Groq whisper-large-v3

text = transcribe("./meeting.m4a", provider="groq")

# Force OpenAI whisper-1

text = transcribe("./lecture.mp3", provider="openai")

```

## Core Architecture and Source Code References

Understanding the internal implementation helps with debugging and customization:

- **[`agent_reach/transcribe.py`](https://github.com/Panniantong/Agent-Reach/blob/main/agent_reach/transcribe.py)**: Contains the complete pipeline implementation including `download_audio`, `compress_audio`, `chunk_audio`, and the public `transcribe` function (lines 207-246) that orchestrates the workflow.
- **[`agent_reach/config.py`](https://github.com/Panniantong/Agent-Reach/blob/main/agent_reach/config.py)**: Defines the configuration schema for API keys and provider settings.
- **[`agent_reach/cli.py`](https://github.com/Panniantong/Agent-Reach/blob/main/agent_reach/cli.py)**: Implements argument parsing and output handling for the CLI interface.

The default provider priority places **Groq** (`groq`) first for its speed and cost efficiency, with **OpenAI** (`openai`) serving as the fallback when the `provider="auto"` parameter is used.

## Summary

- Agent Reach bundles a complete transcription pipeline in [`agent_reach/transcribe.py`](https://github.com/Panniantong/Agent-Reach/blob/main/agent_reach/transcribe.py) that handles downloading, compression, chunking, and API communication.
- The workflow automatically splits files larger than 25 MiB into 10-minute chunks to comply with Whisper API limits.
- **Groq** is the default provider (using `whisper-large-v3`), with automatic fallback to **OpenAI** (`whisper-1`) if the primary provider fails.
- Configuration is stored in `~/.agent-reach/config.yaml` and managed through the `Config` class.
- Both CLI (`agent-reach transcribe`) and Python library interfaces support local files and URLs, with `yt-dlp` handling video platform downloads.

## Frequently Asked Questions

### What audio formats does Agent Reach support for transcription?

Agent Reach accepts any format that ffmpeg can decode, including MP3, WAV, M4A, and OGG. When processing URLs, yt-dlp extracts the best available audio stream automatically. The internal `compress_audio` function standardizes all inputs to mono 16 kHz 32 kbps M4A before sending to the Whisper API.

### How does Agent Reach handle audio files larger than 25 MB?

The `chunk_audio` function in [`agent_reach/transcribe.py`](https://github.com/Panniantong/Agent-Reach/blob/main/agent_reach/transcribe.py) automatically segments audio into 10-minute chunks when files exceed the Whisper API's 25 MiB limit. Each chunk is transcribed separately and concatenated into a single coherent transcript without requiring manual intervention.

### Can I use Agent Reach transcription without Groq or OpenAI API keys?

No, Agent Reach requires valid API credentials for at least one supported provider. The system checks for keys in `~/.agent-reach/config.yaml` or corresponding environment variables (`GROQ_API_KEY`, `OPENAI_API_KEY`). Attempting to transcribe without configured credentials raises a configuration error before any audio processing begins.

### Is it possible to transcribe YouTube videos directly using Agent Reach?

Yes, Agent Reach integrates yt-dlp to extract audio from YouTube URLs and other video platforms automatically. Simply pass the video URL to the `transcribe` function or CLI command, and the pipeline handles download, extraction, and transcription in one operation.