How to Use the Agent-Reach Transcribe Command for YouTube Video Transcription with Groq Whisper
To transcribe YouTube videos with Groq Whisper using Agent-Reach, configure your Groq API key with agent-reach configure groq-key, then run agent-reach transcribe <url> --provider groq to download audio, chunk it, and generate transcripts using the free whisper-large-v3 model.
The Agent-Reach open-source repository (Panniantong/Agent-Reach) provides a unified transcription pipeline that automates the entire workflow from video download to text generation. This guide explains how to leverage the built-in transcribe command with Groq's free Whisper API to convert YouTube videos into accurate text transcripts.
Prerequisites: Configure Your Groq API Key
Before processing videos, you must store your Groq API key in the Agent-Reach configuration. The CLI stores this in ~/.agent-reach/config.yaml or reads from environment variables via agent_reach/config.py.
agent-reach configure groq-key YOUR_GROQ_API_KEY
This one-time setup persists the key for all future transcription requests.
Basic CLI Usage
The transcribe command in agent_reach/cli.py accepts a YouTube URL and optional provider flags. Use the --provider groq flag to explicitly route requests to Groq's Whisper endpoint.
# Transcribe with explicit Groq provider
agent-reach transcribe "https://www.youtube.com/watch?v=abc123DEF" --provider groq
# Output to a file instead of stdout
agent-reach transcribe "https://youtu.be/xyz789" -o transcript.txt
You can also omit the provider flag entirely. The default "auto" setting attempts Groq first, then falls back to OpenAI if Groq returns an error.
How the Transcription Pipeline Works
The Agent-Reach transcription pipeline is implemented across three core modules: agent_reach/cli.py for argument parsing, agent_reach/channels/youtube.py for URL handling, and agent_reach/transcribe.py for the core audio processing logic.
CLI Parsing and Provider Selection
In agent_reach/cli.py (lines 117-124), the transcribe sub-parser accepts the source parameter and --provider flag. Valid provider values are auto, groq, or openai.
The provider selection logic in agent_reach/transcribe.py uses a _provider_order list (lines 99-104) that maps "auto" to ["groq", "openai"]. When you specify groq explicitly, the system looks up the provider configuration in the PROVIDERS dictionary and routes requests to Groq's API endpoint at https://api.groq.com/openai/v1/audio/transcriptions.
Audio Processing and Chunking
The download_audio function in agent_reach/transcribe.py (lines 77-95) calls yt-dlp to extract the best audio stream from the YouTube URL and saves it locally. The pipeline then processes this audio through two critical stages:
-
Compression: The
compress_audiofunction (lines 101-124) re-encodes the audio to mono 16 kHz at 32 kbps in M4A format. This optimizes the file for Whisper's input requirements while minimizing bandwidth. -
Chunking: The
chunk_audiofunction (lines 126-154) checks if the compressed file exceeds 24 MiB. If it does, the audio is split into segments of no more than 10 minutes each to comply with Groq's API limits.
Groq Whisper API Integration
Each audio chunk is processed by the transcribe_chunk function in agent_reach/transcribe.py (lines 63-90). This function sends the binary audio data to Groq's API using the whisper-large-v3 model. The _transcribe_with_fallback_ wrapper (lines 49-61) handles retry logic and provider failover when using auto mode.
Finally, the main transcribe function (lines 107-125) aggregates all chunk transcripts into a single string, handling the concatenation before returning the result to the CLI or writing to the output file specified by -o.
Python API Usage
You can bypass the CLI and use the core transcription functions directly in Python scripts. The agent_reach/transcribe.py module exposes the main transcribe function that accepts URLs and provider strings.
from agent_reach.transcribe import transcribe
text = transcribe(
"https://www.youtube.com/watch?v=abc123DEF",
provider="groq" # Optional: defaults to "auto" which tries Groq first
)
print(text[:200]) # Preview first 200 characters
The agent_reach/channels/youtube.py module (lines 81-90) provides a thin wrapper class YouTubeChannel with a transcribe method that simply forwards to the core transcribe function, allowing seamless integration for channel-specific processing workflows.
Summary
- Configuration: Store your Groq API key using
agent-reach configure groq-keybefore running transcriptions. - CLI Command: Use
agent-reach transcribe <youtube-url> --provider groqto process videos with Groq's freewhisper-large-v3model. - Audio Pipeline: The system automatically downloads audio via
yt-dlp, compresses it to mono 16 kHz 32 kbps M4A, and chunks files larger than 24 MiB into 10-minute segments. - Provider Logic: The
autoprovider setting tries Groq first, then falls back to OpenAI if errors occur. - Core Files:
agent_reach/transcribe.pyhandles the heavy lifting (download, compression, chunking, API calls), whileagent_reach/channels/youtube.pyprovides URL-specific wrappers.
Frequently Asked Questions
What Whisper model does Groq use for transcription?
According to the implementation in agent_reach/transcribe.py, Agent-Reach specifically calls Groq's whisper-large-v3 model endpoint. This is the free, production-ready model hosted on Groq's infrastructure, offering fast transcription speeds without API usage charges.
How does Agent-Reach handle large video files?
The chunk_audio function in agent_reach/transcribe.py (lines 126-154) automatically splits audio files exceeding 24 MiB into chunks of 10 minutes or less. This ensures compliance with Groq's API size limits while maintaining sequential processing. The transcribe function then reassembles these chunks into a single coherent transcript.
Can I use automatic fallback if Groq fails?
Yes. When using --provider auto (the default), the _transcribe_with_fallback_ function in agent_reach/transcribe.py (lines 49-61) implements a retry mechanism that attempts Groq first based on the _provider_order list, then automatically switches to OpenAI if Groq returns an error or is unavailable.
Where are API keys stored in Agent-Reach?
API keys are managed by agent_reach/config.py and stored in a YAML configuration file at ~/.agent-reach/config.yaml. The system also checks for environment variables, allowing keys to be set via export commands or loaded from .env files as fallback mechanisms.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →