Agent Reach Upstream Tools Required Per Platform: Complete Dependency Mapping

Agent Reach requires specific third-party CLI tools for each platform, such as Jina Reader for web pages, yt-dlp for YouTube, and Exa via mcporter for search, with automatic fallback chains defined in individual channel classes.

The Panniantong/Agent-Reach repository implements a thin abstraction layer that lets AI agents consume web content without maintaining dozens of separate binaries. Understanding which upstream tools are required per platform is essential for proper installation and troubleshooting. Each platform maps to a primary tool and optional fallbacks, configured in the channel implementations under agent_reach/channels/.

Platform-to-Tool Mapping

Agent Reach maintains a structured mapping of platforms to their required upstream CLI tools or MCP services. The following table documents the exact dependencies, including primary tools and fallback chains, as defined in the repository README:

Platform Primary Tool Fallback Chain
Web Jina Reader –
YouTube yt-dlp –
RSS feedparser (Python library) –
Full-text search Exa via mcporter –
GitHub GitHub CLI (gh) –
Twitter/X twitter-cli OpenCLI → bird
BiliBili bili-cli OpenCLI → generic search API
Reddit OpenCLI (desktop) rdt-cli
Facebook OpenCLI (desktop) –
Instagram OpenCLI (desktop) Official Graph API
Xiaohongshu OpenCLI (desktop) xiaohongshu-mcp → xhs-cli
LinkedIn linkedin-scraper-mcp Jina Reader
V2EX Built-in HTTP/HTML scraper –
Xueqiu (雪球) Built-in HTTP scraper –
Xiaoyuzhou Built-in HTTP scraper –

Platforms marked with built-in scrapers require no external tool installation, as they rely on native HTTP/HTML parsing implemented directly in the channel classes.

How the Architecture Works

The BaseChannel abstract class in agent_reach/channels/base.py defines the contract that every platform implementation must follow. This architecture enables the dynamic selection of upstream tools at runtime.

BaseChannel Contract

Each platform channel inherits from BaseChannel and implements four core methods:

  • can_handle(url): Determines whether the channel should process a given URL based on domain patterns or content type.
  • read(url): Invokes the primary upstream tool (or fallback) to extract content, typically via subprocess.run().
  • search(query): Delegates to search-specific backends like Exa or platform-specific APIs.
  • check(): Probes the system for the presence and health of each candidate tool, returning the first available implementation.

Runtime Tool Selection

When a user executes a command, the channel class performs the following steps:

  1. URL Classification: can_handle() identifies the appropriate channel for the input.
  2. Health Check: check() verifies which upstream tools are installed and responsive.
  3. Execution: The selected tool runs as a subprocess with platform-specific arguments (e.g., yt-dlp --write-subtitle <url>).
  4. Fallback: If the primary tool fails or is absent, the channel automatically attempts the next tool in the defined chain.

Diagnostic Commands

The doctor command (agent-reach doctor) implemented in agent_reach/doctor.py enumerates which tool is currently active for every platform. This diagnostic tool runs check() on every registered channel and prints remediation steps if a candidate dependency is missing.

Practical Usage Examples

The following commands demonstrate how Agent Reach translates high-level requests into subprocess calls of the respective upstream tools. These examples assume the appropriate dependencies are already installed (use agent-reach install to configure them).

Read a web page using Jina Reader:

agent-reach read https://example.com

Extract YouTube subtitles with yt-dlp:

agent-reach read https://www.youtube.com/watch?v=dQw4w9WgXcQ

Search the web using Exa via mcporter:

agent-reach search "latest LLM frameworks comparison"

Read a GitHub repository using the GitHub CLI:

agent-reach read https://github.com/psf/requests

Process a Twitter/X post using the twitter-cli chain:

agent-reach read https://x.com/username/status/1234567890

Search Reddit using the OpenCLI fallback:

agent-reach search "python web scraping site:reddit.com"

Read an RSS feed via feedparser:

agent-reach read https://news.ycombinator.com/rss

Access BiliBili content using bili-cli:

agent-reach read https://www.bilibili.com/video/BV1xJ411x7dF

Read Xiaohongshu posts via OpenCLI:

agent-reach read https://www.xiaohongshu.com/discovery/item/1234567890abcdef

Access LinkedIn profiles using linkedin-scraper-mcp:

agent-reach read https://www.linkedin.com/in/someone

Key Source Files

Understanding the implementation requires familiarity with these specific files:

  • agent_reach/channels/base.py: Contains the abstract BaseChannel class defining the contract (can_handle, read, search, check).
  • agent_reach/channels/<platform>.py: Concrete implementations for each platform (e.g., twitter.py, youtube.py, github.py). Each file specifies the exact upstream command and fallback logic.
  • agent_reach/doctor.py: Diagnostic engine that executes check() on every channel and reports active tool status.
  • agent_reach/cli.py: Entry point handling subcommands (read, search, install, doctor).
  • README.md: Source of truth for the platform-to-tool mapping table and dependency documentation.

Summary

  • Agent Reach acts as a thin glue layer between AI agents and third-party CLI tools, requiring specific upstream binaries per platform.
  • Primary tools include Jina Reader (Web), yt-dlp (YouTube), gh (GitHub), and mcporter (Search), with defined fallback chains for social platforms.
  • Platform channels in agent_reach/channels/ implement BaseChannel to handle URL detection, tool execution, and health checking.
  • Built-in scrapers handle V2EX, Xueqiu, and Xiaoyuzhou without external dependencies.
  • Diagnostic tools via agent-reach doctor verify which upstream tools are active and available.

Frequently Asked Questions

What upstream tools does Agent Reach require for YouTube?

Agent Reach requires yt-dlp as the primary upstream tool for YouTube content extraction. When you execute agent-reach read <youtube-url>, the channel invokes yt-dlp with the --write-subtitle flag to extract video metadata and captions. There is no fallback chain defined for YouTube, so the tool must be installed and available in the system PATH.

How does Agent Reach handle missing dependencies?

When a required upstream tool is missing, the check() method in the platform's channel class returns a negative status, causing the system to attempt the next tool in the fallback chain. If all candidates fail, the doctor command (agent-reach doctor) identifies the missing dependency and provides installation guidance. The architecture keeps upstream utilities unmodified, so swapping or updating tools only requires editing the specific channel file.

Which platforms use built-in scrapers instead of external tools?

V2EX, Xueqiu (雪球), and Xiaoyuzhou rely on built-in HTTP/HTML scrapers that require no external CLI tools. These platforms implement native parsing logic directly in their respective channel classes under agent_reach/channels/, eliminating the need for third-party binaries while maintaining full functionality.

Where is the platform-to-tool mapping defined?

The definitive mapping of platforms to required upstream tools is documented in the README.md under the section titled "每个平台 = 首选 + 备选的有序后端列表". The concrete implementation logic resides in individual files within agent_reach/channels/, where each platform class defines its primary tool and fallback sequence in the check() and read() methods.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →