# Agent Reach Upstream Tools Required Per Platform: Complete Dependency Mapping

> Discover Agent Reach upstream tools per platform like Jina Reader yt-dlp and Exa. Understand complete dependency mapping for seamless operation and automatic fallback chains.

- Repository: [Pnant/Agent-Reach](https://github.com/Panniantong/Agent-Reach)
- Tags: how-to-guide
- Published: 2026-07-02

---

**Agent Reach requires specific third-party CLI tools for each platform, such as Jina Reader for web pages, yt-dlp for YouTube, and Exa via mcporter for search, with automatic fallback chains defined in individual channel classes.**

The **Panniantong/Agent-Reach** repository implements a thin abstraction layer that lets AI agents consume web content without maintaining dozens of separate binaries. Understanding which **upstream tools are required per platform** is essential for proper installation and troubleshooting. Each platform maps to a primary tool and optional fallbacks, configured in the channel implementations under `agent_reach/channels/`.

## Platform-to-Tool Mapping

Agent Reach maintains a structured mapping of platforms to their required upstream CLI tools or MCP services. The following table documents the exact dependencies, including primary tools and fallback chains, as defined in the repository README:

| Platform | Primary Tool | Fallback Chain |
|----------|--------------|----------------|
| **Web** | Jina Reader | – |
| **YouTube** | `yt-dlp` | – |
| **RSS** | `feedparser` (Python library) | – |
| **Full-text search** | Exa via `mcporter` | – |
| **GitHub** | GitHub CLI (`gh`) | – |
| **Twitter/X** | `twitter-cli` | `OpenCLI` → `bird` |
| **BiliBili** | `bili-cli` | `OpenCLI` → generic search API |
| **Reddit** | `OpenCLI` (desktop) | `rdt-cli` |
| **Facebook** | `OpenCLI` (desktop) | – |
| **Instagram** | `OpenCLI` (desktop) | Official Graph API |
| **Xiaohongshu** | `OpenCLI` (desktop) | `xiaohongshu-mcp` → `xhs-cli` |
| **LinkedIn** | `linkedin-scraper-mcp` | Jina Reader |
| **V2EX** | Built-in HTTP/HTML scraper | – |
| **Xueqiu (雪球)** | Built-in HTTP scraper | – |
| **Xiaoyuzhou** | Built-in HTTP scraper | – |

Platforms marked with built-in scrapers require no external tool installation, as they rely on native HTTP/HTML parsing implemented directly in the channel classes.

## How the Architecture Works

The **BaseChannel** abstract class in [`agent_reach/channels/base.py`](https://github.com/Panniantong/Agent-Reach/blob/main/agent_reach/channels/base.py) defines the contract that every platform implementation must follow. This architecture enables the dynamic selection of upstream tools at runtime.

### BaseChannel Contract

Each platform channel inherits from `BaseChannel` and implements four core methods:

- **`can_handle(url)`**: Determines whether the channel should process a given URL based on domain patterns or content type.
- **`read(url)`**: Invokes the primary upstream tool (or fallback) to extract content, typically via `subprocess.run()`.
- **`search(query)`**: Delegates to search-specific backends like Exa or platform-specific APIs.
- **`check()`**: Probes the system for the presence and health of each candidate tool, returning the first available implementation.

### Runtime Tool Selection

When a user executes a command, the channel class performs the following steps:

1. **URL Classification**: `can_handle()` identifies the appropriate channel for the input.
2. **Health Check**: `check()` verifies which upstream tools are installed and responsive.
3. **Execution**: The selected tool runs as a subprocess with platform-specific arguments (e.g., `yt-dlp --write-subtitle <url>`).
4. **Fallback**: If the primary tool fails or is absent, the channel automatically attempts the next tool in the defined chain.

### Diagnostic Commands

The **`doctor`** command (`agent-reach doctor`) implemented in [`agent_reach/doctor.py`](https://github.com/Panniantong/Agent-Reach/blob/main/agent_reach/doctor.py) enumerates which tool is currently active for every platform. This diagnostic tool runs `check()` on every registered channel and prints remediation steps if a candidate dependency is missing.

## Practical Usage Examples

The following commands demonstrate how Agent Reach translates high-level requests into subprocess calls of the respective upstream tools. These examples assume the appropriate dependencies are already installed (use `agent-reach install` to configure them).

Read a web page using Jina Reader:

```bash
agent-reach read https://example.com

```

Extract YouTube subtitles with yt-dlp:

```bash
agent-reach read https://www.youtube.com/watch?v=dQw4w9WgXcQ

```

Search the web using Exa via mcporter:

```bash
agent-reach search "latest LLM frameworks comparison"

```

Read a GitHub repository using the GitHub CLI:

```bash
agent-reach read https://github.com/psf/requests

```

Process a Twitter/X post using the twitter-cli chain:

```bash
agent-reach read https://x.com/username/status/1234567890

```

Search Reddit using the OpenCLI fallback:

```bash
agent-reach search "python web scraping site:reddit.com"

```

Read an RSS feed via feedparser:

```bash
agent-reach read https://news.ycombinator.com/rss

```

Access BiliBili content using bili-cli:

```bash
agent-reach read https://www.bilibili.com/video/BV1xJ411x7dF

```

Read Xiaohongshu posts via OpenCLI:

```bash
agent-reach read https://www.xiaohongshu.com/discovery/item/1234567890abcdef

```

Access LinkedIn profiles using linkedin-scraper-mcp:

```bash
agent-reach read https://www.linkedin.com/in/someone

```

## Key Source Files

Understanding the implementation requires familiarity with these specific files:

- **[`agent_reach/channels/base.py`](https://github.com/Panniantong/Agent-Reach/blob/main/agent_reach/channels/base.py)**: Contains the abstract `BaseChannel` class defining the contract (`can_handle`, `read`, `search`, `check`).
- **`agent_reach/channels/<platform>.py`**: Concrete implementations for each platform (e.g., [`twitter.py`](https://github.com/Panniantong/Agent-Reach/blob/main/twitter.py), [`youtube.py`](https://github.com/Panniantong/Agent-Reach/blob/main/youtube.py), [`github.py`](https://github.com/Panniantong/Agent-Reach/blob/main/github.py)). Each file specifies the exact upstream command and fallback logic.
- **[`agent_reach/doctor.py`](https://github.com/Panniantong/Agent-Reach/blob/main/agent_reach/doctor.py)**: Diagnostic engine that executes `check()` on every channel and reports active tool status.
- **[`agent_reach/cli.py`](https://github.com/Panniantong/Agent-Reach/blob/main/agent_reach/cli.py)**: Entry point handling subcommands (`read`, `search`, `install`, `doctor`).
- **[`README.md`](https://github.com/Panniantong/Agent-Reach/blob/main/README.md)**: Source of truth for the platform-to-tool mapping table and dependency documentation.

## Summary

- **Agent Reach** acts as a thin glue layer between AI agents and third-party CLI tools, requiring specific upstream binaries per platform.
- **Primary tools** include Jina Reader (Web), `yt-dlp` (YouTube), `gh` (GitHub), and `mcporter` (Search), with defined fallback chains for social platforms.
- **Platform channels** in `agent_reach/channels/` implement `BaseChannel` to handle URL detection, tool execution, and health checking.
- **Built-in scrapers** handle V2EX, Xueqiu, and Xiaoyuzhou without external dependencies.
- **Diagnostic tools** via `agent-reach doctor` verify which upstream tools are active and available.

## Frequently Asked Questions

### What upstream tools does Agent Reach require for YouTube?

Agent Reach requires **`yt-dlp`** as the primary upstream tool for YouTube content extraction. When you execute `agent-reach read <youtube-url>`, the channel invokes `yt-dlp` with the `--write-subtitle` flag to extract video metadata and captions. There is no fallback chain defined for YouTube, so the tool must be installed and available in the system PATH.

### How does Agent Reach handle missing dependencies?

When a required upstream tool is missing, the **`check()`** method in the platform's channel class returns a negative status, causing the system to attempt the next tool in the fallback chain. If all candidates fail, the **`doctor`** command (`agent-reach doctor`) identifies the missing dependency and provides installation guidance. The architecture keeps upstream utilities unmodified, so swapping or updating tools only requires editing the specific channel file.

### Which platforms use built-in scrapers instead of external tools?

**V2EX**, **Xueqiu (雪球)**, and **Xiaoyuzhou** rely on built-in HTTP/HTML scrapers that require no external CLI tools. These platforms implement native parsing logic directly in their respective channel classes under `agent_reach/channels/`, eliminating the need for third-party binaries while maintaining full functionality.

### Where is the platform-to-tool mapping defined?

The definitive mapping of platforms to required upstream tools is documented in the **[`README.md`](https://github.com/Panniantong/Agent-Reach/blob/main/README.md)** under the section titled *"每个平台 = 首选 + 备选的有序后端列表"*. The concrete implementation logic resides in individual files within `agent_reach/channels/`, where each platform class defines its primary tool and fallback sequence in the `check()` and `read()` methods.