How Agent Reach Multi-Backend Routing Handles API Failures
Agent Reach implements fault-tolerant routing by probing an ordered list of backends for each channel, selecting the first healthy candidate while classifying failures as missing, broken, timeout, warning, or error states.
Agent Reach abstracts external platforms—such as YouTube, Twitter, and Reddit—into channel classes that can utilize multiple CLI backends. The routing system, defined primarily in agent_reach/channels/base.py, ensures agents automatically fail over to working backends when APIs are unavailable or misconfigured.
Backend Collection and Ordering
The foundation of resilient routing lies in the ordered_backends() method defined in the Channel base class.
User Override Priority
The ordered_backends() method returns the channel's default backend list (for example, ["twitter-cli", "OpenCLI", "bird CLI (legacy)"] for Twitter) and moves any user-specified override to the front of the list. This guarantees that a stale configuration cannot mask a functional backend.
Candidate Generation
Each concrete channel implements a check() method that iterates over these ordered candidates. The method first clears self.active_backend to None, then probes each backend sequentially until it finds a viable option.
The Probing and Selection Logic
The check() method implements a two-phase scanning strategy to determine backend health.
Phase 1: Probing Candidates
For each backend in the ordered list, the channel invokes a lightweight probe—typically via probe_command (from agent_reach/probe.py) or direct subprocess execution. The probe returns structured status codes:
None: The backend binary is not installed (shutil.whichreturnsNoneor the probe cannot locate the command)."ok": The backend is fully functional and ready for use."warn": The backend is installed but requires user action, such as missing authentication credentials."error": The backend binary is broken, timed out, or returned a non-zero exit code.
Phase 2: Status Selection
After collecting (backend, status, message) tuples, the method scans for the first "ok" entry. If none exists, it selects the first "warn" entry. This prioritization ensures that a partially working backend (e.g., an authenticated CLI) does not block a later, fully functional one.
If only "error" entries are found, the channel returns "error" with a concatenated message containing all failure reasons. If no backends are installed, it returns "warn" with installation instructions.
Handling Specific API Failure Types
The routing system categorizes failures to determine appropriate fallback behavior.
Missing Binaries and Broken Environments
When probe_command detects a missing binary or broken environment (such as a missing Node.js runtime for twitter-cli), it returns status == "missing" or status == "broken". The check() method treats these as probe failures, records the error, and continues to the next candidate.
Timeouts and Non-Zero Exits
Execution timeouts or subprocess failures with non-zero exit codes return status == "timeout" or ok == False. The system treats these as "error" states (or "warn" if the output suggests a credential issue), allowing the loop to proceed to the next backend.
Authentication and Warning States
If a backend runs successfully but outputs authentication error phrases like "not_authenticated", the probe returns "warn". The channel still considers this backend a candidate, selecting it only if no "ok" backends exist, ensuring degraded service beats total failure.
Partial Success Handling
For the OpenCLI shared backend, opencli_status() in agent_reach/backends/opencli.py inspects both daemon and extension state. If the extension exists on disk but the daemon reports "disconnected", the method sets ready=True because the service worker may be sleeping. This allows channels like Twitter and Reddit to treat OpenCLI as an "ok" backend despite transient disconnections.
Code Example: Fallback Routing in TwitterChannel
The TwitterChannel.check() method in agent_reach/channels/twitter.py demonstrates the complete routing logic:
def check(self, config=None):
self.active_backend = None
findings = []
for backend in self.ordered_backends(config):
if backend == "twitter-cli":
result = self._check_twitter_cli()
elif backend == "OpenCLI":
result = self._check_opencli()
elif backend == "bird CLI (legacy)":
result = self._check_bird()
else:
continue
if result is None: # not installed
continue
findings.append((backend, *result))
for wanted in ("ok", "warn"): # prefer ok, then warn
for backend, status, message in findings:
if status == wanted:
self.active_backend = backend
return status, message
if findings:
return "error", "\n".join(m for _, _, m in findings)
return "warn", "Twitter CLI 未安装…"
This implementation ensures that a later, fully functional backend is never hidden by an earlier, merely installed but unauthenticated one.
Practical Usage
To invoke the routing logic and inspect the selected backend:
from agent_reach.channels.twitter import TwitterChannel
from agent_reach.config import Config
cfg = Config() # reads user overrides from env/YAML
tw = TwitterChannel()
status, message = tw.check(cfg) # probes twitter-cli → OpenCLI → bird CLI
print(f"Status: {status}")
print(f"Message: {message}")
print(f"Active backend: {tw.active_backend}")
If twitter-cli is installed but lacks authentication, the method selects OpenCLI when it reports ready, because "ok" takes precedence over "warn".
Summary
- Ordered probing: The
ordered_backends()method inagent_reach/channels/base.pyprioritizes user overrides while maintaining default fallback sequences. - Status hierarchy: The
check()method selects backends in priority order:"ok">"warn">"error", ensuring maximum availability. - Failure classification: The
probe_commandhelper inagent_reach/probe.pynormalizes missing binaries, broken environments, timeouts, and authentication failures to determine whether to skip or select a backend. - OpenCLI resilience: The shared backend in
agent_reach/backends/opencli.pytreats disk-present extensions as ready even when the daemon reports disconnected, preventing false negatives across channels. - Active backend contract: The selected backend name is stored in
self.active_backendfor use by diagnostics and CLI displays.
Frequently Asked Questions
How does Agent Reach prioritize backends when multiple are available?
Agent Reach probes backends in the order returned by ordered_backends(), which places user-specified overrides first, followed by channel defaults. It selects the first backend reporting "ok" status; if none are healthy, it selects the first "warn" status. This ensures fully functional backends take precedence over degraded ones.
What happens if all backends fail the health check?
If all candidates return "error" or None (not installed), the check() method returns "error" with a concatenated string of all failure messages. If no backends are detected at all, it returns "warn" with installation instructions, allowing the application to handle the outage gracefully.
How does the OpenCLI backend handle disconnected states?
According to agent_reach/backends/opencli.py, the opencli_status() function checks whether the extension exists on disk even if the daemon reports "disconnected". If the extension is present, it returns ready=True, allowing channels to use OpenCLI as a valid backend despite transient service worker sleep states.
Can users force a specific backend despite probe failures?
While ordered_backends() moves user-specified backends to the front of the probe list, the check() method still validates the backend's health before selection. If the forced backend returns "error", the system falls back to the next candidate to prevent routing requests to known broken services.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →