How ModLens Handles Multiple Vision Backends with Different Speeds and Costs

ModLens implements a pluggable provider architecture that builds a prioritized fail-over chain, automatically preferring fast API backends over slower CLI agents while allowing users to pin specific providers via command-line flags.

ModLens analyzes images using multiple vision backends that vary significantly in response time and operational cost. The tool dynamically constructs an ordered execution chain that routes requests through the fastest, cheapest providers first, falling back to slower alternatives only when necessary. This speed-aware orchestration is implemented across three core modules that handle provider registration, availability detection, and chain composition.

Provider Registration and Discovery

All vision backends are registered in src/providers/index.ts within the PROVIDERS map. Each entry specifies the canonical provider name, default model, and execution flags that determine whether the provider runs as a subprocess (isolateWorkdir) or an in-process API client (execute).

This centralized registry allows ModLens to treat disparate backends—ranging from cloud APIs to local CLI tools—as interchangeable units. The configuration for each provider includes metadata about its runtime requirements, enabling the system to determine compatibility without executing the binary.

Detecting Available Vision Backends

Before building the execution chain, ModLens filters the global registry to identify which providers are actually usable on the current system. The src/providers/availability.ts module defines PROVIDER_DESCRIPTORS (the provider catalog) and implements the providerAvailable helper function.

Sub-process providers are considered available only when their binary exists on PATH, while API providers require valid configuration fields such as API keys and base URLs. The providerChain(kind, config, env) function then constructs an ordered list of available providers specific to the input type—either local (image files) or remote (URLs).

Speed-Class Regions: Inline vs Agents

ModLens categorizes providers into two distinct speed-class regions to optimize for latency and cost. The classification is implemented as a Set in src/analyzer.ts that defines which providers belong to the inline region versus the agents region.

Inline region providers—including gemini-api, openai, and anthropic—typically return results in 5-10 seconds and consume per-request quotas. These are fast, cost-effective API calls that require no local infrastructure.

Agents region providers—such as antigravity-cli and claude-cli—run as subprocesses and take 15-45 seconds to complete. These often require active subscriptions and consume more resources, making them suitable only as fallbacks when inline providers fail or exhaust quotas.

Composing the Fail-Over Chain

The composeChain function in src/analyzer.ts merges the base provider chain (generated by providerChain) with borrowed routes from other harnesses while preserving speed-class ordering. Inline-region providers are inserted after the user’s own inline providers, while agent-region providers are positioned before claude-cli unless explicitly overridden.

The composition logic also respects user-specified preferences stored in config.provider or passed via the -p flag. When a preference exists, the selected provider moves to the front of its respective region, ensuring it is attempted before other providers in the same speed class.

Handling Failures with Cooldown Logic

When a provider encounters a failure or quota exhaustion, the CooldownView class records a cooldown period for that backend. The reorderByCooldown function then re-sorts the provider chain, moving cooled providers to the back of their own region (inline or agent) while maintaining the relative order of healthy providers.

This ensures that subsequent image analyses skip exhausted backends without breaking the speed-class hierarchy. Fast providers remain at the front of the inline region, while slower agents remain subordinate even when other agents enter cooldown.

User Control: Pinning and Exclusion

Users can override the automatic selection logic through two mechanisms defined in src/providers/availability.ts and src/config.ts.

Pinning a provider via modlens -p <provider> or modlens config set provider <provider> forces exclusive use of that backend, bypassing the fail-over chain entirely. This is useful when a specific model is required for consistent output.

Pin-only providers—such as kimi-cli—are excluded from the default chain entirely and execute only when explicitly requested via the -p flag. This prevents resource-intensive or experimental backends from slowing down standard workflows while keeping them accessible for specialized tasks.


# Use the default fast-first chain (no explicit provider)

modlens -i screenshot.png

# Pin a specific provider (e.g., the Gemini API) – no fail-over

modlens -i screenshot.png -p gemini-api

# Show the provider chain that will be used for the current machine

modlens doctor --json | jq '.chain'

# Set a permanent preference in the config (fast API first)

modlens config set provider gemini-api

# Pin a "pin-only" provider (Kimi Code) – it runs only when requested

modlens -i screenshot.png -p kimi-cli

Summary

  • Provider registration occurs in src/providers/index.ts, where the PROVIDERS map catalogs all available backends with their execution modes.
  • Availability detection in src/providers/availability.ts filters providers based on binary presence or API configuration, creating a base chain via providerChain.
  • Speed-class regions separate providers into inline (fast APIs, 5-10s) and agents (slow CLIs, 15-45s) to prioritize low-latency, low-cost options.
  • Chain composition in src/analyzer.ts merges available providers with borrowed routes while preserving regional ordering and user preferences.
  • Cooldown handling moves failed providers to the back of their region via reorderByCooldown, preventing repeated attempts at exhausted backends.
  • User pinning via -p flags or config settings forces specific providers, while pin-only providers like kimi-cli remain excluded from automatic selection unless explicitly requested.

Frequently Asked Questions

How does ModLens decide which vision backend to use first?

ModLens builds a fail-over chain using providerChain in src/providers/availability.ts, which orders providers first by speed class (inline APIs before CLI agents) and then by availability. The composeChain function in src/analyzer.ts further refines this by applying user preferences and cooldown statuses, ensuring the fastest available provider is attempted before slower fallbacks.

What happens when a vision provider fails or hits a quota limit?

When a provider fails, the CooldownView class records a cooldown period for that backend. During subsequent analyses, reorderByCooldown moves the cooled provider to the back of its respective speed-class region (inline or agents). This maintains the priority of healthy providers while temporarily deprioritizing exhausted ones without removing them from the chain entirely.

Can I force ModLens to use a specific backend instead of the automatic chain?

Yes. Use the -p flag (e.g., modlens -p gemini-api) to pin a specific provider, which bypasses the automatic fail-over logic. You can also set a permanent preference using modlens config set provider <name>. When pinned, ModLens attempts only the specified provider and fails immediately if that backend is unavailable, rather than falling back to alternatives.

What is the difference between inline and agent vision providers in ModLens?

Inline providers are API clients (like gemini-api or openai) that return results in 5-10 seconds and typically use pay-per-request pricing. Agent providers are subprocess CLIs (like claude-cli or antigravity-cli) that spawn external processes, take 15-45 seconds, and often require active subscriptions. ModLens always attempts inline providers before agent providers unless the user explicitly pins an agent or all inline providers enter cooldown.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →