# How ModLens Handles Multiple Vision Backends with Different Speeds and Costs

> Learn how ModLens intelligently manages multiple vision backends. Discover its prioritized fail-over and automatic speed/cost optimization for efficient processing.

- Repository: [liustack/modlens](https://github.com/liustack/modlens)
- Tags: architecture
- Published: 2026-08-25

---

**ModLens implements a pluggable provider architecture that builds a prioritized fail-over chain, automatically preferring fast API backends over slower CLI agents while allowing users to pin specific providers via command-line flags.**

ModLens analyzes images using multiple vision backends that vary significantly in response time and operational cost. The tool dynamically constructs an ordered execution chain that routes requests through the fastest, cheapest providers first, falling back to slower alternatives only when necessary. This speed-aware orchestration is implemented across three core modules that handle provider registration, availability detection, and chain composition.

## Provider Registration and Discovery

All vision backends are registered in [`src/providers/index.ts`](https://github.com/liustack/modlens/blob/main/src/providers/index.ts) within the `PROVIDERS` map. Each entry specifies the canonical provider name, default model, and execution flags that determine whether the provider runs as a subprocess (`isolateWorkdir`) or an in-process API client (`execute`).

This centralized registry allows ModLens to treat disparate backends—ranging from cloud APIs to local CLI tools—as interchangeable units. The configuration for each provider includes metadata about its runtime requirements, enabling the system to determine compatibility without executing the binary.

## Detecting Available Vision Backends

Before building the execution chain, ModLens filters the global registry to identify which providers are actually usable on the current system. The [`src/providers/availability.ts`](https://github.com/liustack/modlens/blob/main/src/providers/availability.ts) module defines `PROVIDER_DESCRIPTORS` (the provider catalog) and implements the `providerAvailable` helper function.

Sub-process providers are considered available only when their binary exists on `PATH`, while API providers require valid configuration fields such as API keys and base URLs. The `providerChain(kind, config, env)` function then constructs an ordered list of available providers specific to the input type—either **local** (image files) or **remote** (URLs).

## Speed-Class Regions: Inline vs Agents

ModLens categorizes providers into two distinct speed-class regions to optimize for latency and cost. The classification is implemented as a `Set` in [`src/analyzer.ts`](https://github.com/liustack/modlens/blob/main/src/analyzer.ts) that defines which providers belong to the **inline** region versus the **agents** region.

**Inline region** providers—including `gemini-api`, `openai`, and `anthropic`—typically return results in **5-10 seconds** and consume per-request quotas. These are fast, cost-effective API calls that require no local infrastructure.

**Agents region** providers—such as `antigravity-cli` and `claude-cli`—run as subprocesses and take **15-45 seconds** to complete. These often require active subscriptions and consume more resources, making them suitable only as fallbacks when inline providers fail or exhaust quotas.

## Composing the Fail-Over Chain

The `composeChain` function in [`src/analyzer.ts`](https://github.com/liustack/modlens/blob/main/src/analyzer.ts) merges the base provider chain (generated by `providerChain`) with borrowed routes from other harnesses while preserving speed-class ordering. Inline-region providers are inserted after the user’s own inline providers, while agent-region providers are positioned before `claude-cli` unless explicitly overridden.

The composition logic also respects user-specified preferences stored in `config.provider` or passed via the `-p` flag. When a preference exists, the selected provider moves to the front of its respective region, ensuring it is attempted before other providers in the same speed class.

## Handling Failures with Cooldown Logic

When a provider encounters a failure or quota exhaustion, the `CooldownView` class records a cooldown period for that backend. The `reorderByCooldown` function then re-sorts the provider chain, moving cooled providers to the **back of their own region** (inline or agent) while maintaining the relative order of healthy providers.

This ensures that subsequent image analyses skip exhausted backends without breaking the speed-class hierarchy. Fast providers remain at the front of the inline region, while slower agents remain subordinate even when other agents enter cooldown.

## User Control: Pinning and Exclusion

Users can override the automatic selection logic through two mechanisms defined in [`src/providers/availability.ts`](https://github.com/liustack/modlens/blob/main/src/providers/availability.ts) and [`src/config.ts`](https://github.com/liustack/modlens/blob/main/src/config.ts).

**Pinning a provider** via `modlens -p <provider>` or `modlens config set provider <provider>` forces exclusive use of that backend, bypassing the fail-over chain entirely. This is useful when a specific model is required for consistent output.

**Pin-only providers**—such as `kimi-cli`—are excluded from the default chain entirely and execute only when explicitly requested via the `-p` flag. This prevents resource-intensive or experimental backends from slowing down standard workflows while keeping them accessible for specialized tasks.

```bash

# Use the default fast-first chain (no explicit provider)

modlens -i screenshot.png

```

```bash

# Pin a specific provider (e.g., the Gemini API) – no fail-over

modlens -i screenshot.png -p gemini-api

```

```bash

# Show the provider chain that will be used for the current machine

modlens doctor --json | jq '.chain'

```

```bash

# Set a permanent preference in the config (fast API first)

modlens config set provider gemini-api

```

```bash

# Pin a "pin-only" provider (Kimi Code) – it runs only when requested

modlens -i screenshot.png -p kimi-cli

```

## Summary

- **Provider registration** occurs in [`src/providers/index.ts`](https://github.com/liustack/modlens/blob/main/src/providers/index.ts), where the `PROVIDERS` map catalogs all available backends with their execution modes.
- **Availability detection** in [`src/providers/availability.ts`](https://github.com/liustack/modlens/blob/main/src/providers/availability.ts) filters providers based on binary presence or API configuration, creating a base chain via `providerChain`.
- **Speed-class regions** separate providers into **inline** (fast APIs, 5-10s) and **agents** (slow CLIs, 15-45s) to prioritize low-latency, low-cost options.
- **Chain composition** in [`src/analyzer.ts`](https://github.com/liustack/modlens/blob/main/src/analyzer.ts) merges available providers with borrowed routes while preserving regional ordering and user preferences.
- **Cooldown handling** moves failed providers to the back of their region via `reorderByCooldown`, preventing repeated attempts at exhausted backends.
- **User pinning** via `-p` flags or config settings forces specific providers, while pin-only providers like `kimi-cli` remain excluded from automatic selection unless explicitly requested.

## Frequently Asked Questions

### How does ModLens decide which vision backend to use first?

ModLens builds a fail-over chain using `providerChain` in [`src/providers/availability.ts`](https://github.com/liustack/modlens/blob/main/src/providers/availability.ts), which orders providers first by speed class (inline APIs before CLI agents) and then by availability. The `composeChain` function in [`src/analyzer.ts`](https://github.com/liustack/modlens/blob/main/src/analyzer.ts) further refines this by applying user preferences and cooldown statuses, ensuring the fastest available provider is attempted before slower fallbacks.

### What happens when a vision provider fails or hits a quota limit?

When a provider fails, the `CooldownView` class records a cooldown period for that backend. During subsequent analyses, `reorderByCooldown` moves the cooled provider to the back of its respective speed-class region (inline or agents). This maintains the priority of healthy providers while temporarily deprioritizing exhausted ones without removing them from the chain entirely.

### Can I force ModLens to use a specific backend instead of the automatic chain?

Yes. Use the `-p` flag (e.g., `modlens -p gemini-api`) to pin a specific provider, which bypasses the automatic fail-over logic. You can also set a permanent preference using `modlens config set provider <name>`. When pinned, ModLens attempts only the specified provider and fails immediately if that backend is unavailable, rather than falling back to alternatives.

### What is the difference between inline and agent vision providers in ModLens?

**Inline providers** are API clients (like `gemini-api` or `openai`) that return results in 5-10 seconds and typically use pay-per-request pricing. **Agent providers** are subprocess CLIs (like `claude-cli` or `antigravity-cli`) that spawn external processes, take 15-45 seconds, and often require active subscriptions. ModLens always attempts inline providers before agent providers unless the user explicitly pins an agent or all inline providers enter cooldown.