CLI-Based Providers in ModLens: How They Work and When to Use Them

CLI-based providers in ModLens execute external command-line tools as subprocesses—specifically antigravity-cli, claude-cli, and kimi-cli—to analyze images and return structured JSON evidence for LLM workflows.

ModLens (liustack/modlens) is an open-source vision-to-JSON tool that supports multiple backends for converting screenshots into structured data. While API-based providers make direct HTTP calls to cloud services, the CLI-based providers in ModLens offer a distinct architecture that invokes locally installed binaries, making them ideal for offline environments or custom toolchains.

What Are CLI-Based Providers in ModLens?

ModLens supports six vision providers total, with three operating as CLI-based providers: Antigravity, Claude CLI, and Kimi CLI. Unlike API-based providers that implement an execute() method for direct network requests, these CLI providers spawn separate subprocesses to run external command-line tools installed on your system. Each provider adheres to a strict contract defined in src/schema.ts, ensuring that every backend returns identically structured JSON regardless of the underlying CLI implementation.

The provider registry in src/providers/index.ts maps short identifiers—antigravity, claude-cli, and kimi-cli—to their concrete implementations, enabling runtime selection via the -p <provider> flag.

How CLI-Based Providers Work

CLI-based providers in ModLens share a unified architecture centered on two abstract methods: buildInvocation() and parseOutput(). This design isolates external process execution from the main application while maintaining a consistent interface for image analysis across all backends.

The Build Invocation Pattern

The buildInvocation() method constructs the command-line string that invokes the external tool. In src/providers/antigravity.ts, this method assembles the antigravity-cli command with the --json-schema flag, ensuring the model outputs raw JSON that conforms to the ModLens schema. Similarly, src/providers/claudeCli.ts builds commands for claude-cli, while src/providers/kimiCli.ts constructs invocations for the Kimi vision CLI.

Output Parsing and Validation

After executing the subprocess, the parseOutput() method extracts the JSON result from stdout. All three CLI providers validate their output against the shared schema defined in src/schema.ts. This guarantees that downstream LLM workflows receive consistently structured evidence, whether the analysis came from Antigravity's zero-config model or Kimi's premium vision endpoint.

The Three CLI Providers

Antigravity (src/providers/antigravity.ts): The default provider that invokes antigravity-cli. It requires zero configuration, passing image data via temporary files or base-64 pipes, and returns JSON matching the ModLens schema immediately.

Claude CLI (src/providers/claudeCli.ts): Wraps the claude-cli command-line tool (Claude Code & Pi), executing Claude's vision capabilities through a local subprocess. It uses the same --json-schema flag to enforce output structure while keeping the CLI isolated from the main process.

Kimi CLI (src/providers/kimiCli.ts): Executes kimi-cli for Kimi vision analysis. This provider is gated behind explicit selection (-p kimi-cli) because it requires a paid subscription, unlike the default Antigravity provider.

CLI vs. API-Based Providers

ModLens distinguishes between CLI-based and API-based providers through their implementation patterns. API providers—Gemini API, OpenAI Compat, and Anthropic API—implement execute() methods that send multipart HTTP requests directly to cloud endpoints. In contrast, CLI-based providers run locally installed binaries, making them suitable for air-gapped environments or scenarios where avoiding network latency is critical.

When selecting a provider programmatically, you can detect CLI backends by checking for the absence of an execute method:

import { providers } from '@/providers';
import { ProviderName } from '@/providers/index';
import { execAsync } from 'child_process';

async function runProvider(
  providerName: ProviderName,
  imagePath: string,
) {
  const provider = providers[providerName];
  if ('execute' in provider) {
    // API-based provider (Gemini, OpenAI, Anthropic)
    return await provider.execute({ imagePath });
  } else {
    // CLI-based provider (Antigravity, Claude, Kimi)
    const cmd = provider.buildInvocation({ imagePath });
    const { stdout } = await execAsync(cmd);
    return provider.parseOutput(stdout);
  }
}

Practical Usage Examples

Invoke CLI-based providers using the ModLens command-line interface:


# Default Antigravity CLI (zero-config)

modlens -i screenshot.png

# Explicitly select the Claude CLI provider

modlens -i screenshot.png -p claude-cli

# Run the Kimi CLI (requires a subscription)

modlens -i screenshot.png -p kimi-cli

Summary

  • CLI-based providers in ModLens execute external binaries (antigravity-cli, claude-cli, kimi-cli) as subprocesses rather than using HTTP APIs
  • They implement buildInvocation() to construct shell commands and parseOutput() to extract JSON from stdout
  • Three CLI providers exist: Antigravity (default, zero-config), Claude CLI, and Kimi CLI (subscription-required)
  • All providers share the JSON schema contract defined in src/schema.ts for consistent LLM evidence output
  • The provider registry in src/providers/index.ts enables backend selection via the -p flag

Frequently Asked Questions

What is the difference between CLI-based providers and API providers in ModLens?

CLI-based providers spawn external command-line processes via buildInvocation() to analyze images locally, while API providers implement execute() methods that send HTTP requests to cloud endpoints like Gemini or OpenAI. CLI providers work offline but require locally installed binaries; API providers need internet connectivity and API keys but no local CLI setup.

How does the Antigravity CLI provider handle image input?

The Antigravity provider in src/providers/antigravity.ts passes image bytes to antigravity-cli either through temporary files or base-64 encoded pipes. The CLI then runs a zero-config vision model and returns JSON via stdout, which parseOutput() extracts and validates against src/schema.ts.

Why does the Kimi CLI provider require explicit selection?

The Kimi CLI provider (-p kimi-cli) requires a paid subscription to access Kimi's vision API. ModLens restricts this provider to explicit opt-in to prevent accidental usage of metered services, unlike the free-to-use Antigravity default that works immediately when you run modlens -i image.png.

How does ModLens ensure consistent JSON output across different CLI tools?

All CLI-based providers append the --json-schema flag (or equivalent) when building invocations in buildInvocation(), forcing the underlying model to return raw JSON. The parseOutput() method in each provider then validates this stdout content against the shared schema defined in src/schema.ts, ensuring every backend delivers identically structured evidence for downstream LLM processing.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →