How ModLens Gives Text-Only LLMs the Ability to Analyze Images

ModLens converts images into structured JSON evidence that text-only LLMs can read, quote, and reason over by using vision-capable providers to process visual inputs and return strictly formatted text output.

ModLens is an open-source CLI tool that bridges the gap between visual data and text-only large language models. By encapsulating image processing within a provider-agnostic pipeline and normalizing results to a strict JSON schema, ModLens enables any text-based LLM to analyze images without native vision capabilities.

The Architecture: Converting Images to Structured Text

ModLens operates as an intermediary layer that handles all visual processing internally, exposing only text-based evidence to downstream consumers.

Input Resolution and Base64 Encoding

The pipeline begins in src/main.ts, which parses the -i/--input flag and determines whether the argument represents a local filesystem path or a remote HTTP URL. This value is handed to src/imageInput.ts, which handles the actual byte retrieval.

For local files, the system uses fs.readFileSync; for remote URLs, it uses fetch. The module determines the MIME type automatically and encodes the raw image data as base64. The resulting string attaches to the request payload under the provider-specific image field, formatted as data:<mime>;base64,<blob>.

The Provider Abstraction Layer

All vision backends implement the Provider contract defined in src/providers/index.ts. This interface standardizes how ModLens communicates with diverse vision services.

Five concrete providers exist:

  • Antigravity CLI (default zero-config subprocess)
  • Claude CLI
  • Kimi CLI
  • Gemini API
  • OpenAI-compatible API

Each provider implements two critical methods: buildInvocation (for subprocess providers) or execute (for in-process API providers) to construct requests, and parseOutput to handle raw responses. This abstraction allows ModLens to switch between local CLI tools and remote APIs without changing the core analysis logic.

The Vision Pipeline: From Request to Validated JSON

Once the image is encoded and a provider selected, ModLens constructs a prompt that forces structured output.

Prompt Construction with Schema Enforcement

The src/prompt.ts module assembles the final payload sent to vision models. This prompt includes:

  • A task description instructing the model to analyze the image
  • The base64-encoded image embedded according to provider expectations
  • A JSON schema reference passed via --json-schema (for Antigravity/Claude) or the responseJsonSchema field (for Gemini)

The canonical schema lives in src/schema.ts and defines exact required fields such as description, tags, and details. Providers enforce this schema strictly; any deviation triggers a validation error surfaced directly to the user.

Execution and Provider Selection

The analyzer (src/analyzer.ts) orchestrates the entire flow. It selects the appropriate provider based on CLI flags, user configuration (src/config.ts), and runtime availability checks (src/providers/availability.ts).

After the provider parses the model's response, the analyzer validates the JSON against src/schema.ts and stringifies the result. The CLI prints this JSON to STDOUT, meaning any downstream text-only LLM consuming the output receives pure text evidence that it can quote verbatim and reason over without image-processing capabilities.

Resilience and Rate Limiting

Production usage requires handling API quotas and failures gracefully.

Cooldown Management for Exhausted Providers

When a provider returns a quota-related error, the cooldown module (src/cooldown.ts) records the failure in ~/.modlens/state.json. Subsequent calls automatically avoid the exhausted provider until the cooldown period expires. This state persistence ensures reliable operation across multiple analysis sessions without manual intervention.

Practical Usage Examples

Analyze a local screenshot using the default provider:

modlens -i screenshot.png

Process a remote image via the Gemini API (fastest free route):

modlens -i https://example.com/photo.jpg -p gemini-api

Output raw JSON evidence for downstream LLM consumption:

modlens -i diagram.png --json

Programmatic usage within TypeScript projects:

import { analyze } from "./analyzer";

const result = await analyze({
  input: "https://example.com/graph.png",
  provider: "anthropic-api",
  json: true,
});

Summary

  • ModLens acts as a vision-to-text bridge by processing images internally and outputting only structured JSON that text-only LLMs can consume.
  • Provider abstraction in src/providers/index.ts enables support for five different vision backends (CLI tools and APIs) through a unified interface.
  • Strict schema enforcement via src/schema.ts ensures all providers return consistent, predictable JSON structures with fields like description and tags.
  • Base64 encoding in src/imageInput.ts handles both local files and remote URLs, making image input source-agnostic.
  • Cooldown tracking in src/cooldown.ts maintains provider reliability by persisting rate-limit state to ~/.modlens/state.json.

Frequently Asked Questions

How does ModLens handle different image formats?

ModLens automatically detects MIME types during the base64 encoding process in src/imageInput.ts. Whether the input is PNG, JPEG, WebP, or other common formats, the system reads the bytes, determines the correct content type, and embeds the data using the standard data:<mime>;base64,<blob> format that vision providers expect.

Can I use ModLens with custom vision models?

Yes, the provider architecture supports custom implementations. By implementing the Provider interface in src/providers/index.ts—specifically the buildInvocation or execute methods for request building and parseOutput for response handling—you can integrate any vision-capable service that accepts base64-encoded images and returns JSON.

What happens when a provider hits a rate limit?

The cooldown system in src/cooldown.ts intercepts quota-related errors and records them in ~/.modlens/state.json. Subsequent analysis calls automatically skip the exhausted provider and fallback to available alternatives. This state persists across CLI sessions, ensuring automatic recovery without manual configuration changes.

How does the JSON schema ensure consistent output?

The schema defined in src/schema.ts specifies mandatory fields and data types that every provider must return. When src/analyzer.ts receives a response, it validates the parsed JSON against this schema. Any missing fields or type mismatches trigger validation errors before the text reaches STDOUT, guaranteeing that downstream text-only LLMs receive well-structured evidence they can reliably parse.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →