How FreeLLMAPI Handles Vision Input for Multimodal LLMs: A Deep Dive into the Routing Pipeline

FreeLLMAPI detects image payloads via a hasImage flag in the proxy layer, validates against model metadata using supports_vision columns and VISION_ID_MARKERS, and explicitly rejects requests when no vision-capable model exists in the active fallback chain.

When integrating multimodal capabilities into a unified LLM gateway, routing vision-enabled requests only to models that can actually process images is critical. FreeLLMAPI implements a strict validation pipeline that inspects incoming payloads, queries model capabilities, and guarantees image data never reaches text-only endpoints. This architecture prevents silent failures and ensures proper handling of base64 or URL-based image inputs across the entire request lifecycle.

Detecting Vision Requirements in Incoming Requests

The vision handling pipeline begins at the edge. In server/src/routes/proxy.ts (lines 1602‑1604), the system scans every incoming JSON payload for image content blocks.

When a request arrives containing an image_url block or base64-encoded image data, the middleware sets a boolean hasImage flag:

// server/src/routes/proxy.ts – Vision detection logic
const hasImage = requestBody.messages?.some(m =>
  m.content?.some(c => c.type === 'image_url')
);

This flag propagates immediately into the internal FusionParams object as the vision property. By separating the vision flag from tool-use flags at the earliest stage, FreeLLMAPI maintains independent gating mechanisms for different modalities. The fusion.ts middleware (lines 560‑570) receives these parameters and treats the vision flag as a first-class constraint during panel selection.

Resolving Vision-Capable Models

Before routing occurs, the model discovery service (server/src/services/model-discovery.ts, lines 48‑92) determines which providers support image processing. The system employs multiple heuristics to identify vision capabilities:

  • Explicit boolean flags: Checks for vision or multimodal properties in provider metadata
  • Catalog column inspection: Queries the supports_vision integer column (0 = no, 1 = yes)
  • ID pattern matching: Scans model identifiers against VISION_ID_MARKERS including llava, internvl, and pixtral

The constants VISION_KEYS, VISION_MODALITIES, and VISION_ID_MARKERS centralize this detection logic, allowing the system to infer vision support even when upstream providers omit dedicated capability flags.

Routing Decisions and Request Validation

The router (server/src/services/router.ts, lines 1970‑2051) implements strict validation before dispatching requests. After building the fallback chain, the system checks each candidate model against the vision requirement:

// server/src/services/router.ts – Vision validation in routing loop
if (requirements.requireVision && !candidate.supportsVision) {
  dropped.push(`${id} (no vision support)`);
  continue;
}

If requireVision is true but no enabled entry has supports_vision === 1, the request fails immediately with a standardized error payload:

{
  "code": "no_vision_model",
  "message": "This request includes an image, but no vision-capable model is enabled."
}

This graceful degradation strategy ensures clients receive explicit feedback rather than silently falling back to text-only models that cannot process image data.

Persistent Model State Management

Vision capabilities persist in the database schema defined in server/src/services/custom-model-register.ts (lines 108‑119). The models table stores vision support as an integer defaulting to 0, requiring explicit opt-in for custom models:

// Custom model registration schema
vision: {
  type: 'integer',
  default: 0  // Must be explicitly set to 1 for vision support
}

When administrators register custom endpoints, they must deliberately set the vision column to 1 to enable image processing. This prevents accidental routing of multimodal traffic to incompatible self-hosted models.

Fallback Chain Pre-Flight Validation

Before executing the main routing logic, the getActiveChain() function performs a pre-flight check (server/src/services/router.ts, lines 2230‑2235). This validation ensures at least one enabled model in the chain advertises supports_vision === 1 for pure-vision requests.

By validating the entire fallback chain upfront, FreeLLMAPI avoids expensive downstream processing when no suitable provider exists. This early-exit pattern reduces latency for invalid requests and preserves computational resources for valid multimodal workloads.

Summary

  • Explicit Detection: The hasImage flag in proxy.ts triggers vision-specific routing logic before any model receives the payload.
  • Multi-Modal Discovery: model-discovery.ts uses VISION_ID_MARKERS and boolean flags to identify vision-capable models even with incomplete provider metadata.
  • Strict Validation: The router rejects requests immediately when requireVision is true but no candidate model has supports_vision === 1.
  • Database Persistence: The supports_vision column defaults to 0 in custom-model-register.ts, requiring explicit configuration for multimodal support.
  • Standardized Errors: All vision-related rejections return code: "no_vision_model", simplifying client-side error handling across /v1/chat/completions and other endpoints.

Frequently Asked Questions

How does FreeLLMAPI detect if a request contains an image?

The system inspects the JSON payload in server/src/routes/proxy.ts for any message containing a content block with type: 'image_url'. When found, it sets the hasImage boolean flag that propagates through the FusionParams object to downstream services.

What happens if no vision-capable models are enabled?

The request is rejected immediately with HTTP error code no_vision_model and the message: "This request includes an image, but no vision-capable model is enabled." This occurs in server/src/services/router.ts during the fallback chain validation phase.

How does the system identify vision-capable models during discovery?

The model discovery service (server/src/services/model-discovery.ts) checks three sources: explicit vision or multimodal boolean flags in provider metadata, the supports_vision database column, and pattern matching against VISION_ID_MARKERS like llava, internvl, or pixtral in the model ID string.

Can custom models be configured for vision support?

Yes. When registering custom models via server/src/services/custom-model-register.ts, administrators must explicitly set the vision column to 1. The schema defaults this value to 0, ensuring custom models do not accidentally receive image traffic unless properly configured.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →