# Characteristics of the Gemini-API Provider in ModLens: In-Process Vision Analysis

> Explore the gemini-api provider in ModLens, an in-process vision solution executing HTTP requests to Google's Gemini API with JSON schema validation and support for local/remote images.

- Repository: [liustack/modlens](https://github.com/liustack/modlens)
- Tags: deep-dive
- Published: 2026-08-25

---

**The gemini-api provider is ModLens's only in-process vision provider, executing HTTP requests directly against Google's Gemini Developer API within the Node.js process while enforcing strict JSON schema validation and supporting both local and remote image sources.**

The gemini-api provider in ModLens offers a tightly integrated, subprocess-free approach to computer vision by communicating directly with Google's Gemini Developer API. Unlike wrapper-based providers that spawn external processes, this implementation handles image encoding, prompt construction, and structured output enforcement entirely within the runtime. According to the liustack/modlens source code, the provider supports advanced configuration options including custom "thinking" parameters, automatic quota error recovery, and zero-latency API key rotation.

## Core Architecture: In-Process Execution Without Subprocesses

The gemini-api provider distinguishes itself from other ModLens providers by implementing the `execute` method directly rather than the `buildInvocation`/`parseOutput` pattern. In [`src/providers/index.ts`](https://github.com/liustack/modlens/blob/main/src/providers/index.ts) (lines 48-56), the `VisionProvider` interface defines two implementation paths: subprocess-based providers return command-line arguments for external execution, while in-process providers like gemini-api handle the entire request lifecycle internally.

This architecture eliminates serialization overhead and process-spawning latency. The `executeGeminiApi` function in [`src/providers/geminiApi.ts`](https://github.com/liustack/modlens/blob/main/src/providers/geminiApi.ts) manages the complete flow from image base64 encoding to HTTP POST submission, running entirely within the current Node.js process.

## Default Model and API Key Management

By default, the provider targets `gemini-3.6-flash` as defined by the `GEMINI_API_DEFAULT_MODEL` constant in [`src/providers/geminiApi.ts`](https://github.com/liustack/modlens/blob/main/src/providers/geminiApi.ts) (lines 17-18). Configuration resolution follows ModLens's layered configuration system, reading from environment variables (`GEMINI_API_KEY`, `GEMINI_BASE_URL`) and the settings object.

API key handling occurs in lines 23-30, where the implementation:

- Validates that at least one key exists in the `apiKey` setting
- Redacts all supplied keys in error messages to prevent leakage
- Throws `ApiKeyFailureError` (defined in [`src/util/apiKeys.ts`](https://github.com/liustack/modlens/blob/main/src/util/apiKeys.ts)) if validation fails

## Dual Image Source Support and Prompt Construction

The provider accepts both local filesystem paths and remote URLs through dedicated utility functions. In [`src/providers/geminiApi.ts`](https://github.com/liustack/modlens/blob/main/src/providers/geminiApi.ts) (lines 35-38), the code branches based on `imageKind`, calling either `readLocalImageBase64` for filesystem access or `fetchRemoteImageBase64` for HTTP retrieval.

Prompt construction happens via `buildVisionPrompt` (lines 40-44), which generates the textual instruction appended to the image payload. This inline prompt construction ensures the vision analysis request contains both the visual data and the specific analysis instructions in a single API call.

## Strict Schema Enforcement for Structured Output

To guarantee type-safe results, the provider enforces `VISION_RESULT_SCHEMA` (defined in [`src/schema.ts`](https://github.com/liustack/modlens/blob/main/src/schema.ts)) through the `generationConfig.responseJsonSchema` parameter. In [`src/providers/geminiApi.ts`](https://github.com/liustack/modlens/blob/main/src/providers/geminiApi.ts) (lines 71-74), the implementation injects this schema into the request body, instructing the Gemini API to return JSON that strictly adheres to ModLens's expected output structure.

This structured output enforcement eliminates parsing ambiguity and ensures the returned object contains the required fields for downstream processing.

## Request Customization and Timeout Handling

Advanced users can inject custom parameters through the `extraBody` configuration option. Lines 56-82 in [`src/providers/geminiApi.ts`](https://github.com/liustack/modlens/blob/main/src/providers/geminiApi.ts) demonstrate how `mergeExtraBody` (from [`src/util/extraBody.ts`](https://github.com/liustack/modlens/blob/main/src/util/extraBody.ts)) merges user-provided configurations while preserving the required schema fields. This enables features like the "thinking knob" for controlling model reasoning depth:

```json
{
  "generationConfig": {
    "thinkingConfig": { "thinkingLevel": "LOW" }
  }
}

```

Timeout enforcement uses native `AbortSignal.timeout` (lines 85-86), respecting the CLI-level deadline without external process management.

## Error Classification and Quota Management

The provider implements sophisticated error handling documented in [`src/providers/geminiApi.test.ts`](https://github.com/liustack/modlens/blob/main/src/providers/geminiApi.test.ts). It distinguishes between generic API failures and specific `ApiKeyFailureError` conditions (lines 90-102). When the Gemini API returns a 402 quota error, the provider classifies this for automatic cooldown handling (lines 124-148), allowing ModLens to retry with exponential backoff or switch to alternative keys.

HTTP status codes and response bodies surface directly to the caller, providing transparent debugging information while maintaining security through key redaction.

## Provider Registration and Aliases

The provider registers itself in [`src/providers/index.ts`](https://github.com/liustack/modlens/blob/main/src/providers/index.ts) (lines 79-81) with the canonical name `gemini-api` and the shorthand alias `gemini`. The exported `VisionProvider` object exposes both the default model configuration and the `execute` function, making it discoverable via `resolveProvider('gemini')` or `resolveProvider('gemini-api')`.

## Usage Examples

### Command-Line Interface

Analyze a local screenshot using the default model:

```bash
modlens -i screenshot.png -p gemini-api

```

Override the model and enable low-level thinking for faster responses:

```bash
modlens -i screenshot.png -p gemini-api \
  --model gemini-1.5-pro \
  --extra-body '{"generationConfig":{"thinkingConfig":{"thinkingLevel":"LOW"}}}'

```

### Programmatic TypeScript Usage

```typescript
import { resolveProvider } from './providers/index.js';
import type { BuildProviderInvocationOptions } from './providers/index.js';

// Resolve using the alias
const provider = resolveProvider('gemini');

const opts: BuildProviderInvocationOptions = {
  imageSource: '/path/to/image.png',
  imageKind: 'local',
  timeoutMs: 15000,
  settings: { 
    apiKey: process.env.GEMINI_API_KEY,
    extraBody: {
      generationConfig: {
        thinkingConfig: { thinkingLevel: 'MEDIUM' }
      }
    }
  },
};

// Execute returns { result, meta }
const { result, meta } = await provider.execute!(opts);
console.log('Analysis:', result);
console.log('Token usage:', meta.usage);

```

## Summary

- **In-process execution**: Implements `execute` directly in [`src/providers/geminiApi.ts`](https://github.com/liustack/modlens/blob/main/src/providers/geminiApi.ts) instead of spawning subprocesses, reducing latency.
- **Default configuration**: Uses `gemini-3.6-flash` by default, configurable via `GEMINI_API_KEY` and `GEMINI_BASE_URL` environment variables.
- **Dual image support**: Handles both local filesystem paths (`readLocalImageBase64`) and remote URLs (`fetchRemoteImageBase64`) with automatic base64 encoding.
- **Schema enforcement**: Injects `VISION_RESULT_SCHEMA` into `generationConfig.responseJsonSchema` to guarantee structured JSON output.
- **Advanced customization**: Supports `extraBody` merging for experimental features like `thinkingConfig` while protecting required schema fields.
- **Resilient error handling**: Classifies 402 quota errors for automatic cooldown and distinguishes `ApiKeyFailureError` from general API failures.

## Frequently Asked Questions

### How does the gemini-api provider differ from other ModLens providers?

Unlike subprocess-based providers that implement `buildInvocation` and `parseOutput`, the gemini-api provider implements `execute` directly in [`src/providers/geminiApi.ts`](https://github.com/liustack/modlens/blob/main/src/providers/geminiApi.ts). This in-process architecture eliminates the overhead of spawning external processes and serializing data through stdin/stdout, resulting in lower latency and simpler error propagation within the Node.js runtime.

### What environment variables does the gemini-api provider support?

The provider recognizes `GEMINI_API_KEY` for authentication and `GEMINI_BASE_URL` for API endpoint customization. These values integrate with ModLens's layered configuration system defined in [`config.ts`](https://github.com/liustack/modlens/blob/main/config.ts), allowing overrides via CLI flags, configuration files, or environment variables.

### How does the provider ensure JSON schema compliance?

The implementation injects `VISION_RESULT_SCHEMA` (defined in [`src/schema.ts`](https://github.com/liustack/modlens/blob/main/src/schema.ts)) into the request's `generationConfig.responseJsonSchema` field at lines 71-74 of [`src/providers/geminiApi.ts`](https://github.com/liustack/modlens/blob/main/src/providers/geminiApi.ts). This instructs the Gemini API to return JSON strictly validated against ModLens's expected structure, eliminating post-processing parsing errors.

### What happens when the Gemini API returns a quota error?

When the API returns a 402 status code, the provider classifies this as a quota exhaustion error in the error handling logic documented in [`src/providers/geminiApi.test.ts`](https://github.com/liustack/modlens/blob/main/src/providers/geminiApi.test.ts) (lines 124-148). ModLens uses this classification to trigger automatic cooldown periods, enabling retry logic with exponential backoff or automatic failover to alternative API keys.