Characteristics of the Gemini-API Provider in ModLens: In-Process Vision Analysis
The gemini-api provider is ModLens's only in-process vision provider, executing HTTP requests directly against Google's Gemini Developer API within the Node.js process while enforcing strict JSON schema validation and supporting both local and remote image sources.
The gemini-api provider in ModLens offers a tightly integrated, subprocess-free approach to computer vision by communicating directly with Google's Gemini Developer API. Unlike wrapper-based providers that spawn external processes, this implementation handles image encoding, prompt construction, and structured output enforcement entirely within the runtime. According to the liustack/modlens source code, the provider supports advanced configuration options including custom "thinking" parameters, automatic quota error recovery, and zero-latency API key rotation.
Core Architecture: In-Process Execution Without Subprocesses
The gemini-api provider distinguishes itself from other ModLens providers by implementing the execute method directly rather than the buildInvocation/parseOutput pattern. In src/providers/index.ts (lines 48-56), the VisionProvider interface defines two implementation paths: subprocess-based providers return command-line arguments for external execution, while in-process providers like gemini-api handle the entire request lifecycle internally.
This architecture eliminates serialization overhead and process-spawning latency. The executeGeminiApi function in src/providers/geminiApi.ts manages the complete flow from image base64 encoding to HTTP POST submission, running entirely within the current Node.js process.
Default Model and API Key Management
By default, the provider targets gemini-3.6-flash as defined by the GEMINI_API_DEFAULT_MODEL constant in src/providers/geminiApi.ts (lines 17-18). Configuration resolution follows ModLens's layered configuration system, reading from environment variables (GEMINI_API_KEY, GEMINI_BASE_URL) and the settings object.
API key handling occurs in lines 23-30, where the implementation:
- Validates that at least one key exists in the
apiKeysetting - Redacts all supplied keys in error messages to prevent leakage
- Throws
ApiKeyFailureError(defined insrc/util/apiKeys.ts) if validation fails
Dual Image Source Support and Prompt Construction
The provider accepts both local filesystem paths and remote URLs through dedicated utility functions. In src/providers/geminiApi.ts (lines 35-38), the code branches based on imageKind, calling either readLocalImageBase64 for filesystem access or fetchRemoteImageBase64 for HTTP retrieval.
Prompt construction happens via buildVisionPrompt (lines 40-44), which generates the textual instruction appended to the image payload. This inline prompt construction ensures the vision analysis request contains both the visual data and the specific analysis instructions in a single API call.
Strict Schema Enforcement for Structured Output
To guarantee type-safe results, the provider enforces VISION_RESULT_SCHEMA (defined in src/schema.ts) through the generationConfig.responseJsonSchema parameter. In src/providers/geminiApi.ts (lines 71-74), the implementation injects this schema into the request body, instructing the Gemini API to return JSON that strictly adheres to ModLens's expected output structure.
This structured output enforcement eliminates parsing ambiguity and ensures the returned object contains the required fields for downstream processing.
Request Customization and Timeout Handling
Advanced users can inject custom parameters through the extraBody configuration option. Lines 56-82 in src/providers/geminiApi.ts demonstrate how mergeExtraBody (from src/util/extraBody.ts) merges user-provided configurations while preserving the required schema fields. This enables features like the "thinking knob" for controlling model reasoning depth:
{
"generationConfig": {
"thinkingConfig": { "thinkingLevel": "LOW" }
}
}
Timeout enforcement uses native AbortSignal.timeout (lines 85-86), respecting the CLI-level deadline without external process management.
Error Classification and Quota Management
The provider implements sophisticated error handling documented in src/providers/geminiApi.test.ts. It distinguishes between generic API failures and specific ApiKeyFailureError conditions (lines 90-102). When the Gemini API returns a 402 quota error, the provider classifies this for automatic cooldown handling (lines 124-148), allowing ModLens to retry with exponential backoff or switch to alternative keys.
HTTP status codes and response bodies surface directly to the caller, providing transparent debugging information while maintaining security through key redaction.
Provider Registration and Aliases
The provider registers itself in src/providers/index.ts (lines 79-81) with the canonical name gemini-api and the shorthand alias gemini. The exported VisionProvider object exposes both the default model configuration and the execute function, making it discoverable via resolveProvider('gemini') or resolveProvider('gemini-api').
Usage Examples
Command-Line Interface
Analyze a local screenshot using the default model:
modlens -i screenshot.png -p gemini-api
Override the model and enable low-level thinking for faster responses:
modlens -i screenshot.png -p gemini-api \
--model gemini-1.5-pro \
--extra-body '{"generationConfig":{"thinkingConfig":{"thinkingLevel":"LOW"}}}'
Programmatic TypeScript Usage
import { resolveProvider } from './providers/index.js';
import type { BuildProviderInvocationOptions } from './providers/index.js';
// Resolve using the alias
const provider = resolveProvider('gemini');
const opts: BuildProviderInvocationOptions = {
imageSource: '/path/to/image.png',
imageKind: 'local',
timeoutMs: 15000,
settings: {
apiKey: process.env.GEMINI_API_KEY,
extraBody: {
generationConfig: {
thinkingConfig: { thinkingLevel: 'MEDIUM' }
}
}
},
};
// Execute returns { result, meta }
const { result, meta } = await provider.execute!(opts);
console.log('Analysis:', result);
console.log('Token usage:', meta.usage);
Summary
- In-process execution: Implements
executedirectly insrc/providers/geminiApi.tsinstead of spawning subprocesses, reducing latency. - Default configuration: Uses
gemini-3.6-flashby default, configurable viaGEMINI_API_KEYandGEMINI_BASE_URLenvironment variables. - Dual image support: Handles both local filesystem paths (
readLocalImageBase64) and remote URLs (fetchRemoteImageBase64) with automatic base64 encoding. - Schema enforcement: Injects
VISION_RESULT_SCHEMAintogenerationConfig.responseJsonSchemato guarantee structured JSON output. - Advanced customization: Supports
extraBodymerging for experimental features likethinkingConfigwhile protecting required schema fields. - Resilient error handling: Classifies 402 quota errors for automatic cooldown and distinguishes
ApiKeyFailureErrorfrom general API failures.
Frequently Asked Questions
How does the gemini-api provider differ from other ModLens providers?
Unlike subprocess-based providers that implement buildInvocation and parseOutput, the gemini-api provider implements execute directly in src/providers/geminiApi.ts. This in-process architecture eliminates the overhead of spawning external processes and serializing data through stdin/stdout, resulting in lower latency and simpler error propagation within the Node.js runtime.
What environment variables does the gemini-api provider support?
The provider recognizes GEMINI_API_KEY for authentication and GEMINI_BASE_URL for API endpoint customization. These values integrate with ModLens's layered configuration system defined in config.ts, allowing overrides via CLI flags, configuration files, or environment variables.
How does the provider ensure JSON schema compliance?
The implementation injects VISION_RESULT_SCHEMA (defined in src/schema.ts) into the request's generationConfig.responseJsonSchema field at lines 71-74 of src/providers/geminiApi.ts. This instructs the Gemini API to return JSON strictly validated against ModLens's expected structure, eliminating post-processing parsing errors.
What happens when the Gemini API returns a quota error?
When the API returns a 402 status code, the provider classifies this as a quota exhaustion error in the error handling logic documented in src/providers/geminiApi.test.ts (lines 124-148). ModLens uses this classification to trigger automatic cooldown periods, enabling retry logic with exponential backoff or automatic failover to alternative API keys.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →