What Makes FreeLLMAPI's Gemini Endpoint Compatible with Vertex AI Patterns
FreeLLMAPI's Gemini endpoint replicates Google's Vertex API request contracts, authentication headers, and response schemas, allowing clients to switch endpoints without code changes.
The open-source FreeLLMAPI project (tashfeenahmed/freellmapi) implements a Gemini-compatible REST surface that translates incoming Vertex AI conventions into internal routing logic. This design ensures that SDKs, CLI tools, and applications built for Google Cloud's Vertex AI can point directly at a FreeLLMAPI instance and receive identical behavior for model selection, streaming, tool calling, and error handling.
REST Path and Routing Conventions
Vertex AI uses a strict URL pattern for content generation: v1beta/models/{model}:generateContent for synchronous requests and …:streamGenerateContent for streaming. In server/src/routes/gemini.ts, FreeLLMAPI exposes identical paths under the /gemini router:
/models/:model:generateContent/models/:model:streamGenerateContent
The router extracts the model identifier using the same actionModel parameter parsing that Vertex AI expects. This means a client sending a request to https://api.freellmapi.com/gemini/models/gemini-1.5-flash:generateContent experiences the same path resolution as it would against Google's endpoint.
Authentication with x-goog-api-key
Vertex AI requires the x-goog-api-key header for authentication. FreeLLMAPI's implementation in server/src/routes/gemini.ts accepts this exact header name to maintain header parity.
The system also supports a ?key= query string parameter as a compatibility escape hatch—matching Vertex AI's own fallback behavior—though it warns about URL leakage risks just as Google's documentation does.
Model Resolution and Family Mapping
Rather than requiring exact model IDs, Vertex AI allows family aliases like auto, pro, or flash. FreeLLMAPI implements this logic in server/src/services/gemini-map.ts, which stores a Gemini model map that resolves family requests (default, pro, flash, flashLite) to concrete catalog IDs.
When the map entry specifies auto, the request remains unpinned—exactly how Vertex AI treats automatic model selection. This abstraction layer ensures that clients using generic model identifiers receive appropriate backend routing without manual configuration.
Generation Configuration Schema
The generationConfig object in FreeLLMAPI mirrors Vertex AI's field names and defaults. The GeminiInboundRequest interface in server/src/lib/gemini-wire.ts accepts:
temperaturetopPmaxOutputTokens(defaults to 8192 when omitted, matching Gemini's behavior)stopSequences
By preserving identical field names and default values, FreeLLMAPI ensures that tuning parameters sent to Vertex AI produce the same results when sent to its own endpoints.
Tool Calling and System Instructions
FreeLLMAPI translates OpenAI-style tool definitions into Gemini's native wire format through utilities in server/src/lib/gemini-wire.ts:
geminiToolsToChatToolsconverts function declarationsgeminiToolChoicehandles thefunctionCallingConfiglogicthoughtSignaturevalidation reproduces the strict tool-call verification required for Gemini 3 models
For system instructions, the endpoint handles the systemInstruction field by folding it into the first user turn for Gemma models—a known Vertex AI quirk—and forwarding it unchanged for other models. This preserves the exact behavior that Vertex AI clients expect when sending system prompts.
Response Structure and Error Handling
Vertex AI returns responses containing candidates, usageMetadata, and modelVersion. FreeLLMAPI constructs identical response bodies in server/src/routes/gemini.ts, ensuring that client-side parsers find the expected fields:
candidatesarray withroleandpartsusageMetadatawithpromptTokenCount,candidatesTokenCount, andtotalTokenCountmodelVersionstring identifying the specific model used
Error handling uses the sendError utility to return Vertex AI-style JSON structures containing code, message, and status fields. This allows existing error-handling logic in client applications to function without modification.
Token Counting Endpoint
Vertex AI provides a dedicated token estimation endpoint at /models/{model}:countTokens. FreeLLMAPI implements this in server/src/routes/gemini.ts via the /models/:model:countTokens route, returning a JSON object with totalTokens that matches the Vertex AI response schema.
Practical Implementation Examples
The following examples demonstrate how existing Vertex AI clients can target FreeLLMAPI without syntax changes.
Generate Content
curl -X POST "https://api.freellmapi.com/gemini/models/gemini-1.5-flash:generateContent?key=YOUR_API_KEY" \
-H "Content-Type: application/json" \
-H "x-goog-api-key: YOUR_API_KEY" \
-d '{
"contents": [
{ "role": "user", "parts": [{ "text": "Explain quantum tunneling in plain language." }] }
],
"generationConfig": {
"temperature": 0.7,
"maxOutputTokens": 1024
}
}'
Streaming Responses
curl -N "https://api.freellmapi.com/gemini/models/gemini-1.5-pro:streamGenerateContent?alt=sse&key=YOUR_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"contents": [{ "role": "user", "parts": [{ "text": "Write a haiku about rain." }] }]
}'
Count Tokens
curl -X POST "https://api.freellmapi.com/gemini/models/gemini-1.5-flash:countTokens?key=YOUR_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"contents": [{ "role": "user", "parts": [{ "text": "Hello world!" }] }]
}'
Summary
- Path parity:
server/src/routes/gemini.tsimplements the exact Vertex AI URL patterns forgenerateContent,streamGenerateContent, andcountTokens. - Auth compatibility: Accepts
x-goog-api-keyheaders and?key=query parameters identical to Google's implementation. - Model abstraction:
server/src/services/gemini-map.tsresolves family aliases (auto,pro,flash) to concrete IDs. - Wire format translation:
server/src/lib/gemini-wire.tshandlesgenerationConfig, tool calling (geminiToolsToChatTools), and system instruction folding for Gemma models. - Response fidelity: Returns
candidates,usageMetadata, andmodelVersionfields with Vertex AI-style error JSON. - Drop-in replacement: No client code changes required when switching base URLs from Vertex AI to FreeLLMAPI.
Frequently Asked Questions
Can I use the official Google Cloud SDK with FreeLLMAPI?
Yes. Because FreeLLMAPI implements the same REST path structure, authentication headers, and response schemas as Vertex AI, you can configure the Google Cloud SDK or Gemini CLI to use a FreeLLMAPI base URL. The SDK will function normally, sending x-goog-api-key headers and parsing the candidates and usageMetadata fields exactly as it would with Google's servers.
How does FreeLLMAPI handle model aliases like "gemini-pro"?
The server/src/services/gemini-map.ts module maintains a mapping between family names (pro, flash, flashLite, default) and specific model catalog IDs. When a request arrives with models/gemini-pro, the resolver looks up the corresponding full model identifier. If the configuration specifies auto, the system leaves the model selection unpinned, matching Vertex AI's automatic routing behavior.
Does tool calling work the same way as in Vertex AI?
Yes. FreeLLMAPI's server/src/lib/gemini-wire.ts file contains translation utilities (geminiToolsToChatTools, geminiToolChoice) that convert OpenAI-style function definitions into Gemini's functionDeclarations format. The endpoint also implements the thoughtSignature validation required for Gemini 3 tool calls, ensuring that multi-turn tool interactions behave identically to Vertex AI.
What happens if I omit the maxOutputTokens parameter?
FreeLLMAPI applies the same default value that Vertex AI uses. According to the generationConfig handling in server/src/lib/gemini-wire.ts, when maxOutputTokens is omitted, the system defaults to 8192 tokens. This ensures that applications relying on Vertex AI's default token limits receive consistent behavior without explicitly setting the parameter in every request.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →