FreeLLMAPI Gemini API Compatible Endpoints: Complete Developer Guide
FreeLLMAPI implements a native Gemini surface that mirrors Google's official Gemini HTTP API, with all compatible endpoints served under /v1beta and handled by the geminiRouter in server/src/routes/gemini.ts.
FreeLLMAPI provides a fully-compatible Gemini API interface that allows developers to call Google's Gemini models—or automatically fall back to equivalent free alternatives—without requiring a Google Cloud project. The Gemini-compatible surface is built on Express routes that normalize requests into the internal chat pipeline while returning responses in Google's native wire format.
Supported Gemini API Endpoints
All Gemini API compatible endpoints in FreeLLMAPI share the base path /v1beta. The router uses regex patterns to capture dynamic model names, allowing flexible routing for any Gemini model identifier.
List Available Models
GET /v1beta/models
Returns all Gemini models plus an auto fallback entry. The implementation in server/src/routes/gemini.ts (line 80) queries the provider registry and injects the synthetic auto model for router-selected fallback:
// server/src/routes/gemini.ts#L80
geminiRouter.get('/models', async (req, res) => { ... })
Get Model Metadata
GET /v1beta/models/{model}
Retrieves configuration and capabilities for a specific model. The regex route geminiRouter.get(/^\/models\/(.+)$/) at line 95 handles both real Gemini models and the auto placeholder.
Generate Content (Synchronous)
POST /v1beta/models/{model}:generateContent
Primary chat completion endpoint for non-streaming requests. The route geminiRouter.post(/^\/models\/(.+):generateContent$/) at line 20 translates Gemini-format requests into the internal pipeline via server/lib/gemini-wire.ts.
Stream Generate Content (Server-Sent Events)
POST /v1beta/models/{model}:streamGenerateContent?alt=sse
Streaming variant required by Gemini CLI and compatible clients. The router at line 24 implements SSE formatting when alt=sse is present in the query string.
Count Tokens
POST /v1beta/models/{model}:countTokens
Estimates token consumption for a given request payload. Implemented at line 28 in gemini.ts using the same translation layer as generation endpoints.
Authentication Methods for Gemini Endpoints
FreeLLMAPI's Gemini API compatible endpoints accept three authentication patterns, implemented in the authenticate function (lines 47–60 of gemini.ts):
| Method | Header/Format | Recommendation |
|---|---|---|
| Preferred | x-goog-api-key: <UNIFIED_KEY> |
Matches Google's official Gemini SDK |
| Bearer token | Authorization: Bearer <UNIFIED_KEY> |
OpenAI-compatible client fallback |
| Query parameter | ?key=<UNIFIED_KEY> |
Not recommended—leaks to logs and browser history |
The query parameter option is restricted to /v1beta routes as a compatibility shim; header-based authentication prevents credential exposure in server logs and shell history.
Request and Response Format
The Gemini API compatible endpoints use Google's standard wire protocol. All fields are defined in server/lib/gemini-wire.ts and serialized through geminiResponseFromResult.
Request Body Structure
{
"contents": [
{
"role": "user",
"parts": [{ "text": "Explain quantum computing" }]
}
],
"generationConfig": {
"maxOutputTokens": 8192,
"temperature": 0.7,
"topP": 0.9,
"stopSequences": ["END"]
},
"tools": [
{
"googleSearch": {}
}
],
"toolConfig": {
"toolChoice": "auto"
}
}
The router normalizes these payloads for the internal provider abstraction while preserving Gemini semantics in responses.
Code Examples for Gemini API Endpoints
List Models with curl
curl "http://localhost:3001/v1beta/models" \
-H "x-goog-api-key: YOUR_UNIFIED_KEY"
Synchronous Generation (Python OpenAI SDK)
import openai
client = openai.OpenAI(
base_url="http://localhost:3001/v1beta",
api_key="YOUR_UNIFIED_KEY" # passed as x-goog-api-key header
)
response = client.chat.completions.create(
model="gemini-2.5-flash", # or "auto" for router-selected model
messages=[{
"role": "user",
"content": "What is the capital of France?"
}]
)
print(response.choices[0].message.content)
Streaming with Server-Sent Events
curl "http://localhost:3001/v1beta/models/gemini-2.5-flash:streamGenerateContent?alt=sse" \
-H "x-goog-api-key: YOUR_UNIFIED_KEY" \
-H "Content-Type: application/json" \
-d '{
"contents":[{
"role":"user",
"parts":[{"text":"Tell me a joke"}]
}]
}'
Token Counting
curl "http://localhost:3001/v1beta/models/gemini-2.5-flash:countTokens" \
-H "x-goog-api-key: YOUR_UNIFIED_KEY" \
-H "Content-Type: application/json" \
-d '{
"contents":[{
"role":"user",
"parts":[{"text":"Hello world"}]
}]
}'
Using Google Search Grounding
response = client.chat.completions.create(
model="gemini-2.5-flash",
messages=[{"role": "user", "content": "Latest AI breakthroughs 2024"}],
tools=[{
"googleSearch": {}
}]
)
Architecture and Fallback Behavior
The Gemini API compatible endpoints leverage FreeLLMAPI's generic provider system. The Google provider implementation in server/src/providers/google.ts (line 20) points to https://generativelanguage.googleapis.com/v1beta for native Gemini calls.
When a model quota is exhausted, the router automatically cascades to the next available free provider while maintaining Gemini-compatible parameter mapping. This fallback chain is documented in docs/architecture.md.
Key architectural files:
server/src/routes/gemini.ts— Route definitions and authentication middlewareserver/src/providers/google.ts— Low-level Google API clientserver/lib/gemini-wire.ts— Payload serialization/deserializationdocs/api.md— User-facing endpoint documentation
Summary
- Five core endpoints:
/v1beta/models,:generateContent,:streamGenerateContent,:countTokens, and model metadata retrieval - Three auth methods:
x-goog-api-keyheader (preferred),Authorization: Bearer, and?key=query parameter - Native wire format: Full compatibility with Google's Gemini request/response schema via
gemini-wire.ts - Automatic fallback: Router chains to alternate providers when primary models hit limits
- Dual compatibility: Same endpoints work with OpenAI SDK clients through payload normalization
Frequently Asked Questions
How does FreeLLMAPI handle Gemini API authentication without a Google Cloud project?
FreeLLMAPI substitutes its own unified API key system for Google's project-based credentials. The authenticate function in server/src/routes/gemini.ts validates keys against the internal registry, then proxies requests through the generic provider layer. Your unified key grants access to Gemini models and the fallback pool without requiring GCP setup.
Can I use the official Google Gemini SDK with FreeLLMAPI?
Yes. Point the SDK's baseUrl to your FreeLLMAPI instance's /v1beta path and provide your unified key as the apiKey. The router's native Gemini surface accepts the same request shapes and returns compatible responses, so existing code requires minimal changes.
What happens when a Gemini model quota is exceeded?
The router's fallback chain activates automatically. According to docs/architecture.md, the system attempts the requested model first, then cascades through equivalent free alternatives while preserving your generation parameters. The response maintains Gemini format regardless of which underlying provider fulfills the request.
Are streaming responses fully compatible with Gemini CLI tools?
Yes. The :streamGenerateContent endpoint recognizes ?alt=sse for Server-Sent Events formatting, which Gemini CLI expects. The implementation in server/src/routes/gemini.ts (line 24) ensures chunk boundaries and event types match Google's specification for drop-in CLI compatibility.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →