FreeLLMAPI Gemini API Compatible Endpoints: Complete Developer Guide

FreeLLMAPI implements a native Gemini surface that mirrors Google's official Gemini HTTP API, with all compatible endpoints served under /v1beta and handled by the geminiRouter in server/src/routes/gemini.ts.

FreeLLMAPI provides a fully-compatible Gemini API interface that allows developers to call Google's Gemini models—or automatically fall back to equivalent free alternatives—without requiring a Google Cloud project. The Gemini-compatible surface is built on Express routes that normalize requests into the internal chat pipeline while returning responses in Google's native wire format.


Supported Gemini API Endpoints

All Gemini API compatible endpoints in FreeLLMAPI share the base path /v1beta. The router uses regex patterns to capture dynamic model names, allowing flexible routing for any Gemini model identifier.

List Available Models

GET /v1beta/models

Returns all Gemini models plus an auto fallback entry. The implementation in server/src/routes/gemini.ts (line 80) queries the provider registry and injects the synthetic auto model for router-selected fallback:

// server/src/routes/gemini.ts#L80
geminiRouter.get('/models', async (req, res) => { ... })

Get Model Metadata

GET /v1beta/models/{model}

Retrieves configuration and capabilities for a specific model. The regex route geminiRouter.get(/^\/models\/(.+)$/) at line 95 handles both real Gemini models and the auto placeholder.

Generate Content (Synchronous)

POST /v1beta/models/{model}:generateContent

Primary chat completion endpoint for non-streaming requests. The route geminiRouter.post(/^\/models\/(.+):generateContent$/) at line 20 translates Gemini-format requests into the internal pipeline via server/lib/gemini-wire.ts.

Stream Generate Content (Server-Sent Events)

POST /v1beta/models/{model}:streamGenerateContent?alt=sse

Streaming variant required by Gemini CLI and compatible clients. The router at line 24 implements SSE formatting when alt=sse is present in the query string.

Count Tokens

POST /v1beta/models/{model}:countTokens

Estimates token consumption for a given request payload. Implemented at line 28 in gemini.ts using the same translation layer as generation endpoints.


Authentication Methods for Gemini Endpoints

FreeLLMAPI's Gemini API compatible endpoints accept three authentication patterns, implemented in the authenticate function (lines 47–60 of gemini.ts):

Method Header/Format Recommendation
Preferred x-goog-api-key: <UNIFIED_KEY> Matches Google's official Gemini SDK
Bearer token Authorization: Bearer <UNIFIED_KEY> OpenAI-compatible client fallback
Query parameter ?key=<UNIFIED_KEY> Not recommended—leaks to logs and browser history

The query parameter option is restricted to /v1beta routes as a compatibility shim; header-based authentication prevents credential exposure in server logs and shell history.


Request and Response Format

The Gemini API compatible endpoints use Google's standard wire protocol. All fields are defined in server/lib/gemini-wire.ts and serialized through geminiResponseFromResult.

Request Body Structure

{
  "contents": [
    {
      "role": "user",
      "parts": [{ "text": "Explain quantum computing" }]
    }
  ],
  "generationConfig": {
    "maxOutputTokens": 8192,
    "temperature": 0.7,
    "topP": 0.9,
    "stopSequences": ["END"]
  },
  "tools": [
    {
      "googleSearch": {}
    }
  ],
  "toolConfig": {
    "toolChoice": "auto"
  }
}

The router normalizes these payloads for the internal provider abstraction while preserving Gemini semantics in responses.


Code Examples for Gemini API Endpoints

List Models with curl

curl "http://localhost:3001/v1beta/models" \
  -H "x-goog-api-key: YOUR_UNIFIED_KEY"

Synchronous Generation (Python OpenAI SDK)

import openai

client = openai.OpenAI(
    base_url="http://localhost:3001/v1beta",
    api_key="YOUR_UNIFIED_KEY"  # passed as x-goog-api-key header

)

response = client.chat.completions.create(
    model="gemini-2.5-flash",  # or "auto" for router-selected model

    messages=[{
        "role": "user",
        "content": "What is the capital of France?"
    }]
)
print(response.choices[0].message.content)

Streaming with Server-Sent Events

curl "http://localhost:3001/v1beta/models/gemini-2.5-flash:streamGenerateContent?alt=sse" \
  -H "x-goog-api-key: YOUR_UNIFIED_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "contents":[{
      "role":"user",
      "parts":[{"text":"Tell me a joke"}]
    }]
  }'

Token Counting

curl "http://localhost:3001/v1beta/models/gemini-2.5-flash:countTokens" \
  -H "x-goog-api-key: YOUR_UNIFIED_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "contents":[{
      "role":"user",
      "parts":[{"text":"Hello world"}]
    }]
  }'

Using Google Search Grounding

response = client.chat.completions.create(
    model="gemini-2.5-flash",
    messages=[{"role": "user", "content": "Latest AI breakthroughs 2024"}],
    tools=[{
        "googleSearch": {}
    }]
)

Architecture and Fallback Behavior

The Gemini API compatible endpoints leverage FreeLLMAPI's generic provider system. The Google provider implementation in server/src/providers/google.ts (line 20) points to https://generativelanguage.googleapis.com/v1beta for native Gemini calls.

When a model quota is exhausted, the router automatically cascades to the next available free provider while maintaining Gemini-compatible parameter mapping. This fallback chain is documented in docs/architecture.md.

Key architectural files:


Summary

  • Five core endpoints: /v1beta/models, :generateContent, :streamGenerateContent, :countTokens, and model metadata retrieval
  • Three auth methods: x-goog-api-key header (preferred), Authorization: Bearer, and ?key= query parameter
  • Native wire format: Full compatibility with Google's Gemini request/response schema via gemini-wire.ts
  • Automatic fallback: Router chains to alternate providers when primary models hit limits
  • Dual compatibility: Same endpoints work with OpenAI SDK clients through payload normalization

Frequently Asked Questions

How does FreeLLMAPI handle Gemini API authentication without a Google Cloud project?

FreeLLMAPI substitutes its own unified API key system for Google's project-based credentials. The authenticate function in server/src/routes/gemini.ts validates keys against the internal registry, then proxies requests through the generic provider layer. Your unified key grants access to Gemini models and the fallback pool without requiring GCP setup.

Can I use the official Google Gemini SDK with FreeLLMAPI?

Yes. Point the SDK's baseUrl to your FreeLLMAPI instance's /v1beta path and provide your unified key as the apiKey. The router's native Gemini surface accepts the same request shapes and returns compatible responses, so existing code requires minimal changes.

What happens when a Gemini model quota is exceeded?

The router's fallback chain activates automatically. According to docs/architecture.md, the system attempts the requested model first, then cascades through equivalent free alternatives while preserving your generation parameters. The response maintains Gemini format regardless of which underlying provider fulfills the request.

Are streaming responses fully compatible with Gemini CLI tools?

Yes. The :streamGenerateContent endpoint recognizes ?alt=sse for Server-Sent Events formatting, which Gemini CLI expects. The implementation in server/src/routes/gemini.ts (line 24) ensures chunk boundaries and event types match Google's specification for drop-in CLI compatibility.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →