# FreeLLMAPI Gemini API Compatible Endpoints: Complete Developer Guide

> Explore FreeLLMAPI's Gemini API compatible endpoints, mirroring Google's official API. Access all features under /v1beta with this comprehensive developer guide.

- Repository: [Tashfeen/freellmapi](https://github.com/tashfeenahmed/freellmapi)
- Tags: developer-guide
- Published: 2026-08-30

---

**FreeLLMAPI implements a native Gemini surface that mirrors Google's official Gemini HTTP API, with all compatible endpoints served under `/v1beta` and handled by the `geminiRouter` in [`server/src/routes/gemini.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/routes/gemini.ts).**

FreeLLMAPI provides a fully-compatible **Gemini API interface** that allows developers to call Google's Gemini models—or automatically fall back to equivalent free alternatives—without requiring a Google Cloud project. The Gemini-compatible surface is built on Express routes that normalize requests into the internal chat pipeline while returning responses in Google's native wire format.

---

## Supported Gemini API Endpoints

All **Gemini API compatible endpoints** in FreeLLMAPI share the base path `/v1beta`. The router uses regex patterns to capture dynamic model names, allowing flexible routing for any Gemini model identifier.

### List Available Models

```http
GET /v1beta/models

```

Returns all Gemini models plus an `auto` fallback entry. The implementation in [`server/src/routes/gemini.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/routes/gemini.ts) (line 80) queries the provider registry and injects the synthetic `auto` model for router-selected fallback:

```typescript
// server/src/routes/gemini.ts#L80
geminiRouter.get('/models', async (req, res) => { ... })

```

### Get Model Metadata

```http
GET /v1beta/models/{model}

```

Retrieves configuration and capabilities for a specific model. The regex route `geminiRouter.get(/^\/models\/(.+)$/)` at line 95 handles both real Gemini models and the `auto` placeholder.

### Generate Content (Synchronous)

```http
POST /v1beta/models/{model}:generateContent

```

Primary chat completion endpoint for non-streaming requests. The route `geminiRouter.post(/^\/models\/(.+):generateContent$/)` at line 20 translates Gemini-format requests into the internal pipeline via [`server/lib/gemini-wire.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/server/lib/gemini-wire.ts).

### Stream Generate Content (Server-Sent Events)

```http
POST /v1beta/models/{model}:streamGenerateContent?alt=sse

```

Streaming variant required by Gemini CLI and compatible clients. The router at line 24 implements SSE formatting when `alt=sse` is present in the query string.

### Count Tokens

```http
POST /v1beta/models/{model}:countTokens

```

Estimates token consumption for a given request payload. Implemented at line 28 in [`gemini.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/gemini.ts) using the same translation layer as generation endpoints.

---

## Authentication Methods for Gemini Endpoints

FreeLLMAPI's **Gemini API compatible endpoints** accept three authentication patterns, implemented in the `authenticate` function (lines 47–60 of [`gemini.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/gemini.ts)):

| Method | Header/Format | Recommendation |
|--------|-------------|----------------|
| **Preferred** | `x-goog-api-key: <UNIFIED_KEY>` | Matches Google's official Gemini SDK |
| Bearer token | `Authorization: Bearer <UNIFIED_KEY>` | OpenAI-compatible client fallback |
| Query parameter | `?key=<UNIFIED_KEY>` | **Not recommended**—leaks to logs and browser history |

The query parameter option is restricted to `/v1beta` routes as a compatibility shim; header-based authentication prevents credential exposure in server logs and shell history.

---

## Request and Response Format

The **Gemini API compatible endpoints** use Google's standard wire protocol. All fields are defined in [`server/lib/gemini-wire.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/server/lib/gemini-wire.ts) and serialized through `geminiResponseFromResult`.

### Request Body Structure

```json
{
  "contents": [
    {
      "role": "user",
      "parts": [{ "text": "Explain quantum computing" }]
    }
  ],
  "generationConfig": {
    "maxOutputTokens": 8192,
    "temperature": 0.7,
    "topP": 0.9,
    "stopSequences": ["END"]
  },
  "tools": [
    {
      "googleSearch": {}
    }
  ],
  "toolConfig": {
    "toolChoice": "auto"
  }
}

```

The router normalizes these payloads for the internal provider abstraction while preserving Gemini semantics in responses.

---

## Code Examples for Gemini API Endpoints

### List Models with curl

```bash
curl "http://localhost:3001/v1beta/models" \
  -H "x-goog-api-key: YOUR_UNIFIED_KEY"

```

### Synchronous Generation (Python OpenAI SDK)

```python
import openai

client = openai.OpenAI(
    base_url="http://localhost:3001/v1beta",
    api_key="YOUR_UNIFIED_KEY"  # passed as x-goog-api-key header

)

response = client.chat.completions.create(
    model="gemini-2.5-flash",  # or "auto" for router-selected model

    messages=[{
        "role": "user",
        "content": "What is the capital of France?"
    }]
)
print(response.choices[0].message.content)

```

### Streaming with Server-Sent Events

```bash
curl "http://localhost:3001/v1beta/models/gemini-2.5-flash:streamGenerateContent?alt=sse" \
  -H "x-goog-api-key: YOUR_UNIFIED_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "contents":[{
      "role":"user",
      "parts":[{"text":"Tell me a joke"}]
    }]
  }'

```

### Token Counting

```bash
curl "http://localhost:3001/v1beta/models/gemini-2.5-flash:countTokens" \
  -H "x-goog-api-key: YOUR_UNIFIED_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "contents":[{
      "role":"user",
      "parts":[{"text":"Hello world"}]
    }]
  }'

```

### Using Google Search Grounding

```python
response = client.chat.completions.create(
    model="gemini-2.5-flash",
    messages=[{"role": "user", "content": "Latest AI breakthroughs 2024"}],
    tools=[{
        "googleSearch": {}
    }]
)

```

---

## Architecture and Fallback Behavior

The **Gemini API compatible endpoints** leverage FreeLLMAPI's generic provider system. The Google provider implementation in [`server/src/providers/google.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/providers/google.ts) (line 20) points to `https://generativelanguage.googleapis.com/v1beta` for native Gemini calls.

When a model quota is exhausted, the router automatically cascades to the next available free provider while maintaining Gemini-compatible parameter mapping. This fallback chain is documented in [`docs/architecture.md`](https://github.com/tashfeenahmed/freellmapi/blob/main/docs/architecture.md).

Key architectural files:

- **[`server/src/routes/gemini.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/routes/gemini.ts)** — Route definitions and authentication middleware
- **[`server/src/providers/google.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/providers/google.ts)** — Low-level Google API client
- **[`server/lib/gemini-wire.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/server/lib/gemini-wire.ts)** — Payload serialization/deserialization
- **[`docs/api.md`](https://github.com/tashfeenahmed/freellmapi/blob/main/docs/api.md)** — User-facing endpoint documentation

---

## Summary

- **Five core endpoints**: `/v1beta/models`, `:generateContent`, `:streamGenerateContent`, `:countTokens`, and model metadata retrieval
- **Three auth methods**: `x-goog-api-key` header (preferred), `Authorization: Bearer`, and `?key=` query parameter
- **Native wire format**: Full compatibility with Google's Gemini request/response schema via [`gemini-wire.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/gemini-wire.ts)
- **Automatic fallback**: Router chains to alternate providers when primary models hit limits
- **Dual compatibility**: Same endpoints work with OpenAI SDK clients through payload normalization

---

## Frequently Asked Questions

### How does FreeLLMAPI handle Gemini API authentication without a Google Cloud project?

FreeLLMAPI substitutes its own unified API key system for Google's project-based credentials. The `authenticate` function in [`server/src/routes/gemini.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/routes/gemini.ts) validates keys against the internal registry, then proxies requests through the generic provider layer. Your unified key grants access to Gemini models and the fallback pool without requiring GCP setup.

### Can I use the official Google Gemini SDK with FreeLLMAPI?

Yes. Point the SDK's `baseUrl` to your FreeLLMAPI instance's `/v1beta` path and provide your unified key as the `apiKey`. The router's native Gemini surface accepts the same request shapes and returns compatible responses, so existing code requires minimal changes.

### What happens when a Gemini model quota is exceeded?

The router's fallback chain activates automatically. According to [`docs/architecture.md`](https://github.com/tashfeenahmed/freellmapi/blob/main/docs/architecture.md), the system attempts the requested model first, then cascades through equivalent free alternatives while preserving your generation parameters. The response maintains Gemini format regardless of which underlying provider fulfills the request.

### Are streaming responses fully compatible with Gemini CLI tools?

Yes. The `:streamGenerateContent` endpoint recognizes `?alt=sse` for Server-Sent Events formatting, which Gemini CLI expects. The implementation in [`server/src/routes/gemini.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/routes/gemini.ts) (line 24) ensures chunk boundaries and event types match Google's specification for drop-in CLI compatibility.