# What Makes FreeLLMAPI's Gemini Endpoint Compatible with Vertex AI Patterns

> Discover how FreeLLMAPI's Gemini endpoint seamlessly integrates with Vertex AI patterns by mirroring request contracts, auth headers, and response schemas, enabling effortless endpoint switching.

- Repository: [Tashfeen/freellmapi](https://github.com/tashfeenahmed/freellmapi)
- Tags: architecture
- Published: 2026-08-31

---

**FreeLLMAPI's Gemini endpoint replicates Google's Vertex API request contracts, authentication headers, and response schemas, allowing clients to switch endpoints without code changes.**

The open-source **FreeLLMAPI** project (`tashfeenahmed/freellmapi`) implements a Gemini-compatible REST surface that translates incoming Vertex AI conventions into internal routing logic. This design ensures that SDKs, CLI tools, and applications built for Google Cloud's Vertex AI can point directly at a FreeLLMAPI instance and receive identical behavior for model selection, streaming, tool calling, and error handling.

## REST Path and Routing Conventions

Vertex AI uses a strict URL pattern for content generation: `v1beta/models/{model}:generateContent` for synchronous requests and `…:streamGenerateContent` for streaming. In [`server/src/routes/gemini.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/routes/gemini.ts), FreeLLMAPI exposes identical paths under the `/gemini` router:

- `/models/:model:generateContent`
- `/models/:model:streamGenerateContent`

The router extracts the model identifier using the same **actionModel** parameter parsing that Vertex AI expects. This means a client sending a request to `https://api.freellmapi.com/gemini/models/gemini-1.5-flash:generateContent` experiences the same path resolution as it would against Google's endpoint.

## Authentication with x-goog-api-key

Vertex AI requires the `x-goog-api-key` header for authentication. FreeLLMAPI's implementation in [`server/src/routes/gemini.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/routes/gemini.ts) accepts this exact header name to maintain header parity. 

The system also supports a `?key=` query string parameter as a compatibility escape hatch—matching Vertex AI's own fallback behavior—though it warns about URL leakage risks just as Google's documentation does.

## Model Resolution and Family Mapping

Rather than requiring exact model IDs, Vertex AI allows family aliases like `auto`, `pro`, or `flash`. FreeLLMAPI implements this logic in [`server/src/services/gemini-map.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/services/gemini-map.ts), which stores a **Gemini model map** that resolves family requests (`default`, `pro`, `flash`, `flashLite`) to concrete catalog IDs.

When the map entry specifies `auto`, the request remains unpinned—exactly how Vertex AI treats automatic model selection. This abstraction layer ensures that clients using generic model identifiers receive appropriate backend routing without manual configuration.

## Generation Configuration Schema

The **generationConfig** object in FreeLLMAPI mirrors Vertex AI's field names and defaults. The `GeminiInboundRequest` interface in [`server/src/lib/gemini-wire.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/lib/gemini-wire.ts) accepts:

- `temperature`
- `topP`
- `maxOutputTokens` (defaults to 8192 when omitted, matching Gemini's behavior)
- `stopSequences`

By preserving identical field names and default values, FreeLLMAPI ensures that tuning parameters sent to Vertex AI produce the same results when sent to its own endpoints.

## Tool Calling and System Instructions

FreeLLMAPI translates OpenAI-style tool definitions into Gemini's native wire format through utilities in [`server/src/lib/gemini-wire.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/lib/gemini-wire.ts):

- **`geminiToolsToChatTools`** converts function declarations
- **`geminiToolChoice`** handles the `functionCallingConfig` logic
- **`thoughtSignature`** validation reproduces the strict tool-call verification required for Gemini 3 models

For **system instructions**, the endpoint handles the `systemInstruction` field by folding it into the first user turn for Gemma models—a known Vertex AI quirk—and forwarding it unchanged for other models. This preserves the exact behavior that Vertex AI clients expect when sending system prompts.

## Response Structure and Error Handling

Vertex AI returns responses containing `candidates`, `usageMetadata`, and `modelVersion`. FreeLLMAPI constructs identical response bodies in [`server/src/routes/gemini.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/routes/gemini.ts), ensuring that client-side parsers find the expected fields:

- `candidates` array with `role` and `parts`
- `usageMetadata` with `promptTokenCount`, `candidatesTokenCount`, and `totalTokenCount`
- `modelVersion` string identifying the specific model used

Error handling uses the **`sendError`** utility to return Vertex AI-style JSON structures containing `code`, `message`, and `status` fields. This allows existing error-handling logic in client applications to function without modification.

## Token Counting Endpoint

Vertex AI provides a dedicated token estimation endpoint at `/models/{model}:countTokens`. FreeLLMAPI implements this in [`server/src/routes/gemini.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/routes/gemini.ts) via the `/models/:model:countTokens` route, returning a JSON object with `totalTokens` that matches the Vertex AI response schema.

## Practical Implementation Examples

The following examples demonstrate how existing Vertex AI clients can target FreeLLMAPI without syntax changes.

### Generate Content

```bash
curl -X POST "https://api.freellmapi.com/gemini/models/gemini-1.5-flash:generateContent?key=YOUR_API_KEY" \
  -H "Content-Type: application/json" \
  -H "x-goog-api-key: YOUR_API_KEY" \
  -d '{
    "contents": [
      { "role": "user", "parts": [{ "text": "Explain quantum tunneling in plain language." }] }
    ],
    "generationConfig": {
      "temperature": 0.7,
      "maxOutputTokens": 1024
    }
  }'

```

### Streaming Responses

```bash
curl -N "https://api.freellmapi.com/gemini/models/gemini-1.5-pro:streamGenerateContent?alt=sse&key=YOUR_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "contents": [{ "role": "user", "parts": [{ "text": "Write a haiku about rain." }] }]
  }'

```

### Count Tokens

```bash
curl -X POST "https://api.freellmapi.com/gemini/models/gemini-1.5-flash:countTokens?key=YOUR_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "contents": [{ "role": "user", "parts": [{ "text": "Hello world!" }] }]
  }'

```

## Summary

- **Path parity**: [`server/src/routes/gemini.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/routes/gemini.ts) implements the exact Vertex AI URL patterns for `generateContent`, `streamGenerateContent`, and `countTokens`.
- **Auth compatibility**: Accepts `x-goog-api-key` headers and `?key=` query parameters identical to Google's implementation.
- **Model abstraction**: [`server/src/services/gemini-map.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/services/gemini-map.ts) resolves family aliases (`auto`, `pro`, `flash`) to concrete IDs.
- **Wire format translation**: [`server/src/lib/gemini-wire.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/lib/gemini-wire.ts) handles `generationConfig`, tool calling (`geminiToolsToChatTools`), and system instruction folding for Gemma models.
- **Response fidelity**: Returns `candidates`, `usageMetadata`, and `modelVersion` fields with Vertex AI-style error JSON.
- **Drop-in replacement**: No client code changes required when switching base URLs from Vertex AI to FreeLLMAPI.

## Frequently Asked Questions

### Can I use the official Google Cloud SDK with FreeLLMAPI?

Yes. Because FreeLLMAPI implements the same REST path structure, authentication headers, and response schemas as Vertex AI, you can configure the Google Cloud SDK or Gemini CLI to use a FreeLLMAPI base URL. The SDK will function normally, sending `x-goog-api-key` headers and parsing the `candidates` and `usageMetadata` fields exactly as it would with Google's servers.

### How does FreeLLMAPI handle model aliases like "gemini-pro"?

The [`server/src/services/gemini-map.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/services/gemini-map.ts) module maintains a mapping between family names (`pro`, `flash`, `flashLite`, `default`) and specific model catalog IDs. When a request arrives with `models/gemini-pro`, the resolver looks up the corresponding full model identifier. If the configuration specifies `auto`, the system leaves the model selection unpinned, matching Vertex AI's automatic routing behavior.

### Does tool calling work the same way as in Vertex AI?

Yes. FreeLLMAPI's [`server/src/lib/gemini-wire.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/lib/gemini-wire.ts) file contains translation utilities (`geminiToolsToChatTools`, `geminiToolChoice`) that convert OpenAI-style function definitions into Gemini's `functionDeclarations` format. The endpoint also implements the `thoughtSignature` validation required for Gemini 3 tool calls, ensuring that multi-turn tool interactions behave identically to Vertex AI.

### What happens if I omit the maxOutputTokens parameter?

FreeLLMAPI applies the same default value that Vertex AI uses. According to the `generationConfig` handling in [`server/src/lib/gemini-wire.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/lib/gemini-wire.ts), when `maxOutputTokens` is omitted, the system defaults to 8192 tokens. This ensures that applications relying on Vertex AI's default token limits receive consistent behavior without explicitly setting the parameter in every request.