# How FreeLLMAPI Emulates the Ollama API for Full Compatibility

> Discover how FreeLLMAPI emulates the Ollama API with a compatible shim. Learn about request validation and payload translation for seamless integration.

- Repository: [Tashfeen/freellmapi](https://github.com/tashfeenahmed/freellmapi)
- Tags: how-to-guide
- Published: 2026-08-31

---

**FreeLLMAPI provides an Ollama-compatible shim via the `ollamaRouter` in [`server/src/routes/ollama.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/routes/ollama.ts) that intercepts standard Ollama HTTP endpoints, validates requests through configurable emulation modes, and translates payloads into the internal FreeLLMAPI service layer.**

FreeLLMAPI emulates the Ollama API to allow existing Ollama clients to communicate with the gateway without code modifications. According to the tashfeenahmed/freellmapi source code, this compatibility layer is centralized in the `ollamaRouter` and leverages the same core services—such as `runInboundChat` and `runEmbeddings`—that power FreeLLMAPI's native OpenAI-compatible endpoints.

## Ollama Router Architecture and Entry Points

The emulation layer is built around the **`ollamaRouter`** defined in [[`server/src/routes/ollama.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/routes/ollama.ts)](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/routes/ollama.ts). This router mounts under the `/api/` path and exposes standard Ollama endpoints including `/api/tags`, `/api/chat`, `/api/generate`, `/api/embed`, and `/api/show`.

Each endpoint performs three operations:

1. **Authorization** via the `authorize` helper function
2. **Request validation** using Zod schemas (e.g., `embedSchema`)
3. **Delegation** to core services like `runInboundChat` (defined in [[`server/src/lib/inbound-chat.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/lib/inbound-chat.ts)](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/lib/inbound-chat.ts)) or `runEmbeddings` (defined in [[`server/src/services/embeddings.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/services/embeddings.ts)](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/services/embeddings.ts))

This architecture ensures that Ollama clients interact with FreeLLMAPI exactly as they would with a native Olloma server, while the gateway routes traffic through its unified model management, quota, and rate-limiting infrastructure.

## Configurable Authorization Modes

FreeLLMAPI supports three distinct emulation modes controlled by the `getOllamaEmulationMode` function. The active mode is stored in the settings database under the key **`ollama_emulation`** and defaults to `'off'` via the migration in [[`server/src/db/migrations/20260727_000001_agent_compat.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/db/migrations/20260727_000001_agent_compat.ts)](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/db/migrations/20260727_000001_agent_compat.ts).

| Mode | Behavior |
|------|----------|
| `off` | All Ollama routes return **404**; emulation is disabled |
| `open-loopback` | Only requests from local loopback addresses (127.0.0.1, ::1) are permitted |
| `key-required` | Requires a valid unified API key in the `Authorization: Bearer <key>` header |

The `authorize` helper checks this setting before every request by querying `getSetting('ollama_emulation')` from [[`server/src/db/index.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/db/index.ts)](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/db/index.ts), ensuring flexible deployment scenarios from fully open local development to secured production environments.

## Model Catalog and Inspection Endpoints

### Listing Available Models (`/api/tags`)

When an Ollama client requests the model catalog via `/api/tags`, FreeLLMAPI calls `buildModelListing()` from [[`server/src/services/model-listing.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/services/model-listing.ts)](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/services/model-listing.ts) to retrieve the available model index. The endpoint returns an "auto" placeholder model plus every configured model, each transformed by the `ollamaModel()` function into the Ollama-expected shape:

- `name` and `model` identifiers
- `modified_at` timestamp
- `size` and `digest` for versioning
- `details` containing architecture family, parameter size, and quantization level

All models report a synthetic `format` value of `freellmapi` to identify the source gateway.

### Model Details (`/api/show`)

The `/api/show` endpoint normalizes incoming model names using `normalizeOllamaModel()` and retrieves detailed metadata. For the special `auto` model, it returns generic capabilities including `completion` and `tools` support with a standard context window. For concrete models, the response includes `modelfile`, `parameters`, `template`, and `model_info` (e.g., `general.architecture`).

## Chat and Generate Request Translation

FreeLLMAPI handles Ollama's conversational endpoints through sophisticated payload transformation before delegating to the core chat orchestration service.

### Message and Tool Conversion

The **`ollamaMessages`** function rewrites Ollama-style message arrays—including multimodal content with vision images and tool call definitions—into the internal `ChatMessage` format. Similarly, **`ollamaTools`** maps Ollama tool definitions to FreeLLMAPI's native `ChatToolDefinition` structures, enabling function calling compatibility.

### Response Wire Formatting

Responses are formatted using **`ollamaWire`** (for `/api/chat`) and **`generateWire`** (for legacy `/api/generate`). These adapters emit newline-delimited JSON (ndjson) streams that match Ollama's streaming protocol exactly:

```typescript
// Conceptual implementation from ollamaWire
{
  model: request.model,
  created_at: new Date().toISOString(),
  message: { role: 'assistant', content: text },
  done: false,
  eval_count: tokenCount,
  // ... timing metrics via ollamaDurations
}

```

The adapters translate internal `finishReason` values to Ollama's `done_reason` field (values: `stop` or `length`) and fabricate timing metrics via `ollamaDurations` to satisfy clients expecting tokens-per-second statistics.

## Embeddings Endpoint Compatibility

Both `/api/embed` and the legacy `/api/embeddings` endpoints share the same implementation flow:

1. Authorization via the `authorize` helper
2. Request validation against `embedSchema`
3. Delegation to `runEmbeddings` for vector generation
4. Response reshaping to match Ollama's embedding schema

This ensures that vector retrieval works identically whether clients use the modern or legacy endpoint naming convention.

## Practical Usage Examples

The following `curl` commands demonstrate interacting with FreeLLMAPI's Ollama-compatible endpoints. Replace `<YOUR_UNIFIED_API_KEY>` with your actual API key when using `key-required` mode.

List available models through the `/api/tags` endpoint:

```bash
curl -X GET http://localhost:3000/api/tags \
  -H "Authorization: Bearer <YOUR_UNIFIED_API_KEY>"

```

Retrieve detailed model information via `/api/show`:

```bash
curl -X POST http://localhost:3000/api/show \
  -H "Content-Type: application/json" \
  -H "Authorization: Bearer <YOUR_UNIFIED_API_KEY>" \
  -d '{"model":"gpt-oss:120b"}'

```

Send a chat completion request to `/api/chat`:

```bash
curl -X POST http://localhost:3000/api/chat \
  -H "Content-Type: application/json" \
  -H "Authorization: Bearer <YOUR_UNIFIED_API_KEY>" \
  -d '{
        "model":"gpt-oss:120b",
        "messages":[{"role":"user","content":"Explain quantum tunnelling"}],
        "stream":false
      }'

```

Use the legacy generate endpoint at `/api/generate`:

```bash
curl -X POST http://localhost:3000/api/generate \
  -H "Content-Type: application/json" \
  -H "Authorization: Bearer <YOUR_UNIFIED_API_KEY>" \
  -d '{
        "model":"gpt-oss:120b",
        "prompt":"Write a haiku about winter",
        "stream":false
      }'

```

Generate embeddings through `/api/embed`:

```bash
curl -X POST http://localhost:3000/api/embed \
  -H "Content-Type: application/json" \
  -H "Authorization: Bearer <YOUR_UNIFIED_API_KEY>" \
  -d '{"model":"gpt-oss:120b","input":"Free LLMAPI"}'

```

## Summary

- **Centralized Router**: The `ollamaRouter` in [`server/src/routes/ollama.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/routes/ollama.ts) implements all Ollama-compatible HTTP endpoints under `/api/`
- **Flexible Security**: Three emulation modes (`off`, `open-loopback`, `key-required`) control access via the `ollama_emulation` setting stored in the database
- **Protocol Translation**: Functions like `ollamaMessages`, `ollamaTools`, and `ollamaWire` convert between Ollama and internal formats while preserving streaming semantics
- **Model Compatibility**: The `buildModelListing()` and `ollamaModel()` adapters in [`server/src/services/model-listing.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/services/model-listing.ts) present FreeLLMAPI's unified model catalog in Ollama-native format
- **Unified Backend**: Despite the Ollama facade, all requests flow through FreeLLMAPI's centralized backend for consistent quota management and rate limiting

## Frequently Asked Questions

### How do I enable Ollama API emulation in FreeLLMAPI?

Set the `ollama_emulation` setting to either `open-loopback` for local-only access or `key-required` for production use with API key authentication. The setting is stored in the database and defaults to `off` via the migration [`20260727_000001_agent_compat.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/20260727_000001_agent_compat.ts) to prevent unauthorized access.

### Which Ollama endpoints does FreeLLMAPI support?

FreeLLMAPI supports the complete Ollama endpoint set including `/api/tags` (model listing), `/api/show` (model details), `/api/chat` (conversational AI), `/api/generate` (legacy completions), and both `/api/embed` and `/api/embeddings` (vector generation), all implemented in [`server/src/routes/ollama.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/routes/ollama.ts).

### Does FreeLLMAPI support tool calling through the Ollama API?

Yes. The `ollamaTools` function translates Ollama tool definitions into FreeLLMAPI's internal `ChatToolDefinition` format, and `ollamaMessages` handles tool response messages. This enables function calling capabilities when using the `/api/chat` endpoint with compatible models.

### Can I use the standard Ollama CLI with FreeLLMAPI?

Yes. Point the Ollama CLI to your FreeLLMAPI instance URL instead of the default `localhost:11434`. When using `key-required` mode, configure the CLI to include the Bearer token in requests, or use environment variables mapped through [`server/src/lib/key-parser.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/lib/key-parser.ts) to authenticate requests.