# How to Use the Ollama‑Compatible Chat Endpoint in FreeLLMAPI

> Connect any Ollama client to FreeLLMAPI's Ollama compatible chat endpoint at /ollama/api/chat for seamless integration and routing through its internal chat pipeline. No code changes needed.

- Repository: [Tashfeen/freellmapi](https://github.com/tashfeenahmed/freellmapi)
- Tags: how-to-guide
- Published: 2026-08-30

---

**FreeLLMAPI exposes a fully Ollama‑compatible API at `/ollama/api/chat`, allowing any Ollama client to connect without code changes while routing requests through its internal chat pipeline.**

The **Ollama‑compatible chat endpoint** in FreeLLMAPI enables seamless integration with existing Ollama tooling. Whether you're using the official `ollama` CLI, third‑party UI widgets, or custom scripts, you can point them at FreeLLMAPI's gateway and immediately access its unified model catalog, rate‑limiting, and usage tracking.

## Where the Ollama Router Lives

The Ollama compatibility layer is implemented in **[`server/src/routes/ollama.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/routes/ollama.ts)**. This `ollamaRouter` is mounted in the main Express application ([`server/src/app.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/app.ts)), exposing all endpoints under the **`/ollama`** path prefix.

When a request arrives at `/ollama/api/chat`, the router executes a six‑step processing pipeline before returning a response.

## Request Processing Pipeline

Each chat request flows through these validated stages:

| Step | Implementation | Source Location |
|------|----------------|---------------|
| **Schema validation** | `chatSchema` (Zod) validates `model`, `messages`, `tools`, `options`, `format`, `stream` | [`ollama.ts:190‑198`](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/routes/ollama.ts#L190-L198) |
| **Model normalization** | `normalizeOllamaModel` strips `ollama:` prefix and validates catalog existence | [`ollama.ts:88‑89`](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/routes/ollama.ts#L88-L89) |
| **Message conversion** | `ollamaMessages` transforms Ollama‑style messages to internal `ChatMessage[]` | [`ollama.ts:200‑214`](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/routes/ollama.ts#L200-L214) |
| **Stream routing** | `stream` flag selects `sendDelta` (SSE lines) or `sendNonStream` (single JSON) | [`ollama.ts:381‑403`](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/routes/ollama.ts#L381-L403) |
| **Output formatting** | `format` parameter triggers `json_object` or `json_schema` wrapping | [`ollama.ts:418‑424`](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/routes/ollama.ts#L418-L424) |
| **Done reason mapping** | `ollamaDoneReason` translates internal states to Ollama‑compatible reasons | [`ollama.ts:70‑72`](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/routes/ollama.ts#L70-L72) |

Invalid payloads return **HTTP 400** at the validation stage, preventing malformed requests from reaching downstream services.

## Making Chat Requests

All examples assume FreeLLMAPI is running locally on port 3000.

### Non‑Streaming Chat

```bash
curl -X POST http://localhost:3000/ollama/api/chat \
  -H "Content-Type: application/json" \
  -d '{
    "model": "gemma:2b",
    "messages": [
      {"role": "system", "content": "You are a helpful assistant."},
      {"role": "user", "content": "What is the capital of France?"}
    ],
    "stream": false
  }'

```

Setting **`stream: false`** returns a single JSON response with the complete generated text.

### Streaming Responses

```bash
curl -N -X POST http://localhost:3000/ollama/api/chat \
  -H "Content-Type: application/json" \
  -d '{
    "model": "gemma:2b",
    "messages": [{"role": "user", "content": "Tell me a short joke."}],
    "stream": true
  }'

```

The `-N` flag disables curl's output buffering, letting you see **Server‑Sent Events (SSE)** lines as they arrive. Each line contains a partial `message` object with incremental `content`.

### Structured JSON Output

```bash
curl -X POST http://localhost:3000/ollama/api/chat \
  -H "Content-Type: application/json" \
  -d '{
    "model": "gemma:2b",
    "messages": [{"role": "user", "content": "Give me a JSON object with fields name and age."}],
    "format": {"name": "string", "age": "integer"}
  }'

```

The **`format`** parameter accepts `json` for basic JSON mode or a JSON schema object for strict structured generation. FreeLLMAPI maps this to either `json_object` or `json_schema` internally.

## Additional Ollama‑Compatible Endpoints

The `ollamaRouter` implements the full Ollama API surface:

| Endpoint | Purpose |
|----------|---------|
| `GET /ollama/api/tags` | List available models |
| `GET /ollama/api/version` | Return server version |
| `POST /ollama/api/generate` | Text completion (non‑chat) |
| `POST /ollama/api/embed` | Batch embedding generation |
| `POST /ollama/api/embeddings` | Single embedding (legacy) |

### Embedding Example

```bash
curl -X POST http://localhost:3000/ollama/api/embeddings \
  -H "Content-Type: application/json" \
  -d '{
    "model": "nomic-embed-text",
    "input": ["FreeLLMAPI provides an Ollama‑compatible chat API."]
  }'

```

## Key Configuration Files

Understanding these files helps with debugging and customization:

| File | Purpose |
|------|---------|
| [[`server/src/routes/ollama.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/routes/ollama.ts)](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/routes/ollama.ts) | Core router with all endpoint handlers |
| [[`server/src/app.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/app.ts)](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/app.ts) | Mounts `ollamaRouter` on `/ollama` |
| [[`shared/types.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/shared/types.ts)](https://github.com/tashfeenahmed/freellmapi/blob/main/shared/types.ts) | Provider enum including `'ollama'` |
| [[`server/src/services/provider-quota.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/services/provider-quota.ts)](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/services/provider-quota.ts) | Quota enforcement for Ollama provider |
| [[`server/src/docs/openapi.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/docs/openapi.ts)](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/docs/openapi.ts) | Auto‑generated OpenAPI documentation |

## Summary

- **Endpoint**: `POST /ollama/api/chat` on your FreeLLMAPI gateway
- **Compatibility**: Drop‑in replacement for native Ollama servers
- **Streaming**: Controlled via `stream: true/false` with proper SSE handling
- **Validation**: Zod schemas enforce payload correctness
- **Extensions**: JSON schema output, tool calls, and embedding endpoints all supported

## Frequently Asked Questions

### What clients work with FreeLLMAPI's Ollama endpoint?

Any client that speaks the Ollama protocol works without modification: the official `ollama` CLI, Ollama web UIs like Open WebUI, LangChain's Ollama integration, and custom HTTP clients. Point them at `http://<your-gateway>/ollama` instead of `http://localhost:11434`.

### Does the Ollama‑compatible endpoint support tool calling?

Yes. The `chatSchema` accepts an optional `tools` array, and `ollamaMessages` handles tool call conversion in [`ollama.ts:200‑214`](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/routes/ollama.ts#L200-L214). Tool responses flow back through the standard Ollama message format.

### How does model naming work with the Ollama adapter?

FreeLLMAPI normalizes model names via `normalizeOllamaModel`, which strips the `ollama:` prefix and validates against the internal catalog. Request `gemma:2b` or `ollama:gemma:2b`—both resolve correctly if the model is configured in your gateway.

### What happens when I hit a rate limit?

The [`provider-quota.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/provider-quota.ts) service enforces quota limits for the Ollama provider. Exceeded limits return appropriate HTTP error codes with Ollama‑compatible error formatting, consistent with other FreeLLMAPI provider adapters.