How to Use the Ollama‑Compatible Chat Endpoint in FreeLLMAPI

FreeLLMAPI exposes a fully Ollama‑compatible API at /ollama/api/chat, allowing any Ollama client to connect without code changes while routing requests through its internal chat pipeline.

The Ollama‑compatible chat endpoint in FreeLLMAPI enables seamless integration with existing Ollama tooling. Whether you're using the official ollama CLI, third‑party UI widgets, or custom scripts, you can point them at FreeLLMAPI's gateway and immediately access its unified model catalog, rate‑limiting, and usage tracking.

Where the Ollama Router Lives

The Ollama compatibility layer is implemented in server/src/routes/ollama.ts. This ollamaRouter is mounted in the main Express application (server/src/app.ts), exposing all endpoints under the /ollama path prefix.

When a request arrives at /ollama/api/chat, the router executes a six‑step processing pipeline before returning a response.

Request Processing Pipeline

Each chat request flows through these validated stages:

Step Implementation Source Location
Schema validation chatSchema (Zod) validates model, messages, tools, options, format, stream ollama.ts:190‑198
Model normalization normalizeOllamaModel strips ollama: prefix and validates catalog existence ollama.ts:88‑89
Message conversion ollamaMessages transforms Ollama‑style messages to internal ChatMessage[] ollama.ts:200‑214
Stream routing stream flag selects sendDelta (SSE lines) or sendNonStream (single JSON) ollama.ts:381‑403
Output formatting format parameter triggers json_object or json_schema wrapping ollama.ts:418‑424
Done reason mapping ollamaDoneReason translates internal states to Ollama‑compatible reasons ollama.ts:70‑72

Invalid payloads return HTTP 400 at the validation stage, preventing malformed requests from reaching downstream services.

Making Chat Requests

All examples assume FreeLLMAPI is running locally on port 3000.

Non‑Streaming Chat

curl -X POST http://localhost:3000/ollama/api/chat \
  -H "Content-Type: application/json" \
  -d '{
    "model": "gemma:2b",
    "messages": [
      {"role": "system", "content": "You are a helpful assistant."},
      {"role": "user", "content": "What is the capital of France?"}
    ],
    "stream": false
  }'

Setting stream: false returns a single JSON response with the complete generated text.

Streaming Responses

curl -N -X POST http://localhost:3000/ollama/api/chat \
  -H "Content-Type: application/json" \
  -d '{
    "model": "gemma:2b",
    "messages": [{"role": "user", "content": "Tell me a short joke."}],
    "stream": true
  }'

The -N flag disables curl's output buffering, letting you see Server‑Sent Events (SSE) lines as they arrive. Each line contains a partial message object with incremental content.

Structured JSON Output

curl -X POST http://localhost:3000/ollama/api/chat \
  -H "Content-Type: application/json" \
  -d '{
    "model": "gemma:2b",
    "messages": [{"role": "user", "content": "Give me a JSON object with fields name and age."}],
    "format": {"name": "string", "age": "integer"}
  }'

The format parameter accepts json for basic JSON mode or a JSON schema object for strict structured generation. FreeLLMAPI maps this to either json_object or json_schema internally.

Additional Ollama‑Compatible Endpoints

The ollamaRouter implements the full Ollama API surface:

Endpoint Purpose
GET /ollama/api/tags List available models
GET /ollama/api/version Return server version
POST /ollama/api/generate Text completion (non‑chat)
POST /ollama/api/embed Batch embedding generation
POST /ollama/api/embeddings Single embedding (legacy)

Embedding Example

curl -X POST http://localhost:3000/ollama/api/embeddings \
  -H "Content-Type: application/json" \
  -d '{
    "model": "nomic-embed-text",
    "input": ["FreeLLMAPI provides an Ollama‑compatible chat API."]
  }'

Key Configuration Files

Understanding these files helps with debugging and customization:

File Purpose
[server/src/routes/ollama.ts](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/routes/ollama.ts) Core router with all endpoint handlers
[server/src/app.ts](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/app.ts) Mounts ollamaRouter on /ollama
[shared/types.ts](https://github.com/tashfeenahmed/freellmapi/blob/main/shared/types.ts) Provider enum including 'ollama'
[server/src/services/provider-quota.ts](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/services/provider-quota.ts) Quota enforcement for Ollama provider
[server/src/docs/openapi.ts](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/docs/openapi.ts) Auto‑generated OpenAPI documentation

Summary

  • Endpoint: POST /ollama/api/chat on your FreeLLMAPI gateway
  • Compatibility: Drop‑in replacement for native Ollama servers
  • Streaming: Controlled via stream: true/false with proper SSE handling
  • Validation: Zod schemas enforce payload correctness
  • Extensions: JSON schema output, tool calls, and embedding endpoints all supported

Frequently Asked Questions

What clients work with FreeLLMAPI's Ollama endpoint?

Any client that speaks the Ollama protocol works without modification: the official ollama CLI, Ollama web UIs like Open WebUI, LangChain's Ollama integration, and custom HTTP clients. Point them at http://<your-gateway>/ollama instead of http://localhost:11434.

Does the Ollama‑compatible endpoint support tool calling?

Yes. The chatSchema accepts an optional tools array, and ollamaMessages handles tool call conversion in ollama.ts:200‑214. Tool responses flow back through the standard Ollama message format.

How does model naming work with the Ollama adapter?

FreeLLMAPI normalizes model names via normalizeOllamaModel, which strips the ollama: prefix and validates against the internal catalog. Request gemma:2b or ollama:gemma:2b—both resolve correctly if the model is configured in your gateway.

What happens when I hit a rate limit?

The provider-quota.ts service enforces quota limits for the Ollama provider. Exceeded limits return appropriate HTTP error codes with Ollama‑compatible error formatting, consistent with other FreeLLMAPI provider adapters.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →