How to Use FreeLLMAPI with Ollama Clients: Complete Integration Guide

FreeLLMAPI includes a built-in Ollama emulation layer that exposes native endpoints (/api/chat, /api/generate, /api/tags) at the router level, allowing any Ollama-compatible client to connect to the FreeLLMAPI base URL and route requests through the smart provider fallback system.

The tashfeenahmed/freellmapi repository provides a unified routing layer for free LLM APIs that natively emulates an Ollama server. By converting Ollama’s wire protocol into the internal InboundChatWire format, FreeLLMAPI enables standard Ollama clients to leverage intelligent rate-limit handling and automatic provider fallback without code changes.

Enabling Ollama Emulation in FreeLLMAPI

The emulation layer is controlled by the ollama_emulation setting stored in the settings table, managed in server/src/routes/settings.ts. This setting supports three operational modes:

  • off – Disables emulation; Ollama endpoints return 404.
  • open-loopback – Accepts unauthenticated requests from localhost.
  • key-required – Requires a valid FreeLLMAPI key for all Ollama endpoint access.

Activate the emulation layer by updating the setting via the REST API:

curl -X POST http://localhost:3001/v1/settings \
  -H "Authorization: Bearer <unified-api-key>" \
  -H "Content-Type: application/json" \
  -d '{"ollamaEmulation":"open-loopback"}'

Once enabled, the server in server/src/routes/ollama.ts initializes the Ollama-compatible route handlers.

Routing Architecture and Endpoint Compatibility

Core Ollama Router Implementation

The file server/src/routes/ollama.ts defines the NDJSON chat and generate endpoints, tag listing, and embedding routes. This router performs bidirectional translation between Ollama’s native request format and FreeLLMAPI’s internal InboundChatWire structure. When a client sends a request to /api/chat, the router:

  1. Parses the Ollama-style JSON payload.
  2. Converts it to InboundChatWire for internal processing.
  3. Selects the best available free provider based on rate limits and quotas.
  4. Translates the provider’s response back into Ollama’s wire format for the client.

Available Endpoints

The emulation layer exposes the following Ollama-compatible HTTP endpoints:

  • POST /api/chat – Streaming chat completions (NDJSON).
  • POST /api/generate – Text generation with streaming support.
  • GET /api/tags – Lists available models cataloged by the router.
  • POST /api/embeddings – Text embedding generation.

Configuring Ollama Clients to Use FreeLLMAPI

Command Line Configuration

Point any Ollama CLI tool at the FreeLLMAPI server by setting the OLLAMA_HOST environment variable:

export OLLAMA_HOST=http://localhost:3001

# Run a completion using the router's model selection

ollama run mistral "Explain quantum entanglement in one sentence."

Direct API Usage

You can interact with the emulated endpoints directly using curl:

List available models:

curl http://localhost:3001/api/tags

Streaming chat completion:

curl -X POST http://localhost:3001/api/chat \
  -H "Content-Type: application/json" \
  -d '{
        "model":"auto",
        "messages":[{"role":"user","content":"Write a haiku about AI."}]
      }' --no-buffer

Provider Routing and Observability

FreeLLMAPI automatically injects the X-Routed-Via header into every response, identifying which upstream provider actually served the request. This enables debugging and cost tracking when using Ollama clients.

Inspect the routing header:

curl -i -X POST http://localhost:3001/api/chat \
  -H "Content-Type: application/json" \
  -d '{"model":"auto","messages":[{"role":"user","content":"Hello"}]}' \
  | grep -i '^x-routed-via'

# Example output: X-Routed-Via: openrouter/gpt-4o

Alternative Integration Methods

Using the Custom Provider

FreeLLMAPI supports a generic custom provider mode described in the README that can forward requests to any OpenAI-compatible endpoint. This allows you to add a locally-running Ollama instance as a backend provider rather than using the emulation layer, useful when you need to chain multiple Ollama instances through the router.

Key Parsing for Ollama Providers

The file server/src/lib/key-parser.ts maps environment variable prefixes (OLLAMA_, OLLAMA_CLOUD_) to the internal platform name ollama. This allows you to register external Ollama endpoints from the Keys page by setting variables like OLLAMA_LOCAL_API_KEY and OLLAMA_LOCAL_BASE_URL, which the parser converts into valid provider configurations.

Summary

  • Enable emulation via the ollamaEmulation setting in server/src/routes/settings.ts using open-loopback or key-required modes.
  • Route implementation resides in server/src/routes/ollama.ts, which translates between Ollama wire format and InboundChatWire.
  • Client configuration requires only setting OLLAMA_HOST to the FreeLLMAPI base URL.
  • Observability is provided through the X-Routed-Via response header showing the actual upstream provider.
  • Environment mapping in server/src/lib/key-parser.ts supports OLLAMA_ prefixed variables for custom provider registration.

Frequently Asked Questions

Can I use FreeLLMAPI with existing Ollama CLI tools?

Yes. Set the OLLAMA_HOST environment variable to your FreeLLMAPI server address (e.g., http://localhost:3001). The Ollama CLI will automatically discover models and send requests through the FreeLLMAPI router without requiring configuration changes.

What Ollama endpoints are supported by the emulation layer?

The router supports /api/chat, /api/generate, /api/tags, and /api/embeddings. These endpoints handle streaming NDJSON responses for chat and generation, plus standard JSON responses for model listing and embeddings, matching native Ollama server behavior.

How does FreeLLMAPI handle authentication for Ollama clients?

Authentication depends on the ollama_emulation setting value. In open-loopback mode, local requests require no authentication. In key-required mode, clients must include a valid FreeLLMAPI key in the Authorization header. In off mode, Ollama endpoints are disabled entirely.

Is streaming supported when using Ollama clients with FreeLLMAPI?

Yes. The /api/chat and /api/generate endpoints return NDJSON streams identical to native Ollama servers. The router streams tokens from the upstream provider through to the client in real-time, handling provider-specific rate limits and fallbacks transparently during the stream.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →